top of page

The Research

Our professors have spent decades in linguistics, researching and publishing findings. This is where FlexSR is born, the next generation of speech recognition.

Most speech recognition works by prediction. A model trained on very large amounts of audio guesses the most probable sequence of words, then writes it down. It has no representation of the sounds a speaker actually produced, which is why it struggles with names, jargon, dialect and any language it has not been trained on. FlexSR works the other way round. Human speech can be described by a small, fixed set of phonological features, and those features are shared across every language. Matching against features rather than predicting words means the system needs no acoustic model per language, no retraining to add vocabulary, and far less compute. It runs offline on hardware as small as a Raspberry Pi. That claim rests on decades of work by our inventor team at Oxford and Frankfurt.

The papers below show examples of the research, before bringing into a FlexSR solution.

Title

Distinctive features: Phonological underspecification in representation and processing

Citation
Lahiri, A. and Reetz, H. Journal of Phonetics, 38(1), 44–59. 2010. doi:10.1016/j.wocn.2010.01.002

What it shows
This is the model FlexSR is built on. Words are stored in the mental lexicon as sparse sets of features rather than full acoustic detail, and incoming speech is matched against them three ways: match, mismatch, or no-mismatch. The point of the third category is that an imperfect signal does not rule out the right word, while a genuinely conflicting one is excluded. Backed by behavioural priming, EEG and MEG evidence that human listeners work this way.

The Paper

https://doi.org/10.1016/j.wocn.2010.01.002

Title

Asymmetric Phonological Representations of Words in the Mental Lexicon

Citation

Lahiri, Aditi (2012). Asymmetric phonological representations of words in the mental lexicon. In A. Cohn et al. (eds) The Oxford Handbook of Laboratory Phonology, 146–161. Oxford University Press.

What it shows

A survey of how lexical and phonological representations are investigated, arguing that a word is best stored as a single abstract underlying form from which every surface pronunciation can be derived. Storing one surface variant instead limits what contrasts can be encoded, because the information that distinguishes words is spread across the variants rather than sitting in any one of them.

The Paper

https://doi.org/10.1093/oxfordhb/9780199575039.013.0008

Title

Height differences in English dialects: consequences for processing and representation

Citation
Scharinger, M. and Lahiri, A. Language and Speech, 53(2), 245-272. 2010.

What it shows

Examines how vowel height differences across English dialects are represented and processed.

The Paper

https://journals.sagepub.com/doi/10.1177/0023830909357154

Title

Crosslinguistic acoustic categorization of sibilants independent of phonological status.​

Citation
Citation  Evers, V., Reetz, H. and Lahiri, A. Journal of Phonetics, 26(4), 345-370. 1998.

What it shows

Cross-language evidence that the same acoustic categories hold regardless of how a language uses them phonologically

The Paper

https://doi.org/10.1006/JPHO.1998.0079

Title

Phonological feature-based speech recognition system for pronunciation training in non-native language learning

Citation

Arora, V., Lahiri, A. and Reetz, H. Journal of the Acoustical Society of America, 143(1), 98-108. 2018.

What it shows

A working feature-based recogniser applied to pronunciation assessment.

The Paper

https://doi.org/10.1121/1.5017834

Title

Phonological feature based mispronunciation detection and diagnosis using multi-task DNNs and active learning

 

Citation

Arora, V., Lahiri, A. and Reetz, H. Interspeech 2017, 1350-1353.

What it shows

Detecting and diagnosing mispronunciation from phonological features.

The Paper

https://doi.org/10.21437/Interspeech.2017-1350

Request the full papers

Several of these papers sit behind publisher paywalls. Contact us for full access, and we will send what we are permitted to share.

© 2035 by Boost360. Powered and secured by Wix

bottom of page