

The Research
Our professors have spent decades in linguistics, researching and publishing findings. This is where FlexSR is born, the next generation of speech recognition.
Most speech recognition works by prediction. A model trained on very large amounts of audio guesses the most probable sequence of words, then writes it down. It has no representation of the sounds a speaker actually produced, which is why it struggles with names, jargon, dialect and any language it has not been trained on. FlexSR works the other way round. Human speech can be described by a small, fixed set of phonological features, and those features are shared across every language. Matching against features rather than predicting words means the system needs no acoustic model per language, no retraining to add vocabulary, and far less compute. It runs offline on hardware as small as a Raspberry Pi. That claim rests on decades of work by our inventor team at Oxford and Frankfurt.
The papers below show examples of the research, before bringing into a FlexSR solution.
Title
Distinctive features: Phonological underspecification in representation and processing
Citation
Lahiri, A. and Reetz, H. Journal of Phonetics, 38(1), 44–59. 2010. doi:10.1016/j.wocn.2010.01.002
What it shows
This is the model FlexSR is built on. Words are stored in the mental lexicon as sparse sets of features rather than full acoustic detail, and incoming speech is matched against them three ways: match, mismatch, or no-mismatch. The point of the third category is that an imperfect signal does not rule out the right word, while a genuinely conflicting one is excluded. Backed by behavioural priming, EEG and MEG evidence that human listeners work this way.
The Paper
https://doi.org/10.1016/j.wocn.2010.01.002
Title
Asymmetric Phonological Representations of Words in the Mental Lexicon
Citation
Lahiri, Aditi (2012). Asymmetric phonological representations of words in the mental lexicon. In A. Cohn et al. (eds) The Oxford Handbook of Laboratory Phonology, 146–161. Oxford University Press.
What it shows
A survey of how lexical and phonological representations are investigated, arguing that a word is best stored as a single abstract underlying form from which every surface pronunciation can be derived. Storing one surface variant instead limits what contrasts can be encoded, because the information that distinguishes words is spread across the variants rather than sitting in any one of them.
The Paper
https://doi.org/10.1093/oxfordhb/9780199575039.013.0008
Title
Height differences in English dialects: consequences for processing and representation
Citation
Scharinger, M. and Lahiri, A. Language and Speech, 53(2), 245-272. 2010.
What it shows
Examines how vowel height differences across English dialects are represented and processed.
The Paper
https://journals.sagepub.com/doi/10.1177/0023830909357154
Title
Crosslinguistic acoustic categorization of sibilants independent of phonological status.
Citation
Citation Evers, V., Reetz, H. and Lahiri, A. Journal of Phonetics, 26(4), 345-370. 1998.
What it shows
Cross-language evidence that the same acoustic categories hold regardless of how a language uses them phonologically
The Paper
https://doi.org/10.1006/JPHO.1998.0079
Title
Phonological feature-based speech recognition system for pronunciation training in non-native language learning
Citation
Arora, V., Lahiri, A. and Reetz, H. Journal of the Acoustical Society of America, 143(1), 98-108. 2018.
What it shows
A working feature-based recogniser applied to pronunciation assessment.
The Paper
https://doi.org/10.1121/1.5017834
Title
Phonological feature based mispronunciation detection and diagnosis using multi-task DNNs and active learning
Citation
Arora, V., Lahiri, A. and Reetz, H. Interspeech 2017, 1350-1353.
What it shows
Detecting and diagnosing mispronunciation from phonological features.
The Paper
https://doi.org/10.21437/Interspeech.2017-1350
Request the full papers
Several of these papers sit behind publisher paywalls. Contact us for full access, and we will send what we are permitted to share.