Systems and methods configured to analyze acoustic parameters of speech to detect, diagnose, predict, and / or monitor the progression of a condition, disorder, or disease
A real-time system for analyzing speech formant frequencies on mobile devices addresses memory and privacy concerns, enabling immediate detection and monitoring of neurological conditions by classifying vowels.
Patent Information
- Application Number
- JP2025512099
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-31
- Filing Date
- 2023-08-30
- Publication Date
- 2025-08-28
AI Technical Summary
Existing systems for analyzing acoustic parameters of speech to detect, diagnose, and monitor neurological and central nervous system conditions face memory depletion, data privacy concerns, and lack of real-time feedback due to processing stored audio files, and are unable to provide immediate results.
A system that extracts formant frequencies from speech in real-time using a mobile device, classifies vowels based on these frequencies, and applies acoustic metrics to generate score data for detecting, diagnosing, and monitoring conditions without recording the entire speech, thereby reducing memory requirements and addressing privacy concerns.
Enables real-time detection, diagnosis, and monitoring of neurological conditions by classifying vowels using formant frequencies, providing immediate feedback and reducing memory and privacy issues.
Smart Images

Figure 2025528447000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to systems and methods configured to analyze acoustic parameters of speech to detect, diagnose, predict, and / or monitor the progression of any of the following conditions, disorders, or diseases, more particularly pediatric and adult neurological and central nervous system conditions, including, but not limited to, back pain, multiple sclerosis, stroke, epileptic seizures, Alzheimer's disease, Parkinson's disease, dementia, motor neuron diseases, muscular atrophy, acquired brain injury, cancers with neurological deficits, childhood developmental conditions, and rare genetic disorders such as spinal muscular atrophy. [Background technology]
[0002] The speech signal contains measurable acoustic parameters related to speech production, and by analyzing the acoustics of the speech, aspects related to motor speech function can be inferred.
[0003] Additionally, acoustic analysis of human speech, and specifically, measuring an individual's vowel articulation dysregulation, can facilitate detecting the presence, severity, and characteristics of motor speech ataxias associated with conditions or diseases such as Parkinson's disease or Alzheimer's disease and / or the pediatric and adult neurological and central nervous system conditions listed above. By analyzing the level of vowel articulation dysregulation in an individual's speech, medical professionals can also monitor the deterioration or improvement of speech due to disease-related progression, recovery, or treatment effects.
[0004] Formants, also known as harmonics, are frequency domain features of speech, which are concentrations of acoustic energy around particular frequencies in a speech signal or waveform. Formants F1 (Hz) and F2 (Hz), and optionally F3 (Hz), are generally considered to be frequency acoustic parameters associated with the perception and production of vowels present in words pronounced by an individual.
[0005] Because individuals with vocal tract narrowing conditions and / or disorders often have impaired range of speech movement (due to the vocal tract narrowing problem), their ability to pronounce vowels in words is often affected. By measuring the frequencies of the F1, F2, and optionally F3 vowel formants in individuals with such conditions and / or disorders and comparing them to the F1 and F2 formant frequencies of individuals who speak normally, the level of this impairment can be measured and inferences can be drawn regarding the severity and treatment progress of individuals with such conditions and / or disorders.
[0006] The frequencies of such formants can be extracted from speech waveforms and measured in a variety of ways, such as using well-known computer software packages for speech analysis in phonetics, such as PRAAT.
[0007] Speech and vowel articulation disorders are therefore well-known digital biomarkers of such conditions and / or diseases.
[0008] However, systems and methods configured to perform such acoustic analysis on human speech suffer from drawbacks, including the need to process saved and therefore stored audio files that encode the human speech to be post-analyzed, which constantly depletes memory resources, raises data privacy concerns, and precludes feedback during training. Furthermore, such systems and methods are unable to provide results in near real time.
[0009] It is therefore an object of the present invention to provide a system and method that overcomes, in at least some way, the above-mentioned problems and / or provides a useful alternative to the public and / or industry.
[0010] Further aspects of the present invention will become apparent from the following description, which is given by way of example only.
[0011] According to the present invention, there is provided a system configured to analyze acoustic parameters of speech to detect, diagnose, predict and / or monitor the progression of a condition, disorder or disease, the system comprising a first computing device configured with means for receiving an audio stream comprising speech data encoding at least one word spoken by an individual, the first computing device comprising: means for converting the audio stream into a speech signal; means for extracting from the speech signal a first formant data set comprising formant frequencies associated with letters in at least one word as the audio stream is received in near real time without recording speech data encoding at least one word spoken by the individual; means for identifying formant frequencies associated with one or more vowel characters in a word from a first formant data set; means for identifying specific vowel character(s) from the identified vowel character formant frequencies; means for recording a second formant data set including at least some of the identified vowel character formant frequencies of the specific vowel character(s); configured to have The system further comprises means executing on the first computing device and / or a second computing device connected to the first computing device by a network for generating score data by applying one or more predetermined acoustic metrics to the recorded second data set of formants, the score data being used to identify a pronunciation level of the vowel letter(s) in at least one word spoken by the individual; The system further comprises means for storing the score data as an output file.
[0012] The present invention is directed to systems and methods for analyzing vowel productions in speech as they relate to clinical manifestations of disorders, conditions, or diseases in individuals suffering from any of the following neurological and central nervous system conditions in children and adults: back pain, multiple sclerosis, stroke, epileptic seizures, Alzheimer's disease, Parkinson's disease, dementia, motor neuron disease, muscular atrophy, acquired brain injury, cancer with neurological deficits, childhood developmental conditions, and rare genetic disorders such as spinal muscular atrophy.
[0013] The present invention provides a computer-implemented system and method for determining whether an individual pronounces vowels in spoken words in order to detect the presence of disease and monitor speech deterioration or improvement due to disease-related progression, recovery, or treatment effects.
[0014] The present invention may be used to predict, diagnose, and / or identify disease progression.
[0015] The system extracts a first formant data set from words spoken by an individual on a first computing device, such as a mobile smartphone equipped with a microphone, into which the individual speaks and uses these to classify vowels within the words, stores at least some of these frequencies of the vowel formants as a recording file in a second formant data set, provides the second formant data set as input to acoustic metrics to generate score data, and evaluates the score data to identify the pronunciation levels of the vowels in the words spoken by the individual, thereby enabling the detection, diagnosis, prediction, and / or monitoring of the progression of a condition, disorder, or disease.
[0016] For example, if an individual says the word "apple" into a microphone of a first computing device, the system is configured to extract, on the first computing device, the vowel formants present in the word spoken by the individual, record at least some of these as a second formant dataset, and use the second formant dataset as input to acoustic metrics to assess whether the individual pronounced the vowels "a" and "e" in the word "apple," which helps to understand whether the individual is saying the important vowel portions of the word.
[0017] Score data is generated as a report on pronunciation levels, which may enable disease detection, remediation, and assessment. The score data may be generated on the first computing device and / or a second computing device connected to the first computing device over a network.
[0018] Thus, the present invention is configured to classify vowels in words spoken by an individual based on formant frequencies in near real time, without the need to record files containing the individual's original speech, since only files containing the necessary vowel formant frequency data are recorded. This configuration significantly reduces memory requirements, delays in providing results, and addresses data privacy concerns, since only vowel formant frequency data extracted from the audio stream of the individual's speech is recorded.
[0019] Preferably, the formant frequencies in the first formant data set include at least formant frequencies F1 (Hz) and F2 (Hz).
[0020] Alternatively, the formant frequencies in the first formant data set include at least formant frequencies F1 (Hz), F2 (Hz), and F3 (Hz).
[0021] Thus, formant frequencies F1 (Hz), F2 (Hz), and optionally F3 (Hz) are used to classify vowels within the word(s) spoken by the individual.
[0022] Preferably, the formant frequencies in the second formant data set include formant frequencies F1 (Hz) and F2 (Hz).
[0023] Thus, using the formant frequencies F1 (Hz) and F2 (Hz) as input to the acoustic metrics, the pronunciation levels of the vowels in the word(s) spoken by the individual are identified.
[0024] Preferably, the first computing device is a mobile computing device, such as a mobile smart phone, having mobile phone and computing capabilities.
[0025] Preferably, the means for extracting the first formant data set comprises: - speech analysis means configured with means for converting the audio signal from a time domain signal to a frequency domain signal, for example by using a Fast Fourier Transform (FFT) algorithm; linear predictive coding means having means for applying an autocorrelation algorithm to estimate a dominant frequency in the frequency domain signal, and means for applying a Levinson-Durbin algorithm to estimate linear prediction parameters for the estimated dominant frequency, compress the frequency domain signal, and identify signal peaks in the compressed frequency domain signal; means for restoring the compressed frequency domain signal; means for extracting the identified signal peaks from the reconstructed frequency domain signal; Including, The extracted signal peaks correspond to formant frequencies in the first formant data set.
[0026] Preferably, the means for identifying formant frequencies associated with one or more vowel characters from the first formant dataset and for identifying the specific vowel characters includes means, at the first computing device, for applying a Mahalanobis distance algorithm to the extracted formant frequencies.
[0027] Preferably, the predetermined metrics include one or more of a Formant Centering Ratio (FCR) algorithm, a Vowel Space Algorithm (VSA), and a Vowel Articulation Index (VAI) algorithm.
[0028] According to a further aspect of the invention, there is provided a method configured to analyze acoustic parameters of speech to detect, predict, and / or monitor progression of a condition, disorder, or disease, the method comprising using a first computing device configured with means for receiving an audio stream comprising speech data encoding at least one word spoken by an individual, the first computing device comprising: converting the audio stream into a speech signal; extracting a first formant data set from the speech signal as the audio stream is received in near real time without recording speech data encoding at least one word spoken by the individual, the first formant data set comprising formant frequencies associated with letters in at least one word; identifying formant frequencies associated with one or more vowel characters in the word from a first formant data set; Identifying specific vowel character(s) from the identified vowel character formant frequencies; recording a second formant data set including at least some of the identified vowel character formant frequencies of the specific vowel character(s); configured to run The method further includes generating score data by applying one or more predetermined acoustic metrics to the recorded second data set of formants at the first computing device and / or a second computing device connected to the first computing device via a network, the score data being used to identify a pronunciation level of the vowel character(s) in at least one word spoken by the individual; The method further includes storing the score data as an output file.
[0029] Preferably, the formant frequencies in the first formant data set include at least formant frequencies F1 (Hz) and F2 (Hz).
[0030] Alternatively, the formant frequencies in the first formant data set include at least formant frequencies F1 (Hz), F2 (Hz), and F3 (Hz).
[0031] Thus, formant frequencies F1 (Hz), F2 (Hz), and optionally F3 (Hz) are used to classify vowels within the word(s) spoken by the individual.
[0032] Preferably, the formant frequencies in the second formant data set include formant frequencies F1 (Hz) and F2 (Hz).
[0033] Thus, using the formant frequencies F1 (Hz) and F2 (Hz) as input to the acoustic metrics, the pronunciation levels of the vowels in the word(s) spoken by the individual are identified.
[0034] Preferably, the first computing device is a mobile computing device, such as a mobile smart phone, having mobile phone and computing capabilities.
[0035] Preferably, the step of extracting the first formant data set comprises: converting the audio signal from a time domain signal to a frequency domain signal, for example by using a Fast Fourier Transform (FFT) algorithm; applying an autocorrelation algorithm to estimate a dominant frequency in the frequency domain signal, and applying a Levinson-Durbin algorithm to estimate linear prediction parameters of the estimated dominant frequency, compressing the frequency domain signal, and identifying signal peaks in the compressed frequency domain signal; restoring the compressed frequency domain signal; extracting the identified signal peaks from the reconstructed frequency domain signal; Including, The extracted signal peaks correspond to formant frequencies in the first formant data set.
[0036] Preferably, the method includes applying, at the first computing device, a Mahalanobis distance algorithm to the extracted formant frequencies to identify formant frequencies associated with one or more vowel characters from the first formant data set and identifying the specific vowel characters.
[0037] Preferably, the predetermined metrics include one or more of a Formant Centering Ratio (FCR) algorithm, a Vowel Space Algorithm (VSA), and a Vowel Articulation Index (VAI) algorithm.
[0038] The invention will be more clearly understood from the following description of some embodiments of the invention, given by way of example only, with reference to the accompanying drawings, in which: [Brief explanation of the drawings]
[0039] [Figure 1] FIG. 2 is a flow diagram illustrating the processing steps of a method performed by a system constructed in accordance with the present invention. [Figure 2] FIG. 2 is a flow diagram illustrating the processing steps performed to extract formant data from an audio signal in accordance with the present invention. [Figure 3]FIG. 3 is a flow diagram showing the processing steps performed to identify and classify vowels in the formant data extracted in FIG. 2. [Figure 4] 1 is a graph of cluster data used to classify vowel characters showing formant frequency F2 (Hz) versus formant frequency F1 (Hz). [Figure 5] FIG. 1 is a flow diagram illustrating processing steps for using acoustic metrics to determine whether an individual has pronounced a vowel letter in a word. DETAILED DESCRIPTION OF THE INVENTION
[0040] Referring to the drawings, and initially to FIG. 1 , there is provided a method configured to analyze acoustic parameters of speech to detect, diagnose, predict, and / or monitor the progression of a condition, ataxia, or disease, including any of pediatric and adult neurological and central nervous system conditions, including, but not limited to, back pain, multiple sclerosis, stroke, epileptic seizures, Alzheimer's disease, Parkinson's disease, dementia, motor neuron disease, muscular atrophy, acquired brain injury, cancer with neurological deficits, childhood developmental conditions, and rare genetic ataxias, such as spinal muscular atrophy.
[0041] The method includes using a first computing device 1 configured with means, e.g., a microphone, for receiving an audio stream including speech data encoding at least one word spoken by an individual 2.
[0042] The first computing device 1 is provided as a mobile smartphone configured to perform the step of converting a received audio stream into an audio signal 3 .
[0043] A formant data extraction process 4 is performed to extract a first formant data set 5 from the acquired speech signal 3. The first formant data set 5 comprises formant frequencies associated with letters in the word(s) contained in the audio stream speech signal 3.
[0044] As shown in FIG. 2, the step for extracting the first formant data set 5 involves converting the speech signal 3 from a time domain signal to a frequency domain signal 11, which may be achieved, for example, by using a Fast Fourier Transform (FFT) algorithm 10 before applying various Linear Predictive Coding (LPC) filters.
[0045] The Linear Predictive Coding (LPC) step involves applying an autocorrelation algorithm 12 to the frequency domain signal 11 to estimate dominant frequencies in the frequency domain signal 11. Then, a Levinson-Durbin algorithm 13 is applied to estimate linear prediction parameters for the estimated dominant frequencies, compress the frequency domain signal, and identify signal peaks in the compressed frequency domain signal.
[0046] Next, in step 14, the compressed frequency domain signal is decompressed and the identified signal peaks are extracted from the decompressed frequency domain signal using peak detection means.
[0047] The extracted signal peaks correspond to formant frequencies in the first formant data set 5 .
[0048] In the example shown in FIG. 2, the extracted first formant data set 5 includes formant frequencies F1 (Hz), F2 (Hz), and F3 (Hz), where F1=574.21 Hz, F2=1210.93 Hz, and F3=1898.43 Hz.
[0049] However, although the extracted first formant data set 5 of FIG. 2 shows formant frequencies F1 (Hz), F2 (Hz), and F3 (Hz), it will be appreciated that extraction of only formant frequencies F1 (Hz) and F2 (Hz) may be used to identify the vowels spoken by an individual.
[0050] In step 6 of FIG. 1, the vowels spoken or pronounced by the individual are identified by analyzing the formant frequencies in the first formant data set 5.
[0051] As shown in FIG. 3, this stage includes applying a Mahalanobis distance algorithm to the extracted formant frequencies at a first computing device to identify formant frequencies associated with one or more vowel characters from the first formant data set and to identify the specific vowel characters.
[0052] As shown in Figure 4, in a graph 40 of formant frequency F2 (Hz) against formant frequency F1 (Hz) data, each vowel letter typically has frequencies clustered or distributed around a particular zone or region. For example, the cluster of frequencies for the vowel letter "A" is generally designated by reference numeral 41, the cluster of frequencies for the vowel letter "E" is generally designated by reference numeral 42, etc. Analyzing which vowel letters an individual spoke using the distributions shown in Figure 4 involves steps 30, 31 of classifying the formant frequencies 5 according to these known vowel frequency clusters or distributions 41, 42.
[0053] In step 32, the Mahalanobis distance algorithm is then used to measure the distance between the extracted formant frequencies F1 (Hz) and F2 (Hz) and optionally F3 (Hz) and the cluster distribution of Figure 4, and the output from the Mahalanobis distance algorithm makes it possible to identify each vowel in step 33 and to calculate the individual's mean formant frequency values for each vowel, for example, Vowel A - average: F1: 800Hz, F2: 1450Hz, and Vowel A - average: F1: 875Hz, F2: 1490Hz, and similarly for all other vowel letters.
[0054] Formant frequency data not associated with a vowel letter is discarded.
[0055] As shown in FIG. 5, as the individual practices saying or pronouncing selected words into the microphone of the first computing device, a second formant data set is recorded at step 50, which includes the average formant frequencies of the individual's specific vowel letter(s) (calculated in step 33 of FIG. 3).
[0056] In step 51, a recorded second formant data set 50 is recorded, which includes the individual's mean formant frequency values F1 (Hz) and F2 (Hz) for each vowel.
[0057] Next, the step of generating score data by applying one or more predetermined acoustic metrics to the recorded second formant data set 50 is performed on the first computing device 1 (i.e., the personal mobile phone) and / or on a second computing device 52 communicatively connected to the first computing device 1 and providing back-end computing capabilities.
[0058] As shown in step 7 of Figure 1, the predetermined metrics include one or more of a formant centering ratio (FCR) algorithm, a vowel space algorithm (VSA), and a vowel pronunciation index (VAI) algorithm. In the example shown in Figure 5, the back-end second computing device 52 applies a vowel space algorithm (VSA) 53 and a vowel pronunciation index (VAI) algorithms 54, 55.
[0059] 1, an evaluation or inference may be made (automatically or via manual clinician interpretation) based on the generated score data to verify and check whether the individual is producing the expected vowels. The score data, if calculated on the second computing device 52, may be transmitted as an output file over the network to the first computing device 1, where feedback may be provided to the individual.
[0060] The score data may be used to generate reports on pronunciation levels, which may enable detection, diagnosis, prediction, and / or monitoring of symptoms, disorders, or disease progression.
[0061] It will be understood that the invention is not limited to the specific details set forth herein, which are given by way of example only, and that various modifications and alterations are possible without departing from the scope of the invention.
Claims
1. 1. A system configured to analyze acoustic parameters of speech to detect, diagnose, predict, and / or monitor the progression of a condition, disorder, or disease, the system comprising: a first computing device configured with means for receiving an audio stream including speech data encoding at least one word spoken by an individual, the first computing device comprising: means for converting the audio stream into a speech signal; means for extracting from the speech signal a first formant data set comprising formant frequencies associated with letters in the at least one word as the audio stream is received in near real time without recording the speech data encoding the at least one word spoken by the individual; and means for identifying formant frequencies associated with one or more vowel characters in the word from the first formant data set; means for identifying specific vowel character(s) from the identified vowel character formant frequencies; means for recording a second formant data set including at least some of the identified vowel character formant frequencies of the specific vowel character(s); configured to have the system further comprises means executing on the first computing device and / or a second computing device connected to the first computing device by a network for generating score data by applying one or more predetermined acoustic metrics to the recorded second data set of formants, the score data being used to identify a pronunciation level of the vowel character(s) in the at least one word spoken by the individual; The system further comprises means for storing the score data as an output file. The system.
2. The system of claim 1 , wherein the first computing device is a mobile computing device such as a mobile smart phone having cellular and computing capabilities.
3. The means for extracting the first formant data set comprises: - speech analysis means configured to have means for transforming said audio signal from a time domain signal to a frequency domain signal, for example by using a Fast Fourier Transform (FFT) algorithm; linear predictive coding means having means for applying an autocorrelation algorithm to estimate a dominant frequency in the frequency domain signal, and means for applying a Levinson-Durbin algorithm to estimate linear prediction parameters for the estimated dominant frequency, compressing the frequency domain signal, and identifying signal peaks in the compressed frequency domain signal; means for decompressing the compressed frequency domain signal; means for extracting the identified signal peaks from the reconstructed frequency domain signal; Including, the extracted signal peaks correspond to the formant frequencies in the first formant data set. A system according to any one of the preceding claims.
4. 10. The system of claim 1, wherein the means for identifying formant frequencies associated with the one or more vowel characters from the first formant data set and for identifying the specific vowel characters comprises means, at the first computing device, for applying a Mahalanobis distance algorithm to the extracted formant frequencies.
5. 10. The system of claim 1, wherein the predetermined metrics include one or more of a Formant Centering Ratio (FCR) algorithm, a Vowel Space Algorithm (VSA), and a Vowel Articulation Index (VAI) algorithm.
6. 10. The system of claim 1, wherein the first computing device is a mobile phone.
7. 10. A system according to any one of the preceding claims, wherein the formant frequencies in the first and second formant data sets include at least formant frequencies F1 (Hz) and F2 (Hz).
8. 10. The system of claim 1, wherein the formant frequencies in the first formant data set include at least formant frequencies F1 (Hz), F2 (Hz), and F3 (Hz), and the formant frequencies in the second formant data set include at least formant frequencies F1 (Hz) and F2 (Hz).
9. 1. A method configured to analyze acoustic parameters of speech to detect, diagnose, predict, and / or monitor the progression of a condition, disorder, or disease, the method comprising using a first computing device configured with means for receiving an audio stream including speech data encoding at least one word spoken by an individual, the first computing device comprising: converting the audio stream into a speech signal; extracting a first formant data set from the speech signal as the audio stream is received in near real time without recording the speech data encoding the at least one word spoken by the individual, the first formant data set comprising formant frequencies associated with letters in the at least one word; identifying formant frequencies associated with one or more vowel characters in the word from the first formant data set; Identifying specific vowel character(s) from the identified vowel character formant frequencies; recording a second formant data set including at least some of the identified vowel character formant frequencies of the specific vowel character(s); configured to run the method further comprising applying, at the first computing device and / or a second computing device connected to the first computing device via a network, one or more predetermined acoustic metrics to the recorded second data set of formants to generate score data, the score data being used to identify a pronunciation level of the vowel character(s) in the at least one word spoken by the individual; The method further includes storing the score data as an output file. The method.
10. The method of claim 9 , wherein the first computing device is a mobile computing device, such as a mobile smart phone, having cellular and computing capabilities.
11. The step of extracting the first formant data set comprises: converting said audio signal from a time domain signal to a frequency domain signal, for example by using a Fast Fourier Transform (FFT) algorithm; applying an autocorrelation algorithm to estimate a dominant frequency in the frequency domain signal, and applying a Levinson-Durbin algorithm to estimate linear prediction parameters of the estimated dominant frequency, compressing the frequency domain signal, and identifying signal peaks in the compressed frequency domain signal; decompressing the compressed frequency domain signal; extracting the identified signal peaks from the reconstructed frequency domain signal; Including, the extracted signal peaks correspond to the formant frequencies in the first formant data set.
11. The method of claim 9 or 10.
12. 12. The method of claim 9, further comprising applying, at the first computing device, a Mahalanobis distance algorithm to the extracted formant frequencies to identify formant frequencies associated with the one or more vowel characters from the first formant data set, and identifying the specific vowel characters.
13. 13. The method of any one of claims 9 to 12, wherein the predetermined metrics include one or more of a Formant Centering Ratio (FCR) algorithm, a Vowel Space Algorithm (VSA), and a Vowel Articulation Index (VAI) algorithm.
14. The method of any one of claims 9 to 13, wherein the first computing device is a mobile phone.
15. 10. A method according to any one of the preceding claims, wherein the formant frequencies in the first and second formant data sets include at least formant frequencies F1 (Hz) and F2 (Hz).
16. 10. A method according to any one of the preceding claims, wherein the formant frequencies in the first formant data set comprise at least formant frequencies F1 (Hz), F2 (Hz), and F3 (Hz), and wherein the formant frequencies in the second formant data set comprise at least formant frequencies F1 (Hz) and F2 (Hz).