Quantifying symptoms of neurological voice disorders using a multimodal approach
A multimodal model combining phonetic and acoustic analysis addresses the challenge of quantifying neurological voice disorders, offering accurate and efficient symptom detection and progression monitoring.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2026-03-26
AI Technical Summary
Current automated systems struggle to accurately quantify symptoms of neurological voice disorders, particularly in free or unguided speech, due to noise interference and complexity of natural speech patterns, leading to prolonged and subjective medical professional analysis.
A multimodal model that combines phonetic and acoustic content analysis of patient speech to identify and quantify symptoms, using a neural network architecture with separate branches for acoustic and phonetic features, enabling accurate symptom detection and severity prediction.
The multimodal model effectively quantifies neurological voice disorder symptoms, facilitating disease progression monitoring and diagnosis support by providing objective and efficient symptom analysis.
Smart Images

Figure US2025046893_26032026_PF_FP_ABST
Abstract
Description
NONPROVISIONAL APPLICATIONQUANTIFYING SYMPTOMS OF NEUROLOGICAL VOICE DISORDERS USING A MULTIMODAL APPROACHGOVERNMENT SUPPORT
[0001] This invention was made with government support under 1 P50DC019900-03 awarded by the National Institutes of Health - National Institute on Deafness and Other Communication Disorders (NIH-NIDCD). The government has certain rights in the invention.Related Applications
[0002] This application claims priority to U.S. Provisional Application Serial No. 63 / 695,994, filed September 18, 2024, entitled “METHOD FOR QUANTIFYING SYMPTOMS OF NEUROLOGICAL VOICE DISORDERS USING MULTIMODAL TRANSFORMER-BASED APPROACH”. The entirety of this provisional application is hereby incorporated by reference for all purposes.Technical Field
[0003] This disclosure relates generally to quantifying symptoms of neurological voice disorders, and more specifically to systems and methods that can quantify symptoms of neurological voice disorders from an audio signal of a patient speaking one or more sentences using a multimodal model.Background
[0004] Neurological voice disorders can weaken or impair the control of a person’s voice by affecting the nerves and / or muscles associated with production of the person’s voice. Traditionally, in order to diagnose a neurological vocal disorder, an audio recording is made of the person speaking various vocal sounds, words, phrases, and / or sentences. However, these audio recordings are quite noisy. While current automated system processing of the audio recording can remove some of the noise, significant noise remains on the recordings. The symptoms can be detected in the processed audio recording through the remaining noise such that a medical professional may quantify the symptoms and make a diagnosis. However,the quantifying by the medical professional often takes a prohibitively long time and can be negatively affected by the remaining noise and by personal interpretation.
[0005] Automated systems can theoretically quantify the voice symptoms if only simplistic pre-determined vocal sounds are used. Current automated systems struggle and fail to quantify free or unguided speech and / or sentences spoken by patients. In other cases, a patient can be instructed to engage in speech and / or sentences with sustained vowels. Using sustained vowels can help to avoid language-related noise in audio recordings, but the sustained vowels can cause the automated systems to miss dynamic changes that happen during speech. Scripted symptom-eliciting sentences (also referred to as eliciting sentences) show promise but are not commonly used due to phonetic and acoustic complexity compared to the processing capabilities of current automated systems.Summary
[0006] Described herein are systems and methods that can quantify symptoms of neurological voice disorders from an audio signal of a patient speaking one or more sentences using a multimodal model. The multimodal model can employ both phonetic content of one or more vocal sounds, words, phrases, and / or sentences (e.g., symptom-eliciting sentences) and acoustic content of an audio signal of the patient speaking the one or more vocal sounds, words, phrases, and / or sentences (e.g., symptom-eliciting sentences, sentences with sustained vowels, etc.). The multimodal model can predict symptom severity, which in turn can be used for disease progression monitoring and decision support for diagnosis.
[0007] In an aspect, the present disclosure can include a system that can quantify symptoms of neurological voice disorders using both phonetic content and acoustic content of an audio signal of a patient speaking one or more sentences. Steps of the method can be performed by a system comprising a processor. An audio signal comprising speech of one to N sentences (N is an integer value) by a patient can be received. One or more different voice symptoms exhibited by the patient and evident in the audio signal can be detected by: analyzing a phonetic content of the one to N sentences; and analyzing an acoustic content of the speech of the one to N sentences. Amounts of the one or more different voice symptoms exhibited by the patient in the audio signal can be measured.
[0008] In another aspect, the present disclosure can include a method for quantifying symptoms of neurological voice disorders using both phonetic content and acoustic content of an audio signal of a patient speaking one or more sentences. The system includes a recording device and a computer in communication with the recording device. The recording device can be configured to record an audio signal comprising speech of one to N sentences by a patient, where N is an integer value. The computing device includes a non-transitory memory configured to store instructions and the audio signal; and a processor configured to access the memory and execute the instructions to at least: receive the audio signal and store the audio signal in the non-transitory memory; detect one or more different voice symptoms exhibited by the patient and evident in the audio signal by analyzing: a phonetic content of the one to N sentences; and an acoustic content of the speech of the one to N sentences; and measure amounts of the one or more different voice symptoms exhibited by the patient in the audio signal.Brief Description of the Drawings
[0009] The foregoing and other features of the present disclosure will become apparent to those skilled in the art to which the present disclosure relates upon reading the following description with reference to the accompanying drawings, in which:
[0010] FIG. 1 is a block diagram of a system that can quantify symptoms of neurological voice disorders from an audio signal of a patient speaking one or more sentences using a multimodal model;
[0011] FIG. 2 is a block diagram of an example of the computing device of FIG. 1 ;
[0012] FIG. 3 is a block diagram of an example of the model of FIG. 3;
[0013] FIG. 4 is a process flow diagram of a method for quantifying symptoms of neurological voice disorders from an audio signal of a patient speaking one or more sentences using a multimodal model;
[0014] FIG. 5 is a process flow diagram of a method for detecting one or more different voice symptoms exhibited by a patient;
[0015] FIG. 6 includes different plots of a dataset used in the experimental;
[0016] FIG. 7 includes different plots showing how symptoms were found; and
[0017] FIG. 8 includes model architecture diagrams.Detailed DescriptionI. Definitions
[0018] Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure pertains.
[0019] As used herein, the singular forms “a,” “an,” and “the” can also include the plural forms, unless the context clearly indicates otherwise.
[0020] As used herein, the terms “comprises” and / or “comprising,” can specify the presence of stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups.
[0021] As used herein, the term “and / or” can include any and all combinations of one or more of the associated listed items.
[0022] As used herein, the terms “first,” “second,” etc. should not limit the elements being described by these terms. These terms are only used to distinguish one element from another. Thus, a “first” element discussed below could also be termed a “second” element without departing from the teachings of the present disclosure. The sequence of operations (or acts / steps) is not limited to the order presented in the claims or figures unless specifically indicated otherwise.
[0023] It will be understood that when an element is referred to as being "on," "attached" to, "connected" to, "coupled" with, "contacting," etc., another element, it can be directly on, attached to, connected to, coupled with or contacting the other element or intervening elements may also be present. In contrast, when an element is referred to as being, for example, "directly on," "directly attached" to, "directly connected" to, "directly coupled" with or "directly contacting" another element, there are no intervening elements present. It will also be appreciated by those of skill in the art that references to a structure or feature that is disposed “adjacent” another feature may have portions that overlap or underlie the adjacent feature.
[0024] As used herein, the term “neurological voice disorder” can refer to a condition that affects the muscles and / or nerves involved in voice production (e.g., speech, sound formation, singing, etc.). A neurological voice disorder may be referred to herein as a “voice disorder”. Examples of neurological voice disorderscan include, but are not limited to, spasmodic dysphonia, vocal tremor, vocal cord paralysis, disorders associated with Parkinson’s disease, and dysarthria.
[0025] As used herein, the term “speech” can refer to the ability to communicate ideas through vocal production, the use of articulate vocal sounds, words, phrases, and sentences to convey thoughts and meaning.
[0026] As used herein, the term “audio signal” can refer to an electrical representation of sound (which can be visualized as a waveform). For example, the audio signal can transform sound into an electrical voltage, with characteristics including amplitude (corresponding to loudness), frequency (corresponding to pitch), and timing of sounds. It should be noted that an “audio recording” or “recording” includes the audio signal and that the audio signal can be referred to as an audio recording.
[0027] As used herein, the term “noise profile” can refer to a distinctive statistical description of ambient background noise of an audio signal that is not produced by the subject’s vocal chords. For example, a few seconds of silence can be included during the recording of an audio signal (e.g., a few seconds at the beginning of the recording) so that the noise profile of the audio signal can be identified and filtered out.
[0028] As used herein, the term “spectrogram” can refer to a visual representation of a spectrum. As an example, an audio signal recording can include a spectrum of frequencies as the recorded audio signal varies with time and the visual representation of the spectrum of frequencies can be referred to as a spectrogram.
[0029] As used herein, the term “multimodal model” refers to an artificial intelligence model trained to process and find relationships between different types of data (e.g., phonetic content of one or more of the sounds, words, phrases, and / or sentences (e.g., symptom-eliciting sentences) spoken by a patient and acoustic content of an audio signal of the patient speaking the one or more of the sounds, words, phrases, and / or sentences). For example, the multimodal model can include at least two different modes, which can be (but are not limited to) audio content analysis and phonetic content analysis. Multimodal models can rely, for example, on deep neural networks, hidden Markov models (HMM), Restricted Boltzmann Machines (RBM), or the like.
[0030] As used herein, the multimodal model can be “tuned” with historical data used for the same purpose (e.g., historical data from the same patient, historical data from a population of other patients undergoing a similar test / with similar demographic information, historical data from a segment of the population of other patients, historical data from populations with specific voice disorder(s), etc.).
[0031] As used herein, the terms “subject” and “patient” can refer to a human being. However, the systems, methods, and techniques described herein can also be used with respect to other warm-blood organisms that can produce vocal sounds (with any necessary small modifications).II. Overview
[0032] As many as one in five Americans self-report as having suffered from a neurological voice disorder at some point in their life. Neurological voice disorders can affect the muscles and / or nerves involved in voice production, causing symptoms including, but not limited to, harshness, strained or breathy speech, vocal tremor, and a noticeable number of sentence-level breaks / stops. Generally, an audio recording can be created of a patient speaking various vocal sounds, words, phrases, and / or sentences. Traditionally, a professional can analyze the audio recording and attempt to identify symptoms (e.g., through system noise, background noise, overlapping symptoms, limited vocalization data, or the like). Automated techniques have been developed to analyze audio recordings to identify symptoms faster and more accurately than the traditional professional, but suffer from a variety of problems (e.g., inability to filter enough noise, simplistic processing abilities, limited vocalizations that can be processed, or the like). In both case, audio recordings tend to be noisy, so the content of the audio recording (e.g., the type of vocal sounds, words, phrases, and / or sentences spoken by the patient) can be muddled. It is important to have clean recordings (without significant noise) so that symptoms of a neurological voice disorder can be properly identified. Audio recordings can range from basic vowel and consonant sounds to sustained vowels and one or two words up to one or more sentences. A patient may be given written instructions for vocalizations and / or may free speak (e.g., unguided speech). The more complex and unguided the audio recordings are the more difficult symptom quantification is (for a professional or an automated system). A professional may be able to separate free speech from the noise in some cases, but automatedtechniques struggle to separate the free speech from the noise or understand spoken words. Automated technology has been shown to have more luck when a patient is instructed to engage in shorter bursts of speech with sustained vowels. Using sustained vowels can help to avoid language-related noise but automated systems can miss dynamic changes within sustained vowels. Additionally, more simplistic cues (e.g., sustained vowels, vocalization without language, short words or sentences, or the like) may not show one or more symptoms and / or symptom severity. Scripted symptom-eliciting sentences (also referred to as eliciting sentences) show promise for improved quantification of symptoms and / or symptom severity but are not commonly used due to complexity of sentences and processing of those sentences (e.g., language, phonetics, acoustics, etc.). While the complexity has been prohibitive before, automated systems have gained the processing power to analyze complex matters.
[0033] Multi-modal models, as described herein, can handle the complexity of eliciting sentences, as well as other types of patient speech. Recently, automated systems using multimodal models have shown promise in other fields, such as detecting Parkinson’s disease, but multimodal models have not been utilized to identify symptoms of neurological voice disorders within an audio signal. However, automated systems have not utilized multimodal models to identify symptoms of neurological voice disorders within an audio signal.
[0034] Described herein are systems and methods that can quantify (also referred to as quantitate) symptoms of neurological voice disorders using a multimodal model. The multimodal model can use both phonetic content one or more vocal sounds, words, phrases, and / or sentences (e.g., sustained vowels, symptomeliciting sentences, etc.) and acoustic content of an audio signal of the patient speaking one or more vocal sounds, words, phrases, and / or sentences (e.g., symptom-eliciting sentences). In some instances, demographic information about the patient can also be considered. The systems and methods described herein can be used to predict symptom severity, which in turn can be used for disease progression monitoring and decision support for diagnosis.III. Systems
[0035] To diagnose a patient with a neurological voice disorder (also referred to as a voice disorder), a test is conducted in which an audio recording is made of thepatient speaking (e.g., various vocal sounds, words, phrases, and / or sentences). The recording transforms the patient’s speech, as well as any ambient background noise, into an audio signal. The patient’s speech can include one or more complex symptom-eliciting sentences designed to trigger one or more speech symptoms (if present). It should be understood that the system 100 is not limited to symptomeliciting sentences and instead could be any type of sentence (with sustained vowels, for example) or other voice input. The system 100 of FIG. 1 can identify and quantify (also referred to as quantitate) one or more symptoms (also referred to as voice symptoms or vocal symptoms) of one or more neurological voice disorders within the audio recording of the patient’s speech. Examples of the one or more different symptoms exhibited by the patient and evident in the audio signal can include but are not limited to, harshness, breathiness, tremor, and / or an abnormal number of speech breaks. The system 100 can be used to facilitate the diagnosis of one or more various neurological voice disorders, including severity of the one or more neurological voice disorders, and is not limited as to the type of neurological voice disorder that can be identified.
[0036] Notably, the system 100 can employ a multi-modal model that inputs at least phonetic content of the patient’s speech in the audio recording and acoustic content of an audio signal of the audio recording, analyzes the input, and identifies and quantifies the one or more symptoms. The system 100 can output the identification and / or the quantification of the one or more symptoms to predict symptom severity. Quantification and identification of the one or more symptoms of a patient over time (e.g., weeks, months, years, etc.) can facilitate disease progression monitoring and decision support for diagnosis. The system 100 can combine acoustic analysis and language analysis to determine the identification and / or quantification of one or more symptoms and / or one or more neurological voice disorders. It should be understood that the example system shown in FIG. 1 is simply for ease of illustration and other examples / configurations are contemplated. It should be understood that unless otherwise stated, elements of the systems described herein generally operate as widely known in related fields.
[0037] The system 100 can include a recording device 102 and a computing device 104. The recording device 102 can be configured to record an audio signal that can include at least speech by a patient (also referred to as desired audio). The speech can include various vocal sounds, words, phrases, and / or sentences, whichmay be pre-instructed by a professional administering the test but need not be instructed (e.g., can be free speech). In some instances, the speech can include from one to N sentences (e.g., complex symptom-eliciting sentences), where N is an integer value. For example, the one to N sentences can be bound only by the complexity of the neural vocal signal being diagnosed and / or the processing power / limits of the computing device 104. The recorded audio signal can also include noise, any unwanted sound or interference that degrades the quality of the desired audio, making the desired audio harder to hear or understand. As an example, noise can come from external sources (e.g., environmental or ambient noise) and / or can be generated internally by one or more elements of the recording device 102. The recording device 102 can include any number of sub-devices including, but not limited to, a microphone and a transmission means (e.g., for wired and / or wireless communication with at least computing device 104). In some instances, the recording device 102 can also include one or more processing components and / or memory components (not shown).
[0038] The computing device 104 can be in wired and / or wireless communication with the recording device 102. In some instances, the recording device 102 can be linked to the computing device 104 by the wired and / or wireless connection. In other instances, not shown, the recording device 102 and the computing device 104 can be at least partially embodied in a single device (e.g., computing device with recording capabilities or a recording device with a memory and / or processor). The computing device 104 is an example of one or more hardware devices that can be part of the system 100 and capable of implementing at least the data input, data pre-processing, the multi-modal model, and the output described herein. The computing device 104 can include various systems and subsystems and can be a personal computer, a laptop computer, a workstation, a computer system, an application-specific integrated circuit (ASIC), a server BladeCenter, a server farm, or the like. Although not illustrated, it should be understood (for example) that the computing device 104 can include a system bus, the processor 108 can include one or more processing units, the memory 106 can include a system memory and / or memory devices, a communication interface (e.g., a network interface, which may be in communication with a network (public or private) that the recording device 102 may be on), or the like commonly found in computing devices.
[0039] As illustrated, in any implementation, the computing device 104 can include at least a memory 106 (e.g., a non-transitory memory) and a processor 108. In some instances, the functionality of the memory 106 and the processor 108 can be implemented by a microprocessor. The memory 106 can store instructions and audio signals recorded by the recording device (e.g., instructions and / or audio recordings can be permanently or temporarily saved in RAM and / or ROM). In some instances, the symptoms identified and / or any analysis can be stored and accessible with the corresponding audio signal. In some instances, the memory 106 can be implemented, for example, as computer-readable media (integrated or removable), such as a memory card, disk drive, compact disk (CD), or server accessible over a network. The processor 108 can access the memory 106 and execute the instructions to perform actions including: input at least phonetic content spoken by the patient and acoustic content of an audio signal of the audio recording, analyze the input, and identify and quantify the one or more symptoms found in the input. Examples of the processor 108 can include an application-specific integrated circuit (ASIC), processing core, or the like. While not shown, it should be understood that the computing device 104 can also include at least one user interface (e.g., mouse, keyboard, knob(s), button(s), touch screen, or the like) and / or display (visual, audio, haptic, or the like)
[0040] The system 100 can also include, in some instances, remote storage 110 (e.g., a remote server or the like) that can store one or more types of information and / or data and communicate (with a wired and / or wireless connection) with the computing device 104. For instance, the remote storage 110 can include demographic information from a plurality of patients and / or populations, information related to one or more symptoms and / or vocal disorders, or the like. As an example, the remote storage 110 can store at least a portion of demographic information about the patient and send the demographic information to the computing device 104. The demographic information can include age, sex, race, ethnicity, income, education, location, native language, or the like. The computing device 104 can use at least a portion of the demographic information to improve symptom identification and quantification and / or the diagnosis. For example, the processor 108 can receive demographic information of the patient from the remote storage 110 (or a user interface (not shown) and use the demographic information when associating the one or more different voice symptoms exhibited by the patient and evident in theaudio signal with the one or more neurological voice disorders. In other words, the processor 108 can consider a likelihood a symptom contributes to a diagnosis rather than is characteristic of speech of a demographic that the user fits into.
[0041] FIG. 2 shows an example computing device 104 that can be implemented by the system 100. It should be noted that the example computing device 104 may be associated with more / different components than are shown in FIG. 1 . As shown, the computing device 104 can include the memory 106 and processor 108 as described previously. The processor 108 can access the memory 106 to retrieve executable instructions and data. The executable instructions, upon execution, cause the processor 108 to receive 202 at least audio signal recordings, detect 204 by implementing a multi-modal model (model 206), and measure 208 one or more symptoms. The receive 202 instruction can cause the processor 208 to receive at least a portion of the audio recording (including the audio signal) and store the audio recording in the non-transitory memory 106. The entire audio recording can be stored in the memory 106 (audio signal + noise) in some instances. In other instances, the audio signal and at least a portion of the noise (after filtering) can be stored in the memory 106. In still other instances, the audio signal (after filtering) can be stored alone (with minimal noise and / or a not significant amount of noise) in the memory 106. The detect 204 instruction can cause the processor 106 to detect one or more different voice symptoms exhibited by the patient and evident in the audio signal by analyzing: a phonetic content of the audio signal, and an acoustic content of the speech in the audio signal. As an example, phonetic differences can be extracted from the phonetic content and acoustic differences can be extracted from the acoustic content. Notably, the model 206 can consider the language content of the audio (phonetic content) and the acoustic content of the audio (acoustic content) when determining the number (or measure) of symptoms. The phonetic content / phonetic differences and the acoustic signal / acoustic differences can be input into the model 208 (shown in FIG. 3). For example, the model 206 can be a tuned model (e.g., tuned based on historical data for a population (or portion of the population), historical data for the patient, or the like). The measure 208 instruction can cause the processor 106 to measure amounts of the one or more different voice symptoms exhibited by the patient in the audio signal.
[0042] Additionally, the example computing device 104 can include at least a communication component 210, an input 212 (e.g., for a user interface, a data inputfrom a near or remote storage or devices, or the like), and an output 214 (e.g., visual display, audio output, haptic output, or the like). The communication component 210 can enable wired and / or wireless communication and may include at least a portion of the devices / components to engage in wired and / or wireless communication. The input 212 can enable data to be input into the memory 106 and / or the processor 108 (e.g., the data can include the audio signal). The output 214 can enable other data and / or a suggestion about the data to be output from the memory 106 and / or the processor 108 (e.g., the other data can include one or more symptoms, a likelihood of one or more neurological voice disorders, etc.). As an example, the input 212 can communicate with the recording device (e.g., recording device 102) and the output 214 can communicate with an output device (not shown, such as a display device).
[0043] Referring now to FIG. 3, illustrated is an example execution of the multimodal model 206. The multimodal model 206 can be, for instance, a neural network model with at least two branches (e.g., audio branch 306 (also called vocal branch) and language branch 308). An audio recording 300 (as described above) can be input into the model 206. In some instances, the audio recording 300 be prefiltered (e.g., by the recording device or the processor prior to being input into the model) and can include only the audio signal or the audio signal and an insignificant amount of noise. In other instances, the audio recording 300 can include the audio signal and the noise and the noise can be filtered out from the audio recording within the model 206. The audio signal (within the audio recording 300) can include speech of a patient with 1 to N various vocal sounds, words, phrases, and / or sentences. The model 206 can split the audio signal of audio recording 300 into acoustic content 302 (e.g., timing, frequency, amplitude, etc.) and phonetic content 304 (e.g., speech sounds) for analysis in the audio branch 306 and / or the language branch 308. As noted, the model 206 can consider elements of both the language and the audio components of the audio signal when determining the number (or measure) of one or more symptoms (e.g., harshness, breathiness, tremor, number of (abnormal) sentence-level breaks, or the like).
[0044] Prior to entering the branches of the model 206 the processor (e.g., processor 108) can separate the audio recording 300 into the acoustic content 302 and the phonetic content 304 by segmenting the audio signal to identify sentences, words, specific vowel or consent sounds, or the like. For instance, the 1 to N various vocal sounds, words, phrases, and / or sentences can be segmented. The segmentscan be identified based on a pre-instructed 1 to N various vocal sounds, words, phrases, and / or sentences. Optionally the processor (e.g., processor 108) can remove noise (if not previously removed). The noise can be removed, for example, with a within-sample noise profile based on several seconds of audio recording without patient speech (e.g., at the beginning or the end), such that the noises that match the profile are filtered out to leave the audio signal. Then, the processor (e.g., processor 108) can transform the segmented audio signal into one or more spectrograms (e.g., Log-Mel spectrogram) that can be input into the model 206. A Log-Mel spectrogram is a Mel spectrogram where a logarithmic transformation is applied to the amplitude of frequency components). Log-Mel is an example of an audio feature extraction technique used in machine learning; it should be understood that other audio feature extraction techniques can be used.
[0045] The audio branch 306 can utilize the pre-processed acoustic content 302 while the language branch 308 can utilize a combination of the acoustic content and the phonetic content 304. The acoustic content 302 of a given audio recording 300 can be input into audio branch 306 and / or can be retrieved from memory by audio branch 306. The audio branch 306 can utilize one or more transformer encoders that can understand basic audio data structures (such as the Whisper encoder by OpenAI) as a software backbone. The audio branch 306 can be trained to multitask to notice the acoustic content 302 of multiple different symptoms (e.g., breathiness, harshness, tremor, sentence-level breaks, etc.) (e.g., by comparison with acoustic content with no symptoms or combinations of one or more symptoms). The audio branch 306 can transform the one or more spectrogram of acoustic content 302 (e.g., with one dimensional convolutions and deep learning error correction, or the like). The transformed acoustic content can be combined with sinusoidal positional encoding and fed into one or more transformer encoder blocks (not shown for ease of illustration) of the audio branch 306. The transformer encoder blocks can weigh (e.g., utilizing self-attention coding) the importance of pre-determined parts of the input (e.g., related to vowel(s), consonant(s), word(s), etc.). The transformer encoder blocks can also utilize multilayer perceptron (MLP) as part of a feedforward for the model 206 when the model comprises a neural network consisting of fully connected neurons with nonlinear activation functions. The audio branch 306 can output one or more context aware embeddings that can be sent to scorer 310. Each of theembeddings can contain information about one or more acoustic properties and / or speaker characteristics.
[0046] Both the acoustic content 302 and the phonetic content 304 of audio recording 300 can be input into and / or retrieved by the language branch 308 of model 206. In some instances, the acoustic content 302 and / or the phonetic content 304 can be pre-processed or can be in a single combined audio signal (that has been filtered as discussed above but not separated into spectrogram(s)). The language branch 308 can also retrieve (or receive) target text 314 that can include a vocal target for comparison with the audio signal (e.g., acoustic content 302 and phonetic content 304). The phonetic content 304 can be determined by the model 206 converting speech into the international phonetic alphabet (I PA). The target text 314 (which can be input from memory or by a user-interface) can also be converted into IPA by the model 206 (either before or within the language branch 308). The language branch 308 can compare the IPA version of the target text and the IPA version of the phonetic content 304 of the speech in the audio recording. For example, the target text 314 can be “Sam has a rabbit in his hat” which in IPA is “saem haez e raebet in hiz hast” and the phonetic content 308 of the audio signal can be “Dan had a rabbit in his head” which in IPA is “daen haed e raebet in hiz hsd”.The language branch 306 can include at least one embedding model (e.g., with learned semantic and syntactic information) that can transform the IPA of the target text 314 and the phonetic content 308 to determine the differences (e.g., the deviation from the IPA of the target text). The embedded phonetic difference (e.g., deviation from actual and target IPA) can be fed into the scorer 310. The scorer 310 can analyze the differences and output symptoms, a number of times a symptom appears in the sentence or sentences, a severity of each symptom, a potential diagnosis / likelihood of the potential diagnosis, or the like.IV. Methods
[0047] Another aspect of the present disclosure can include methods (e.g., methods 400 and 500 of FIGS. 4 and 5) The methods 400 and 500 for quantifying symptoms of neurological voice disorders from an audio signal of a patient speaking one or more sentences using a multimodal model. Methods described herein can be used to predict symptom severity, which in turn can be used for disease progression monitoring and decision support for diagnosis. The system 100 of FIG. 1 , and the examples shown in FIGS. 2 and 3, for example, can be used to execute the methods400 and 500. Unless otherwise stated, the methods described herein generally follow the usual practices widely known in related fields. Moreover, examples of how elements of the method can operate and / or what the elements can be are described in the Systems section above (e.g., recording device 102, computing device 104, etc.). Additionally, as stated above, one or more steps of methods 400 and 500 can be executed by a computing device comprising a processor.
[0048] Referring now to FIG. 4, illustrated is a method 400 for quantifying symptoms of neurological voice disorders from an audio signal of a patient speaking one or more sentences using a multimodal model. The multimodal model can utilize both phonetic content of one or more vocal sounds, words, phrases, and / or sentences (e.g., symptom-eliciting sentences, sustained vowels, etc.) and acoustic content of an audio signal of the patient speaking one or more symptom-eliciting sentences. The multimodal model can predict symptom severity, which in turn can be used for disease progression monitoring and decision support for diagnosis.
[0049] At 402, an audio signal comprising speech by a patient (also referred to as desired audio) (e.g., recorded by a recording device 102) can be received (e.g., by computing device 104). The speech can include various (from 1 to N) vocal sounds, words, phrases, and / or sentences (which may be pre-instructed by a professional administering the test but need not be instructed (e.g., can be free speech)). In some instances, the speech can include one or more symptom-eliciting sentences. A recording of the audio signal also includes noise, any unwanted sound or interference that degrades the quality of the desired audio, making the desired audio harder to hear or understand. As an example, noise can come from external sources (e.g., environmental noise) and / or can be generated internally by one or more elements of the recording device.
[0050] At 404, one or more different voice symptoms (e.g., tremor, harshness, breathiness, abnormal sentence-level breaks, etc.)) exhibited by the speech of a patient can be detected from an audio recording. The speech in the recorded acoustic signal can have a phonetic content (e.g., speech sounds) and an acoustic content (e.g., timing, frequency, amplitude, etc. of the sounds). As an example, phonetic differences can be extracted from the phonetic content and acoustic differences can be extracted from the acoustic content. Notably, a multimodal model (described with respect to FIG. 5) can consider the language content of the audio (phonetic content) and the acoustic content of the audio (acoustic content) whendetermining the number (or measure) of symptoms. The phonetic content and the acoustic content can be input into the model and the model can detect elements (phonetic or acoustic) of the one or more different voice symptoms (if present). For example, the model can be a tuned model (e.g., tuned based on historical data for a population (or portion of the population), historical data for the patient, or the like). At 406, amounts of the one or more different voice symptoms exhibited by the patient in the audio signal can be measured and the measured amounts can be output. The output, in some instances, can further include symptom combinations, symptom severity combinations, and / or patterns with potential diseases / severity and likelihood (e.g., based at least in part on demographic information). In other instances, one or more neurological vocal disorders can be predicted based on the quantification of the one or more symptoms of the patient’s speech.
[0051] FIG. 5 is an illustration of a method 500 for detecting one or more different voice symptoms exhibited by a patient. In some instances, method 500 can be a part of method 400 (step 404). The one or more different voice symptoms can be detected and analyzed by a multimodal model (e.g., shown in FIG. 3, also referred to as “model”). The multimodal model can be, for example, a neural network model with at least two branches (e.g., an audio branch (e.g., 306) (also called vocal branch) and language branch (e.g., 308)). A recording (e.g., audio recording 300) can be input into the model and / or retrieved by the model (e.g., from memory). In some instances, the recording can be pre-filtered (e.g., by the recording device or the processor prior to being input into the model) and can include only the audio signal or the audio signal and an insignificant amount of noise. In other instances, the recording can include the audio signal and the noise and the noise can be filtered out from the audio recording within the model. The audio signal (within the recording) can include speech of a patient with 1 to N various vocal sounds, words, phrases, and / or sentences. The model can split the audio signal of the recording into acoustic content (e.g., acoustic content 302) (e.g., timing, frequency, amplitude, etc.) and phonetic content (e.g., phonetic content 304) (e.g., speech sounds) for analysis by branches of the model (e.g., the audio branch and / or the language branch). As noted, the model can consider elements of both the language and the audio components of the audio signal when determining the number (or measure) of one or more symptoms (e.g., harshness, breathiness, tremor, number of (abnormal) sentence-level breaks, or the like).
[0052] Prior to entering the branches of the model the audio recording can be separated (e.g., by processor 108) into the acoustic content and the phonetic content by segmenting the audio signal to identify sentences, words, specific vowel or consent sounds, or the like. For instance, the 1 to N various vocal sounds, words, phrases, and / or sentences can be segmented. The segments can be identified based on a pre-instructed 1 to N various vocal sounds, words, phrases, and / or sentences. Optionally noise can be removed (if not previously removed, or not enough was previously removed) by the processor (e.g., processor 108). The noise can be removed, for example, with a within-sample noise profile based on several seconds of audio recording without patient speech (e.g., at the beginning or the end), such that the noises that match the profile are filtered out to leave the audio signal. Then, the segmented audio signal can be transformed (e.g., by processor 108) into one or more spectrograms (e.g., Log-Mel spectrogram) that can be input into the model. A Log-Mel spectrogram is a Mel spectrogram where a logarithmic transformation is applied to the amplitude of frequency components). Log-Mel is an example of an audio feature extraction technique used in machine learning; it should be understood that other audio feature extraction techniques can be used.
[0053] At 502, the phonetic content (combined with the acoustic content) of the from one to N sentences can be analyzed (through a language branch of the model). The language branch can utilize a combination of the acoustic content and the phonetic content. Both the acoustic content and the phonetic content of the audio recording can be input into and / or retrieved by the language branch of the model. In some instances, the acoustic content and / or the phonetic content can be pre- processed or can be in a single combined audio signal (that has been filtered as discussed above but not separated into spectrogram(s)). The language branch can also retrieve (or receive) target text that can include a vocal target for comparison with the audio signal (e.g., acoustic content and phonetic content). The phonetic content can be determined by the model converting speech into the international phonetic alphabet (IPA). The target text (which can be input from memory or by a user- in terface) can also be converted into IPA by the model (either before or within the language branch). The language branch can compare the IPA version of the target text and the IPA version of the phonetic content of the speech in the audio recording. For example, the target text can be “Sam has a rabbit in his hat” which in IPA is “seem haez e raebet in hiz haet” and the phonetic content of the audio signalcan be “Dan had a rabbit in his head” which in I PA is “daen haed e raebet in hiz hed”. The language branch can include at least one embedding model (e.g., with learned semantic and syntactic information) that can transform the I PA of the target text and the phonetic content to determine the differences (e.g., the deviation from the IPA of the target text). The embedded phonetic difference (e.g., deviation from actual and target IPA) can be fed into the scorer.
[0054] At 504, the acoustic content of the from one to N sentences can be analyzed (through an audio branch of the model). The acoustic content of a given audio recording can be input into audio branch and / or can be retrieved from memory by audio branch. The audio branch can utilize one or more transformer encoders that can understand basic audio data structures (such as the Whisper encoder by OpenAI) as a software backbone. The audio branch can be trained to multitask to notice the acoustic content of multiple different symptoms (e.g., breathiness, harshness, tremor, sentence-level breaks, etc.) (e.g., by comparison with acoustic content with no symptoms or combinations of one or more symptoms). The audio branch can transform the one or more spectrogram of acoustic content (e.g., with one dimensional convolutions and deep learning error correction, or the like). The transformed acoustic content can be combined with sinusoidal positional encoding and fed into one or more transformer encoder blocks of the audio branch. The transformer encoder blocks can weigh (e.g., utilizing self-attention coding) the importance of pre-determined parts of the input (e.g., related to vowel(s), consonant(s), word(s), etc.). The transformer encoder blocks can also utilize multilayer perceptron (MLP) as part of a feedforward for the model when the model comprises a neural network consisting of fully connected neurons with nonlinear activation functions. The audio branch can output one or more context aware embeddings that can be sent to scorer. Each of the embeddings can contain information about one or more acoustic properties and / or speaker characteristics. While not shown in FIG. 5, the analyzed outputs (e.g., embeddings) of the language branch and the audio branch of the model can be fed into a scorer (e.g., scorer 310 also called a concatenator). The scorer can output symptoms, a number of times a symptom appears in the sentence or sentences, a severity of each symptom, a potential diagnosis / likelihood of the potential diagnosis, or the like.V. Experimental
[0055] Neurological voice disorders such as laryngeal dystonia (LD) and voice tremor (VT) often manifest through perceptual symptoms including breathiness, tremor, harshness, and increased voice breaks. An automated method capable of accurately detecting and quantifying these symptoms from voice recordings have been developed and is described herein, enabling more scalable and consistent clinical assessment.Methods
[0056] An artificial neural network architecture called NeuroVoiceNet for the four associated learning tasks is proposed, which includes finetuning a Whisper transformer encoder model, and attention-based artificial neural network modules. This architecture has two branches, one for processing acoustic features and another one for phonetic features, making it a multimodal approach. For training ordinal regression tasks, the Condition-Aware Tiered Ordinal Loss or CATOL is proposed for separating symptomatic and asymptomatic cases, while considering noisy clinical labels.
[0057] The proposed multi-task model, NeuroVoiceNet, would accept one or more recorded sentences, task target sentences and demographic information as input, and predict breathiness, harshness and tremor scores for the whole input, and number of voice breaks for each of the sentences used. The architecture of NeuroVoiceNet enables it to utilize the acoustic information of voice recordings and the language information of the speech task targets.
[0058] Study data
[0059] This study is a retrospective analysis of audio data of ADLD, ABLD and VT patients collected between 2011 and 2021 . These studies investigated the change in voice symptoms related to corresponding research topics, therefore for this study only baseline recordings were used for analysis and modeling. Each patient and healthy volunteer had to complete from 8 to 20 speech tasks, dictated by standardized sentences designed to induce ABLD or ADLD -related symptoms. These sentences were repeated until the speaker achieved a proper recording.
[0060] The audio data for all samples was collected in the same environment. At first the record setup used a ADInstruments PowerLab 8 / 30 ML870, TDT MA3 Microphone Amplifier, Yamaha Q2031 B and a Shure WH20 microphone headset. This setup was later replaced with a Zoom H6 Handy Recorder and the WH20 headset. Voice quality scores were annotated by two experienced clinicians,specializing in voice disorders. Diagnosis information was used from previously reported studies. Individual sentences were identified and timestamped manually. As for demographics, the age variable was normalized using a min-max scaler so that the age variable would be suitable for neural network modeling. Both age and gender variables were transformed with a label encoder and used as tabular input into the NeuroVoiceNet model.
[0061] While preparing the audio dataset, the preprocessing pipeline was done in Python 3.12. The audio files were imported; individual sentence segments were identified and sliced and lastly resampled to 16 kHz using librosa 0.10.2. After this, spectral gating from noisereduce 3.0.3 was used to filter out background noise. Dynamic range compression with a lenient threshold of -6 dB and a ratio of 4 was then performed. This smoothens the dynamics of the audio, which enhances consistency and clarity. Then, root mean square normalization was used with a lenient reference value of -20 dBFS to normalize loudness. Finally, all sentences were normalized between [-1 ,1] value range, padded to 30 seconds and transformed into log-magnitude Mel spectrograms to be compliant with the foundation model used by NeuroVoiceNet, Whisper-small. The preprocessing pipeline is depicted FIG. 8, subplot a.
[0062] Learning tasks
[0063] The method contained a total of 5 learning tasks that were trained. As recording-level results, the three voice quality scores breathiness, harshness and tremor were predicted. The model also predicts the sentence-level number of breaks. The fifth learning task was a diagnosis vector, which encoded the presence of ABLD, ADLD and VT, and the patient can have any combination of the three. Because of this, differentiation from voice modality alone can be incredibly difficult, resulting in sub-optimal classification accuracy. However, the fifth learning task was used in an auxiliary learning manner, meaning that the fifth learning task helped the model perform in other primary tasks by providing additional supervisory signals during training. Specifically, by exposing the model to diagnostic labels that reflected the underlying etiology of voice symptoms, the shared representations learned become more robust and symptom-aware. This auxiliary signal encouraged the model to disentangle overlapping acoustic cues that may be indicative of multiple conditions. Detailed implementation and learning logic for this are described in the model training section.
[0064] Proposed model
[0065] The proposed model, NeuroVoiceNet, is comprised of three branches (FIG. 8, subplot a). The acoustic branch consists of a pretrained Whisper-small encoder model that was fine-tuned during training. With its decoder counterpart, this model has been originally pretrained with multi-task training data of over 600 thousand hours. Every log-Mel spectrogram tensor was fed individually to this transformer encoder, and the resulting acoustic embeddings were then used as inputs to an attention mechanism, Acoustic Attention Module (FIG. 8, subplot d). Because of the multi-task learning problem, this branch will learn to identify influential areas in the spectrogram for all prediction variables.
[0066] In addition to acoustic and demographic features, the model utilized target language information in a novel way. Because the symptoms in this study affect the pronunciation of the spoken words, the phonetic dimension of this problem can be inspected to improve the prediction model performance. A target sentence was processed into phonetic letter array t by using the International Phonetic Alphabet and eng_to_ipa 0.02 python package. This would represent the target phonetic information the speaker would be striving to produce during speech tasks. Spoken audio matching that target sentence would also be transcribed into a phonetic letter array s using a pretrained Wav2Vec2 transformer model trained for phonetic prediction. This would represent the phonetic information produced by the speaker.
[0067] A method was developed to compare the two arrays, to assess how off- target the phonetic information is. By tokenizing both inputs, embeddingsand X that share the same feature space can be derived, using a small neural network called Phoneme Embedding Module (FIG. 8, subplot b). While Wav2Vec2 performed well on predicting phonetic transcription from asymptomatic speech data, the performance was reduced with symptomatic audio where the assumed voice quality was not met. This resulted in missing or added tokens in the prediction. To account for this, the two embedding tensorsand X^ were aligned using a fast dynamic time warping algorithm. Finally, the two embedding tensors were subtracted, which resulted in a phonetic difference embedding tensor(d). If the speaker was phonetically on target and produced acceptable quality of voice, X would be a zero-valued tensor. However, if speaker exhibited symptoms during speech, X(d)would contain information of this off-target speech production. This method was integrated into the proposed NeuroVoiceNet architecture as the phonetic computational branch, and theembeddings were fed to the Phonetic Attention Module (FIG. 8, subplot c).
[0068] The model further included two attention modules, one for acoustic embeddings and one for phonetic difference embeddings. Because voice quality scores are derived by clinicians assessing a global phenomenon over all sentences, the task for these attention mechanisms was to mimic this behavior. The embedding input tensor was defined asXe^B xSxLxD where B is batch size, S is number of sentences, L is sequence length and D is the feature dimension. The attention weights were computed using a small neural network: ab s=softmax Linear2GeLU Linear1xb s)')')') where ab se RL. Then, the weighted aggregation of features was done withwhere zb se IRD. Lastly, the terms were aggregated along the sentence dimension with s zb 'zb,s- s=l
[0069] This way, the attention mechanism computed a weighted feature representation per sentence, aggregated them across sentences, and ultimately output a batch-wise feature summary. Both individually attended acoustic and the phonetic tensors were then flattened and concatenated together with the demographics variables and fed into the final linear layer Linear^ which output the predictions for breathiness, harshness and tremor scores. The ordinal output was achieved by rounding to zero decimals, clipping negative floats to zero, and transforming the floats to ordinal integers. The granularity for the prediction task of the number of voice breaks was different, as the model is supposed to give a local ordinal prediction for each sentence. To achieve this, the output of acoustic attention mechanisms Linear2layer was also fed into a separate Linear^ layer, that could thenoutput number of voice breaks predictions. The complete architecture of NeuroVoiceNet and its modules are depicted in FIG. 8.
[0070] Cross validation
[0071] To validate the architecture of NeuroVoiceNet within the dataset, a fivefold cross-validation experiment was conducted that compared the full model against two alternative configurations. The first was an ablated version of NeuroVoiceNet in which the phonetic computation branch was removed. This modification allowed direct assessment of the contribution of the language-aware processing pathway to overall performance across tasks, particularly in symptom differentiation and sentence-level predictions. The second comparison model was a convolutional neural network (CNN) that utilized the same input features and was trained to solve the same multi-task prediction problem. By benchmarking against these two variants, the aim was to evaluate both the architectural advantage of incorporating phonetic representations and the broader effectiveness of the model design relative to conventional deep learning approaches.
[0072] Model training
[0073] For training the final NeuroVoiceNet model, data was segmented into 80% training, 10% validation, and 10% test datasets. This would amount to 540 observations for training, 67 observations for validation, and 68 observations for testing. The split was stratified by diagnosis, so that in each dataset the distribution of healthy controls and patients would remain the same. Models were trained using an NVIDIA H100 GPU with CUDA 12.5, PyTorch 2.5.0 and Transformers 4.44.2.
[0074] The ordinal regression learning tasks required a loss function that appropriately accounted for the clinical context. The ordinal scale between 0 and 10 scores was inherently asymmetric in terms of importance: distinguishing between symptomatic (scores > 1) and asymptomatic (score = 0) voice was more critical than differentiating between the remaining scores. This asymmetry makes commonly used regression loss functions, such as Mean Squared Error (MSE) or Huber Loss, unsuitable without modifications. One alternative is to restructure the learning tasks as multi-class classification and apply class weighting, a widely adopted technique to address class imbalance. However, these classification-based loss functions do not account for the numerical hierarchy of ordinal score values.
[0075] To address these challenges, Condition-Aware Tiered Ordinal Loss or CATOL, which combines binary classification to separate symptomatic andasymptomatic cases, ordinal bucket classification to capture severity levels, and intra-bucket regression to fine-tune predictions within each group, was proposed. The first of these components is the binary classification loss L1;defined asLr= BCE(sigmoid 2(p - O. ,ybinary) where p is the predicted score and ybinaryis a binary indicator for symptomatic and asymptomatic cases. The second component is a Cross-Entropy or CE loss for the separation of symptomatic bucketized scores and are defined as [1 ,2], [3,4], [5,6], [7,8] and [9,10]. True and predicted bucket indices are computed as b =0,4) respectively, while the loss is defined asL2= CE(b, b~) using the CE loss. The last component is an intra-bucket regression loss to capture fine-grained variation. Intra-bucket regression loss is defined asL3= MSE(p,y')
[0076] The final loss function is expressed as a weighted sum of the three components, defined aswhere a, (i and are weight hyperparameters. L, was weighted as 1 .0, being the most critical loss for the clinical task. L2was weighted as 0.5, and L3as 0.25. The weights were empirically chosen. The last loss drives the adaptation of separating the within-bucket values from each other, while considering the continuous hierarchy of the ordinal scoring. With bucketized L2and coarse separation of L3, CATOL addresses the problem of noisy labels, which are prominent in similar clinical learning problems.
[0077] Breathiness, harshness and tremor tasks were learned using CATOL, while number of voice breaks was learned using Root Mean Squared Error (RMSE). The reason for this was the fact that number of voice breaks does not have a numerical positive bound of 10, so this task was trained more like a traditional regression task.
[0078] The following key techniques were utilized during model training: data augmentation, node dropout and auxiliary learning. For data augmentation, the nlpaug 1.1.11 python package was used to add acceptable variance into the training dataset. Vocal tract length perturbation, loudness adjustment, noise injection, and time shifting were used to quadruple the training dataset to 2144 recordings byapplying augmentations randomly with 50% probability. Also, the sentence ordering within a recording was shuffled randomly, so that the model would not mistakenly learn to associate sentence ordering with the prediction tasks.
[0079] Node dropout regularization was used between Linear2and Linear^ with a probability of 30% during training. To further regularize learning and incorporating domain knowledge, auxiliary learning of the diagnosis vectors was used during model training. The vectors can be defined as d = [dd, db, dt] e {0,1} where d denotes ADLD, b ABLD and t voice tremor. For learning, Linear^ layer with a final Sigmoid activation was added after Linear2which would output the predicted diagnosis vectors. Multi-label Hamming loss was be used as the loss function, defined aswhere N is the total number of samples, L is the total number of classes,is the target, is the prediction and xor(-) is the exclusive operator that returns zero when the target and the prediction are identical. Hamming loss was added to the existing multi-task loss. This way the model could learn shared representations with the voice score tasks, which then enabled per diagnosis logic to be learned. This is especially helpful in cases where one type of voice symptom, such as tremor of voice has a significant influence over the whole signal, making the interpretation of other symptoms difficult. With auxiliary learning, NeuroVoiceNet was made aware of the different diagnosis combinations that will help NeuroVoiceNet generalize better to unseen data.
[0080] Model evaluation
[0081] After model training, the performance of the trained model was investigated using an independent test dataset. This evaluation focused on four primary learning tasks: the three recording-level voice quality scores: breathiness, harshness, and tremor, and the sentence-level number of voice breaks. Each task was evaluated separately using appropriate metrics to assess the model's generalization ability. Mean Absolute Error (MAE) and Mean Square Error (MSE) are routinely used regression metrics that were reported for all tasks. In addition to this, Pairwise Accuracy was used to report the performance of the learning tasks, as it isdeemed appropriate for ordinal regression. Performance on the auxiliary diagnosis vector task was not used for final evaluation but served to improve learning during training. This structured evaluation allowed validation of how well the model performed on clinically relevant voice assessments across different symptom dimensions.Results
[0082] This cross-validation experiment showcases how using an attentionbased transformer architecture is favorable over more traditional convolutional neural networks, and the addition of the phonetic branch in our architecture improves the result even further. For the independent test dataset, pairwise accuracies of 90.8%, 92.2%, 89.2 and 88.9% were reported for voice breathiness, tremor, harshness and number of breaks, respectively.
[0083] Patient population
[0084] A total of 672 voice recordings were collected retrospectively from multiple studies conducted between 2011 and 2021 . These recordings included 207 unique patients. From the recordings, 43 were diagnosed as ADLD, 30 as ABLD, 11 as mixed LD, 53 ADLD with VT, 21 ABLD with VT, 24 mixed LD with VT, and 25 as healthy controls. In addition to a voice corpus, demographic information, such as gender and age, was collected. Voice quality scores were gathered as ordinal values between 0 and 10, while number of voice breaks is the count of such events.
[0085] For every recording, the number of collected standardized sentences would vary between 8 and 20 due to study protocol differences. These Rainbow Passage sentences are designed to induce ADLD and ABLD symptoms. The total number of sentences in the dataset was 8925, and the uncut duration of the audio was 25.1 hours. The data would have its voice quality scores annotated by two voice pathology clinicians. Due to the rarity of LD, compiling such a large, well- characterized audio dataset is exceptionally uncommon. This is one of the most comprehensive clinical voice corpora available for studying ADLD, ABLD, and VT. FIG. 6, a-b showcases two recording examples, the distribution of recording durations in corpus, the uniformity of the corpus and the distribution of diagnosis outcomes.
[0086] Cross validation of candidate models
[0087] For comparing the neural network-based models, three candidate architectures were tested: CNN, NeuroVoiceNet without the phonetic branch, and fullNeuroVoiceNet. These models represent a progression from conventional deep learning approaches to more specialized transformer-based architectures that incorporate multi-modal information. The 5-fold cross validation results in FIG. 7, element a show how in terms of symptoms, the number of voice breaks task performed the best and benefited the most from a full NeuroVoiceNet architecture. After this, tremor was the best performing and showed benefits of using a transformer model over the CNN. Utilizing the full NeuroVoiceNet architecture gave the best result here. The breathiness task was the third best performing task, showcasing benefits of using a transformer over a CNN, however the inclusion of the phonetic branch in full NeuroVoiceNet did not yield statistically significant differences. The last task, harshness, performed the worst and demonstrated no benefit from using a transformer over a CNN model. Overall, the full NeuroVoiceNet model, which includes the phonetic branch, achieved the highest performances across these tasks. Based on these results, the full NeuroVoiceNet was selected for final model training and evaluation on a held-out test set.
[0088] Training and model evaluation
[0089] For the five learning tasks, NeuroVoiceNet model was fine-tuned with the training data for 40 epochs, which ran for approximately 7 hours. During this time, the validation data loss descended and ultimately plateaued after 20 epochs (FIG. 7, element b). The best model checkpoint was selected based on combined task validation error, which reached its global minima after 17.68 epochs. After this point, the evaluation error started accumulating again while training loss kept decreasing, indicating over-fitting.
[0090] Pairwise accuracy results of predicting the 4 symptom types from an independent test dataset were 90.8%, 89.2%, 92.2% and 93.0% for breathiness, harshness, tremor and number of breaks, respectively. These results align with the five-fold cross-validation experiment, where consistent performance trends across symptom types were observed. The results also highlight the model’s strong generalization across all tasks.
[0091] From the above description, those skilled in the art will perceive improvements, changes, and modifications. Such improvements, changes and modifications are within the skill of one in the art and are intended to be covered by the appended claims.
Claims
The following is claimed:1 . A method comprising: receiving, by a system comprising a processor, an audio signal comprising speech of one to N predefined sentences by a patient, where N is an integer value; detecting, by the system, one or more different voice symptoms exhibited by the patient by: analyzing a phonetic content of the one to N sentences; and analyzing an acoustic content of the speech of the one to N sentences; and outputting, by the system, measured amounts of the one or more different voice symptoms exhibited by the patient in the audio signal.
2. The method of claim 1 , further comprising associating, by the system, one or more neurological voice disorders with the one or more different voice symptoms exhibited by the patient and evident in the audio signal.
3. The method of claim 2, further comprising receiving, by the system, demographic information of the patient and using the demographic information of the patient when associating the one or more neurological voice disorders with the one or more different voice symptoms exhibited by the patient and evident in the audio signal.
4. The method of claim 1 , wherein the one or more different voice symptoms exhibited by the patient and evident in the audio signal are one or more of harshness, breathiness, tremor, or number of uncontrolled sentence-level breaks.
5. The method of claim 1 , wherein: the analyzing the phonetic content of the one to N sentences comprises extracting one or more phonetic difference features from the one to N sentences; and the analyzing the acoustic component comprises extracting one or more audio features from the one to N sentences.
6. The method of claim 1 , wherein, after the receiving the audio signal, the method further comprises preprocessing the audio signal, wherein preprocessing comprises: segmenting and identifying the one to N sentences; removing noise from the one to N sentences with a within-sample noise profile; and transforming the one to N sentences into spectrograms.
7. The method of claim 6, wherein the spectrograms are Log-Mel spectrograms.
8. The method of claim 6, wherein the measuring further comprises comparing the spectrograms to a tuned model to determine the one or more symptoms of the voice disorder, wherein the tuned model is tuned based on historical data for a population or historical data for the patient.
9. The method of claim 8, wherein the tuned model reveals predicted results, the spectrograms reveal actual results, and a difference between the predicted results and the actual results reveals the one or more different symptoms exhibited by the patient.
10. A system comprising: a recording device configured to record an audio signal comprising speech of one to N predefined sentences by a patient, where N is an integer value; a computing device in communication with the recording device and comprising: a non-transitory memory configured to store instructions and the audio signal; and a processor configured to access the memory and execute the instructions to at least: receive the audio signal and store the audio signal in the non- transitory memory; detect one or more different voice symptoms exhibited by the patient by analyzing: a phonetic content of the one to N sentences; andan acoustic content of the speech of the one to N sentences; and output measured amounts of the one or more different voice symptoms exhibited by the patient in the audio signal.11 . The system of claim 10, further comprising an output device configured to output the one or more different voice symptoms exhibited by the patient and / or a symptom severity of the one or more different voice symptom exhibited by the patient based on the amounts measured.
12. The system of claim 10, wherein the processor associates the one or more different voice symptoms exhibited by the patient and evident in the audio signal with one or more neurological voice disorders.
13. The system of claim 12, wherein the processor receives demographic information of the patient and uses the demographic information when associating the one or more different voice symptoms exhibited by the patient and evident in the audio signal with the one or more neurological voice disorders.
14. The system of claim 10, wherein the one or more different voice symptoms exhibited by the patient and evident in the audio signal are one or more of harshness, breathiness, tremor, or number of breaks.
15. The system of claim 10, wherein the processor analyzes the phonetic content of the one to N sentences by extracting phonetic difference features from the one to N sentences and the processor analyzes the acoustic component by extracting audio features from the one to N sentences.
16. The system of claim 10, wherein the processor is configured to execute the instructions to preprocess the audio signal by: segmenting and identifying the one to N sentences; removing noise from the one to N sentences with a within-sample noise profile; and transforming the one to N sentences into spectrograms.
17. The system of claim 16, wherein the spectrograms are Log-Mel spectrograms.
18. The system of claim 16, wherein the processor measures the amounts of the one or more different voice symptoms exhibited by the patient by comparing the spectrograms to a tuned model to determine the one or more symptoms of the voice disorder.
19. The system of claim 18, wherein the tuned model reveals predicted results, the spectrograms reveal actual results, and a difference between the predicted results and the actual results reveals the one or more symptoms.
20. The system of claim 19, wherein the tuned model is tuned based on historical data for a population or historical data for the patient.
Citation Information
Patent Citations
Language disorder diagnosis / screening
US20200160881A1
Mispronunciation detection with phonological feedback
US20210319786A1
Automated assessment of cognitive and speech motor impairment
US20230172526A1
Speech analysis for monitoring or diagnosis of a health condition
US20230255553A1