Speech function evaluation
The diagnostic device uses a smartphone to measure digital biomarkers from speech audio data to assess speech function, addressing the challenge of remote SMA monitoring by accurately tracking bulbar function and disease progression.
Patent Information
- Application Number
- JP2025519564
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-07
- Filing Date
- 2023-10-06
- Publication Date
- 2025-09-29
AI Technical Summary
Existing methods fail to effectively assess speech function in individuals with spinal muscular atrophy (SMA) due to the inability to measure speech level variations accurately in remote settings, which are indicative of bulbar function and potential disease progression.
A diagnostic device and method using a smartphone to measure digital biomarkers from speech audio data, including speaking rate, pronunciation accuracy, pauses, and speech level changes, applying algorithms to assess speech function and map it to bulbar function, enabling remote monitoring of SMA progression.
Enables effective tracking of SMA progression through remote speech assessments, correlating with clinical outcomes, facilitating early detection and personalized treatment.
Smart Images

Figure 2025532348000001_ABST
Abstract
Description
[Technical Field]
[0001] Technical field of the invention The present invention relates to a diagnostic device and computer-implemented method for assessing a subject's speech function. [Background technology]
[0002] Background of the Invention Spinal muscular atrophy (SMA) is associated with bulbar weakness. People living with SMA report difficulty speaking loudly (e.g., to be heard in noisy environments) and may experience shortness of breath while speaking. It is important to note that speech impairment does not appear to be a meaningful aspect of health—in a Roche-sponsored qualitative study using structured interviews, 0% of SMA patients and their families mentioned speech ability among their greatest difficulties. However, of seven healthcare professionals (HCPs) surveyed, three ranked speech impairment as important. More commonly, HCPs rated the need to measure bulbar function more than patients and their caregivers.
[0003] Furthermore, the Scientific Advisory Working Group (SAWG) recommended that combining measures from speech and respiratory assessments may help detect deterioration of bulbar function, which may be a precursor to serious events (such as aspiration). In other words, measures derived from speech-based assessments can serve as a leading indicator for hospitalization.
[0004] Passage reading is often used in the assessment of motor speech disorders as a controlled, repeatable approximation of spontaneous contextual speech. 1 Early signs of bulbar dysfunction in patients with amyotrophic lateral sclerosis (ALS) may be associated with slower speech rate, more frequent speech interruptions, and reduced articulatory accuracy. 2,3 . TIFF2025532348000002.tif108170
[0005] In addition, people with spinal muscular atrophy complain of difficulty speaking loudly, so the sound pressure level of speech 4 It is hypothesized that may be a further outcome measure.
[0006] In a remote patient monitoring setting using a smartphone, the mouth-to-microphone distance is unknown, and therefore absolute speech level cannot be inferred from the recorded audio signal. However, assuming that the mouth-to-microphone distance remains constant during the task, changes in speech level can be measured over the duration of the task. This is based on the idea that during prolonged reading efforts, people with SMA may experience fatigue or shortness of breath, which may manifest as a gradual decrease in speech level. Summary of the Invention
[0007] Summary of the Invention The present invention provides a diagnostic device and computer-implemented method for assessing a subject's speech function, and the output may be useful in assessing the subject's bulbar function and tracking the status or progression of diseases affecting bulbar function, such as (but not limited to) SMA.
[0008] More specifically, a first aspect of the present invention provides a diagnostic device configured to assess a speech function of a subject, the diagnostic device comprising at least one processor; a microphone; and a memory storing computer-readable instructions that, when executed by the at least one processor, cause the diagnostic device to: prompt the subject to perform a speaking aloud diagnostic task; receive, via the microphone, audio data associated with the diagnostic task; extract digital biomarker data associated with the subject's speech function from the audio data; and apply a speech function assessment model to the digital biomarker data, the speech function assessment model being configured to generate an output indicative of the subject's speech function based on the digital biomarker data.
[0009] By measuring speech function using the diagnostic device according to the first aspect of the present invention, it may be possible to effectively track the progression of various muscular disorders, such as SMA, in a subject through active testing of the subject. In particular, the computer-readable instructions, when executed by the processor, may further be configured to cause the diagnostic device to map the output indicative of the subject's speech function to a bulbar function assessment grade indicative of the subject's bulbar function. As described in detail later in this application, the diagnostic device according to the first aspect of the present invention may use the output indicative of the speech function and / or bulbar function assessment grade to indicate and / or track the presence or progression of a muscular disorder, such as SMA, in a subject or user.
[0010] In a preferred implementation, the device is or comprises a smartphone. This is advantageous since virtually everyone owns a smartphone these days. By performing a computer-implemented process such as that described on a smartphone, a user does not need to travel to, for example, a hospital or other clinical setting to have their speech function measured. Other types of diagnostic devices, such as tablets, laptop computers, desktop computers, etc., may also be used. Alternatively, the diagnostic device may be a dedicated speech function assessment device.
[0011] The computer-readable instructions, when executed by a processor, may cause the diagnostic device to prompt the subject to perform a diagnostic task by displaying text to be spoken by the subject and / or by outputting a portion of the text to be spoken by the subject. In some cases, the computer-readable instructions, when executed by at least one processor, may cause the diagnostic device to further display the text for a preview period before receiving audio data, to better inform the user of the text to be spoken before the user reads it aloud. During that time, the diagnostic device may be configured to indicate that audio data has not yet been received. This ensures that the subject does not begin reading aloud until a dedicated time. In some cases, the diagnostic device may include a display component, such as a touchscreen, including one or more sensors. Specifically, in response to a user tapping or touching the touchscreen at a location corresponding to a portion of the text to be spoken, the diagnostic device may be configured to generate an audio output corresponding to the text in the area touched or tapped by the subject. This may help the user pronounce the word displayed in that portion. This is particularly useful because reduced pronunciation accuracy may be a sign of reduced bulbar muscle function, which will be described in more detail shortly. In response to user input (via one or more sensors in the touchscreen or otherwise), the subject may be able to change the font size in which the spoken text is displayed. The computer-readable instructions, when executed by the processor, may further cause the diagnostic device to prompt the user to speak with as few interruptions as possible (e.g., to read the displayed text aloud). This may be particularly important because an increase in the number and / or duration of interruptions in the subject's speech may be indicative of decreased bulbar muscle function, which will be described in more detail shortly.
[0012] The nature of digital biomarker data and its extraction will now be described in more detail. Various types of digital biomarker data may be extracted from recorded speech data, and the list of examples provided below is by no means exhaustive. Essentially, the digital biomarker types parameterize different aspects of a subject's speech function, which may be affected by, for example, reduced bulbar muscle function as a result of SMA.
[0013] In some cases, the digital biomarker data may include the subject's speaking rate, which may be expressed in syllables, words, or phonemes per unit of time (e.g., per minute or per second). In such cases, extracting the digital biomarker data may include applying a speaking rate determination algorithm to the recorded audio data, the speaking rate determination algorithm configured to determine the subject's speaking rate. The algorithm may be configured to detect words, syllables, and / or phonemes in the recorded audio data and divide the number of detected words, syllables, and / or phonemes by the total duration of active speech (i.e., excluding speech pauses).
[0014] In some cases, digital biomarkers may include parameterized pronunciation accuracy, for example, in the form of the distance between the acoustic features of the subject's speech in the recorded speech data and a reference (i.e., highly intelligible) speech recording. A suitable measure of pronunciation accuracy is described in Janbakhshi et al. (2019) 5 , Bartelds et al. (2020) 6 , and Ullmann et al. (2015) 7 In such cases, extracting the digital biomarker data may include applying a pronunciation accuracy determination algorithm to the recorded speech data, the pronunciation accuracy determination algorithm being configured to determine the pronunciation accuracy of the subject. The algorithm may be an algorithm according to one of the above-cited references. TIFF2025532348000003.tif103170
[0015] In some cases, the digital biomarker data may include information about pauses in the subject's speech in the recorded audio data. For example, the digital biomarker data may include the total duration of pauses in the recorded audio data and / or a pause ratio (i.e., the total duration of pauses divided by the total duration of the recorded audio data, or more specifically, the total duration of speech in the recorded audio data). In such cases, extracting the digital biomarker data may further include applying a pause identification algorithm to the recorded audio data, the pause identification algorithm configured to identify pauses in the recorded audio data. When the digital biomarker data includes the total duration of pauses, extracting the digital biomarker may further include calculating the total duration of pauses in the recorded audio data based on the pauses identified by the pause identification algorithm. The pause identification algorithm may be configured to generate multiple timestamps indicating the start and end of each respective pause, and extracting the digital biomarker data may include calculating the total duration of pauses in the recorded audio data based on the generated multiple timestamps.
[0016] When the digital biomarker data includes a pause ratio, extracting the digital biomarker may further include calculating a total duration of pauses in the recorded audio data based on the pauses identified by the pause identification algorithm and dividing it by a total duration of speech in the recorded audio data. The pause identification algorithm may be configured to generate a plurality of timestamps indicating a start and an end of each respective pause, and extracting the digital biomarker data may include calculating a total duration of pauses in the recorded audio data based on the generated plurality of timestamps. Extracting the digital biomarker data may further include dividing the total duration of pauses by a total duration of speech in the recorded audio data.
[0017] In some cases, the digital biomarker data may include data indicating changes in a subject's speech level over time. This may include the slope of a linear fit to speech level over the duration of the task. The term "level" is used herein to refer to the absolute acoustic intensity of speech. Such intensity may be measured as either a physical quantity such as sound pressure level (i.e., local pressure deviation from ambient atmospheric pressure), a voltage fluctuation measured by a microphone, or a measure of acoustic intensity that mimics the characteristics of human hearing, such as A-weighted decibels (denoted "dB(A)") or perceived loudness. Because people with SMA may experience fatigue or shortness of breath during prolonged reading efforts, which may manifest as a gradual decrease in speech level, it is desirable to assess, track, or measure changes in level. In implementations in which speech level is monitored, the mouth-to-microphone distance is preferably constant. Thus, when executed by the processor, the computer readable instructions may cause the diagnostic device to prompt the user to maintain a certain distance between their mouth and the microphone when performing a diagnostic task, or to place the diagnostic device in a predetermined position.
[0018] In the cases outlined in the previous paragraph, extracting the digital biomarker data may include applying a speech level determination algorithm to the recorded audio data. In those cases, extracting the digital biomarker data may include determining changes in the subject's level of speech over time based on the output of the speech level determination algorithm. In particular, extracting the digital biomarker data may include plotting speech level against time, generating a linear fit, and extracting, deriving, determining, or otherwise calculating the slope of the linear fit.
[0019] In some cases, the digital biomarker data may include one or more Mel-Frequency Cepstral Coefficient (MFCC) values associated with the audio data. In particular, the digital biomarker data may include one or more low-order MFCC values associated with the audio data, for example, the digital biomarker data may include one or more first Mel-Frequency Cepstral Coefficient (MFCC1) values associated with the audio data. Thus, extracting the digital biomarker data may include calculating one or more MFCC values, low-order MFCC values, or MFCC1 values associated with the audio data.
[0020] When extracting digital biomarker data, it is desirable to extract data from relevant portions of the recorded audio data. For example, there may be a beginning, end, or middle portion of the audio recording corresponding to background noise before the subject begins speaking, after the subject has finished speaking, or between two segments of the subject's speech. In other words, the recorded audio data may include multiple segments, and the diagnostic device may be configured to perform a pre-processing step including applying an active speech detection algorithm to the audio data, the active speech detection algorithm being configured to classify the segments of the audio data into active speech segments and background noise segments. Classifying the segments of the audio data into active speech segments and background noise segments includes generating timestamps indicating the start and end times of each respective active speech segment and background noise segment.
[0021] Extraction of digital biomarker data (described in the previous paragraphs) may then proceed only on active speech segments.
[0022] In some examples, the computer-readable instructions, when executed by a processor, may cause the diagnostic device to perform a preprocessing step that includes resampling received audio data to a predetermined sampling rate. This preprocessing step may be performed before the preprocessing step of applying an active speech detection algorithm to the audio data, such that applying the active speech detection algorithm to the audio data may include applying the active speech detection algorithm to the resampled audio data. The predetermined sampling rate may be at least 10 kHz and no more than 20 kHz; for example, the predetermined sampling rate may be 16 kHz. In some cases, an active speech segment may include a voiced subsegment in which the vocal fold or vocal cord is actually vibrating, and an unvoiced subsegment in which the vocal cords are not vibrating. For example, a subject's vocal cords may vibrate while producing the sound "a" but not while producing the sound "sh." Typically, a voiced subsegment may be associated with a vowel. This is in contrast to a background noise segment in which the subject is not speaking. The diagnostic device may then be configured to apply a voiced subsegment detection algorithm to the active speech segments of the audio data, the voiced subsegment detection algorithm being configured to classify the subsegments into voiced and unvoiced speech subsegments. Classifying the subsegments of the active speech segments of the audio data into voiced and unvoiced speech segments includes generating timestamps indicating the start and end times of each respective voiced and unvoiced speech subsegment. Extraction of digital biomarker data, for example, relating to speech level or pronunciation accuracy, or one or more MFCC values, may then be performed only on the voiced speech subsegments to, for example, assess vocal cord phonation function, or vowel pronunciation accuracy, or articulatory range, respectively.
[0023] When the digital biomarker data includes one or more MFCC values, extracting the digital biomarker data may include extracting one or more lower-order MFCC values, such as one or more MFCC1 values, from each voiced speech sub-segment. In this case, the speech function assessment model applied to the digital biomarker data (i.e., the MFCC values) may be configured to calculate a variance of the MFCC values. An output indicative of the subject's speech function may correspond to the variance of the MFCC values. For example, the output indicative of the subject's speech function may correspond to the variance of the lower-order MFCC values, such as the variance of the MFCC1 values.
[0024] Lower-order MFCCs, such as MFCC1, may describe the overall shape of the speech spectrum. During vocalizations, the overall shape of the speech spectrum may be determined by the articulatory apparatus and resonances within the vocal tract. Therefore, lower MFCC variance values may indicate a decrease in articulatory range, such as may be caused by bulbar weakness. Therefore, the variance of lower-order MFCC values may be used to indicate the presence and / or progression of muscular disorders such as SMA, as described in more detail below.
[0025] In some cases, the computer-readable instructions, when executed by the processor, may further cause the device to receive noise data via a microphone, calculate background noise from the noise data, and apply a correction to the audio data using the background noise.
[0026] In some examples, the output indicative of the subject's speech function may correspond to digital biomarker data. For example, the output indicative of the speech function may correspond to a change in speaking rate, pronunciation accuracy, total duration of pauses, pause ratio, or level of speech of the subject.
[0027] We now describe how the output indicative of speech function may be used to indicate the presence or progression of a muscular disorder such as SMA. The computer-readable instructions, when executed by at least one processor, may cause the diagnostic device to apply a clinical interpretation model to the output indicative of speech function. The clinical interpretation model may be configured to output an indication of the presence or absence of a muscular disorder such as SMA in the user, or an indication of the progression of the muscular disorder in the user. The clinical interpretation model may be configured to compare the output indicative of speech function to a predetermined value and, based on the comparison, output an indication of the presence or absence of a muscular disorder such as SMA. In particular, the clinical interpretation model may be configured to determine whether the output indicative of speech function is greater than a predetermined threshold. In some examples, the clinical interpretation model may be configured to output an indication that a muscular disorder is present (e.g., the user has PlwSMA) if the output indicative of speech function is determined to be greater than the predetermined threshold, and / or to output an indication that a muscular disorder is not present if the output indicative of speech function is determined to be equal to or less than the predetermined threshold. In another example, the clinical interpretation model may be configured to output an indication that a muscle disorder is present (e.g., the user has PlwSMA) if the output indicative of the speech function is determined to be less than a predetermined threshold, and / or to output an indication that a muscle disorder is not present if the output indicative of the speech function is determined to be equal to or greater than the predetermined threshold. This may be the case when the output indicative of the speech function is a variance of the low-order MFCC values. When the output indicative of the speech function is a variance of the low-order MFCC values, the predetermined threshold may be less than or equal to 10,000. The predetermined threshold may be less than or equal to 9,000 or less than or equal to 8,000. The predetermined threshold may be greater than or equal to 6,000. The predetermined threshold may be greater than or equal to 7,000 or greater than or equal to 8,000.
[0028] For example, PlwSMA may exhibit significantly lower MFCC1 variance values compared to healthy subjects due to bulbar weakness in PlwSMA, as explained above.
[0029] A second aspect of the present invention provides a computer-implemented method for assessing a subject's speech function, the computer-implemented method comprising the steps of prompting the subject to perform a speaking aloud diagnostic task; receiving, via a microphone, audio data associated with the diagnostic task; extracting from the audio data digital biomarker data associated with the subject's speech function; and applying a speech function assessment model to the digital biomarker data, the speech function assessment model being configured to generate an output indicative of the subject's speech function based on the digital biomarker data. Preferably, the computer-implemented method of the second aspect of the present invention is executed by a processor of a diagnostic device, such as the diagnostic device of the first aspect of the present invention. It will be understood that any feature described above with respect to the first aspect of the present invention applies equally well to the second aspect of the present invention, unless the context clearly dictates otherwise, or unless combinations of such features are clearly technically incompatible.
[0030] A third aspect of the present invention provides a computer program comprising instructions which, when executed by a processor of a computer (or other suitable data processing device), cause the processor to perform the computer-implemented method of the second aspect of the invention. A further aspect of the present invention provides a computer-readable storage medium having stored thereon the computer program of the third aspect of the invention.
[0031] The present invention includes combinations of the described embodiments and preferred features except where such combinations are clearly unacceptable or explicitly avoided. [Brief explanation of the drawings]
[0032] Embodiments of the present invention will now be described with reference to the accompanying drawings.
[0033] [Figure 1] 1 is a diagram of an exemplary environment in which a diagnostic device for assessing a subject's speech function is provided. [Figure 2]1 is a flow diagram of a computer-implemented method for assessing a user's speech function. [Figure 3] 1 is a flow diagram of a computer-implemented method for determining an indication of the presence or absence of a muscle disorder such as SMA. [Figure 4] Plot showing MFCC1 variance values calculated for PlwSMA and healthy individuals. [Figure 5] FIG. 1 illustrates an example of a network architecture and data processing device that may be used to implement one or more example aspects described herein. DETAILED DESCRIPTION OF THE INVENTION
[0034] Detailed Description of the Drawings Aspects and embodiments of the present invention will now be described with reference to the accompanying figures. Further aspects and embodiments will be apparent to those skilled in the art. All documents mentioned in this text are incorporated herein by reference.
[0035] In the following description of various aspects, reference is made to the accompanying drawings which form a part hereof, and which show, by way of illustration, various embodiments in which the aspects described herein may be practiced. It is to be understood that other aspects and / or embodiments may be utilized and structural and functional modifications may be made without departing from the scope of the described aspects and embodiments.
[0036] The aspects described herein are capable of other embodiments and of being practiced or carried out in various ways. It is also to be understood that the terminology and phraseology used herein is for the purpose of description and should not be regarded as limiting. Rather, the words and terms used herein should be given their broadest interpretation and meaning. The use of "including" and "comprising" and variations thereof is intended to encompass the items listed thereafter and their equivalents as well as additional items and their equivalents. The use of the terms "mounted," "connected," "coupled," "disposed," "engaged," and similar terms is intended to encompass both direct and indirect mounting, connecting, coupling, disposing, and engaging.
[0037] The systems, methods, and devices described herein provide diagnostic devices and computer-implemented methods for evaluating, measuring, or determining the speech function of patients, particularly those suffering from muscular disorders such as SMA. In some cases, the diagnostic device may be in the form of a mobile device, particularly a smartphone, with a specific software application installed. The software application may be configured to execute (or cause a processor of the mobile device to execute) a corresponding computer-implemented method.
[0038] In some cases, the diagnostic device acquires or receives sensor data from one or more sensors associated with the mobile device when the subject uses the mobile device to interact with the software application. In some cases, the sensors may be within the mobile device. In some cases, data indicative of the subject's speech function is derived, calculated, or extracted from the received or acquired sensor data. In some cases, an assessment of the severity and progression of symptoms of muscle disorders, particularly SMA, in the subject may be determined based on the extracted sensor features.
[0039] In an embodiment of the present invention, a diagnostic device may prompt a subject to perform a diagnostic task. In some cases, the diagnostic task is fixed or modeled after an established method and standardized test. In some cases, in response to the subject performing the diagnostic task, the diagnostic device acquires or receives sensor data via one or more sensors. In some cases, the sensor may be in a mobile device or a wearable sensor worn by the subject. In some cases, sensor features associated with symptoms of myopathy, particularly SMA, are extracted from the received or acquired sensor data. In some cases, an assessment of the severity and progression of symptoms of myopathy, particularly SMA, in the subject is determined based on the extracted features of the sensor data.
[0040] Assessment of the severity and progression of symptoms of myopathies, particularly SMA, using a diagnostic according to the present disclosure correlates well with assessments based on clinical outcomes and may therefore replace clinical subject monitoring and testing. Exemplary diagnostics according to the present disclosure may be used outside of a clinical setting, thus offering advantages in cost, ease of subject monitoring, and convenience for the subject. This facilitates frequent, particularly daily, subject monitoring and testing, resulting in a better understanding of disease stage and providing insights into the disease that are useful to both the clinical and research communities. Exemplary diagnostic devices according to the present disclosure may be able to detect even small changes in a subject's speech function early, which may indicate the presence and progression of myopathies, particularly SMA, in the subject, and thus can be used for better disease management, including personalized treatment.
[0041] FIG. 1 is a diagram of an exemplary environment in which a diagnostic device 105 for assessing a subject's 110 speech function is provided. In some cases, the device 105 may be a smartphone, smartwatch, or other mobile computing device. The device 105 includes a display screen 160. In some cases, the display screen 160 may be a touchscreen. The device 105 includes at least one processor 115 and a memory 125 that stores computer instructions for a symptom monitoring application 130 that, when executed by the at least one processor 115, causes the device 105 to assess the subject's speech function. The device 105 receives a plurality of sensor data via one or more sensors associated with the device 105. In some cases, the one or more sensors associated with the device are at least one of sensors disposed within the device or sensors worn by the subject and configured to communicate with the device. In FIG. 1, the sensors associated with the device 105 include a first sensor 120, such as a microphone disposed within the device 105.
[0042] The device 105 extracts digital biomarker data from the received first sensor data that can be used to determine the subject's speech function.
[0043] The device 105 determines the speech function of the subject 110 based on the extracted features. In some cases, the device 105 transmits the extracted features to the server 150 via the network 180. In some cases, the device 105 transmits first sensor data to the server 150 via the network 180. The server 150 includes at least one processor 155 and a memory 161 that stores computer instructions for a symptom assessment application 170 that, when executed by the server's processor 155, causes the processor 155 to determine the speech function of the subject 110 based on the extracted features received by the server 150 from the device 105. In some cases, the symptom assessment application 170 may cause the processor 115 to extract features from the sensor data received from the device 105. In some cases, the symptom assessment application 170 may determine the speech function of the subject 110 based on the extracted features of the sensor data, which may be received from the device 105, and a subject database 175 stored in the memory 160. In some cases, subject database 175 may include subject data and / or clinical data. In some cases, subject database 175 may include in-clinic and sensor-based measures of speech function. In some cases, subject database 175 may be separate from server 150. In some cases, server 150 transmits the determined speech function of subject 110 to device 105. In some cases, device 105 may output the speech function of subject 110. In some cases, device 105 may communicate information to subject 110 based on the assessment. In some cases, the assessment of subject 110's speech function may be communicated to a clinician, who may determine a personalized treatment for subject 110 based on the assessment.
[0044] In some cases, the computer instructions for the symptom monitoring application 130, when executed by the at least one processor 115, cause the device 105 to determine the speech function of the subject 110 based on active testing of the subject 110. The device 105 prompts the subject 110 to perform one or more tasks. In some cases, prompting the subject to perform the one or more diagnostic tasks includes prompting the subject to make as long a continuous "aaah" sound as possible.
[0045] In response to the subject 110 performing one or more diagnostic tasks, the diagnostic device 105 receives a plurality of sensor data via one or more sensors associated with the device 105. The device 105 extracts various digital biomarker data from the received sensor data from which an assessment of the speech function of the subject 110 may be made. Symptoms of muscle disorders, particularly SMA, in the subject 110 may include symptoms that affect the speech function of the subject 110.
[0046] Figure 2 shows an exemplary method for assessing speech function of a subject 110 based on active testing of the subject using the exemplary device 105 of Figure 1. While Figure 2 is described with reference to Figure 1, it should be noted that the method steps of Figure 2 may be performed by other systems. The computer-implemented method includes, in step 205, prompting the subject to perform the diagnostic tasks outlined above. The method includes receiving a plurality of sensor data (step 210), for example via a microphone, in response to the subject performing one or more tasks.
[0047] Then, in step 215, digital biomarker data is extracted from the sensor data and a speech function assessment model is applied to the digital biomarker data.
[0048] In step 220, data indicative of the subject's speech function is output, for example, by processor 107 generating instructions that, when executed by display component 160 of device 105, cause display component 160 to display output indicative of the subject's speech function. Alternatively, as outlined elsewhere in this application, the calculated data indicative of the subject's speech function may be transmitted to server 150.
[0049] As noted above, assessment of the severity and progression of symptoms of muscle disorders, particularly SMA, using diagnostics according to the present disclosure correlates well with clinical outcome-based assessments and may therefore replace clinical subject monitoring and testing.
[0050] FIG. 3 illustrates an exemplary method for determining an indication of the presence or absence of SMA in a subject based on active testing of the subject using the exemplary device 105 of FIG. 1. While FIG. 3 is described with reference to FIG. 1, it should be noted that the method steps of FIG. 3 may be performed by other systems. The computer-implemented method includes, at step 235, calculating a variance value of the MFCC1 values of the audio data. This variance value corresponds to the output indicative of speech function in the method described with reference to FIG. 2. At step 240, the computer-implemented method includes determining whether the calculated variance value is greater than a predetermined threshold. If it is determined that the calculated variance value is greater than the predetermined threshold, then at step 250, the computer-implemented method includes outputting an indication that SMA is absent. If it is determined that the calculated variance value is equal to or less than the predetermined threshold, then at step 245, the computer-implemented method includes outputting an indication that SMA is present.
[0051] For example, the predetermined threshold may be approximately 9,000. That is, a variance value >9,000 may indicate that the user is a healthy individual, and a variance value ≦9,000 may indicate that the user has PlwSMA. This may be explained with reference to FIG. 4, which is a plot showing variance values of MFCC1 values calculated for PlwSMA and healthy individuals. The plot shows that a majority of PlwSMA individuals may have a test result of a variance value of MFCC1 values ≦9,000, while a majority of healthy individuals may have a test result of a variance value of MFCC1 values >9,000.
[0052] 5 illustrates an example of a network architecture and data processing devices that may be used to implement one or more exemplary embodiments described herein, such as those described in FIGS. 1 and 2. Various network nodes 303, 305, 307, and 309 may be interconnected via a wide area network (WAN) 301, such as the Internet. Other networks, including private intranets, corporate networks, LANs, wireless networks, personal networks (PANs), etc., may also or instead be used. Network 301 is for illustrative purposes and may be replaced by fewer or additional computer networks. The local area network (LAN) may have one or more of any known LAN topologies and may use one or more of a variety of different protocols, such as Ethernet. Devices 303, 305, 307, 309, and other devices (not shown) may be connected to one or more of the networks via twisted pair wire, coaxial cable, optical fiber, radio waves, or other communication media.
[0053] As used herein and depicted in the drawings, the term "network" refers not only to a system in which remote storage devices are coupled together via one or more communication paths, but also to stand-alone devices that may be coupled to such a system from time to time with storage capabilities. Consequently, the term "network" includes not only a "physical network" but also a "content network" consisting of data—attributable to a single entity—that resides across all physical networks.
[0054] The components may include a data server 303, a web server 305, and client computers 307, 309. The data server 303 provides overall access, control, and management of the databases and control software for implementing one or more exemplary embodiments described herein. The data server 303 may be connected to a web server 305 through which users interact and obtain requested data. Alternatively, the data server 303 may itself function as a web server and be directly connected to the Internet. The data server 303 may be connected to the web server 305 via a network 301 (e.g., the Internet) via a direct or indirect connection, or via some other network. Users may interact with the data server 303 using remote computers 307, 309, for example, using a web browser, to connect to the data server 303 through one or more public-facing websites hosted by the web server 305. The client computers 307, 309 may be used in conjunction with the data server 303 to access data stored in the data server 303, or may be used for other purposes. For example, from client device 307, a user may access web server 305 using an internet browser, as known in the art, or by executing a software application that communicates with web server 305 and / or data server 303 over a computer network (such as the internet). In some cases, client computer 307 may be a smartphone, smartwatch, or other mobile computing device and may implement a diagnostic device such as device 105 shown in FIG. 1. In some cases, data server 303 may implement a server such as server 150 shown in FIG. 1.
[0055] The servers and applications may be combined on the same physical machine, maintain separate virtual or logical addresses, or reside on separate physical machines. Figure 1 illustrates only one example of a network architecture that may be used, and those skilled in the art will appreciate that the particular network architecture and data processing devices used may vary and are secondary to the functions they provide, as described further herein. For example, the services provided by web server 305 and data server 303 may be combined on a single server.
[0056] Each of the components 303, 305, 307, and 309 may be any type of known computer, server, or data processing device. The data server 303 may include, for example, a processor 311 that controls the overall operation of the rate server 303. The data server 303 may further include RAM 313, ROM 315, a network interface 317, an input / output interface 319 (e.g., a keyboard, a mouse, a display, a printer, etc.), and memory 321. The I / O 319 may include various interface units and drivers for reading, writing, displaying, and / or printing data or files. The memory 321 may further store operating system software 323 for controlling the overall operation of the data processing device 303, control logic 325 for directing the data server 303 to perform aspects described herein, and other application software 327 that provides secondary support and / or other functions that may or may not be used in conjunction with other aspects described herein. The control logic may also be referred to herein as data server software 325. The functionality of data server software may refer to actions or decisions that are performed automatically based on rules coded into control logic, actions or decisions that are performed manually by a user providing input to the system, and / or a combination of automated processing based on user input (e.g., queries, data updates, etc.).
[0057] Memory 321 may store data used to perform one or more aspects described herein, including first database 329 and second database 331. In some cases, the first database may include a second database (e.g., as a separate table, report, etc.). That is, information may be stored in a single database or separated into different logical, virtual, or physical databases, depending on the system design. Devices 305, 307, and 309 may have an architecture similar to or different from that described with respect to device 303. Those skilled in the art will understand that the functionality of data processing device 303 (or devices 305, 307, and 309) described herein may be distributed across multiple data processing devices, for example, to distribute processing load across multiple computers, segregate transactions based on geographic location, user access level, quality of service (QoS), etc.
[0058] One or more aspects described herein may be embodied in computer-usable or readable data and / or computer-executable instructions, such as in one or more program modules executed by one or more computers or other devices described herein. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types when executed by a processor within a computer or other device. Modules may be written in source code programming languages that are later compiled for execution, or may be written in scripting languages such as (but not limited to) HTML or XML. Computer-executable instructions may be stored in a computer-readable medium such as a hard disk, optical disk, removable storage media, solid-state memory, RAM, etc. As will be appreciated by those skilled in the art, the functionality of the program modules may be combined or distributed as desired in various embodiments. Additionally, the functionality may be embodied, in whole or in part, in firmware or hardware equivalents, such as integrated circuits, field programmable gate arrays (FPGAs), etc. Particular data structures may be used to more efficiently implement one or more aspects, and such data structures are contemplated within the scope of computer-executable instructions and computer-usable data described herein.
[0059] The features disclosed in the foregoing description, or the following claims, or the accompanying drawings, and expressed in their specific form or as means for performing a disclosed function or as methods or processes for obtaining a disclosed result, may be utilized, individually or in any combination of such features, as appropriate, to realize the invention in diverse forms thereof.
[0060] While the present invention has been described in conjunction with the exemplary embodiments set forth above, many equivalent modifications and variations will be apparent to those skilled in the art given this disclosure. Accordingly, the exemplary embodiments of the present invention set forth above are considered to be illustrative and not limiting. Various changes may be made to the described embodiments without departing from the spirit and scope of the invention.
[0061] To avoid any misunderstanding, any theoretical explanations provided herein are provided for the purpose of enhancing the understanding of the reader, and the inventors do not wish to be bound by any of these theoretical explanations.
[0062] Any section headings used herein are for organizational purposes only and should not be construed as limiting the subject matter described.
[0063] Throughout this specification, including the claims which follow, unless the context requires otherwise, the words "comprises" and "includes," and variations such as "comprises," "comprising," and "including," will be understood to imply the inclusion of stated integers or steps or groups of integers or steps, but not the exclusion of any other integers or steps or groups of integers or steps.
[0064] It should be noted that, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein as from "about" one particular value and / or to "about" another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by use of the antecedent "about," it will be understood that the particular value forms another embodiment. The term "about" with respect to numerical values is arbitrary and means, for example, + / - 10%.
Claims
1. 1. A diagnostic device configured to assess speech function of a subject, comprising: at least one processor; A microphone and a memory that, when executed by the at least one processor, prompting the subject to perform a speak-aloud diagnostic task; receiving, via the microphone, audio data associated with the diagnostic task; extracting digital biomarker data associated with the speech function of the subject from the audio data; applying a speech function assessment model to the digital biomarker data, the speech function assessment model configured to generate an output indicative of the speech function of the subject; a memory storing computer readable instructions that cause the diagnostic device to perform the steps of: A diagnostic device comprising:
2. the computer-readable instructions, when executed by the processor, cause the diagnostic device to prompt the subject to perform the diagnostic task by displaying text spoken by the subject or by outputting aloud a portion of the text spoken by the subject. The diagnostic device of claim 1 .
3. the digital biomarker data includes the subject's speaking rate in units of syllables, words, or phonemes per unit time; extracting the digital biomarker data includes applying a speaking rate determination algorithm to the recorded audio data, the speaking rate determination algorithm configured to determine the speaking rate of the subject. A diagnostic device according to claim 1 or claim 2.
4. the digital biomarker data includes pronunciation accuracy; 4. The diagnostic device of claim 1, wherein extracting the digital biomarker data comprises applying a pronunciation accuracy determination algorithm to the recorded audio data, the pronunciation accuracy determination algorithm being configured to determine the subject's pronunciation accuracy.
5. 5. The diagnostic device of claim 1, wherein extracting the digital biomarker data comprises applying a pause identification algorithm to the recorded audio data, the pause identification algorithm configured to identify pauses in the recorded audio data.
6. the digital biomarker data includes a total duration of interruptions in the recorded audio data; and extracting the digital biomarker data further comprises calculating a total duration of pauses in the recorded audio data based on the pauses identified by the pause identification algorithm. The diagnostic device of claim 5.
7. The diagnostic device of any one of claims 1 to 6, wherein extracting the digital biomarker data comprises applying a speech level determination algorithm to the recorded audio data.
8. the digital biomarker data includes data indicative of a change in the subject's level of speech over time; extracting the digital biomarker data includes determining a change in the level of speech of the subject over time based on the output of the speech level determination algorithm. The diagnostic device of claim 7.
9. 9. The diagnostic device of claim 1, wherein the computer-readable instructions, when executed by the processor, cause the diagnostic device to perform a pre-processing step, the pre-processing step comprising applying an active speech detection algorithm to the audio data, the active speech detection algorithm being configured to classify segments of the audio data into active speech segments and background noise segments.
10. 10. The diagnostic device of claim 9, wherein the computer-readable instructions, when executed by the processor, cause the diagnostic device to apply a voiced subsegment detection algorithm to the active speech segments of the audio data, the voiced speech subsegment detection algorithm being configured to classify the subsegments into voiced and unvoiced speech subsegments.
11. 11. The diagnostic device of claim 10, wherein extracting the digital biomarker data comprises extracting one or more low-order Mel-Frequency Cepstral Coefficient (MFCC) values from each voiced speech sub-segment.
12. The diagnostic device of claim 11, wherein the speech function assessment model is configured to calculate the variance of the low-order MFCC values so that the output indicating the speech function of the subject corresponds to the variance of the low-order MFCC values.
13. 13. The diagnostic device of claim 1, wherein the computer readable instructions, when executed by the at least one processor, cause the diagnostic device to apply a clinical interpretation model to the output indicative of the speech function, the clinical interpretation model outputting an indication of the presence or absence of a muscle disorder.
14. 14. The diagnostic device of claim 13, wherein the clinical interpretation model is configured to compare the output indicative of the speech function with a predetermined value and to output an indication of the presence or absence of the myopathy based on the comparison.
15. the clinical interpretation model: determining whether the output indicative of the speech function is greater than a predetermined threshold; outputting an indication that a muscle disorder exists when the output indicating the speech function is determined to be greater than the predetermined threshold; When it is determined that the output indicating the speech function is equal to or less than the predetermined threshold, an indication that the muscle disorder does not exist is output. The diagnostic device of claim 14 configured to:
16. the clinical interpretation model: determining whether the output indicative of the speech function is less than a predetermined threshold; outputting an indication that a muscle disorder exists when the output indicating the speech function is determined to be smaller than the predetermined threshold; When it is determined that the output indicating the speech function is equal to or greater than the predetermined threshold, a message indicating that the muscle disorder does not exist is output. The diagnostic device of claim 14 configured to:
17. 17. The diagnostic device of claim 16 when dependent on claim 12, wherein the predetermined threshold is less than or equal to 10,000 and greater than or equal to 6,000.
18. 1. A computer-implemented method for assessing speech function of a subject, comprising: prompting the subject to perform a speak-aloud diagnostic task; receiving, via a microphone, audio data associated with the diagnostic task; extracting digital biomarker data associated with the speech function of the subject from the audio data; applying a speech function assessment model to the digital biomarker data, the speech function assessment model configured to generate an output indicative of the speech function of the subject based on the digital biomarker data; 11. A computer-implemented method comprising:
19. 20. The computer-implemented method of claim 18, further comprising applying a clinical interpretation model to the output indicative of the speech function, the clinical interpretation model outputting an indication of the presence or absence of a myopathy or an indication of the progression of a myopathy.
20. A computer-implemented method according to claim 18 or claim 19, wherein the computer-implemented method is executed by a processor of a diagnostic device according to any one of claims 1 to 17.
21. The steps of prompting the subject and receiving the speech data are performed by a processor of a diagnostic device, and the steps of extracting digital biomarker data and applying the speech function assessment model are performed by a processor of a server, and the diagnostic device is configured to transmit the speech data to the server, and the diagnostic device: at least one processor; A microphone and a memory that, when executed by the at least one processor, prompting the subject to perform a speak-aloud diagnostic task; receiving, via the microphone, audio data associated with the diagnostic task; a memory storing computer readable instructions that cause the diagnostic device to perform the steps of:
20. The computer-implemented method of claim 18 or claim 19, comprising: