Speech function assessment

By using diagnostic devices and computer-implemented methods, digital biomarkers in subject audio data are extracted and analyzed, and their speech functions are evaluated, which solves the problem of difficulty in evaluating and monitoring speech functions in patients with muscular disability in the prior art, and effectively tracking speech function changes and providing individualized treatment plans.

CN119998879APending Publication Date: 2025-05-13F HOFFMANN LA ROCHE & CO AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380070670.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-07
Filing Date
2023-10-06
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

It is difficult to effectively assess and monitor the speech function of patients with muscle disabilities such as spinal muscular atrophy (SMA). Especially in noisy environments or during prolonged reading, patients may experience speech disorders and shortness of breath.

Method used

It provides a diagnostic device and computer implementation method, which receives the loud speaking audio data of the subject through a microphone, extracts digital biomarker data, including indicators such as speech speed, pronunciation accuracy and speech pause, and generates evaluation output using a speech function evaluation model.

Benefits of technology

Effectively track changes in subjects’ speech function through active testing, help diagnose and monitor symptoms of muscle disability, provide individualized treatment options, and replace traditional clinical subject monitoring and testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119998879A_ABST
    Figure CN119998879A_ABST
Patent Text Reader

Abstract

The present invention provides a diagnostic device configured to assess speech function of a subject, the diagnostic device comprising: at least one processor; a microphone; and a memory storing computer readable instructions that, when executed by the at least one processor, cause the diagnostic device to: prompt the subject to perform a diagnostic task of speaking loud; receiving audio data associated with the diagnostic task via the microphone; extracting digital biomarker data associated with the speech function of the subject from the audio data; and applying a verbal function assessment model to the digital biomarker data, the verbal function assessment model configured to generate an output indicative of the verbal function of the subject.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a diagnostic apparatus and a computer-implemented method for assessing the speech function of a subject. Background Art

[0002] Spinal muscular atrophy (SMA) is associated with bulbar weakness. People with SMA are known to have difficulty speaking loudly (e.g., being heard in a noisy environment) and may experience shortness of breath when speaking. Notably, speech impairment does not appear to be a meaningful aspect of well-being – in a qualitative study sponsored by Roche using structured interviews, 0% of SMA patients and their families listed speech impairment as their top difficulty. However, 3 of 7 healthcare professionals (HCPs) surveyed listed speech difficulties as an important issue. More generally, HCPs were perceived to have a greater need to measure bulbar abilities than patients and their caregivers.

[0003] In addition, the Scientific Advisory Working Group (SAWG) suggests that combining measurements from speech and respiratory assessments could help detect worsening bulbar function that could foreshadow critical events such as aspiration. In other words, measurements derived from speech-based assessments could serve as leading indicators for hospitalization.

[0004] Paragraph reading is commonly used to assess motor speech disorders as a controlled and reproducible approximation of spontaneous, contextual speech. 1 Early signs of bulbar dysfunction in patients with amyotrophic lateral sclerosis (ALS) may be associated with slower speech rate, more frequent speech pauses and reduced articulatory accuracy, research suggests. 2,3

[0005]

[0006] In addition, since people with spinal muscular atrophy are known to have difficulty speaking loudly, the sound pressure level of speech was estimated 4 Possibly further outcome measures.

[0007] In a remote patient monitoring setting using a smartphone, the distance from the mouth to the microphone is unknown, so the absolute speech level cannot be inferred from the recorded audio signal. However, assuming that the distance from the mouth to the microphone remains constant during the task, we can measure the change in speech level over the duration of the task. This is based on the idea that during long reading sessions, people with SMA may experience fatigue or shortness of breath, which may manifest as a gradual decrease in speech level. Summary of the invention

[0008] The present invention provides a diagnostic device and computer-implemented method for assessing speech function in a subject. The output can be used to assess bulbar function in a subject and track the status or progression of a condition affecting bulbar function, such as (but not limited to) SMA.

[0009] More specifically, the first aspect of the present invention provides a diagnostic device configured to assess the speech function of a subject, the diagnostic device comprising: at least one processor; a microphone; and a memory storing computer-readable instructions which, when executed by the at least one processor, cause the diagnostic device to: prompt the subject to perform a diagnostic task of speaking out loud; receive audio data associated with the diagnostic task via the microphone; extract digital biomarker data associated with the speech function of the subject from the audio data; and apply a speech function assessment model to the digital biomarker data, the speech function assessment model being configured to generate an output indicative of the speech function of the subject based on the digital biomarker data.

[0010] By measuring speech function using a diagnostic device according to the first aspect of the present invention, the progress of various muscle disabilities (such as SMA) in the subject can be effectively tracked by actively testing the subject. Specifically, the computer-readable instructions, when executed by the processor, can be further configured to cause the diagnostic device to map the output indicating the speech function of the subject to the medullary function assessment grade indicating the medullary function of the subject. As described in detail later in this application, the diagnostic device according to the first aspect of the present invention can use the output indicating speech function and / or medullary function assessment grade to indicate and / or track the presence or progress of muscle disabilities (such as SMA) in a subject or user.

[0012] In a preferred embodiment, the device is or includes a smartphone. This is advantageous because almost everyone has a smartphone today. By implementing a computer-implemented process (such as the process described on a smartphone), the user does not need to go to, for example, a hospital or other clinical setting in order to measure speech function. Other types of diagnostic devices may be used, such as tablet computers, laptop computers, desktop computers, etc. Alternatively, the diagnostic device may be a dedicated speech function assessment device.

[0013] Computer readable instructions, when executed by a processor, cause the diagnostic device to prompt the subject to perform a diagnostic task by displaying the text to be spoken aloud by the subject and / or audibly outputting a part of the text to be spoken aloud by the subject. In some cases, in order to allow the user to better familiarize himself with the text to be spoken aloud before reading the text, the computer readable instructions, when executed by at least one processor, can further cause the diagnostic device to display the text in a preview period before receiving audio data. During this time, the diagnostic device can be configured to indicate that the audio data has not yet been received. This ensures that the subject does not start reading until the specified time. In some cases, the diagnostic device may include a display component, such as a touch screen, which includes one or more sensors. Specifically, in response to a user tapping or touching the touch screen at a position corresponding to a part of the text to be spoken aloud, the diagnostic device may be configured to generate an audio output corresponding to the text in the area touched or tapped by the subject. This can help the user pronounce the words displayed at the part. This is particularly useful because a decrease in pronunciation accuracy may be a symptom of decreased bulbar muscle function, which will be discussed in more detail later. In response to user input (via one or more sensors in the touch screen or otherwise), the subject can also change the font size of the text to be spoken aloud. The computer readable instructions, when executed by the processor, may further cause the diagnostic device to prompt the user to speak aloud (e.g., read aloud the displayed text) with as few interruptions as possible. This may be particularly important because an increase in the number and / or duration of pauses in the subject's speech may be a symptom of decreased bulbar muscle function, which will be discussed in more detail later.

[0014] We now discuss the nature of digital biomarker data and its extraction in more detail. Various types of digital biomarker data can be extracted from recorded audio data, and the list of examples listed below is by no means exhaustive. Essentially, the types of digital biomarkers parameterize various aspects of a subject's speech function that may be affected by decreased bulbar muscle function, for example, as a result of SMA.

[0015] In some cases, the digital biomarker data may include the subject's speech rate. This can be expressed in syllables, words, or phonemes per unit time (e.g., per minute or per second). In such cases, extracting the digital biomarker data may include applying a speech rate determination algorithm to the recorded audio data, the speech rate determination algorithm being configured to determine the subject's speech rate. The algorithm may be configured to detect words, syllables, and / or phonemes within the recorded audio data, and the number of detected words, syllables, and / or phonemes divided by the total duration of valid speech (i.e., excluding speech pauses).

[0016] In some cases, the digital biomarker may include pronunciation accuracy, which is parameterized, for example, in the form of the distance between the acoustic features of the subject's speech in the recorded audio data and a reference (i.e., highly intelligible) speech recording. Suitable pronunciation accuracy measures are listed in: Janbakhshi et al. (2019) 5 , Bartelds et al. (2020) 6 , and Ullmann et al. (2015) 7 In such cases, extracting the digital biomarker data may include applying a pronunciation accuracy determination algorithm to the recorded audio data, the pronunciation accuracy determination algorithm being configured to determine the pronunciation accuracy of the subject. The algorithm may be an algorithm according to one of the previously cited references.

[0017] In some cases, the digital biomarker data may include information about pauses in the subject's speech in the recorded audio data. For example, the digital biomarker data may include the total duration of pauses in the recorded audio data, and / or a pause rate (i.e., the total duration of pauses divided by the total duration of the recorded audio data or, more specifically, the total duration of speech in the recorded audio data). In such cases, extracting the digital biomarker data may further include applying a pause recognition algorithm to the recorded audio data, the pause recognition algorithm being configured to identify pauses within the recorded audio data. When the digital biomarker data includes the total duration of pauses, extracting the digital biomarker may further include calculating the total duration of pauses in the recorded audio data based on the pauses identified by the pause recognition algorithm. The pause recognition algorithm may be configured to generate multiple timestamps indicating the start and end of each respective pause, and

[0018]

[0019] And extracting the digital biomarker data may include calculating a total duration of pauses in the recorded audio data based on the generated plurality of timestamps.

[0020] When the digital biomarker data includes a pause rate, extracting the digital biomarker may further include calculating the total duration of pauses in the recorded audio data based on the pauses identified by the pause recognition algorithm and dividing it by the total duration of speech in the recorded audio data. The pause recognition algorithm may be configured to generate a plurality of timestamps indicating the start and end of each respective pause, and extracting the digital biomarker data may include calculating the total duration of pauses in the recorded audio data based on the generated plurality of timestamps. Extracting the digital biomarker data may further include dividing the total duration of pauses by the total duration of speech in the recorded audio data.

[0021] In some cases, the digital biomarker data may include data indicating the level of speech of the subject over time. This may include the slope of a linear fit to the speech level over the duration of the task. In this article, the term "level" is used to refer to the absolute sound intensity of speech. This intensity can be measured as a physical quantity, such as sound pressure level (i.e., local pressure deviation from ambient atmospheric pressure), a voltage change measured by a microphone, or a sound intensity measurement result that simulates the characteristics of human hearing, such as A-weighted decibels (expressed as "dB (A)") or perceived loudness. It is desirable to assess, track or measure the change in level because during a long reading process, people with SMA may experience fatigue or shortness of breath, which may be manifested as a gradual decrease in speech level. In the specific implementation of monitoring speech level, the distance from the mouth to the microphone is preferably constant. Therefore, when the computer-readable instructions are executed by the processor, the diagnostic device can prompt the user to maintain a constant distance between their mouth and the microphone when performing a diagnostic task, or place the diagnostic device in a predetermined position.

[0022] In the cases outlined in the previous paragraph, extracting digital biomarker data can include applying a speech level determination algorithm to the recorded audio data. In these cases, extracting digital biomarker data can include determining the level of the subject's speech over time based on the output of the speech level determination algorithm. Specifically, extracting digital biomarker data can include plotting speech level versus time, generating a linear fit, and extracting, deriving, determining, or otherwise calculating the slope of the linear fit.

[0023] In some cases, the digital biomarker data may include one or more Mel-frequency cepstral coefficient (MFCC) values ​​associated with the audio data. Specifically, the digital biomarker data may include one or more low-order MFCC values ​​associated with the audio data, for example, the digital biomarker data may include one or more first Mel-frequency cepstral coefficient (MFCC 1) values ​​associated with the audio data. Therefore, extracting the digital biomarker data may include calculating one or more MFCC values, low-order MFCC values, or MFCC 1 values ​​associated with the audio data.

[0024] When extracting digital biomarker data, it is desirable to extract data from relevant parts of the recorded audio data. For example, there may be parts corresponding to background noise at the beginning, end, or middle of the recorded data, which are from before the subject starts speaking, after the subject has finished speaking, or between two segments of the subject's speech. In other words, the recorded audio data may include multiple segments, and the diagnostic device may be configured to perform a pre-processing step, which includes applying an effective speech detection algorithm to the audio data, and the effective speech detection algorithm is configured to classify the segments of the audio data into effective speech segments and background noise segments. Classifying the segments of the audio data into effective speech segments and background noise segments includes generating timestamps indicating the start time and end time of each corresponding effective speech segment and background noise segment.

[0025] Then, extraction of digital biomarker data (as explained in the previous paragraphs) can be performed only on valid speech segments.

[0026] In some examples, the computer-readable instructions, when executed by a processor, may cause the diagnostic device to perform a preprocessing step, the preprocessing step comprising resampling the received audio data to a predetermined sampling rate. The preprocessing step may be performed before the preprocessing step of applying the effective speech detection algorithm to the audio data, so that applying the effective speech detection algorithm to the audio data may include applying the effective speech detection algorithm to the resampled audio data. The predetermined sampling rate may be at least 10kHz and not greater than 20kHz, for example, the predetermined sampling rate may be 16kHz. In some cases, the effective speech segment may include a phonic subsegment in which the vocal cords or vocal folds actually vibrate and a non-phonic subsegment in which the vocal cords do not vibrate. For example, the vocal cords of a subject may vibrate when an "a" sound is uttered, but not when an "sh" sound is uttered. Typically, the phonic subsegment may be associated with a vowel. This is in contrast to a background noise segment in which the subject is not speaking. The diagnostic device may then be configured to apply the phonic subsegment detection algorithm to the effective speech segment of the audio data, and the phonic speech subsegment detection algorithm is configured to classify the subsegments into phonic speech subsegments and non-phonic speech subsegments. Classifying sub-segments of the valid speech segment of the audio data into voiced speech segments and unvoiced speech segments includes generating timestamps indicating the start time and the end time of each corresponding voiced speech sub-segment and unvoiced speech sub-segment. Then, extraction of digital biomarker data related to, for example, speech level or pronunciation accuracy or one or more MFCC values ​​can be performed only on the voiced speech sub-segments, for example, to assess the vocal function of the vocal cords, or the accuracy of vowel pronunciation, or the pronunciation range, respectively.

[0027] When the digital biomarker data includes one or more MFCC values, extracting the digital biomarker data may include extracting one or more MFCC values, such as one or more low-order MFCC values, for example, one or more MFCC 1 values, from each voiced speech sub-segment. In this case, the speech function assessment model applied to the digital biomarker data (i.e., MFCC values) can be configured to calculate the variance of the MFCC values. The output indicating the speech function of the subject may correspond to the variance of the MFCC values. For example, the output indicating the speech function of the subject may correspond to the variance of the low-order MFCC values, such as the variance of the MFCC 1 values.

[0028] Low-order MFCCs, such as MFCC 1, can describe the overall shape of the speech spectrum. The overall shape of the speech spectrum can be determined by the resonances within the articulatory organs and vocal tract during the phonation interval. Therefore, a lower MFCC variance can indicate a reduced range of articulation, which may be caused by bulbar weakness. In this way, the variance of low-order MFCC values ​​can be used to indicate the presence and / or progression of muscle disability (such as SMA), as discussed in further detail below.

[0029] In some cases, the computer readable instructions, when executed by the processor, may further cause the apparatus to: receive noise data via a microphone; calculate background noise from the noise data; and apply a correction to the audio data using the background noise.

[0030] In some examples, the output indicative of the subject's speech function may correspond to digital biomarker data. For example, the output indicative of speech function may correspond to a change in speech rate, pronunciation accuracy, total duration of pauses, pause rate, or level of speech of the subject.

[0031] We now discuss how to use the output indicating speech function to indicate the presence or progress of muscle disability (such as SMA). Computer readable instructions can cause the diagnostic device to apply the clinical interpretation model to the output indicating speech function when executed by at least one processor. The clinical interpretation model can be configured to output an indication of the presence or absence of muscle disability (such as SMA) in the user or an indication of the progress of muscle disability in the user. The clinical interpretation model can be configured to compare the output indicating speech function with a predetermined value, and output an indication of the presence or absence of muscle disability (such as SMA) based on the comparison. Specifically, the clinical interpretation model can be configured to determine whether the output indicating speech function is greater than a predetermined threshold. In some examples, the clinical interpretation model can be configured to: if it is determined that the output indicating speech function is greater than a predetermined threshold, an indication of the presence of muscle disability (e.g., the user is PlwSMA) is output, and / or if it is determined that the output indicating speech function is less than or equal to a predetermined threshold, an indication of the absence of muscle disability is output. In other examples, the clinical interpretation model can be configured to output an indication of the presence of muscle disability (e.g., the user is PlwSMA) if it is determined that the output indicating speech function is less than a predetermined threshold, and / or output an indication of the absence of muscle disability if it is determined that the output indicating speech function is greater than or equal to a predetermined threshold. This may be the case when the output indicating speech function is the variance of low-order MFCC values. When the output indicating speech function is the variance of low-order MFCC values, the predetermined threshold may be 10,000 or less. The predetermined threshold may be 9,000 or less, or 8,000 or less. The predetermined threshold may be 6,000 or more. The predetermined threshold may be 7,000 or more, or 8,000 or more.

[0032] For example, as mentioned above, PlwSMA may exhibit significantly lower MFCC 1 variance compared to healthy subjects due to their bulbar weakness.

[0033] The second aspect of the present invention provides a computer-implemented method for assessing the speech function of a subject, the computer-implemented method comprising the following steps: prompting the subject to perform a diagnostic task of speaking loudly; receiving audio data associated with the diagnostic task via a microphone; extracting digital biomarker data associated with the speech function of the subject from the audio data; and applying a speech function assessment model to the digital biomarker data, the speech function assessment model being configured to generate an output indicating the speech function of the subject based on the digital biomarker data. Preferably, the computer-implemented method of the second aspect of the present invention is performed by a processor of a diagnostic device such as the diagnostic device of the first aspect of the present invention. It should be understood that the optional features set forth above with respect to the first aspect of the present invention are equally applicable to the second aspect of the present invention, unless the context clearly indicates otherwise, or whether the combination of such features is obviously technically incompatible.

[0034] A third aspect of the invention provides a computer program comprising instructions which, when executed by a processor of a computer (or other suitable data processing apparatus), cause the processor to perform the computer-implemented method of the second aspect of the invention. Another aspect of the invention provides a computer-readable storage medium having the computer program of the third aspect of the invention stored thereon.

[0035] The present invention includes any combination of described aspects and preferred features unless such a combination is expressly impermissible or explicitly avoided. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Embodiments of the present invention will now be described with reference to the accompanying drawings, in which:

[0037] - Figure 1 is a diagram of an exemplary environment in which a diagnostic apparatus for assessing speech function of a subject is provided.

[0038] - Figure 2 is a flow chart of a computer-implemented method for assessing a user's speech function.

[0039] - Figure 3 is a flow chart of a computer-implemented method for determining an indication of the presence or absence of muscle disability, such as SMA.

[0040] - Figure 4 is a graph showing the MFCC 1 variance calculated for PlwSMA and for healthy individuals.

[0041] - Figure 5 An example of a network architecture and data processing device that can be used to implement one or more illustrative aspects described herein is shown. DETAILED DESCRIPTION

[0042] Aspects and embodiments of the present invention will now be discussed with reference to the accompanying drawings. Other aspects and embodiments will be apparent to those skilled in the art. All documents mentioned in this text are incorporated herein by reference.

[0043] In the following description of the various aspects, reference is made to the accompanying drawings which form a part hereof and in which are shown by way of illustration various embodiments in which the aspects described herein may be practiced. It should be understood that other aspects and / or embodiments may be utilized and structural and functional modifications may be made without departing from the scope of the described aspects and embodiments.

[0044] Aspects described herein can be used for other embodiments and can be practiced or executed in various ways. In addition, it should be understood that the wording and terminology used herein are for illustrative purposes and should not be considered as restrictive. On the contrary, the phrases and terms used in this article will be given their broadest interpretation and meaning. The use of "include" and "comprise" and its variants is meant to cover the projects listed thereafter and their equivalents, as well as other projects and their equivalents. The use of the terms "install", "connect", "couple", "locate", "engage" and similar terms is intended to include direct and indirect installation, connection, coupling, positioning and engagement.

[0045] The systems, methods, and devices described herein provide diagnostic devices and computer-implemented methods for assessing, measuring, or determining speech function of a patient (e.g., a patient with muscle disability, such as, in particular, SMA). In some cases, the diagnostic device may be in the form of a mobile device, in particular a smartphone, on which a specific software application is installed. The software application may be configured to execute (or cause a processor of the mobile device) a corresponding computer-implemented method.

[0046] In some cases, when the subject uses the mobile device to interact with the software application, the diagnostic device obtains or receives sensor data from one or more sensors associated with the mobile device. In some cases, the sensor can be within the mobile device. In some cases, data indicating the subject's speech function is derived, calculated, or extracted from the received or obtained sensor data. In some cases, an assessment of the severity and progression of the subject's muscle disability, particularly SMA, can be determined based on the extracted sensor features.

[0047] In a specific implementation of the present invention, the diagnostic device can prompt the subject to perform a diagnostic task. In some cases, the diagnostic task is anchored in a given method and standardized test, or the diagnostic task is modeled after a given method and standardized test. In some cases, in response to the subject performing the diagnostic task, the diagnostic device obtains or receives sensor data via one or more sensors. In some cases, the sensor may be in a mobile device or a wearable sensor worn by the subject. In some cases, sensor features associated with muscle disability, particularly symptoms of SMA, are extracted from the received or obtained sensor data. In some cases, an assessment of the severity and progression of the symptoms of muscle disability, particularly SMA, of the subject is determined based on the extracted features of the sensor data.

[0048] Assessment of symptom severity and progression of muscle disability, particularly SMA, using diagnostics according to the present disclosure is well correlated with assessments based on clinical outcomes and can therefore replace clinical subject monitoring and testing. Exemplary diagnostic methods according to the present disclosure can be used outside of a clinical setting and therefore have advantages in terms of cost, ease of monitoring a subject, and convenience for the subject. This facilitates frequent, particularly daily, subject monitoring and testing, thereby providing a better understanding of the disease stage and providing disease insights useful to both the clinical and research communities. Exemplary diagnostic devices according to the present disclosure can provide earlier detection of even small changes in a subject's speech function, which can indicate the presence or progression of muscle disability, particularly SMA, in a subject, and can therefore be used for better disease management, including personalized treatment.

[0049] Figure 1 is an exemplary environment diagram, in which a diagnostic device 105 for assessing the speech function of a subject 110 is shown. In some cases, device 105 can be a smart phone, smart watch, or other mobile computing device. Device 105 includes a display screen 160. In some cases, display screen 160 can be a touch screen. Device 105 includes at least one processor 115 and a memory 125 storing computer instructions for a symptom monitoring application 130, which when executed by at least one processor 115, causes device 105 to assess the speech function of the subject. Device 105 receives a plurality of sensor data via one or more sensors associated with device 105. In some cases, the one or more sensors associated with the device are at least one of the following: a sensor disposed within the device or a sensor worn by the subject and configured to communicate with the device. Figure 1 In FIG. 1 , the sensors associated with the device 105 include a first sensor 120 , such as a microphone located in the device 105 .

[0050] The device 105 extracts digital biomarker data from the received first sensor data, which digital biomarker data can be used to determine the subject's speech function.

[0051] The device 105 determines the speech function of the subject 110 based on the extracted features. In some cases, the device 105 sends the extracted features to the server 150 via the network 180. In some cases, the device 105 sends the first sensor data to the server 150 via the network 180. The server 150 includes at least one processor 155 and a memory 161 storing computer instructions for a symptom assessment application 170, which, when executed by the server processor 155, causes the processor 155 to determine the speech function of the subject 110 based on the extracted features received by the server 150 from the device 105. In some cases, the symptom assessment application 170 may cause the processor 115 to extract features from the sensor data received from the device 105. In some cases, the symptom assessment application 170 may determine the speech function of the subject 110 based on the extracted features of the sensor data that may be received from the device 105 and a subject database 175 stored in the memory 160. In some cases, the subject database 175 may include subject data and / or clinical data. In some cases, subject database 175 may include clinical and sensor-based measurements of speech function. In some cases, subject database 175 may be independent of server 150. In some cases, server 150 sends the determined speech function of subject 110 to device 105. In some cases, device 105 may output the speech function of subject 110. In some cases, device 105 may communicate information to subject 110 based on the assessment. In some cases, the assessment of the speech function of subject 110 may be communicated to a clinician, who may determine an individualized therapy for subject 110 based on the assessment.

[0052] In some cases, the computer instructions of symptom monitoring application 130, when executed by at least one processor 115, cause device 105 to determine the speech function of subject 110 based on active testing of subject 110. Device 105 prompts subject 110 to perform one or more tasks. In some cases, prompting the subject to perform the one or more diagnostic tasks includes prompting the subject to make a continuous "ah" sound for as long as possible.

[0053] In response to subject 110 performing one or more diagnostic tasks, diagnostic device 105 receives a plurality of sensor data via one or more sensors associated with device 105. Device 105 extracts various digital biomarker data from the received sensor data, and the speech function of subject 110 can be assessed based on the digital biomarker data. Muscle disability in subject 110, in particular, symptoms of SMA, may include symptoms that affect the speech function of subject 110.

[0054] Figure 2 Shows the use of Figure 1 The exemplary apparatus 105 of the exemplary method of assessing the speech function of the subject 110 based on the subject's active testing. Figure 2 refer to Figure 1 However, it should be noted that Figure 2 The method steps may be performed by other systems. The computer-implemented method includes, in step 205, prompting the subject to perform a diagnostic task as described above. The method includes, in response to the subject performing one or more tasks, receiving a plurality of sensor data via a microphone (step 210).

[0055] Then, in step 215 , digital biomarker data is extracted from the sensor data, and the speech function assessment model is applied to the digital biomarker data.

[0056] In step 220, data indicative of the subject's speech function is output, for example, by processor 107 generating instructions that, when executed by display component 160 of apparatus 105, cause display component 160 to display output indicative of the subject's speech function. Alternatively, the calculated data indicative of the subject's speech function may be transmitted to server 150, as outlined elsewhere in this application.

[0057] As described above, assessment of symptom severity and progression of muscle disability, particularly SMA, using diagnostics according to the present disclosure correlates well with assessments based on clinical findings and, therefore, can replace clinical subject monitoring and testing.

[0058] Figure 3 Shows the use of Figure 1 The exemplary apparatus 105 of the present invention provides an exemplary method for determining an indication of the presence or absence of SMA in a subject based on active testing of the subject. Figure 3 refer to Figure 1 However, it should be noted that Figure 3 The method steps may be performed by other systems. The computer-implemented method includes, in step 235, calculating the variance of the MFCC 1 values ​​of the audio data. The variance corresponds to the reference Figure 2The computer-implemented method includes outputting an indication of speech function in the method described herein. In step 240, the computer-implemented method includes determining whether the calculated variance is greater than a predetermined threshold. If the calculated variance is determined to be greater than the predetermined threshold, then in step 250, the computer-implemented method includes outputting an indication that SMA is not present. If the calculated variance is determined to be less than or equal to the predetermined threshold, then in step 245, the computer-implemented method includes outputting an indication that SMA is present.

[0059] For example, the predetermined threshold may be approximately 9,000. That is, a variance of >9,000 may indicate that the user is a healthy individual, and a variance of ≤9,000 may indicate that the user is a PlwSMA. This can be referred to Figure 4 To explain. Figure 4 is a graph showing the variance of the calculated MFCC 1 values ​​for PlwSMA and for healthy individuals. The graph shows that the test results of the variance of the MFCC 1 values ​​of most PlwSMA may be ≤9,000, while the test results of the variance of the MFCC 1 values ​​of most healthy individuals may be >9,000.

[0060] Figure 5 A method is shown that can be used to implement one or more illustrative aspects described herein (such as Figure 1 and Figure 2 301 ) and an example of a network architecture and data processing device. Various network nodes 303, 305, 307, and 309 may be interconnected via a wide area network (WAN) 301 (such as the Internet). Other networks may also or alternatively be used, including private intranets, corporate networks, LANs, wireless networks, personal networks (PANs), etc. Network 301 is for illustrative purposes and may be replaced with fewer or other computer networks. A local area network (LAN) may have one or more of any known LAN topologies and may use one or more of a variety of different protocols, such as Ethernet. Devices 303, 305, 307, 309 and other devices (not shown) may be connected to one or more networks via twisted pair, coaxial cable, optical fiber, radio waves, or other communication media.

[0061] The term "network" as used herein and depicted in the accompanying drawings refers not only to a system in which remote storage devices are coupled together via one or more communication paths, but also to independent devices that may occasionally be coupled to a system having storage functions. Thus, the term "network" includes not only a "physical network" but also a "content network", which consists of data (attributed to a single entity) residing in all physical networks.

[0062] Components may include a data server 303, a web server 305, and client computers 307, 309. The data server 303 provides overall access, control, and management of the database and control software for performing one or more illustrative aspects described herein. The data server 303 may be connected to the web server 305, through which users interact and obtain data upon request. Alternatively, the data server 303 itself may act as a web server and be directly connected to the Internet. The data server 303 may be connected to the web server 305 via a network 301 (e.g., the Internet), via a direct or indirect connection, or via some other network. Users may interact with the data server 303 using a remote computer 307, 309 (e.g., using a web browser) via one or more externally disclosed websites hosted by the web server 305 to connect to the data server 303. The client computers 307, 309 may be used together with the data server 303 to access data stored therein, or may be used for other purposes. For example, as is known in the art, a user may access the web server 305 from a client device 307 using an Internet browser, or by executing a software application that communicates with the web server 305 and / or data server 303 over a computer network, such as the Internet. In some cases, the client computer 307 may be a smartphone, smart watch, or other mobile computing device, and may implement a diagnostic device, such as a Figure 1 The device 105 shown. In some cases, the data server 303 can implement a server such as Figure 1 Server 150 is shown.

[0063] The server and application can be combined on the same physical computer and retain separate virtual or logical addresses, or they can reside on separate physical computers. Figure 1 Only one example of a network architecture that can be used is shown, and those skilled in the art will appreciate that the specific network architecture and data processing devices used can vary and assist in the functions they provide, as further described herein. For example, the services provided by the network server 305 and the data server 303 can be combined on a single server.

[0064] Each component 303, 305, 307, 309 can be any type of known computer, server or data processing device. The data server 303 may include, for example, a processor 311 that controls the overall operation of the rate server 303. The data server 303 may further include a RAM 313, a ROM 315, a network interface 317, an input / output interface 319 (e.g., a keyboard, a mouse, a display, a printer, etc.) and a memory 321. The I / O 319 may include various interface units and drivers for reading, writing, displaying and / or printing data or files. The memory 321 may further store operating system software 323 for controlling the overall operation of the data processing device 303, control logic 325 for instructing the data server 303 to perform aspects described herein, and other application software 327 that provides assistance, support and / or other functions, which may or may not be used in conjunction with other aspects described herein. The control logic may also be referred to as data server software 325 in this article. The functionality of the data server software may refer to a combination of operations or decisions made automatically based on rules encoded into the control logic, operations made manually by users providing input to the system, and / or automatic processing based on user input (e.g., queries, data updates, etc.).

[0065] The memory 321 may also store data for performing one or more aspects described herein, including a first database 329 and a second database 331. In some cases, the first database may include the second database (e.g., as a separate table, report, etc.). That is, depending on the system design, information may be stored in a single database or divided into different logical, virtual or physical databases. The devices 305, 307, 309 may have an architecture similar to or different from that described with respect to the device 303. Those skilled in the art will appreciate that the functionality of the data processing device 303 (or devices 305, 307, 309) as described herein may be distributed across multiple data processing devices, for example, to distribute processing loads between multiple computers to separate processing performed based on geographic location, user access level, quality of service (QoS), etc.

[0066] One or more aspects described herein may be embodied in computer-usable or readable data and / or computer-executable instructions executed by one or more computers or other devices described herein, such as in one or more program modules. Typically, program modules include routines, programs, targets, components, data structures, etc., which perform specific tasks or implement specific abstract data types when executed by a processor in a computer or other device. Modules may be written in a source code programming language and then compiled to execute the module, or modules may be written in a scripting language (such as, but not limited to, HTML or XML). Computer-executable instructions may be stored on a computer-readable medium, such as a hard disk, an optical disk, a removable storage medium, a solid-state memory, a RAM, etc. As will be appreciated by those skilled in the art, the functions of the program modules may be combined or distributed as needed in various embodiments. In addition, the functions may be embodied in firmware or equivalent hardware (such as an integrated circuit, a field programmable gate array (FPGA), etc.) in whole or in part. Specific data structures may be used to more effectively implement one or more aspects, and such data structures are included within the scope of computer-executable instructions and computer-usable data described herein.

[0067] The features disclosed in the foregoing description, or in the following claims, in terms of the manner of expressing or implementing the disclosed functions in their specific forms, or in terms of the methods or processes for obtaining the disclosed results, may be appropriately used alone in their various forms, or in any combination to implement the present invention.

[0068] Although the present invention has been described in conjunction with the above exemplary embodiments, many equivalent modifications and variations will be apparent to those skilled in the art when this disclosure is given. Therefore, the above exemplary embodiments of the present invention are considered to be illustrative rather than restrictive. Various changes may be made to the described embodiments without departing from the spirit and scope of the present invention.

[0069] For the avoidance of any doubt, any theoretical explanations provided herein are intended to improve the reader's understanding. The inventors do not wish to be bound by any of these theoretical explanations.

[0070] Any section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.

[0071] Throughout the specification, including the following claims, unless the context requires otherwise, the words "comprise" and "include" and variations such as "comprises and comprising" and "including", will be understood to imply the inclusion of stated integers or steps or groups of integers or steps but not the exclusion of any other integers or steps or groups of integers or steps.

[0072] It must be noted that, as used in this specification and the appended claims, the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein as from "about" one particular value and / or to "about" another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when values ​​are expressed as approximations, by use of the antecedent "about", it will be understood that the particular value forms another embodiment. The term "about" in relation to a numerical value is optional and means, for example, + / - 10%.

Claims

1. A diagnostic device configured to assess a subject's speech function, the diagnostic device comprising: at least one processor; microphone; as well as a memory storing computer readable instructions that, when executed by the at least one processor, cause the diagnostic device to: prompting the subject to perform a speaking-out-loud diagnostic task; receiving, via the microphone, audio data associated with the diagnostic task; extracting digital biomarker data associated with the speech function of the subject from the audio data; as well as A speech function assessment model is applied to the digital biomarker data, the speech function assessment model being configured to generate an output indicative of the speech function of the subject.

2. The diagnostic device according to claim 1, wherein: The computer readable instructions, when executed by the processor, cause the diagnostic device to prompt the subject to perform the diagnostic task by displaying text to be spoken aloud by the subject or audibly outputting a portion of the text to be spoken aloud by the subject.

3. The diagnostic device according to claim 1 or claim 2, wherein: The digital biomarker data includes the subject's speaking rate in syllables, words, or phonemes per unit time; and Extracting the digital biomarker data includes applying a speaking rate determination algorithm to the recorded audio data, the speaking rate determination algorithm configured to determine a speaking rate of the subject.

4. The diagnostic device according to any one of claims 1 to 3, wherein: The digital biomarker data includes pronunciation accuracy; and Extracting the digital biomarker data includes applying a pronunciation accuracy determination algorithm to the recorded audio data, the pronunciation accuracy determination algorithm configured to determine the subject's pronunciation accuracy.

5. The diagnostic device according to any one of claims 1 to 4, wherein: Extracting the digital biomarker data includes applying a pause recognition algorithm to the recorded audio data, the pause recognition algorithm configured to recognize pauses within the recorded audio data.

6. The diagnostic device according to claim 5, wherein: The digital biomarker data comprises a total duration of pauses in the recorded audio data; and Extracting the digital biomarker data further includes calculating the total duration of pauses in the recorded audio data based on the pauses identified by the pause identification algorithm.

7. The diagnostic device according to any one of claims 1 to 6, wherein: Extracting the digital biomarker data includes applying a speech level determination algorithm to the recorded audio data.

8. The diagnostic device according to claim 7, wherein: The digital biomarker data includes data indicative of changes in the level of speech of the subject over time; and Extracting the digital biomarker data includes determining a level of speech of the subject over time based on an output of the speech level determination algorithm.

9. A diagnostic device according to any of the preceding claims, wherein the computer-readable instructions, when executed by the processor, cause the diagnostic device to perform a preprocessing step, the preprocessing step comprising applying an effective speech detection algorithm to the audio data, the effective speech detection algorithm being configured to classify segments of the audio data into effective speech segments and background noise segments.

10. A diagnostic device according to claim 9, wherein the computer-readable instructions, when executed by the processor, cause the diagnostic device to apply a voiced sub-segment detection algorithm to the valid speech segment of the audio data, and the voiced speech sub-segment detection algorithm is configured to classify sub-segments into voiced speech sub-segments and unvoiced speech sub-segments.

11. The diagnostic apparatus of claim 10, wherein extracting the digital biomarker data comprises extracting one or more low-order Mel-Frequency Cepstral Coefficient (MFCC) values ​​from each voiced speech sub-segment. 12 . The diagnostic apparatus according to claim 11 , wherein the speech function assessment model is configured to calculate a variance of the low-order MFCC values ​​so that the output indicative of the speech function of the subject corresponds to the variance of the low-order MFCC values.

13. A diagnostic device according to any of the preceding claims, wherein the computer-readable instructions, when executed by the at least one processor, cause the diagnostic device to apply a clinical interpretation model to the output indicative of the speech function, wherein the clinical interpretation model outputs an indication of the presence or absence of muscle disability.

14. The diagnostic apparatus of claim 13, wherein the clinical interpretation model is configured to compare the output indicative of the speech function with a predetermined value and output an indication of the presence or absence of the muscle disability based on the comparison.

15. The diagnostic apparatus according to claim 14, wherein the clinical interpretation model is configured to: determining whether the output indicative of the speech function is greater than a predetermined threshold; and If it is determined that the output indicative of the speech function is greater than the predetermined threshold, outputting an indication that muscle disability exists; and If it is determined that the output indicative of the speech function is less than or equal to the predetermined threshold, an indication that the muscle disability is absent is output.

16. The diagnostic apparatus according to claim 14, wherein the clinical interpretation model is configured to: determining whether the output indicative of the speech function is less than a predetermined threshold; and If it is determined that the output indicative of the speech function is less than the predetermined threshold, outputting an indication that muscle disability exists; and If it is determined that the output indicative of the speech function is greater than or equal to the predetermined threshold, an indication that the muscle disability is absent is output.

17. The diagnostic apparatus of claim 16 when dependent on claim 12, wherein the predetermined threshold is 10,000 or less and 6,000 or more.

18. A computer-implemented method for assessing a subject's speech function, the computer-implemented method comprising the following steps: prompting the subject to perform a speaking-out-loud diagnostic task; receiving, via a microphone, audio data associated with the diagnostic task; extracting digital biomarker data associated with the speech function of the subject from the audio data; as well as A speech function assessment model is applied to the digital biomarker data, the speech function assessment model being configured to generate an output indicative of the speech function of the subject based on the digital biomarker data.

19. The computer-implemented method of claim 18, wherein the computer-implemented method further comprises the steps of: A clinical interpretation model is applied to the output indicative of the speech function, wherein the clinical interpretation model outputs an indication of the presence or absence of muscle disability or an indication of the progression of muscle disability.

20. A computer-implemented method according to claim 18 or claim 19, wherein: The computer-implemented method is executed by a processor of a diagnostic device according to any one of claims 1 to 17.

21. The computer-implemented method of claim 18 or claim 19, wherein the steps of prompting the subject and receiving the audio data are performed by a processor of a diagnostic device, and wherein the steps of extracting the digital biomarker data and applying the speech function assessment model are performed by a processor of a server, wherein the diagnostic device is configured to transmit the audio data to the server, and wherein the diagnostic device comprises: at least one processor; microphone; as well as a memory storing computer readable instructions that, when executed by the at least one processor, cause the diagnostic device to: prompting the subject to perform a speaking-out-loud diagnostic task; Audio data associated with the diagnostic task is received via the microphone.