Automated generation of targeted feedback using speech characteristics extracted from audio samples to address speech defects

The digital therapeutic application addresses the challenges of treating speech disorders by providing real-time, targeted feedback using machine learning models, offering an objective and personalized solution that improves therapy adherence and efficiency.

JP2025093900APending Publication Date: 2025-06-24CLICK THERAPEUTICS INC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
JP2024216304
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-12-03
Filing Date
2024-12-11
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Current methods for treating speech disorders are often ineffective due to the subjective nature of evaluation and treatment, and the difficulty in providing individualized solutions that account for the unique characteristics of each patient's speech.

Method used

A digital therapeutic application that enables users to record their voice samples and receive real-time, targeted feedback with corrective actions, using machine learning models to classify speech features and provide user-specific interventions.

Benefits of technology

This approach allows for objective and personalized feedback, improving user adherence to therapy, saving computational resources, and enabling self-practice anywhere, thus addressing the limitations of traditional speech disorder treatments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025093900000001_ABST
    Figure 2025093900000001_ABST
Patent Text Reader

Abstract

To provide a system and method for providing speech instructions based on classification of speech of language communication from a user.SOLUTION: A computing system can generate speech characteristics of a first verbal communication using a first audio sample from a user. The computing system can determine a first speech classification of the first verbal communication based on the speech characteristics from a plurality of speech classifications. The computing system can select one action from a plurality of actions that includes modifying one or more speech characteristics that define a user's speech based on the first speech classification. The computing system can provide instructions to present a message prompting the user to perform the speech defined by this one action selected from the plurality of actions. The efficacy of the medication that the user is taking to address the condition can be increased.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit and priority of U.S. Non - Provisional Application No. 18 / 967,074, filed on December 3, 2024, entitled "Automated Generation of Targeted Feedback Using Speech Characteristics Extracted from Audio Samples to Address Speech Defects". This U.S. Non - Provisional Application claims the benefit and priority of U.S. Provisional Application No. 63 / 609,270, filed on December 12, 2023, entitled "Automated Generation of Targeted Feedback Using Speech Characteristics Extracted from Audio Samples to Address Speech Defects", which is hereby incorporated by reference in its entirety.

Background Art

[0002] Speech disorders include impairments or deviations in the production of speech sounds that can affect an individual's effective communication ability. These disorders can manifest in various forms and can affect the clarity, fluency, or overall comprehensibility of speech. Speech disorders are classified in various ways. In articulation disorders, sounds are distorted, replaced, or omitted, reducing the intelligibility of speech. In fluency disorders, the natural flow of speech is disrupted, resulting in stuttering, (low and unclear) mumbling, repetition, prolongation, breaks, etc. when speaking. Voice disorders can affect the quality, pitch, or volume of the voice. Such vocal impairments can have a wide range of causes, including various types of neurological disorders such as alogia (poverty of speech) and blunted emotional expression.

[0003] The daily life of an individual with a speech disorder may be adversely affected by the condition. For example, speech disorders can result in unclear speech or distorted pronunciation, making it difficult for others to understand, prone to misunderstanding, and unable to effectively convey information to others. Furthermore, individuals with speech disorders may avoid social situations or distance themselves from conversations due to fear of being misunderstood or judged negatively. Additionally, continued speech disorder may lead to feelings of mental stress and anxiety. The fear of negative judgment, ridicule, and rejection can increase emotional pain and further exacerbate communication problems.

[0004] In the treatment of speech disorders, there may be cases where speech intervention techniques are used (for example, under the guidance of a speech therapist). Due to the existence of various speech disorders, there may be difficulties in treating the speech disorders of patients. Each speech disorder has its own unique characteristics, and each patient's voice has its own unique intonation, making treatment difficult. For example, the treatment for a person with a monotonous way of speaking may be different from that for a person with stuttering. Furthermore, if individualized solutions are not used, the differences in the speech of each patient may cause the treatment for an individual to be prolonged. Therefore, without a treatment method tailored to an individual's speech, there may be no effect in treating individuals with speech disorders.

[0005] Furthermore, the subjectivity in the evaluation and treatment of speech disorders and language disorders can also be an issue in providing effective treatment to people with speech disorders and language disorders. For example, in the evaluation of speech and language disorders, even in the case of computer-assisted treatment methods, subjective judgments by speech therapists are often involved. If the speech therapists are different, the interpretation of speech samples and performances in treatment sessions will be different, and there may be variations in diagnosis and treatment plans. Moreover, there may be cases where it is difficult to secure a speech therapist for treatment.

Summary of the Invention

[0006] To address the above and other technical challenges, the digital therapeutic applications described herein can be provided to a user, where the user is enabled to record their own voice samples of language communication, and then, from the digital therapeutic application, receive direct, targeted, real-time feedback with corrective actions to address specific speech and language disorders, enabling the user to address speech and language disorders. In the context of a digital therapeutic application, an output that includes targeted therapy and instructions to correct the user's utterances can provide user-specific interventions in real time that improve the user's adherence or compliance to the therapy over time. Further, the digital therapeutic application can provide a more targeted, individualized response aimed at addressing the user's specific utterances, thereby saving the user's time, effort, and computational resources (e.g., processing and memory) that would otherwise be used in speech interventions that may not be very effective. Additionally, this digital therapeutic application enables the user to repeat self-practice anywhere by themselves, and further enables the user to receive objective measurements without being influenced by the subjective and inconsistent opinions of multiple different pathologists.

[0007] This application can prompt the user to record language communication through the microphone of the user device. For example, the user is instructed to utter a series of words or phrases displayed on the graphical user interface of the application on the device. When an audio sample of the language communication is obtained, the application (or a service associated with the application) can process the audio sample and identify a set of utterance features of the user's language communication. These utterance features can include, for example, respiration, phonation, articulation, resonance, prosody, pitch, jitter, shimmer, rhythm, etc. For each utterance feature, the application can calculate a score indicating the objective severity of the corresponding utterance feature with respect to an understandable utterance. The application can also use a video recording obtained simultaneously with the audio sample to identify the non-verbal features of the user associated with the language communication. The non-verbal features can include gestures, facial expressions, or eye contact, etc.

[0008] Using a set of speech features, this digital therapy application can classify a user's speech communication. The classification can identify whether the speech communication is understandable or not, and can identify speech as having (low and unclear) mumbling, lisp, dysarthria, etc. The classification may be based on any number of functions of a set of speech features identified from an audio sample, and may be augmented by non-verbal features identified from a video recording. For example, the application can apply the speech features to a machine learning model to determine the classification of spoken utterances. The model may be trained using a training dataset according to a supervised learning technique. The training dataset can include a set of examples, each example including a sample audio recording of a speech communication by another individual and an annotation indicating the classification of that speech communication. Other functions can also be used to classify a user's speech.

[0009] Based on the classification of language communication, this digital therapy application can select the user's actions regarding the utterance of words by the user in language communication. When the classification indicates that the language communication is understandable, the application can identify that the user should maintain their word utterance method. In contrast, when the classification indicates that the language communication is unintelligible, the application can select a corrective action to correct the user's utterance. The corrective action can be associated with one or more of the utterance features and the severity regarding the utterance features. For example, if the user's language communication is classified as whispering, the application can select actions including increasing clarity and pacing during speech to address the whispering. Further, the application can identify the causal factor (or diagnosis) of the classification based on the utterance features. For example, if the user's language communication is classified as unintelligible, the application can identify the utterance feature with the highest severity as part of the causal factor leading to that classification.

[0010] The digital therapy application can generate an instruction that includes a message instructing the user to speak for classification purposes. This instruction can identify not only the classification itself but also the score of each utterance feature. In this way, the user can visually confirm the degree of each utterance feature related to the classification of the user's utterance. Further, the application can generate a corrected version of the audio sample from the user to which a corrective action has been applied. The corrected version can include the voice of language communication such as how the pronunciation of the words should sound so that the words included can be understood by others. The application can use text-to-speech (TTS) technology to process the audio sample from the user and output the corrected version. The application can give an instruction to the user by displaying a message via the graphical user interface on the user's device. The application can also play the corrected version of the audio sample. In this way, after hearing their own voice including the corrective action specified for the user's utterance classification, the user can adjust the pronunciation of their words.

[0011] A digital therapy application can repeat this process by receiving additional audio samples of the user's speech communication and providing instructions to correct the user's utterances. The application can identify the user's progress metric, i.e., an indicator, in multiple audio samples over a period of time. Also, the application can display the progress metric to the user to provide an indication of whether the user's utterances have improved over time. By repeatedly collecting audio samples and providing instructions regarding the utterances, the user can learn about the corrective measures to take to improve their utterances. Further, by listening to a corrected version of their own utterances, the user can gain a sense of how their voice should sound in order to be understood by others. Additionally, the digital therapy application can provide feedback beyond the linguistic aspects of the utterance, for example, regarding the context of the utterance (e.g., providing different feedback depending on whether the user is speaking to a friend or a supervisor) or the emotion (e.g., identifying a flat affect and providing feedback on how to express emotions more strongly).

[0012] In this way, digital therapeutic applications enable users to directly record audio samples of language communication and quickly receive direct feedback on corrective actions to address their speech disorders. This functionality can improve the quality of human-computer interaction (HCI) between the user and the device by imparting additional usefulness to the device. In the context of digital therapeutic applications, instructions regarding the user's utterances lead to providing user-specific interventions in real time, improving the user's adherence to the treatment. Further, since digital therapeutic applications can provide more targeted responses aimed at addressing the user's specific utterances, computational resources (e.g., processing and memory) that would have been consumed by relatively less effective computer-assisted speech interventions can be saved.

[0013] Aspects of the present disclosure relate to systems and methods for providing utterance instructions based on the classification of utterances of language communication from a user. One or more processors coupled to a memory can identify a first audio sample of a first language communication from the user. The one or more processors can generate a plurality of first utterance features of the first language communication using the first audio sample. The one or more processors can identify a first utterance classification of the first language communication based on the plurality of first utterance features from a plurality of utterance classifications. The one or more processors can select one action including modifying one or more of the utterance features defining the user's utterance based on the first utterance classification from a plurality of actions. The one or more processors can provide an instruction to present a message prompting the user to perform the utterance defined by the one action selected from the plurality of actions.

[0014] In some embodiments, the one or more processors can determine that the language communication is unintelligible. In some embodiments, the one or more processors can select the one action such that the user modifies at least one of the plurality of first utterance features in the utterance. In some embodiments, the one or more processors can determine that the language communication is intelligible. In some embodiments, the one or more processors can select the one action such that the user maintains one or more of the plurality of first utterance features in the utterance.

[0015] In some embodiments, the one or more processors can generate a score indicative of the severity of at least one of the plurality of first utterance features. In some embodiments, the one or more processors can provide the instruction including the message identifying the score for presentation to the user. In some embodiments, the plurality of first utterance features can include a corresponding plurality of scores. Each of the plurality of scores can be defined along a scale of the corresponding utterance feature.

[0016] In some embodiments, the one or more processors can identify the first utterance classification based on at least one of (i) an average of the plurality of scores, (ii) a weighted combination of the plurality of scores, (iii) a comparison with a dataset consisting of a plurality of second scores, (iv) a neural network model, or (v) a generative transformer model. In some embodiments, the one or more processors can apply a machine learning (ML) model to the plurality of first utterance features. The machine learning (ML) model can be configured using a training dataset including a plurality of examples, each of the plurality of examples identifying (i) a respective second audio sample of a second language communication and (ii) a respective second classification from the plurality of utterance classifications.

[0017] In some embodiments, the one or more processors can provide the instruction including the message that identifies at least one of (i) one or more of the plurality of first utterance features and (ii) the one action to modify the utterance. In some embodiments, the one or more processors can identify, based on at least one of the plurality of first utterance features, one factor from a plurality of factors as the cause of the first utterance classification. In some embodiments, the one or more processors can provide the instruction including the message that identifies the one factor as the cause of the first utterance classification.

[0018] In some embodiments, the one or more processors can generate a second audio sample for playback to the user by modifying the first audio sample according to the one action. In some embodiments, the one or more processors can apply a speech synthesis model to the first audio sample and the one action to generate the second audio sample. In some embodiments, the one or more processors can identify a second utterance classification of the user's second language communication based on a plurality of second utterance features. The plurality of second utterance features can be generated from the second audio sample identified at a time point after the provision of the instruction. In some embodiments, the one or more processors can determine a progress metric based on a comparison between the first utterance classification prior to the instruction and the second utterance classification after the submission of the instruction.

[0019] In some embodiments, the plurality of first speech features may further include at least one of (i) breathing, (ii) phonation, (iii) articulation, (iv) resonance, (v) prosody, (vii) pitch, (viii) jitter, (ix) shimmer, or (x) rhythm. In some embodiments, the plurality of speech classifications may include at least one of (i) whispering, (ii) tongue-twisting, (iii) paralytic dysarthria, (iv) stuttering, or (v) intelligibility. In some embodiments, the user may suffer from either a speech disorder or a language disorder, and may be receiving speech therapy at least partially simultaneously with the providing of the instructions. In some embodiments, the user may suffer from a disease associated with the speech disorder, and may be receiving pharmacotherapy for the disease at least partially simultaneously with the providing of the instructions.

[0020] In some embodiments, the one or more processors can identify a first video sample of a first non-verbal communication from the user at least partially simultaneously with the first verbal communication. In some embodiments, the one or more processors can identify a plurality of first non-verbal features of the first non-verbal communication using the first video sample. In some embodiments, the plurality of first non-verbal features can include at least one of the user's gestures or eye contact. In some embodiments, the one or more processors can identify the first speech classification based on the plurality of first non-verbal features.

[0021] Aspects of the present disclosure relate to systems and methods for providing utterance instructions based on characteristics of a user's language communication. One or more processors coupled to a memory can identify a first audio sample of a first language communication from the user. The one or more processors can generate a plurality of first utterance characteristics of the first language communication using the first audio sample. The one or more processors can select one action from a plurality of actions to modify one or more of the plurality of utterance characteristics that define the user's utterance. The one or more processors can provide an instruction to present a message prompting the user to perform the utterance defined by the one action selected from the plurality of actions.

[0022] In some embodiments, the one or more processors can determine that the language communication is unintelligible. In some embodiments, the one or more processors can select the one action such that the user modifies at least one of the plurality of first utterance characteristics in the utterance. In some embodiments, the one or more processors can determine that the language communication is intelligible. In some embodiments, the one or more processors can select the one action such that the user maintains one or more of the plurality of first utterance characteristics in the utterance.

[0023] In some embodiments, the one or more processors can generate a score indicative of the severity of at least one of the plurality of first utterance features. In some embodiments, the one or more processors can provide the instruction including the message identifying the score for presentation to the user. In some embodiments, the plurality of first utterance features can include a corresponding plurality of scores. Each of the plurality of scores can be defined along a scale of the corresponding utterance feature. In some embodiments, the plurality of first utterance features can include a corresponding plurality of scores. Each of the plurality of scores can be defined along a scale of the corresponding utterance feature.

[0024] In some embodiments, the one or more processors can identify the first utterance classification based on at least one of (i) an average of the plurality of scores, (ii) a weighted combination of the plurality of scores, (iii) a comparison with a dataset consisting of a plurality of second scores, (iv) a neural network model, or (v) a generative transformer model. In some embodiments, the one or more processors can apply a machine learning (ML) model to the plurality of first utterance features. The ML model can be configured using a training dataset including a plurality of examples. Each of the plurality of examples can identify (i) a respective second audio sample of a second language communication and (ii) a respective second classification from the plurality of utterance classifications.

[0025] In some embodiments, the one or more processors can provide the instruction including the message that identifies at least one of (i) one or more of the plurality of first utterance features and (ii) the one action to modify the utterance. In some embodiments, the one or more processors can identify one factor from a plurality of factors based on at least one of the plurality of first utterance features and provide the message that identifies the one factor. In some embodiments, the one or more processors can generate a second audio sample for playback to the user by modifying the first audio sample according to the one action. In some embodiments, the one or more processors can apply a speech synthesis model to the first audio sample and the one action to generate the second audio sample.

[0026] In some embodiments, the one or more processors can identify a second utterance classification of the user's second language communication based on a plurality of second utterance features. The plurality of second utterance features can be generated from a second audio sample identified at a time point after the provision of the instruction. In some embodiments, the one or more processors can determine a progress metric based on a comparison between the first utterance classification prior to the instruction and the second utterance classification after the submission of the instruction.

[0027] In some embodiments, the plurality of first speech characteristics may further include at least one of (i) breathing, (ii) phonation, (iii) articulation, (iv) resonance, (v) prosody, (vii) pitch, (viii) jitter, (ix) shimmer, or (x) rhythm, (xi) pacing, or (xii) pause. In some embodiments, the user may suffer from either a speech disorder or a language disorder, and may receive speech therapy at least partially simultaneously with the provision of the instruction. In some embodiments, the user may suffer from a disease associated with the speech disorder, and may receive drug therapy for the disease at least partially simultaneously with the provision of the instruction.

[0028] In some embodiments, the one or more processors can identify a first video sample of a first non-verbal communication from the user at least partially simultaneously with the first verbal communication. In some embodiments, the one or more processors can use the first video sample to identify a plurality of first non-verbal characteristics of the first non-verbal communication. The plurality of first non-verbal characteristics can include at least one of the user's gestures, facial expressions, or eye contact. In some embodiments, the one or more processors can identify the one action based on the plurality of first non-verbal characteristics.

[0029] Other aspects of the disclosure relate to systems and methods that provide relief for deficiencies in a user's speech expressiveness. One or more processors coupled to a memory can obtain a first metric associated with the user before at least one completion of a plurality of sessions. The one or more processors can repeat providing the plurality of sessions to the user. Each of the plurality of sessions can identify a first audio sample of a first language communication from the user; generate a plurality of first utterance features of the first language communication using the first audio sample; identify one action from a plurality of actions to modify one or more of the plurality of first utterance features that define the user's utterance; and provide an instruction to present a message prompting the user to perform the utterance defined by the one action selected from the plurality of actions. The one or more processors can obtain a second metric associated with the user after at least one completion of the plurality of sessions. Relief of the deficiency in speech expressiveness can be provided to the user when the second metric (i) decreases from the first metric by a first predetermined margin or (ii) increases from the first metric by a second predetermined margin.

[0030] In some embodiments, the user can be diagnosed with a condition including at least one of speech disorder, autism spectrum disorder (ASD), multiple sclerosis, mood disorder, neurodegenerative disease, Alzheimer's disease, dementia, Parkinson's disease, or schizophrenia. In some embodiments, the user can receive treatment at least partially concurrently with the at least one of the plurality of sessions. The treatment can include psychosocial intervention or medication to address the condition.

[0031] In some embodiments, the defect in the speech expressiveness may be caused by the pathological condition. In some embodiments, the user may be an adult who is at least 18 years old or older. In some embodiments, the plurality of sessions may be provided over a period ranging from 3 days to 6 months. In some embodiments, the first language communication of the first audio sample may include the utterance of one or more words by the user. In some embodiments, the plurality of first speech features may further include at least one of (i) respiration, (ii) phonation, (iii) articulation, (iv) resonance, (v) prosody, (vii) pitch, (viii) jitter, (ix) shimmer, or (x) rhythm, (xi) pacing, or (xii) pause.

[0032] In some embodiments, at least one of the plurality of sessions includes identifying a first speech classification of the first language communication based on the plurality of first speech features from a plurality of speech classifications, and selecting one action including modifying one or more of the speech features that define the user's utterance based on the first speech classification from a plurality of actions. The plurality of speech classifications may include at least one of (i) whispering, (ii) tongue-twisting, (iii) paralytic dysarthria, (iv) stuttering, or (v) intelligibility.

[0033] In some embodiments, the remission of the defect in the utterance expressiveness of the user having a speech disorder can be brought about when the second metric decreases from the first metric by the first predetermined margin, or when the second metric increases from the first metric by the second predetermined margin, and the first metric and the second metric are: Goldman-Fristoe Test of Articulation (GFTA-3) value, Arizona Articulation Proficiency Scale (Arizona-3) value, Speech Intelligibility Index (SII) value, Percentage of Intelligible Words (PIW) value, Percentage of Intelligible Utterances (PIU) value, Percentage of Intelligible Syllables (PIS) value, Percentage of Correct Consonants (PCC) value, Percentage of Correct Vowels (PVC) value, Percentage of Correct Vowels and Diphthongs (PVC-R) value, Stuttering Severity Instrument (SSI-4) value, Overall Assessment of the Speaker's Experience of Stuttering (OASES) value, Maximum Phonation Time (MPT) value, GRBAS scale value, Voice Range Profile (VRP) value, Voice Handicap Index (VHI) value, Voice-Related Quality of Life (V-RQOL) value, Consensus Auditory-Perceptual Evaluation of Voice (CAPE-V) value, Diadochokinetic Rate (DDK) value, Prosodic Voice Screening Profile (PVSP) value, Bzoch Hypernasality Scale value, Resonance Severity Index value, Nasalization Rate value, Western Aphasia Battery (WAB) value, Boston Diagnostic Aphasia Examination (BDAE) value, Communication Effectiveness Index (CETI) value, Adult Aphasia Battery (ABA-2) value, DDK rate, Revised Percentage of Correct Consonants (PCC-R) value, Frenchay Dysarthria Assessment (FDA-2) value, or Dysarthria Impact Profile (DIP) value, at least one of which is used.

[0034] In some embodiments, the remission of the user's utterance expressiveness having ASD can be brought about when the second metric decreases by the first predetermined margin from the first metric, or when the second metric increases by the second predetermined margin from the first metric, and the first metric and the second metric are at least one of an Autism Diagnostic Observation Schedule (ADOS) value, a Pragmatic Language Skills Test (TOPL-2) value, a CETI value, a Social Responsiveness Scale, Second Edition (SRS-2) value, a Comprehensive Assessment of Spoken Language-2 (CASL-2) value, a Functional Communication Profile-Revised (FCP-R) value. In some embodiments, the remission of the user's utterance expressiveness having multiple sclerosis can be brought about when the second metric decreases by the first predetermined margin from the first metric, or when the second metric changes by the second predetermined margin from the first metric, and the first metric and the second metric are at least one of a Western Aphasia Battery (WAB) value, a Boston Diagnostic Aphasia Examination (BDAE) value, a CETI value, an Assessment of Basic Language and Learning Skills-Revised (ABA-2) value, a DDK rate value, a Profile of Communicative Competence-Revised (PCC-R) value, a Frenchay Dysarthria Assessment-2 (FDA-2) value, or a Dysarthria Impact Profile (DIP) value.

[0035] In some embodiments, the remission of the user's utterance expressiveness with an affective disorder is brought about when the second metric decreases by the first predetermined margin from the first metric, or when the second metric changes by the second predetermined margin from the first metric, and the first metric and the second metric are at least one of the Hamilton Depression Rating Scale (Ham-D) values. In some embodiments, the remission of the user's utterance expressiveness with schizophrenia is brought about when the second metric decreases by the first predetermined margin from the first metric, or when the second metric changes by the second predetermined margin from the first metric, and the first metric and the second metric are at least one of the Motivation and Pleasure Scale - Self-Report (MAP-SR) value, the Social Effort and Conscientiousness Scale (SEACS) social effort value, and the SEACS social conscientiousness value. In some embodiments, the first metric can be identified based on the corresponding utterance classification of a plurality of utterance classifications in the first session of the plurality of sessions. The second metric can be identified based on the corresponding first utterance classification in the second session of the plurality of sessions.

Brief Description of the Drawings

[0036] The above-described objects, other objects, aspects, features, and advantages of the present disclosure should be better understood with reference to the following description together with the accompanying drawings.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

[0037] In reading the following descriptions of the various embodiments, the explanations and their respective contents listed following each section of the specification should be helpful.

[0038] In Section A, a system and method for providing utterance instructions based on the utterance classification of a user's language communication are described.

[0039] In Section B, a method for providing relief for deficiencies in a user's utterance expressiveness is described.

[0040] Section C is an explanation of networks and computing environments that may be useful in implementing the embodiments described herein.

[0041] A. Systems and Methods for Providing Utterance Instructions Based on Utterance Classification of User Language Communication Referring now to FIG. 1, a block diagram of a system 100 for providing utterance instructions based on utterance classification of user language communication is shown. Generally speaking, system 100 can include at least one session management service 105 and a set of user devices 110A-N (hereinafter generally referred to as user devices 110) communicatively coupled to each other via at least one network 115. At least one of the user devices 110 (e.g., the illustrated first user device 110A) can include at least one application 125. Application 125 can include or provide at least one user interface 130 having one or more user interface (UI) elements 135A-N (hereinafter generally referred to as UI elements 135). Session management service 105 can include at least one session handler 140, at least one utterance classifier 150, at least one feature extractor 145, at least one feedback generator 155, at least one performance evaluator 160, and at least one classification function 165, among others. Session management service 105 can include or have access to at least one database 170. Database 170 can store, maintain, or otherwise include at least one user profile 175A-N (hereinafter generally referred to as user profiles 175) and a training data set 180. The functions of application 125 on user device 110 can be partially executed on session management service 105, and vice versa. Each component of system 100 can be implemented using the computing systems described in Section C.

[0042] More specifically, the session management service 105 (which may generally be referred to herein as a message communication service) is a computing device that includes one or more processors coupled with memory and software, and may be any computing device capable of executing the various processes and tasks described herein. The session management service 105 can communicate with one or more user devices 110 and a database 170 via a network 115. The session management service 105 may be located, disposed, or otherwise associated with at least one computer system. This computer system may correspond to a data center, a branch office, or a site where one or more computers corresponding to the session management service 105 are located.

[0043] Within the session management service 105, the session handler 140 can manage sessions and can receive audio samples of language communication from a user. The feature extractor 145 can generate a set of utterance features from the audio samples. The utterance classifier 150 can identify the classification of the user's language communication using the utterance features. The feedback generator 155 can generate a modified version of the audio sample for providing to the user. The performance evaluator 160 can track the user's progress over time across multiple sessions.

[0044] The classification function 165 can be used to identify the classification of an utterance from a user based on a set of utterance features. In some embodiments, the classification function 165 can include a combination (e.g., an average) of scores corresponding to a set of utterance features. In some embodiments, the classification function 165 can include a weighted combination of scores corresponding to a set of utterance features. In some embodiments, the classification function 165 can be a mapping between a combination of scores of utterance features and an utterance classification.

[0045] In some embodiments, the classification function 165 can include a set of weights arranged across a set of layers according to a machine learning (ML) model. The architecture of the machine learning model can include, for example, a deep learning neural network (e.g., a convolutional neural network architecture), a regression model (e.g., a linear regression model or a logistic regression model), a random forest, a support vector machine (SVM), a clustering algorithm (e.g., k-nearest neighbors), or a naive Bayes model, among others. Generally, the classification function 165 can have at least one input and output. The input and output can be associated via a set of weights. The input can, in particular, be an audio sample, a set of utterance classifications, or a set of acoustic features, among others. The output is at least one of a classification or an action. The machine learning model of the classification function 165 can be trained using a training data set 180. The training data set 180 can include a set of examples. Each example can include an input (e.g., an audio sample, a set of utterance classifications, or a set of acoustic features) and an expected output (e.g., a classification or an action).

[0046] In some embodiments, the classification function 165 may include a set of weights arranged across a set of layers according to a transformer architecture. The transformer architecture can receive an input in the form of a set of character strings (e.g., from a text input) and output content in one or more modalities (e.g., in the form of text strings, audio content, images, videos, or multimedia content). The generative transformer model can be a machine learning model according to a transformer model (e.g., a generative pre-trained model or bidirectional encoder representations from a transformer). Under this architecture, the classification function 165 can include, in particular, at least one tokenization layer (also sometimes referred to herein as a tokenizer), at least one input embedding layer, at least one positional encoder, at least one encoder stack, at least one decoder stack, and at least one output layer, interconnected with each other (e.g., via forward, reverse, or skip connections). This generative transformer model can be, for example, a large language model (LLM), a model from text to image, a model from text to audio, or a model from text to video. This generative transformer model can be trained using a training dataset 180. The training dataset 180 can include a set of examples (e.g., in the form of a corpus). Each example can include an input (e.g., an audio sample, a set of utterance classifications, or a set of acoustic features) and an expected output (e.g., a classification or an action).

[0047] The user device 110 (which may also be referred to herein as an end-user computing device) is a computing device that includes one or more processors coupled with memory and software, and may be any computing device capable of executing the various processes and tasks described herein. The user device 110 can communicate with the session management service 105 and the database 170 via the network 115. The user device 110 may be a smartphone, other mobile phone, tablet computer, wearable computing device (e.g., smartwatch, glasses), or laptop computer. The application 125 can be accessed using the user device 110. In some embodiments, the application 125 can be downloaded (e.g., via a digital distribution platform) and installed on the user device 110. In some embodiments, the application 125 may be a web application with resources accessible via the network 115.

[0048] The application 125 executed on the user device 110 can be a digital therapeutic application. The application 125 can provide a session (which may also be referred to as a therapeutic session herein) for addressing at least one language disorder of the user. The user's language disorders can include, for example, articulation disorders (e.g., substitutions or omissions), fluency disorders (e.g., (low and unclear) muttering or stuttering), voice disorders (e.g., hoarseness, nasal voice, pitch disorder, or volume disorder), resonance disorders (e.g., hypernasality or hyponasality), apraxia of speech, aphasia, or paralytic dysarthria, etc. The user of the application 125 can be an individual who has been diagnosed with some condition or is at risk of it. The condition can include any number of disorders that cause a language disorder in the user. Such conditions can include, for example, language disorder or speech disorder, neurological disorders (e.g., schizophrenia with positive or negative symptoms, mild cognitive impairment, autism spectrum disorder (ASD), neurodegenerative diseases (e.g., Alzheimer's disease, dementia, Parkinson's disease, or multiple sclerosis), mood disorders (e.g., major depressive disorder, anxiety disorder, bipolar disorder, or post-traumatic stress disorder (PTSD)), or substance abuse, etc.

[0049] Users of the application may suffer from at least one of speech disorders (e.g., dysarthria, fluency disorder, voice disorder, resonance disorder, aphemia, aphasia, or paralytic dysarthria) or language disorders (e.g., specific language impairment (SLI), expressive language disorder, receptive language disorder, and language processing disorder). The user may be undergoing speech therapy (e.g., articulation therapy, phonological awareness activities, voice therapy, fluency shaping therapy) at least partially simultaneously with the use of application 125. The user may be undergoing pharmacotherapy to address disorders related to language or speech disorders at least partially simultaneously with the use of application 125. Such medications may be administered at least orally, intravenously, or topically. For example, for schizophrenia, there are typical antipsychotics (e.g., haloperidol, chlorpromazine, fluphenazine, perphenazine, loxitane, thioridazine, or trifluoperazine) or atypical antipsychotics (e.g., aripiprazole, risperidone, clozapine, quetiapine, olanzapine, ziprasidone, lurasidone, paliperidone, or iloperidone). For mood disorders (e.g., PTSD or depression), selective serotonin reuptake inhibitors (SRI) or mood stabilizers (e.g., lithium, valproic acid, divalproex sodium, carbamazepine, lamotrigine) can be used. Application 125 can enhance the efficacy of the medications the user is taking to address their condition. Such interventions include psychosocial interventions such as psychoeducation, group therapy, cognitive behavioral therapy (CBT), early intervention in first-episode psychosis (FEP), cognitive rehabilitation, and educational programs.

[0050] Application 125 can include, present, or otherwise provide to a user of user device 110 one or more user interface elements 135A-N (hereinafter generally referred to as UI elements 135) according to a configuration on application 125. UI elements 135 can correspond to visual components of user interface 130, such as command buttons, text boxes, check boxes, radio buttons, menu items, and sliders. In some embodiments, application 125 can be a digital therapy application and can provide sessions (sometimes referred to herein as therapy sessions) for addressing speech disorders via user interface 130.

[0051] Database 170 can store and maintain various resources and data associated with session management service 105 and application 125. Database 170 can include a database management system (DBMS) for organizing and compiling data maintained on the database as user profiles 175 and the like. Database 170 can communicate with session management service 105 and one or more user devices 110 via network 115. During the execution of various operations, session management service 105 and application 125 can access database 170 and retrieve specified data therefrom. Also, session management service 105 and application 125 can write data to database 170 from the execution of such operations.

[0052] In database 170, each user profile 175 (which may be referred to herein as a user account, user information, or subject profile) can store and maintain information related to a user of application 125 via user device 110. Each user profile 175 can be associated with or correspond to a user of application 125. User profile 175 can identify various information about the user, such as a user identifier, a medical condition to be addressed, information regarding sessions conducted by the user (e.g., completed activities or lessons), message preferences, user characteristic information, and the progress of addressing the medical condition (e.g., completion of endpoints). Session information can include various parameters of previous sessions executed by the user and may initially be blank. Message preferences can include treatment preferences and user input preferences, such as the type of message preferred and the timing of the message. Message preferences can also include preferences determined by session management service 105, such as the type of message to which the user can respond. The progress can be initially set to a starting value (e.g., blank or "0") and can correspond to the alleviation, reduction, or treatment of the medical condition. User profile 175 may be continuously updated by application 125 and session management service 105.

[0053] In some embodiments, the user profile 175 can identify or include information regarding a treatment regimen implemented by the user, such as the type of treatment (e.g., therapy, pharmacotherapy, or psychotherapy), duration (e.g., days, weeks, or years), and frequency (e.g., daily, weekly, quarterly, annually). The user profile 175 can include at least one activity log regarding messages provided to the user, the user's interactions that identify the user's performance, and responses from the user device 110 associated with the user. The user profile 175 can be stored and maintained in the database 160 using one or more files (e.g., Extensible Markup Language (XML), Comma Separated Values (CSV) delimited text file, or Structured Query Language (SQL) file). The user profile 175 can be updated iteratively as the user executes additional sessions or responds to additional messages.

[0054] Referring now to FIG. 2, a block diagram of a process 200 for parsing an audio sample in a system 100 for providing instructions is depicted. This process can include or correspond to operations performed by the system 100 for parsing an audio sample. Under process 200, a session handler 140 executing on the session management service 105 can send or transmit a request 205 to the application 125. The session handler 140 can generate the request 205 based on the user profile 175. The request 205 can take various forms related to the application 125. For example, the request 205 may be directly displayed as a notification on the application 125. In another example, the request 205 can also be displayed as an electronic communication (e.g., an email) on the user 210. In another example, the request 205 may be displayed as a Short Message Service (SMS) (e.g., a text message) or a Multimedia Messaging Service (MMS) (e.g., an audio message, a video message).

[0055] Requirement 205 can include a message having an instruction for user 210 of application 125 via user device 110. These instructions can provide the next steps for advancing the treatment plan based on user profile 175. For example, the instructions can indicate the actions that user 210 of application 125 should take for the next stage of the treatment plan. The message of requirement 205 can identify or include one or more words that user 210 should record or dictate via application 125. The one or more words can take the form of a set of text strings presented on user interface 130 of application 125. Session handler 140 can select or identify one or more words to be provided via requirement 205 based on user profile 175. In some embodiments, session handler 140 can select one or more words based on the treatment plan designated for user 210 as defined in user profile 175. In some embodiments, requirement 205 can include or identify a context for the one or more words recorded via application 125. The context can define the environmental settings in which user 210 records the words. For example, the context can include a role-play scenario that user 210 should imagine when speaking the words directed by application 125.

[0056] Upon receiving request 205, application 125 on user device 110 can display, render, or otherwise present the message of request 205 via user interface 130. The message can prompt user 210 to record the utterance of one or more specified words. Application 125 can obtain, acquire, or otherwise record at least one language communication 212A via a microphone on user device 110. Language communication 212A can correspond to or include the utterance of one or more words by user 210. The utterance can correspond to an action by user 210 of orally or verbally speaking one or more words. To record language communication 212A from user 210, application 125 can include a recording button on user interface 130. For example, user 210 can press "Record" on user interface 130 to activate application 125 to create a recording for capturing language communication 212A by user 210.

[0057] From the recording of language communication 212A, application 125 can output, generate, or otherwise produce at least one audio sample 215A. Audio sample 215A can correspond to or include one or more audio files. The audio files can be generated according to any number of audio formats, such as waveform audio file format (WAV), audio interchange file format (AIFF), MPEG format (MPEG), or Ogg Vorbis (Ogg). Upon obtaining audio sample 215A, application 125 can provide, transmit, or otherwise send audio sample 215A to session management service 105.

[0058] In some embodiments, application 125 can obtain, acquire, or otherwise record at least one non-verbal communication from user 210 via the camera of user device 110, at least partially simultaneously with verbal communication 212A. The non-verbal communication can correspond to or include actions of user 210 during the utterance of one or more words for verbal communication 212A, such as gestures, facial expressions, eye contact, etc. From the recording of non-verbal communication 212A, application 125 can output, generate, or otherwise produce at least one video sample (or image). The video sample can be generated according to any number of video formats, such as MPEG format (MPEG), Windows® Media Video (WMV), QuickTime movie (MOV), or Audio Video Interleave (AVI). The non-verbal communication can be used to reinforce verbal communication 212A when evaluating the utterances of user 210. Other input / output (I / O) devices of user device 110 can also be used to record non-verbal communication.

[0059] Session handler 140 can obtain, receive, or identify audio sample 215A of verbal communication 212A from user 210 transmitted by user device 110. In some embodiments, session handler 140 can obtain, receive, or identify video samples of non-verbal communication. Session handler 140 can identify audio sample 215A (and video samples) associated with one or more words selected for the user. Session handler 140 can store and maintain audio sample 215A (and video samples) in database 170. Audio sample 215A can be stored as being associated with user profile 175.

[0060] The feature extractor 145 that runs on the session management service 105 can process and parse the audio sample 215A. The feature extractor 145 can extract, determine, or otherwise generate a set of acoustic features from the parsed audio sample 215A. For example, the feature extractor 145 can extract information or data from the audio sample 215 during the parsing process. A set of acoustic features can include spectral characteristics, temporal patterns, or pitch information, etc. Temporal patterns can include, for example, zero-crossing rate, root mean square (RMS) energy, and temporal centroid, etc. Spectral characteristics can include, for example, Mel-frequency cepstral coefficients, spectral centroid, spectral flux, and spectral roll-off, etc. A set of acoustic characteristics can be determined using a number of speech processing algorithms, such as speech recognition algorithms, formant analysis algorithms (e.g., linear predictive coding), or Mel-frequency cepstral coefficient (MFCC) extraction algorithms.

[0061] The feature extractor 145 can generate, identify, or otherwise specify a set of utterance features 220A-N (hereinafter generally referred to as utterance features 220) using the parsed audio sample 215A. The utterance features 220 can define or specify various aspects of the user 210's language communication 212A identified from the audio sample 215A. In some embodiments, the feature extractor 145 can specify a set of utterance features 220 using the audio sample 215A according to any number of speech processing algorithms. A set of utterance features 220 can be specified using the audio sample 215A itself and one or more words designated to be uttered by the user 210. The algorithms can include, for example, speech recognition algorithms (e.g., deep learning algorithms), pitch estimation algorithms, speech rate measurement algorithms, prosody detection algorithms, speech quality analysis (e.g., jitter and shimmer analysis), or rhythm pattern recognition, etc.

[0062] A set of utterance features 220 can identify or include breathing, phonation, articulation, resonance, prosody, pitch, jitter, shimmer, rhythm, pacing, or pauses. Breathing can correspond to the duration of inhalation in the speech communication 212A. Phonation can correspond to the vibration of the vocal folds when uttering one or more words in the speech communication 212A. Articulation can correspond to the clarity and accuracy with which speech sounds are produced in the speech communication 212A. Resonance can correspond to the vibration of sound waves in the oral and nasal cavities in the oral communication 212A. Prosody can correspond to the variations in pitch, duration, and intensity in the utterance of one or more words. Pitch can correspond to the frequency (or pitch) in the utterance of one or more words in the speech communication 212A. Jitter can correspond to the frequency variations (or pitch) in the speech communication 212A. Shimmer can correspond to the degree of amplitude or intensity variations. Rhythm can correspond to the patterns and timing of speech sounds and pauses. Pacing can correspond to the progress or speed at which utterances are made in the speech communication 212A. Pauses can correspond to interruptions or silences during an utterance. In some embodiments, the utterance features 220 can include or specify the number of words within the speech communication 212A and the length (e.g., duration) of the words within the speech communication 212A.

[0063] For each utterance feature 220, the feature extractor 145 can calculate, generate, or otherwise determine a score for the utterance feature 220. Each score can be determined or defined along a scale of the corresponding utterance feature 220. In some embodiments, the score can be generated based on the severity of one or more utterance features 220A-N. The severity may be calculated based on the deviation from a standard utterance feature. The standard utterance feature can be stored in the database 170, and the feature extraction unit 145 can compare the generated utterance feature 220 with the standard utterance feature 220. For example, the standard speaking speed is 150 words per minute, but the speaking speed of the user 210 may be 100 words per minute. To indicate that the user is 50 words away from the standard utterance feature 220, a value of 50 can be assigned to this deviation. In another example, the standard speaking rhythm is from 3 to 8 syllables per second, but the speaking rhythm of the user 210 may be 1 syllable per second. To indicate that the user is 2 syllables away from the standard utterance feature 220, a value of 2 can be assigned to the deviation. In some embodiments, the standard utterance feature may depend on the context (e.g., a role-play scenario) given in the request 205 to the user 210. For example, the standard utterance feature can specify that for a scenario of asking a person for directions, the expected number of words is 100-250 words over a time frame of 1-2 minutes.

[0064] In some embodiments, the feature extractor 145 can process and analyze video samples of non-verbal communication from the user 210. The feature extractor 145 can use the video samples to generate, identify, or otherwise specify a set of non-verbal features. The set of non-verbal features can include, for example, gestures (e.g., hand gestures, body poses, or head movements), facial expressions indicating emotions (e.g., happiness, sadness, anger, fear, surprise, or contempt), or eye contact (e.g., gaze). The feature extractor 145 can use any number of algorithms to extract the set of non-verbal features. The algorithms can include, for example, computer vision algorithms (e.g., deep learning artificial neural networks, dynamic time warping (DTW), or scale-invariant feature transform (SIFT)), or eye gaze algorithms (e.g., eye gaze tracking algorithms, pupil detection algorithms, iris-based methods, or gaze estimation). In some embodiments, the feature extractor 145 can specify or calculate a score for each non-verbal feature. Each score can be specified or defined along a scale of the corresponding non-verbal feature.

[0065] Next, referring to FIG. 3, a block diagram of a process 300 for applying utterance classification in a system 100 for providing instructions is depicted. This process can include or correspond to operations performed by the system 100 to apply utterance classification in the system 100 for providing instructions. Under process 300, an utterance classifier 150 running on a session management service 105 can generate, identify, or otherwise specify at least one utterance classification 305 from a set of utterance classifications based on a set of utterance features 220. In some embodiments, the utterance classifier 150 may determine the classification 305 based on a set of non-verbal features along with the set of utterance features 220. The set of utterance classifications can include, for example, (low and unclear) muttering, lisp, dysarthria, stuttering, or understandable (or clear), etc. (Low and unclear) muttering corresponds to an utterance with a low, unclear, or ambiguous way of speaking and insufficient articulation or pronunciation. Lisp means an error in the pronunciation of a particular sound (e.g., fricatives such as "s" and "z"). Dysarthria corresponds to an unclear, slow, or inaccurate utterance. Stuttering corresponds to a sound, syllable, word, word repetition, sound elongation, or involuntary pause. Understandable (or clear) refers to an utterance that can be understood by others, and there may be no other defects such as (low and unclear) muttering, lisp, dysarthria, stuttering, etc.

[0066] The utterance classifier 150 can determine an utterance classification 305 according to a classification function 165. The classification function 165 can be used to identify or determine the utterance classification 305 based on utterance features 220 and non-verbal features. The classification function 165 can include, for example, the average of scores corresponding to the utterance features 220 (and non-verbal features); a weighted combination of scores corresponding to the utterance features 220; a mapping of scores to different utterance classifications; a machine learning model; and a generative transducer model, etc. In some embodiments, the utterance classifier 150 can identify the average of scores corresponding to the utterance features 220 (and non-verbal features). The classification function 165 can define a mapping between the average score and the utterance classification. Using the average value, the utterance classifier 150 can identify the utterance classification 305 defined by the classification function 165.

[0067] In some embodiments, the utterance classifier 150 can determine the classification 305 using a weighted combination of scores corresponding to the utterance features 220 (and non-verbal features). The classification function 165 can define the weight of each score and the mapping between the weighted value and one of a set of utterance classifications. The utterance classifier 150 can identify the utterance classification defined by the classification function 165 based on the weighted combination. For example, the user 210 can have breathing characteristics, articulation characteristics, and pitch characteristics with scores of 1, 26, and 79 respectively. The utterance classifier 150 can calculate the weighted combination of these scores and determine that the classification 305 is tongue-tied based on the weighted combination and the mapping defined by the classification function 165. In some embodiments, the utterance classifier 150 may determine the classification 305 based on a comparison between a set of utterance features 220 and the mapping of scores to different utterance classifications defined by the classification function 165. In some embodiments, the utterance classifier 150 may determine the classification 305 based on the number of spoken words identified in the utterance features 220 and the length of the duration of the utterance.

[0068] In some embodiments, the utterance classifier 150 can apply a set of utterance features 220 (and non-verbal characteristics) to the machine learning model of the classification function 165 to determine the utterance classification 305. To apply the set of utterance features 220, the utterance classifier 150 can provide the set of utterance features 220 as an input to the machine learning model. When the set of utterance features 220 is input into the machine learning model, the utterance classifier 150 can process the input according to a set of weights defined by the machine learning model and generate the utterance classification 305 as the output of the machine learning model. In some embodiments, the utterance classifier 150 may apply a set of utterance features 220 (and non-verbal characteristics) to the generative transformer of the classification function 165 to determine the utterance classification 305. To apply the set of utterance features 220, the utterance classifier 150 can generate an input prompt using the set of utterance features 220 according to a template. The utterance classifier 150 can input the prompt into the generative transformer model. When the prompt is input into the generative transformer model, the utterance classifier 150 can process the input according to a set of weights defined by the generative transformer model and generate the utterance classification 305 as the output of the generative transformer model.

[0069] The machine learning model (or transformer model) of the classification function 165 can be initialized, configured, and trained using the training dataset 180 on the database 170. The training dataset can include a set of examples. Each example can include a set of sample utterance features 220'A-N (hereinafter generally referred to as features 220'), a predicted classification 305', and one or more predicted actions 310'A-N (hereinafter generally referred to as actions 310'). In some embodiments, each example can include a sample audio recording from which the sample utterance features 220' are derived. To train, the utterance classifier 150 (or another computing device) can apply the sample utterance features 220' to the machine learning model of the classification function 165 to generate predicted classifications and actions. The utterance classifier 150 can compare the predicted classifications and actions output from the machine learning model with the predicted classifications 305' and actions 310' of the examples in the training dataset 180. Based on the comparison, the utterance classifier 150 may determine a loss metric (i.e., an indicator) according to a loss function (e.g., mean squared error, cross-entropy loss, hinge loss, or Huber loss). The utterance classifier 150 can use the loss metric to update one or more metrics of the machine learning model of the classification function 165.

[0070] Based on the utterance classification 305, the utterance classifier 150 can identify or select at least one action 310 from a set of action candidates. The action 310 can indicate or identify a modification to one or more of the utterance features 220 that define the user 210's utterance. The set of actions 310 can identify or include changes in any of the utterance features 220, such as breathing, voicing, articulation, resonance, prosody, pitch, jitter, shimmer, rhythm, pacing, or pauses in the utterance of words by the user 210. When the utterance classification 305 indicates that the language communication 212A is understandable, the utterance classifier 150 can select an action 310 for the user 210 to maintain one or more of the set of utterance features 220 of the utterance. Conversely, when the utterance classification 305 indicates that the language communication 212A is not understandable, the utterance classifier 150 can select an action 310 for the user 210 to modify one or more of the set of utterance features 220 of the utterance. In some embodiments, the set of actions 310 can identify or include a target number (or range) of words or a target length (e.g., duration) of the utterance.

[0071] In some embodiments, the utterance classifier 150 can select an action 310 based on a set of utterance features 220 (and non - verbal features). For example, if it is determined that the user 210 has a prosody problem, the action 310 can specify a change in pitch during the utterance of one or more words. The selection of the action 310 may be based on a defined mapping between the scores of the utterance features 220 and different action candidates. In some embodiments, the speech classifier 150 can use a classification function 165 to determine the action 310. For example, the utterance classifier 150 can apply a set of utterance features 220 to a machine - learning model of the classification function 165. By processing the input utterance features 220 according to the weights of the model, the utterance classifier 150 can generate an action 310 that the user 210 should take.

[0072] In some embodiments, the utterance classifier 150 can select or identify at least one factor from a set of factor candidates as the cause of the utterance classification 305 based on a set of utterance features 220. The factor candidates can identify potential causes contributing to the utterance classification 305. The set of factor candidates 220 can include utterance features such as respiration, phonation, articulation, resonance, prosody, pitch, jitter, shimmer, rhythm, pacing, or pauses. To identify, the utterance classifier 150 can compare the score of each utterance feature 220 with an expected range of values of the utterance feature for an intelligible utterance. If the score is outside the expected range, the utterance classifier 150 can select that utterance feature as a factor. Otherwise, if the score is within the expected range, the utterance classifier 150 can exclude that utterance feature as a factor. For example, the utterance classifier 150 can determine that the articulation score is low compared to a reference range and identify articulation as a factor causing the (low and unclear) murmur identified as the utterance classification 305.

[0073] The feedback generator 155 executed on the session management service 105 can write, output, or otherwise generate an audio sample 215B for playback for the user 210. The audio sample 215B can be a modified or corrected version of the original audio sample 215A according to the action 310. In some embodiments, when generating the audio sample 215B, the feedback generator 155 can use one or more words specified for the user 210 to speak together with the original audio sample 215A and the action 310. To generate the second audio sample 215B, the feedback generator 155 can change, modify, or otherwise correct a set of acoustic features extracted from the original audio sample 215A based on the action 310 and the words. For example, if the utterance of the original utterance sample 215A is characterized as a tongue twister, the feedback generator 155 can insert a frequency component into the audio sample 215B to add a fricative consonant (corresponding to a word having such a consonant), thereby deleting instances of tongue twisters in the utterance. The feedback generator 155 can use any number of utterance or voice processing algorithms, such as text-to-speech (TTS) algorithms, pitch shifting algorithms, voice conversion models, formant correction algorithms, or spectral envelope correction.

[0074] In some embodiments, the feedback generator 155 can apply a speech synthesis model to the audio sample 215A and the action 310. The speech synthesis model can include any number of machine learning models, such as a deep learning speech synthesis model or a generative transformer model. The speech synthesis model may include a text analyzer that converts text into linguistic features, an acoustic model that extracts features from the original recording or utterance (recording) based on the linguistic features, and a vocoder that creates the waveform of the corrected utterance. The speech synthesis model can be trained using a training data set that includes a set of examples. Each example can include a sample audio recording, one or more words used in the utterance of the sample audio recording, the action applied (e.g., action 310), and the output expected from applying the action to the sample audio recording. The feedback generator 155 can input the audio sample 215A, one or more words, and the action 310 into the speech synthesis model. Upon receiving the input, the feedback generator 155 can process the input audio sample 215A and the action 310 according to a set of weights of the speech synthesis model. From the processing, the feedback generator 155 can generate a corrected audio sample 215B.

[0075] Next, referring to FIG. 4, a block diagram of a process 400 for providing feedback within system 100 to provide an instruction is depicted. Process 400 can include or correspond to operations performed by system 100 to provide feedback within system 100 to provide an instruction 405. Under process 400, feedback generator 155 can write, output, or otherwise generate an instruction 405 for user 210. Instruction 405 can include a message prompting user 210 to make an utterance defined by action 310. In some embodiments, feedback generator 155 can generate instruction 405 to include a message for identifying classification 305. In some embodiments, feedback generator 155 can generate instruction 405 that includes a message for identifying any one or more of utterance features 220 (e.g., including a score), utterance classification 305, and action 310 that user 210 should take.

[0076] In some embodiments, feedback generator 155 can generate instruction 405 to include a message for identifying, for example, a factor as a cause of utterance classification 305. In some embodiments, feedback generator 155 can generate instruction 405 to include modified audio sample 205B. In some embodiments, feedback generator 155 can generate a command for user 210 to execute to include instruction 405. For example, this command can include "Please practice pursing your lips every 8 hours for the next 3 days." Along with the generation of instruction 405, session handler 140 can send, transmit, or otherwise provide instruction 405 to user device 110.

[0077] Upon receiving the indication 405, the application 125 on the user device 110 can render, display, or otherwise present the message of the indication 405. The message can prompt the user 210 to make utterances defined by the action 310. The message can identify any one or more of a set of utterance features 220 (including scores), classification 305, action 310, and factors via the UI element 135 of the user interface 130. In some embodiments, the application 125 can play back 410 the audio sample 215B. The playback 410 can be video or audio. In some embodiments, the application 125 can present the action 310, particularly as a tutorial video, practice session, interactive practice, animation, and guided practice. For example, the practice of taking a sip and speaking can be presented on the application 125 in a video where a speech therapist guides the practice. In another example, the UI element 135 can be used to place an animation of an individual's lip shape as part of the application 125.

[0078] Referring now to FIG. 5, a block diagram of a process 500 for tracking the progress of the user 210 in a system for providing an indication is depicted. The process 500 can include or correspond to operations performed by the system 100 to track the progress of the user 210. The process 500 can include any number of operations of the processes 200, 300, and 400. Under the process 200, the session handler 140 can start a subsequent session by sending a request. The request can include a message having an indication for the user 210 of the application 125 via the user device 110. The message of the request 205 can identify or include one or more words that the user 210 should record via the application 125. The one or more words can be different from the one or more words the user 210 was instructed to utter in the previous session.

[0079] The application 125 on the user device 110 can display, render, or otherwise present the message of the request 205 via the user interface 130. The message can prompt the user 210 to record the utterance of one or more specified words. The application 125 can obtain, acquire, or otherwise record at least one language communication 212A via the microphone on the user device 110. The language communication 212B can correspond to or include the utterance of one or more words by the user 210. The utterance can correspond to an action by the user 210 of orally or verbally speaking one or more words. From the recording of the language communication 212B, the application 125 can output, generate, or otherwise cause at least one audio sample 215C to occur. By obtaining the audio sample 215C, the application 125 can provide, transmit, or otherwise send the audio sample 215C to the session management service 105. In some embodiments, the application 125 can obtain, acquire, or otherwise record at least one non-verbal communication from the user 210 via the camera of the user device 110 at least partially simultaneously with the language communication 215C.

[0080] The session handler 140 can receive the audio sample 215C (and video samples) from the user device 110. Upon receiving, the feature extractor 145 can generate, identify, or otherwise specify a set of utterance features 505A-N (hereinafter generally referred to as features 505) using the parsed audio sample 215C. The utterance features 505 can be specified in a similar manner as the utterance features 220. The utterance features 505 can define or specify various aspects of the user 210's language communication 212B identified from the audio sample 215C. For each of the utterance features 220, the feature extractor 145 can calculate, generate, or otherwise determine a score for the utterance feature 505. Each score can be determined or defined along a scale of the corresponding utterance feature 505. The feature extractor 145 can generate, identify, or otherwise specify a set of non-verbal features using the video sample. In some embodiments, the feature extractor 145 can specify or calculate a score for each non-verbal feature. Each score can be specified or defined along a scale of the corresponding non-verbal feature.

[0081] The utterance classifier 150 can generate, identify, or otherwise specify at least one utterance classification 510 from a set of utterance classifications based on a set of utterance features 505. The utterance classification 510 can be specified in a similar manner as the utterance classification 305. In some embodiments, the utterance classifier 150 may determine the classification 510 based on a set of utterance features 505 and a set of non-verbal features. In some embodiments, the utterance classifier 150 can determine the utterance classification 505 according to the classification function 165. The utterance classifier 150 can identify or select at least one action from a set of action candidates based on the utterance classification 505. This action can indicate or specify a modification to one or more of the utterance features 505 that define the user 210's utterance.

[0082] The performance evaluator 160 that executes on the session management service 105 can calculate, determine, or otherwise generate at least one progress metric 515 based on a comparison between the classification 510 of the current language communication 212B and the classification 305 of the previous language communication 212A. In some embodiments, the performance evaluator 160 can determine the progress metric 515 based on a comparison between the utterance features 505 of the current language communication 212B and the utterance features 220 of the previous language communication 212A. The progress metric 515 can correspond to or identify a quantification of the improvement or deterioration between the current language communication 212B and one or more previous language communications (e.g., language communication 212A). For example, if the progress metric 515 for the classification 510 is low, it may indicate a low deviation from the classification 305 and a slow improvement. In another example, if the progress metric 515 for the classification 510 is high, it may indicate a high deviation from the classification 305 and a rapid improvement. In some embodiments, the progress metric 515 can generate a function of the level of deviation between the classification 305 and the classification 510. The function can include percent improvement, relative improvement, improvement ratio, or improvement index. The level of deviation can indicate how fast or slow the improvement of the user 210 is based on the execution of the instruction 405. By determining the progress metric 515, the performance evaluator 160 can store and maintain the progress metric 515 in the user profile 175 of the user 210.

[0083] Next, referring to FIGS. 6A and 6B, a set of screenshots of user interfaces 600 in a system for providing an utterance indication based on the classification of an utterance of a language communication from a user are depicted. The set of user interfaces 600 may be part of the application 125 and is presented via the user interface 130. The user interface 605 may prompt the user to record a sentence. The user interface 610 may be a waiting screen while an audio sample from the user is being processed and analyzed. The user interface 615 can display a classification of a speech disorder (e.g., (low and unclear) mumbles), along with scores of various utterance features (e.g., breathing, voicing, articulation, resonance, and prosody) from the user's utterance sample. Also, the user interface 615 can include a button for the user to press to listen to the playback of the corrected speech utterance. Further, the user interface 620 can provide feedback in text form, for example, to instruct the user to improve clarity, and can include a button for recording the utterance sample again. The user interface 625 can provide context-based feedback and can notify the user that the user's utterance may not be appropriate in a particular setting, and can include a button for recording the utterance sample again. The user interface 630 can provide text feedback notifying the user that the length of the recorded sentence is shorter than the expected time, and can include a button for recording the utterance sample again.

[0084] Next, referring to FIG. 6C, a screenshot of another set of user interfaces 600 in a system for providing utterance instructions based on the classification of utterances of user language communication is depicted. User interfaces 635 - 645 may be part of a role-play practice session. The session can include, for example, asking whether there is an empty seat on the bus. User interface 635 can indicate the percentage of the session completed by the user. User interfaces 640 and 645 may be introduction screens to the role-play practice session.

[0085] Next, referring to FIGS. 6D and 6E, a screenshot of another set of user interfaces 600 in a system for providing utterance instructions based on the classification of utterances of user language communication is depicted. User interface 650 may be for prompting the user to record a voice recording when given a role-play scenario (e.g., a conversation with a doctor). One role-play scenario is provided to the user. The user can make one voice recording per session, and the user is provided with feedback regarding the length of the response during recording by a dial-in indicator. User interface 655 may be for instructing the user to select the response most similar to their own. User interface 660 can indicate the response selected by the user. User interface 665 can provide feedback regarding the clarity of the user's utterance in the voice recording. User interface 660 can provide feedback regarding the length of the user's utterance in the voice recording. User interface 675 can prompt the user to select the usefulness of the feedback to the application.

[0086] Next, referring to FIG. 6F, a screenshot of another set of user interfaces 600 in a system for providing an utterance indication based on the classification of an utterance of a user's language communication is depicted. The user interface 680 can prompt the user to imagine that they are participating in a given role-play scenario (e.g., a conversation with a volunteer supervisor). When interacting with the "Continue" button, the user interface 685 can be displayed, prompting the user to record a conversation where the number of words in the record is within a target range. When the recording is complete, the user interface 690 can be displayed, providing feedback indicating the number of words spoken by the user. The user interface 695 can include additional feedback in text form indicating that the user should have spoken additional (more) words.

[0087] Referring now to FIG. 7, a flowchart of a method 700 for providing an utterance indication based on the classification of an utterance of a user's language communication is shown. The method 700 can be implemented or executed using any of the components detailed herein, such as the session management service 105 or the user device 110, or any combination thereof. Under the method 700, a computing system (e.g., the session management service 105 or the user device 110) can identify an audio sample (705). The computer can generate a set of utterance features (710). The computer may apply the set of utterance features to a classification function (715). The computer can determine an utterance classification based on the application of the classification function (720). The computer can select an action based on the utterance classification (725). The computer can generate a modified audio sample using the action (730). The computer can provide feedback (735).

[0088] B. A method for providing relief for deficiencies in the user's utterance expressiveness Next, referring to FIG. 8, a flowchart of a method 800 that provides relief for deficiencies in a user's speech expressiveness is depicted. Method 800 can be performed by any component or actor described herein, such as session management service 105, user device 110, or user 210. Method 800 can be used in combination with any of the functions or actions described in Section A of this document. Briefly described, method 800 can include obtaining a baseline metric (805). Method 800 can include receiving an audio sample (810). Method 800 can include generating speech features (815). Method 800 can include determining an action (820). Method 800 can include providing an instruction (825). Method 800 can include obtaining a session metric (830). Method 800 can include determining whether to continue (835). Method 800 can include determining whether the session metric is an improvement over the baseline metric (840). Method 800 can include determining that relief has been shown if it is determined that the session metric is an improvement over the baseline metric (845). Method 800 can include determining that no relief has been shown if it is determined that the session metric is not an improvement over the baseline metric (850).

[0089] More specifically, method 800 can include taking, identifying, or otherwise obtaining a baseline metric (805). The baseline metric can be associated with a user (e.g., user 210) having a defect in speech expressiveness. The baseline metric can be obtained (e.g., by a computing system such as session management service 105 or user device 110 or both) before providing any session to the user via a digital therapy application (e.g., application 125 described herein). The baseline metric can indicate the severity of the user's speech expressiveness. This baseline measure may depend on the type of condition and may include those detailed in Examples 1-10 herein.

[0090] Defects in speech expressiveness can be caused by alogia or blunting of vocal affect. Alogia may correspond to or mean a decrease in the variety of speech uttered by the user and a decrease in the quality of the speech. For example, users with alogia may exhibit a decrease in speech fluency, lack of response detail, and lack of meaningful content in communication. Blunting of vocal affect may mean a decrease in the range or intensity of emotional expression in the user's voice. For example, the voice of a user with blunted vocal affect is monotone and typically lacks the normal variations in pitch, tone, rhythm, etc. associated with various emotions.

[0091] The user can have any demographic characteristics, such as by age group (e.g., adult (18 years and older) or late adolescence (between 18 - 24 years old)) or gender (e.g., male, female, or non - binary). In some embodiments, a user with a speech expressiveness defect may be diagnosed with, or have the potential to have, a pathological condition. The pathological condition can include any number of disorders that cause this speech disorder in the user. Such pathological conditions can include, for example, speech disorders, neurological disorders (e.g., schizophrenia with positive or negative symptoms, mild cognitive impairment, autism spectrum disorder (ASD), neurodegenerative diseases (e.g., Alzheimer's disease, dementia, Parkinson's disease, or multiple sclerosis), mood disorders (e.g., major depressive disorder, anxiety disorder, bipolar disorder, or post - traumatic stress disorder (PTSD)), etc. The user's schizophrenia may further be accompanied by positive symptoms including hallucinations and delusions, or negative symptoms including reduced motivation or emotional expression.

[0092] The user may be receiving treatment at least partially simultaneously with at least one of a plurality of sessions provided to the user. The treatment can include at least one of psychosocial interventions or medications for dealing with schizophrenia. The user may be receiving treatment for schizophrenia at least partially simultaneously with one or more sessions. The treatment includes psychosocial interventions or medications for dealing with schizophrenia. Psychosocial interventions include, for example, psychoeducation, group therapy, cognitive - behavioral therapy (CBT), or early intervention in first - episode psychosis (FEP), etc. For example, medications include typical antipsychotics (e.g., haloperidol, chlorpromazine, fluphenazine, perphenazine, loxitane, thioridazine, or trifluoperazine) or atypical antipsychotics (e.g., aripiprazole, risperidone, clozapine, quetiapine, olanzapine, ziprasidone, lurasidone, paliperidone, or iloperidone), etc. The treatment can enhance the efficacy of the medications the user is taking to deal with the pathological condition.

[0093] Method 800 can include retrieving, identifying, or receiving an audio sample of a first language communication from a user (810). The computing system can obtain, acquire, or otherwise record at least one language communication (e.g., language communication 212A or 212B) via a microphone. The language communication can correspond to or include the utterance of one or more words by the user. The utterance can correspond to an act by the user of speaking one or more words orally or verbally. In some embodiments, the computing system can obtain, acquire, or otherwise record at least one non - language communication from the user via a camera, at least partially simultaneously with the language communication. The non - language communication can correspond to or include the user's actions during the utterance of one or more words for the language communication, such as gestures, facial expressions, eye contact, etc.

[0094] Method 800 can include identifying, discriminating, or generating a set of utterance features of a language communication using a first audio sample (815). The computing system can identify a set of utterance features (e.g., utterance features 220 or 505) based on an audio sample of the language communication. The set of utterance features can identify or include, for example, breathing, voicing, articulation, resonance, prosody, pitch, jitter, shimmer, rhythm, pacing, or pauses. For each utterance feature, the computing system can calculate, generate, or otherwise determine a score of the utterance feature. Each score can be determined or defined along a scale of the corresponding utterance feature. In some embodiments, the computing system can process and analyze a video sample of a non-verbal communication. The computing system can use the video sample to generate, identify, or otherwise specify a set of non-verbal features. The set of non-verbal features can identify or include, for example, gestures (e.g., hand gestures, body poses, or head movements), facial expressions indicating emotions (e.g., happiness, sadness, anger, fear, surprise, or contempt), or eye contact (e.g., gaze).

[0095] Method 800 can include selecting, identifying, or otherwise specifying an action that modifies one or more of the utterance features that define the user's utterance (820). The action (e.g., action 310) can indicate or identify a modification to one or more of the utterance features that define the user's utterance. The computing system can identify an action based on a set of utterance features. In some embodiments, the computing system can identify at least one utterance classification from a set of utterance classifications based on a set of utterance features (or non-verbal features). The set of utterance classifications can include, for example, (soft and unclear) muttering, lisping, paralytic dysarthria, stuttering, or understandable. The computing system can identify the utterance classification using a classification function (e.g., classification function 165).

[0096] Method 800 can include providing an instruction to present a message that prompts the user to make an utterance defined by an action selected from a plurality of actions (825). The computing system can write, output, or otherwise generate an instruction for the user (e.g., instruction 405). The instruction can include a message that prompts the user to make an utterance defined by an action. In some embodiments, the computing system can generate an instruction that includes a message for identifying any one or more of utterance features (e.g., including a score), utterance classification, and actions the user should take. In some embodiments, the computing system can generate an instruction to include an altered audio sample of the original audio recording (e.g., audio sample 205B). The message, when provided, can be presented to the user.

[0097] Method 800 can include obtaining, identifying, or otherwise acquiring session metrics (835). Session metrics can be acquired (e.g., by a computing system) after providing at least one session to the user via a digital therapy application. This session metric can indicate the severity of the user's utterance expressiveness after at least one session is provided. This session metric may depend on the type of condition and may include those detailed in Examples 1-10 herein. The session metric can be the same type of metric or measure as the baseline metric.

[0098] Method 800 can include a stage of determining or judging whether to continue (840). This determination may be based on a set test length (e.g., days, weeks, or years), a set number of time instances for running one or more sessions, or a set number of sessions provided to the user. For example, the set number of time instances can range from 3 days to 6 months relative to the acquisition of baseline metrics or the start of the initial session by the user (e.g., the test ends 6 months from the start). If the amount of time from the acquisition of baseline metrics exceeds the set length, this determination may be to stop providing additional tasks. In contrast, if the amount of time does not exceed the set length, this determination can continue to provide additional tasks and repeat from stage (810).

[0099] Method 800 can include a stage of determining or judging whether the session metric is an improvement over the baseline metric (845). The improvement can correspond to the remission of the severity of speech expression defects in the user. This improvement may be considered to have occurred when the session metric has increased by a first predetermined margin compared to the baseline metric, or when the session metric has decreased by a second predetermined margin compared to the baseline metric. For example, for certain types of metrics, the improvement is indicated by an increase in the score between the baseline and the session. For other types of metrics, the improvement is indicated by a decrease in the score between the baseline and the session. The margin may also depend on the type of metric used and generally may correspond to a difference in values that shows a significant difference to a clinician or user, or a difference in values that shows a statistically significant result in the difference between the baseline metric and the session metric.

[0100] Method 800 can include a stage of determining that remission has been shown when it is determined that the session metric is an improvement over the baseline metric (850). In some embodiments, it can be determined (e.g., by a clinician examining the computing system or user) that remission occurred when the session metric increased by a second predetermined margin from the baseline metric. In some embodiments, it can be determined (e.g., by a clinician examining the computing system or user) that remission occurred when the session metric decreased by a first predetermined margin from the baseline metric. Method 800 can include a stage of determining that no remission has been shown when it is determined that the session metric is not an improvement over the baseline metric (855). In some embodiments, it can be determined (e.g., by a clinician examining the computing system or user) that no remission has occurred when the session metric has not increased by a second predetermined margin from the baseline metric. In some embodiments, it can be determined (e.g., by a clinician examining the computing system or user) that no remission has occurred when the session metric has not increased by a second predetermined margin from the baseline metric.

[0101] Referring now to FIG. 9, a block diagram of a research design for testing an application for remitting a user's speech expressiveness defect is depicted. Briefly, at the screening visit, eligible participants download (e.g., application 125) and install the application on the user device. At the baseline visit, the participants launched the application. During the engagement period, the participants were required to use the application daily (e.g., between 3 days and 6 months). At the follow-up visit, the user uninstalled the application. The participants (i.e., users) in this study numbered from 10 to 300.

[0102] Screening Period (Day A - Day B): All participants who have given informed consent enter the screening period (e.g., up to 7 - 21 days before participation) to determine their eligibility. Assessments and activities are conducted. At the screening visit, the facility staff assist the participants in downloading and installing the application.

[0103] Baseline Survey (Day B): At the baseline visit on Day B, the eligibility of the participants is confirmed. Assessments and activities are conducted. The participants are considered eligible to launch the application. Once registration is complete, the facility staff assist the participants to launch Application 125 and complete the practice activities. The content of the application is only available after launch at baseline.

[0104] Engagement Period (Day B - Day C): The participants enter the engagement period (e.g., 5 days - 6 months), during which they interact with the application and complete the assessments and activities.

[0105] End - of - Study Visit (Day C - Day D): At the end - of - study (EOS) visit (e.g., Day 7 - Day 14 after the end of engagement) between Day C and Day D, the participants uninstall the application. Assessments and activities are conducted.

[0106] Example 1: Use of the application for those with speech disorders in general In one example, an application (e.g., application 125) is provided to individuals with various speech disorders according to the International Classification of Diseases, 11th Revision (ICD-11) or the Diagnostic and Statistical Manual of Mental Disorders, 5th Edition (DSM-5). Speech disorders include, for example, articulation disorders (e.g., dysarthria, phonological disorder, childhood apraxia of speech, paralytic dysarthria, and speech sound disorder (SSD)), fluency disorders (e.g., stuttering, and cluttering), voice disorders (e.g., spasmodic dysphonia, vocal cord paralysis, laryngitis, dysphonia, etc.), resonance disorders (hypernasality, hyponasality, velopharyngeal insufficiency, etc.), or neurological speech disorders (aphasia, apraxia of speech, paralytic dysarthria, etc.). The speech disorder can be mild to moderate. The treatment period is from 5 days to 6 months. Improvement of the defect in the speech disorder is measured by either a metric based on speech classification (e.g., classification 305) or the metric of Example 2-6 and is shown over the usage period of the application.

[0107] Example 2: Use of the application for persons with articulation disorders In one example, an application (e.g., application 125) is provided to an individual having a speech disorder according to the International Classification of Diseases, 11th Revision (ICD-11) or the Diagnostic and Statistical Manual of Mental Disorders, 5th Edition (DSM-5). The speech disorder can include, for example, dysarthria, phonological disorder, apraxia of speech, paralytic dysarthria, and speech sound disorder (SSD). The speech disorder can be from mild to moderate. The treatment period is from 5 days to 6 months. Improvement in the speech disorder is measured by a metric based on the Goldman-Fristoe Test of Articulation (GFTA-3), Arizona Articulation Proficiency Scale (Arizona-3), Speech Intelligibility Index (SII), Percentage of Intelligible Words (PIW), Percent Intelligible Utterances (PIU), Percentage of Intelligible Syllables (PIS), Percentage of Consonants Correct (PCC), Percentage of Vowels Correct (PVC), Percentage of Vowels and Diphthongs Correct (PVC-R), or speech classification (e.g., classification 305), and is shown over the period of use of the application. Other metrics such as those detailed in other examples (e.g., examples 3-6) can also be used to show improvement.

[0108] The Goldman-Fristoe Test of Articulation (GFTA-3) is an assessment tool used to evaluate an individual's ability to pronounce consonants. This test is designed to evaluate pronunciation at various word positions (initial, medial, final). There are two main subtests in this test. Namely, there are two subtests: Sounds-in-Words and Sounds-in-Sentences, and a Stimulability Assessment to determine the ability of the participant to produce the target sound when prompted. The raw score is calculated by summing the number of correct responses for each subtest (Sounds-in-Words and Sounds-in-Sentences). The raw score is converted to a standard score using the standard data described in the GFTA-3 manual. The lower the score, the more severe the condition.

[0109] The Arizona Articulation Proficiency Scale (Arizona-3) is an evaluation method for testing articulation proficiency. It measures the pronunciation of consonants in various contexts and provides detailed information on sound substitutions, omissions, and distortions. The total score measures the number of correct responses and indicates the severity of articulation deviations.

[0110] The Speech Intelligibility Index (SII) quantifies the potential intelligibility of speech based on acoustic characteristics such as frequency and intensity. This index helps predict how well speech is understood under various listening conditions, taking into account the impact of background noise on speech perception. The SII score ranges from 0.0 to 1.0, where 0.0 indicates that no speech information is heard by the listener, and 1.0 indicates that all speech information is heard.

[0111] The Percentage of Intelligible Words (PIW) is a measure for evaluating the comprehensibility of speech by calculating the proportion of words in an utterance sample that the listener can understand. PIW is expressed as a percentage and can indicate how clear an individual's communication is. The Percentage of Intelligible Utterances (PIU) measures the proportion of utterances within an utterance sample that are understood by the listener. This metric focuses on complete phrases or sentences rather than individual words.

[0112] The Percentage of Intelligible Syllables (PIS) evaluates the proportion of syllables within an utterance sample that are clearly pronounced and correctly understood by the listener. This criterion is particularly useful when evaluating language sounds in more detail. The Percentage of Correct Consonants (PCC) is widely used as a metric for evaluating articulation by calculating the proportion of correctly and clearly pronounced consonants compared to the total number of consonants uttered in an utterance sample.

[0113] The Percentage of Correct Vowels (PVC) is a criterion used to evaluate the accuracy of vowel pronunciation in an utterance sample. It is calculated by determining the proportion of correctly and clearly pronounced vowels relative to the total number of vowels uttered. The Percentage of Correct Vowels and Diphthongs (PVC-R) extends the PVC criterion by including both vowels and diphthongs in the evaluation. This criterion evaluates the accuracy of both types of sounds in an utterance sample and provides a more comprehensive overview of the speaker's vowel pronunciation.

[0114] Example 3: Use of the Application for Persons with Fluency Disorders In one example, an application (e.g., Application 125) is provided to individuals with various speech disorders according to the 11th Revision of the International Classification of Diseases (ICD-11) or the Diagnostic and Statistical Manual of Mental Disorders, 5th Edition (DSM-5). Such fluency disorders can include, for example, stuttering and rapid speech. These speech disorders can range from mild to moderate. The treatment period is from 5 days to 6 months. Improvement in fluency disorders is measured by the Stuttering Severity Instrument, 4th Edition (SSI-4), the Overall Assessment of the Speaker's Experience of Stuttering (OASES), or a metric based on speech classification (e.g., Classification 305), and is shown over the period of use of the application. Other metrics, such as those detailed in other examples (e.g., Examples 2 and 4-6), can also be used to show improvement.

[0115] The Stuttering Severity Instrument, 4th Edition (SSI-4) is an assessment tool used to evaluate the severity of stuttering. It measures stuttering across the following four areas: (1) Frequency (e.g., the proportion of syllables that stutter, calculated from a speech sample), (2) Duration (e.g., the average length of the three longest stuttering events, rounded to the nearest 0.1 second), (3) Physical concomitants (e.g., observations of secondary behaviors related to stuttering, such as facial grimaces, head movements, and distracting sounds), (4) Naturalness of speech (e.g., an evaluation of how well the speaker's speech sounds like that of a typical person compared to others). Each element is scored independently and contributes to an overall severity assessment that ranges from very mild to severe.

[0116] The Overall Assessment of the Speaker's Experience of Stuttering (OASES) is a comprehensive self-report tool designed to evaluate the impact of stuttering on an individual's life. The OASES assesses various aspects, such as awareness of stuttering (how the individual perceives their stuttering behavior and emotional reactions), communication in daily life (to what extent stuttering affects social interaction, academic performance, and overall quality of life), and reactions to stuttering (how the individual copes with their stuttering in different situations).

[0117] Example 4: Use of the Application for Persons with Speech Abnormalities In one example, an application (e.g., Application 125) is provided to an individual with a speech abnormality according to the 11th Revision of the International Classification of Diseases (ICD-11) or the Diagnostic and Statistical Manual of Mental Disorders, 5th Edition (DSM-5). The speech abnormalities can include, for example, spasmodic dysphonia, vocal cord paralysis, laryngitis, phonation disorders, etc. These speech disorders can be from mild to moderate. The treatment period is from 5 days to 6 months. The improvement in speech abnormalities is measured by, for example, Maximum Phonation Time (MPT), GRBAS scale, Vocal Range Profile (VRP), Vocal Handicap Index (VHI), Voice Related Quality of Life (V-RQOL), Consensus Auditory-Perceptual Evaluation of Voice (CAPE-V), Diadochokinetic Rate (DDK), Prosody-Voice Screening Profile (PVSP), or by metrics based on speech classification (e.g., Classification 305), and is shown over the period of application use. To show improvement, other metrics such as those detailed in other examples (e.g., Examples 2, 3, 5, and 6) can also be used.

[0118] The maximum phonation time (MPT) is a measure of the longest time a person can sustain a vowel in one breath at a comfortable pitch and loudness. When evaluated using a vowel (e.g., "ah"), the MPT is used to assess voice quality and monitor changes over time. The MPT is timed for the subject's trials, and the longest of three trials is used as the score. In adults, the normal MPT is 15 - 25 seconds for women and 25 - 35 seconds for men.

[0119] The GRBAS scale is used to evaluate and score the perceptual quality of a participant's voice. This scale assesses five main aspects of the voice, and each aspect is rated on a scale from 0 to 3. A 0 is normal (no perceivable deviation), and a 3 is severe. The components of the GRBAS scale include the following: Grade (G) measures the overall severity of the voice disorder. Roughness (R) measures the irregularity of vocal fold vibration and reflects the recognition of a harsh, irregular sound of the voice, often associated with non-uniform vocal fold vibration. Breathiness (B) measures the audible escape of air from the voice, indicating how much extra air escapes during phonation, and often results in a soft or weak voice quality due to incomplete vocal fold closure. Asthenia (A) means weakness or lack of power in the voice and captures the weakness of phonation or reduction in vocal energy that can affect voice loudness and clarity. Strained voice (S) is associated with the recognition of effort or tension in phonation and measures the degree of vocal effort or tension, often resulting from excessive phonation effort or hyperfunction.

[0120] The voice range profile (VRP) (also called the voice range diagram) objectively evaluates an individual's phonatory ability and measures the range of pitch and sound pressure level (SPL) that can be produced. This profile is created by mapping the fundamental frequency (F0) against intensity across the voice range and comprehensively represents the dynamics of pitch and loudness.

[0121] The Voice Handicap Index (VHI) is a self-assessment tool with proven effectiveness designed to quantify the psychological and functional impact of voice disorders on an individual's quality of life. It consists of 30 items across three domains: functional, physical, and emotional, and comprehensively measures the severity of perceived voice disorders. By scoring each item on a Likert scale, the VHI provides a quantitative measure for evaluating patients' subjective experiences of voice disorders.

[0122] Voice-Related Quality of Life (V-RQOL) is a voice-specific patient-reported outcome measure developed to evaluate the impact of voice disorders on the quality of life perceived by patients, and its effectiveness has been confirmed. The V-RQOL questionnaire contains 10 items, evaluates the functional, emotional, and social aspects of voice-related quality of life, and provides a score reflecting the degree of impact of voice disorders on daily life and well-being.

[0123] The Consensus Auditory-Perceptual Evaluation of Voice (CAPE-V) is a standardized tool for perceptually evaluating voice quality. It assesses the main parameters of voice, including overall severity, roughness, breathiness, strain, pitch, and loudness. The CAPE-V uses a visual analog scale and evaluates the severity of each parameter based on a structured speech task.

[0124] The Diadochokinetic Rate (DDK) refers to a measure of an individual's ability to rapidly alternate speech sounds. This evaluation typically involves asking the subject to repeat syllables such as "pa·ta·ka" as quickly as possible within a specified time. The DDK rate quantifies the speed of oral motor skills and the coordination of muscle movements essential for fluent speech. It serves as an indicator of motor control related to speech.

[0125] The Prosodic Vocal Screening Profile (PVSP) is a standardized assessment tool designed to evaluate the prosody and vocal characteristics of conversational speech. The PVSP incorporates perceptual judgments across seven suprasegmental areas: phrasing, rate, stress, loudness, pitch, laryngeal quality, and resonance. Each area is typically rated on a scale of 0 - 5 or similar, where 0 indicates no problems or typical performance, and increasing scores towards 5 indicate a higher degree of prosodic or phonatory abnormality. The scores for each area are summed to calculate an overall score that reflects an individual's prosodic and vocal characteristics. The higher the overall score, the more difficult the prosody and voice are indicated to be, and the lower the score, the more typical the performance is indicated to be.

[0126] Example 5: Use of the Application for Persons with Resonance Abnormalities In one example, an application (e.g., Application 125) is provided to an individual with resonance abnormalities according to the 11th Revision of the International Classification of Diseases (ICD - 11) or the Diagnostic and Statistical Manual of Mental Disorders, 5th Edition (DSM - 5). Resonance abnormalities include, among others, hyponasality, hypernasality, velopharyngeal dysfunction, etc. These speech disorders can range from mild to moderate. The treatment period is from 5 days to 6 months. Improvement in resonance abnormalities is measured by the Bzoch Hypernasality Scale, the Resonance Severity Index, the nasalization rate, or by metrics based on speech classification (e.g., Classification 305) and is shown over the period of application use. Other metrics, such as those detailed in other examples (e.g., Examples 2 - 4 and 6), can also be used to show improvement.

[0127] The Bzoch Hypernasality Scale is a perceptual evaluation tool for assessing the severity of hypernasality in speech production. This scale provides a systematic method for clinicians to evaluate the degree of hypernasality based on auditory and perceptual judgments. This scale typically ranges from normal resonance (0) to severe hypernasality (4), enabling standardized documentation of speech production characteristics related to velopharyngeal dysfunction.

[0128] The Resonance Severity Index is a quantitative scale used to evaluate resonance disorders, particularly the severity of hypernasality and hyponasality. This index combines various evaluation tools, including perceptual ratings and instrumental measurements such as nasometry, to calculate a comprehensive score reflecting the severity of resonance problems.

[0129] The nasalization rate is a numerical value obtained from nasometry and quantifies the relative amount of acoustic energy of nasal sounds included in speech production. The nasalization rate is calculated as the ratio of nasal energy to total acoustic energy (nasal and oral) and is expressed as a percentage. The higher the nasalization rate, the greater the nasal resonance, which is often associated with hypernasality. The lower this value, the smaller the nasal resonance, indicating the possibility of hyponasality.

[0130] Example 6: Use of the Application for Persons with Neurological Speech Disorders In one example, an application (e.g., application 125) is provided to individuals having various speech disorders according to the 11th Revision of the International Classification of Diseases (ICD-11) or the Diagnostic and Statistical Manual of Mental Disorders, 5th Edition (DSM-5). Such neurological speech disorders can include aphasia, apraxia of speech, paralytic dysarthria, etc. These speech disorders can range from mild to moderate. The treatment period is from 5 days to 6 months. Improvement of the deficits in neurological speech disorders is measured by the Western Aphasia Battery (WAB), the Boston Diagnostic Aphasia Examination (BDAE), the Communicative Effectiveness Index (CETI), the Apraxia Battery for Adults (ABA-2), DDK rate, the Percentage of Consonants Correct-Revised (PCC-R), the Frenchay Dysarthria Assessment (FDA-2), the Dysarthria Impact Profile (DIP), or by metrics based on speech classification (e.g., classification 305), etc., and is shown over the usage period of the application. Other metrics such as those detailed in other examples (e.g., examples 2-5) can also be used to show improvement.

[0131] The Western Aphasia Battery (WAB) is a comprehensive standardized assessment that evaluates language functions across areas such as spontaneous speech, auditory comprehension, repetition, and naming. The WAB calculates an Aphasia Quotient (AQ) and a Cortical Quotient (CQ). The AQ is an overall score that quantifies the language abilities of participants across major areas including spontaneous speech, auditory comprehension, repetition, and naming. The score ranges from 0 - 100, with higher scores indicating better language function. The CQ is a broader overall score that assesses general cortical function. The CQ is derived by combining scores from language subtests and non - language subtests, providing a measure of overall cortical or cognitive impairment. Also, the score ranges from 0 - 100, with higher scores indicating better cognitive function.

[0132] The Boston Diagnostic Aphasia Examination (BDAE) evaluates fluency, comprehension, repetition, and other language abilities, creating a detailed overall picture of aphasic symptoms. The BDAE enables differential diagnosis of aphasia types and tracks language recovery, thus supporting targeted therapeutic intervention. The BDAE is divided into several subtests, each focusing on a specific language modality. Fluency is evaluated through spontaneous speech, looking at features such as melody line, sentence length, articulatory agility, grammatical form, paraphasia, and word - finding ability. Auditory comprehension is evaluated based on the ability to understand spoken words through various tasks. Oral expression assesses naming ability, sentence completion, use of automated sequences, etc. In reading and writing, there are tasks such as reading aloud words and sentences, and writing tasks that evaluate spelling and narrative ability. The BDAE includes an Aphasia Severity Rating Scale, which gives an overall score reflecting the severity of the language disorder.

[0133] The Communication Effectiveness Index (CETI) is an evaluation scale devised to assess the functional communication ability of aphasic patients, especially those who have suffered a stroke. CETI comprehensively evaluates the ability to convey meaning and understand the intentions of others in daily situations using all available communication methods. CETI consists of a series of items that reflect common communication scenarios, enabling important others such as caregivers and family members to evaluate the individual's performance in such situations. CETI is usually evaluated on a Likert scale from 1 (unable to perform communication) to 10 (able to perform very well) based on the effectiveness of communication observed or reported in each scenario.

[0134] The Adult Aphasia Battery - 2 (ABA - 2) is an assessment tool designed to evaluate the presence and severity of apraxia of speech in adolescents and adults. This battery consists of six sub - tests that evaluate various aspects of speech, including articulation accuracy, phoneme sequencing, and the ability to perform automatic and spontaneous speech tasks.

[0135] The Revised Percentage of Correct Consonants (PCC - R) quantifies the accuracy of speech sounds by measuring the correct consonant production in a speech sample. A low PCC - R score in apraxia of speech reflects a disorder in motor planning and contributes to overall articulation inaccuracy.

[0136] The Frenchay Dysarthria Assessment (FDA - 2) is an assessment tool that measures speech production across areas such as reflexes, respiration, phonation, and articulation. The FDA - 2 score classifies the type and severity of flaccid dysarthria and aids in differential diagnosis and treatment planning.

[0137] The Dysarthria Impact Profile (DIP) is a self-report measure that assesses the physical, emotional, and social impact of paralytic dysarthria on an individual's quality of life. The DIP consists of multiple items classified into specific areas, such as self-awareness, social interaction, and the overall impact of paralytic dysarthria in daily life. It aims to quantify how paralytic dysarthria affects an individual's self-esteem, self-concept, and interpersonal relationships.

[0138] The improvement of the above-described diagnostic value takes into account the user's articulatory expressiveness, and the improvement of the user's articulatory expressiveness becomes the improvement of the diagnostic value. Therefore, if any one of the diagnostic values after treatment according to the protocol described in this specification is improved, it means that the improvement of the user's articulatory expressiveness is shown.

[0139] Example 7: Use of the application for persons with autism spectrum disorder (ASD) In one example, an application (e.g., Application 125) is provided to an individual with an autism spectrum disorder (ASD) according to the 11th Revision of the International Classification of Diseases (ICD-11) or the Diagnostic and Statistical Manual of Mental Disorders, 5th Edition (DSM-5). The severity of the ASD can range from mild to moderate. The treatment period is from 5 days to 6 months. The improvement in expressive language deficits is measured by the Autism Diagnostic Observation Schedule (ADOS), the Test of Pragmatic Language (TOPL-2), CETI, the Social Responsiveness Scale, 2nd Edition (SRS-2), the Comprehensive Assessment of Spoken Language (CASL-2), the Functional Communication Profile-Revised (FCP-R), or by metrics based on speech classification (e.g., Classification 305), and is shown over the period of use of the application. Other metrics, such as those detailed in other examples, can also be used to show improvement.

[0140] The Autism Diagnostic Observation Schedule (ADOS) is a standardized assessment tool used for the diagnosis of ASD. The ADOS consists of a series of structured and semi-structured tasks that facilitate the direct observation of social interaction, communication, and play behaviors. It consists of five modules tailored to an individual's developmental and language levels, and allows an examiner to evaluate behaviors relevant to an autism diagnosis. Behaviors observed during the assessment are scored on a 0-3 scale, where 0 indicates typical development and 3 indicates severe impairment. These scores are summed across multiple domains to calculate a total score that is compared to the ASD diagnostic cut-off value. The higher the total score, the higher the severity of ASD symptoms.

[0141] The Pragmatic Language Test (TOPL-2) is a standardized tool for assessing pragmatic language ability. TOPL-2 evaluates the ability to understand the norms of conversation, which are important for social communication, the ability to recognize context, and the ability to adjust language based on social interaction, etc. Responses are scored based on accuracy and appropriateness, and each item is scored. These scores are summarized as the total score of pragmatic language and then compared with standard data. Percentile ranks and standard scores are useful for interpreting performance compared to age-matched peers, and the lower the score, the more significant the difficulties in pragmatic language use.

[0142] The Communication Effectiveness Index (CETI) is an evaluation scale devised to assess the functional communication ability of aphasic patients, especially those after stroke. CETI comprehensively evaluates the ability to convey meaning and understand the intentions of others in daily situations using all available communication methods. CETI consists of a series of items that reflect common communication scenarios, enabling the evaluation of the patient's performance in such situations by important others such as caregivers and family members. CETI is evaluated on a Likert scale, generally ranging from 1 (low performance) to 10 (high performance).

[0143] The Social Responsiveness Scale - Second Edition (SRS-2) is an evaluation scale for assessing social communication, social awareness, and repetitive behaviors related to autism spectrum disorder. Each item is evaluated on a 4-point Likert scale from 1 (not true) to 4 (almost always true) to capture the frequency of observed social and communication behaviors. The scores are summed, and a T-score is calculated to reflect the severity of social impairment. The higher the T-score, the more significant the social problems.

[0144] The Comprehensive Assessment of Spoken Language - Second Edition (CASL-2) is a standardized test that evaluates language abilities across areas such as comprehension, expression, syntax, semantic knowledge, and pragmatic language. The CASL-2 assesses an individual's ability to use spoken language in a structured format to facilitate the diagnosis of language disorders and the development of intervention plans. The score for each subtest is based on the number of correct responses and is converted into a standard score. By combining the scores of the subtests, a composite score for a broader language domain (e.g., syntax, semantics) can be calculated. These composite scores are interpreted using standard data, with lower scores indicating a greater language disorder.

[0145] The Functional Communication Profile - Revised (FCP-R) is an assessment tool that evaluates an individual's functional communication abilities in various contexts. It focuses not only on formal language abilities but also on how effectively a person can communicate in real-life situations. The FCP-R assesses receptive and expressive communication abilities through observation and caregiver reports, providing insights into an individual's strengths and challenges in functional communication settings. Each communication area of the FCP-R is evaluated based on observed or reported communication behaviors, typically using a rating scale from 1 (not able) to 5 (able independently). The scores are aggregated across areas to create a profile of communication abilities, with higher scores indicating a greater degree of functional independence in communication.

[0146] The improvement of the above-mentioned diagnostic values takes into account the user's expressive ability, and the improvement of the user's expressive ability results in the improvement of the diagnostic values. Therefore, if any one of the diagnostic values after treatment according to the protocol described in this specification is improved, it indicates an improvement in the user's expressive ability.

[0147] Example 8: Use of the Application for Persons with Multiple Sclerosis In one example, the application is provided to an individual with multiple sclerosis according to the International Classification of Diseases, 11th Revision (ICD-11) or the Diagnostic and Statistical Manual of Mental Disorders, 5th Edition (DSM-5). The treatment period is from 5 days to 6 months. The improvement in the defect in expressive language ability is measured by a metric based on expressive language classification (e.g., classification 305) and is shown over the period of use of the application. To show improvement, other metrics such as those detailed in other examples (e.g., Example 6) can also be used. The improvement in the diagnostic value described herein takes into account the user's expressive language ability, and the improvement in the user's expressive language ability results in the improvement of the diagnostic value. Therefore, if any one of the diagnostic values after treatment according to the protocol described herein is improved, it will indicate an improvement in the user's expressive language ability.

[0148] Example 9: Use of the application for persons with mood disorders In one example, the application (e.g., application 125) is provided to an individual with a mood disorder according to the International Classification of Diseases, 11th Revision (ICD-11) or the Diagnostic and Statistical Manual of Mental Disorders, 5th Edition (DSM-5). The treatment period is from 5 days to 6 months. The improvement in the defect in expressive language ability is measured by the Hamilton Depression Rating Scale (Ham-D) or a metric based on expressive language classification (e.g., classification 305) and is shown over the period of use of the application. To show improvement, other metrics such as those detailed in other examples can also be used.

[0149] The Hamilton Depression Rating Scale (Ham-D) is an assessment scale implemented by clinicians to evaluate the severity of depressive symptoms, such as changes in language expression related to depression. The HAM-D assesses a certain range of depressive characteristics across areas such as mood, psychomotor activity, and cognitive function, and includes specific items that evaluate speech-related symptoms such as psychomotor retardation and decreased language output. These speech characteristics (e.g., often a decrease in speech rate, a decrease in vocal intonation, a decrease in spontaneous participation or engagement through words, etc.) are scored on a Likert scale from 0 (none) to 4 (severe) and contribute to the overall severity score of depression.

[0150] The improvement in the above-mentioned diagnostic values takes into account the user's speech expression ability, and the improvement of the user's speech expression ability becomes the improvement of the diagnostic values. Therefore, if any one of the diagnostic values after treatment according to the protocol described in this specification is improved, it means that the user's speech expression ability has improved.

[0151] Example 10: Use of the application for persons with neurodegenerative diseases In one example, an application (e.g., application 125) is given to an individual with a neurodegenerative disease such as dementia, Alzheimer's disease, or Parkinson's disease according to the International Classification of Diseases, 11th Revision (ICD-11) or the Diagnostic and Statistical Manual of Mental Disorders, 5th Edition (DSM-5). The treatment period is from 5 days to 6 months. The improvement in the defect in speech expression ability is measured by a metric based on speech classification (e.g., classification 305) and is shown over the period of use of the application. To show improvement, other metrics such as those detailed in other examples (e.g., example 6) can also be used. The improvement in the diagnostic values described in this specification takes into account the user's speech expression ability, and the improvement of the user's speech expression ability becomes the improvement of the diagnostic values. Therefore, if any one of the diagnostic values after treatment according to the protocol described in this specification is improved, it means that the user's speech expression ability has improved.

[0152] Example 11: Use of the application for persons with schizophrenia In one example, an application (e.g., application 125) is provided to an individual who has schizophrenia according to the International Classification of Diseases, 11th Revision (ICD-11) or the Diagnostic and Statistical Manual of Mental Disorders, 5th Edition (DSM-5), and who has experienced mild to moderate functional impairment as indicated by the WHO-DAS 2.0 and is prescribed an antipsychotic. The treatment period is from 5 days to 6 months. Improvement in speech expressive ability is measured by the Motivation and Pleasure Scale - Self Report value (MAP-SR value), the SEACS metric for social effort value, the SEACS metric for social conscientiousness value, or a metric based on speech classification (e.g., classification 305), and is shown over the period of application use. Other metrics, such as those detailed in other examples, can also be used to show improvement.

[0153] The Motivation and Pleasure Scale - Self Report (MAP-SR) is a self-report tool derived from the Clinical Assessment Interview for Negative Symptoms and assesses the motivation and pleasure domains of negative symptoms in patients with psychotic disorders. This scale includes 15 items that record motivation, effort, interest, and pleasure in various areas of life. All items are rated on a 5-point Likert scale, with lower scores indicating higher severity.

[0154] The Social Effort and Conscientiousness Scale (SEACS) is a self-report measure of effortful behavior for the formation and maintenance of social bonds. The questionnaire has 17 items and is rated on a 6-point Likert scale. The SEACS is divided into two subscales. The social effort subscale reflects the tendency to strive to obtain social connections for one's own purposes, while the social conscientiousness subscale reflects the tendency to strive to comply with social norms.

[0155] The improvement of the above-described diagnostic value takes into account the user's expressive ability, and the improvement of the user's expressive ability results in the improvement of the diagnostic value. Therefore, if any one of the diagnostic values after treatment according to the protocol described in this specification is improved, it indicates that the user's expressive ability has been improved.

[0156] C. Network and Computing Environment The various operations described in this specification can be implemented on a computer system. FIG. 10 shows a simplified block diagram of a representative server system 1000, a client computer system 1014, and a network 1026 that can be used to implement a particular embodiment of the present disclosure. In various embodiments, the server system 1000 or a similar system can implement the services or servers or portions thereof described herein. The client computer system 1014 or a similar system can implement the clients described herein. The system 100 described in this specification may be similar to the server system 1000. The server system 1000 can have a modular design incorporating a number of modules 1002 (e.g., blades in a blade server embodiment), and although two modules 1002 are shown, any number can be provided. Each module 1002 can include one or more processing devices 1004 and a local storage device 1006.

[0157] One or more processing devices 1004 can include a single processor having one or more cores, or multiple processors. In some embodiments, one or more processing devices 1004 can also include a general-purpose primary processor in addition to one or more dedicated coprocessors such as a graphics processor, a digital signal processor, etc. In some embodiments, some or all of the processing devices 1004 can be implemented using custom circuits such as application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs) that can be written by users. In some embodiments, such integrated circuits execute instructions stored in the circuit itself. In other embodiments, one or more processing devices 1004 can execute instructions stored in a local storage device 1006. Any combination of any type of processor can be included in one or more processing devices 1004.

[0158] The local storage device 1006 can include a volatile storage medium (e.g., DRAM, SRAM, SDRAM, etc.) and / or a non-volatile storage medium (e.g., magnetic disk or optical disk, flash memory, etc.). The storage medium incorporated in the local storage device 1006 can be made fixable, removable, or upgradable as desired. The local storage device 1006 can be physically or logically divided into various subunits such as a system memory, a read-only memory (ROM), a permanent storage device, etc. The system memory can be a read-write memory device or a volatile read-write memory such as a dynamic random access memory. The system memory can store some or all of the instructions and data required by one or more processing devices 1004 during execution. The ROM can store the static data and instructions required by one or more processing devices 1004. The permanent storage device can be a non-volatile read-write memory device that can store instructions and data even when the power of the module 1002 is turned off. As used herein, the term "storage medium" includes any medium that can store data indefinitely (although due to overwriting, electrical interference, power loss, etc.), and does not include carrier waves propagated by wireless or wired connections or transient electronic signals.

[0159] In some embodiments, the local storage device 1006 can store one or more software programs executed by one or more processing devices 1004 such as an operating system and / or the functions of the system 100 or any other system described herein or programs implementing various server functions such as any other one or more servers associated with the system 100 or any other system described herein.

[0160] "Software" generally refers to a sequence of instructions that, when executed by one or more processing devices 1004, cause the server system 1000 (or a part thereof) to perform various operations. Thus, it defines an embodiment of one or more specific machines that execute and carry out the operations of a software program. The instructions can be stored as firmware resident in read-only memory for execution by one or more processing devices 1004 and / or as program code stored in a non-volatile storage medium that can be loaded into volatile working memory. Software can be implemented as a single program or, if desired, as an aggregate of separate programs or program modules that interact with each other. To perform the various operations described above, one or more processing devices 1004 can retrieve the program instructions to be executed and the data to be processed from the local storage device 1006 (or a non-local storage device described later).

[0161] In some server systems 1000, a plurality of modules 1002 can be interconnected via a bus or other interconnect 1008 to form a local area network that supports communication between the modules 1002 and other components of the server system 1000. The interconnect 1008 can be implemented using various technologies including server racks, hubs, routers, and the like.

[0162] The wide area network (WAN) interface 1010 can implement a data communication function between a local area network (e.g., via the interconnect 1008) and a network 1026 such as the Internet. Other technologies including wired technologies (e.g., Ethernet, IEEE 802.3 standard) and / or wireless technologies (e.g., Wi-Fi, IEEE 802.11 standard) can be used to communicatively couple the server system to the network 1026.

[0163] In some embodiments, the local storage device 1006 is intended to provide working memory for one or more processing devices 1004, providing fast access to the programs and / or data being processed while reducing traffic on the interconnect 1008. Storage devices for larger amounts of data can be provided on a local area network by one or more mass storage subsystems 1012 connectable to the interconnect 1008. The mass storage subsystem 1012 can be based on magnetic, optical, semiconductor, or other data storage media. Direct attached storage, storage area networks, network attached storage, etc. can be used. Any data storage mechanism or other aggregation of data described herein as being generated, consumed, or maintained by a service or server can be stored in the mass storage subsystem 1012. In some embodiments, additional data storage resources are accessible via the WAN interface 1010 (although potentially with increased latency).

[0164] The server system 1000 can operate in response to requests received via the WAN interface 1010. For example, one of the plurality of modules 1002 can implement a management function and, in response to a received request, can assign individual tasks to other modules 902. Work assignment techniques can be used. When a request is processed, the results can be returned to the requester via the WAN interface 1010. Generally, such operations can be automated. Further, in some embodiments, the WAN interface 1010 can provide an extensible system that can connect multiple server systems 1000 to each other and manage high volumes of activity. Other techniques for managing server systems and server farms (aggregations of server systems that cooperate with each other) can be used, including dynamic resource allocation and reallocation.

[0165] Server system 1000 can interact with various user-owned devices or user-operated devices via a wide-area network such as the Internet. An example of a device operated by a user is shown as client computing system 1014 in FIG. 10. Client computing system 1014 can be implemented as a consumer device such as, for example, a smartphone, other mobile phone, tablet computer, wearable computing device (e.g., smartwatch, glasses), desktop computer, laptop computer, etc.

[0166] For example, client computing system 1014 can communicate via WAN interface 1010. Client computing system 1014 can include computer components such as one or more processing devices 1016, a storage device 1018, a network interface 1020, a user input device 1022, and a user output device 1024. Client computing system 1014 can be a computing device implemented in various form factors such as a desktop computer, laptop computer, tablet computer, smartphone, other mobile computing device, wearable computing device, etc.

[0167] Processing device 1016 and storage device 1018 can be similar to the one or more processing devices 1004 and local storage device 1006 described above. Appropriate devices can be selected based on the requirements imposed on client computing system 1014. For example, client computing system 1014 can be implemented as a "thin" client with limited processing power or as a high-performance computing device. Client computing system 1014 can be provided with program code executable by one or more processing units 1016 to enable various interactions with server system 1000.

[0168] The network interface 1020 can provide a connection to a network 1026 such as a wide area network (e.g., the Internet) to which the WAN interface 1010 of the server system 1000 is also connected. In various embodiments, the network interface 1020 can include a wired interface (e.g., Ethernet) and / or a wireless interface implementing various wireless data communication standards such as Wi-Fi, Bluetooth, or cellular data network standards (e.g., 3G, 4G, LTE, etc.).

[0169] The user input device 1022 can include any device (or devices) through which a user can send a signal to the client computing system 1014, and the client computing system 1014 can interpret the signal as indicating a specific user request or information. In various embodiments, the user input device 1022 can include any or all of a keyboard, touchpad, touch screen, mouse or other pointing device, scroll wheel, click wheel, dial, button, switch, keypad, microphone, etc.

[0170] The user output device 1024 can include any device through which the client computing system 1014 can provide information to the user. For example, the user output device 1024 can include a display sharing image generated by or transmitted to the client computing system 1014. The display can incorporate various image generation techniques such as, for example, a liquid crystal display (LCD), a light-emitting diode (LED) display including an organic light-emitting diode (OLED), a projection system, a cathode ray tube (CRT), etc., along with auxiliary electronic devices (such as a digital-to-analog converter or an analog-to-digital converter, a signal processor, etc.). Some embodiments can include devices such as a touch screen that functions as both an input and output device. In some embodiments, in addition to or instead of the display, other user output devices 1024 can be provided. Examples include indicator lights, speakers, tactile "display" devices, printers, etc.

[0171] Some embodiments include electronic components such as a microprocessor, a storage device, a memory, etc., that store computer program instructions on a computer-readable storage medium. Many of the features described herein can be implemented as a process specified as a set of program instructions encoded on a computer-readable storage medium. When these program instructions are executed by one or more processing units, the various processes indicated by the program instructions are caused to be executed by the one or more processing units. Examples of program instructions or computer code include machine code generated by a compiler and files containing higher-level code that is executed by a computer, an electronic component, or a microprocessor using an interpreter. Through appropriate programming, one or more processing devices 1004 and 1016 can provide various functions, including any of the functions described herein as being executed by a server or a client, or other functions, to the server system 1000 and the client computing system 1014.

[0172] Server system 1000 and client computing system 1014 are exemplary and it will be understood that variations and modifications are possible. The computer systems used in connection with the embodiments of the present disclosure may have other functions not specifically described herein. Further, although server system 1000 and client computing system 1014 are described with reference to specific blocks, it should be understood that these blocks are defined for convenience of explanation and are not intended to imply a particular physical arrangement of components. For example, different blocks may be located within the same facility, within the same server rack, or on the same motherboard, but need not be so located. Further, these blocks need not correspond to physically distinct components. The blocks can be configured to perform various operations, for example, by programming a processor or providing appropriate control circuitry, and various blocks may be reconfigurable or non-reconfigurable depending on how the initial configuration is obtained. Embodiments of the present disclosure can be implemented in various devices including electronic devices implemented using any combination of circuitry and software.

[0173] Although the present disclosure has been described with respect to specific embodiments, those skilled in the art will recognize that numerous modifications are possible. Embodiments of the present disclosure can be implemented using a variety of computer systems and communication technologies, including but not limited to the specific examples described herein. Embodiments of the present disclosure can be implemented using any combination of dedicated components and / or programmable processors and / or other programmable devices. The various processes described herein can be implemented using any combination of the same processor or different processors. When a component is described as being configured to perform a particular operation, such a configuration can be achieved, for example, by designing an electronic circuit to perform the operation, programming a programmable electronic circuit (such as a microprocessor) to perform the operation, or any combination thereof. Further, while the above-described embodiments may refer to specific hardware components and software components, those skilled in the art will understand that different combinations of hardware components and / or software components may also be used, that a particular operation described as being implemented in hardware may also be implemented in software, or vice versa.

[0174] Computer programs incorporating various features of the present disclosure can be encoded and stored on various computer-readable storage media. Suitable media include magnetic disks or tapes, optical storage media such as compact discs (CDs) or digital versatile discs (DVDs), flash memory, and other non-transitory media. A computer-readable medium encoded with program code may be packaged together with a compatible electronic device, or the program code may be provided separately from the electronic device (e.g., via an Internet download or as a separately packaged computer-readable storage medium).

[0175] Thus, while the present disclosure has been described with respect to particular embodiments, it will be understood that the present disclosure is intended to cover all modifications and equivalents within the scope of the following claims.

Claims

1. 1. A method for providing speech instructions based on a speech classification of a verbal communication from a user, the method comprising: identifying, by the one or more processors, a first audio sample of a first language communication from a user; generating, by the one or more processors, a plurality of first speech features of the first language communication using the first audio samples; identifying, by the one or more processors, a first speech classification of the first language communication based on the plurality of first speech features from a plurality of speech classifications; selecting, by the one or more processors, an action from a plurality of actions, the action including modifying one or more of the speech features defining the user's speech based on the first speech classification; and providing, by the one or more processors, instructions to present a message prompting the user to make the utterance defined by the one action selected from the plurality of actions.

2. determining the first speech classification further comprises determining that the linguistic communication is unintelligible; The method of claim 1 , wherein selecting the one action further comprises the user selecting the one action to modify at least one of the first plurality of speech features in the utterance.

3. determining that the first speech classification further comprises determining that the linguistic communication is understandable; The method of claim 1 , wherein selecting the one action comprises the user selecting the one action to maintain one or more of the first plurality of speech features in the utterance.

4. wherein identifying the first plurality of speech features further comprises generating a score indicative of a severity of at least one of the first plurality of speech features; The method of claim 1 , wherein providing the instructions further comprises providing the instructions including the message identifying the score for presentation to the user.

5. The method of claim 1 , wherein the first plurality of speech features further comprises a corresponding plurality of scores, each of the plurality of scores being defined along a scale for the respective speech feature.

6. 6. The method of claim 5, wherein identifying the first speech classification further comprises identifying the first speech classification based on at least one of: (i) an average of the plurality of scores; (ii) a weighted combination of the plurality of scores; (iii) a comparison to a dataset of a plurality of second scores; (iv) a neural network model; or (v) a generative transformer model.

7. 2. The method of claim 1, wherein identifying the first speech classification further comprises applying a machine learning (ML) model to the first plurality of speech features, the ML model being configured using a training dataset including a plurality of examples, each of the plurality of examples identifying (i) a respective second audio sample of a second language communication and (ii) a respective second classification from the plurality of speech classifications.

8. 2. The method of claim 1 , wherein providing the instructions further comprises providing the instructions including the message identifying at least one of: (i) one or more of the first plurality of speech features; and (ii) the one action to modify the utterance.

9. identifying, by the one or more processors, a factor from a plurality of factors as a cause of the first speech classification based on at least one of the first plurality of speech features; The method of claim 1 , wherein providing the instruction further comprises providing the message identifying the one factor as the cause of the first speech classification.

10. 2. The method of claim 1, further comprising generating, by the one or more processors, a second audio sample for playback to the user by modifying the first audio sample in accordance with the one action.

11. 11. The method of claim 10, wherein generating the second audio sample further comprises applying a speech synthesis model to the first audio sample and the one action to generate the second audio sample.

12. identifying, by the one or more processors, a second speech classification of the user's second language communication based on a plurality of second speech features, the plurality of second speech features being generated from a second audio sample identified at a time after providing the instruction; 13. The method of claim 1, further comprising determining, by the one or more processors, a progress metric based on a comparison of the first speech classification prior to the instruction and the second speech classification after the instruction is provided.

13. 2. The method of claim 1, wherein the first plurality of speech features further comprises at least one of: (i) respiration, (ii) voicing, (iii) articulation, (iv) resonance, (v) prosody, (vii) pitch, (viii) jitter, (ix) shimmer, (x) rhythm, (xi) pacing, or (xii) pauses.

14. 2. The method of claim 1, wherein the plurality of speech classifications includes at least one of: (i) muttering (low and slurred), (ii) lisp, (iii) paralytic dysarthria, (iv) stuttering, or (v) intelligible.

15. identifying, by the one or more processors, a first video sample of a first non-verbal communication from the user at least partially contemporaneously with the first verbal communication; and identifying, by the one or more processors, a first plurality of non-verbal features of the first non-verbal communication using the first video sample, the first plurality of non-verbal features including at least one of a gesture, a facial expression, and an eye contact of the user; The method of claim 1 , wherein identifying the first speech classification further comprises identifying the first speech classification based on the first plurality of non-linguistic features.

16. 10. The method of claim 1, wherein the user suffers from either a speech or language disorder and is receiving speech therapy at least partially contemporaneously with the providing of the instructions.

17. The method of claim 1 , wherein the user is suffering from a disorder associated with the speech disorder and is receiving medication for the disorder at least partially contemporaneously with the providing of the instructions.

18. 1. A system for providing speech instructions based on a speech classification of a verbal communication from a user, the system comprising: One or more processors coupled to a memory, the processors comprising: Identifying a first audio sample of a first language communication from the user; generating a first plurality of speech features of the first language communication using the first audio samples; identifying a first speech classification of the first language communication based on the plurality of first speech features from a plurality of speech classifications; selecting an action from a plurality of actions that includes modifying one or more of the speech features defining the user's speech based on the first speech classification; The system is configured to provide instructions to present a message prompting the user to make the utterance defined by the one action selected from the plurality of actions.

19. The one or more processors: determining that the verbal communication is unintelligible; 20. The system of claim 18, further configured for the user to select the one action to modify at least one of the first plurality of speech features in the utterance.

20. The one or more processors: determining that the verbal communication is understandable; 20. The system of claim 18, further configured for the user to select the one action to maintain one or more of the first plurality of speech features in the utterance.

21. The one or more processors: generating a score indicative of a severity of at least one of the first plurality of speech features; 20. The system of claim 18, further configured to provide the instructions, the instructions including the message identifying the score for presentation to the user.

22. 20. The system of claim 18, wherein the first plurality of speech features further includes a corresponding plurality of scores, each of the plurality of scores defined along a scale for the respective speech feature.

23. 23. The system of claim 22, wherein the one or more processors are further configured to determine the first utterance classification based on at least one of: (i) an average of the plurality of scores; (ii) a weighted combination of the plurality of scores; (iii) a comparison to a dataset of a plurality of second scores; (iv) a neural network model; or (v) a generative transformer model.

24. 20. The system of claim 18, wherein the one or more processors are further configured to apply a machine learning (ML) model to the first plurality of speech features, the ML model being configured using a training dataset including a plurality of examples, each of the plurality of examples identifying (i) a respective second audio sample of a second language communication and (ii) a respective second classification from the plurality of speech classifications.

25. 20. The system of claim 18, wherein the one or more processors are further configured to provide the instructions including the message identifying at least one of: (i) one or more of the first plurality of speech features; and (ii) the one action to modify the utterance.

26. The one or more processors: identifying a factor from a plurality of factors as a cause of the first speech classification based on at least one of the first plurality of speech features; 20. The system of claim 18, further configured to provide the instruction comprising the message identifying the one factor as the cause of the first speech classification.

27. 20. The system of claim 18, wherein the one or more processors are further configured to generate a second audio sample for playback to the user by modifying the first audio sample according to the one action.

28. 20. The system of claim 18, wherein the one or more processors are further configured to apply a speech synthesis model to the first audio sample and the one action to generate the second audio sample.

29. the one or more processors are further configured to determine a second speech classification of the user's second language communication based on a plurality of second speech features generated from a second audio sample determined at a time subsequent to providing the instruction; 20. The system of claim 18, wherein the one or more processors are further configured to determine a progress metric based on a comparison of the first speech classification prior to the instruction and the second speech classification after submission of the instruction.

30. 20. The system of claim 18, wherein the first plurality of speech features further comprises at least one of: (i) respiration, (ii) voicing, (iii) articulation, (iv) resonance, (v) prosody, (vii) pitch, (viii) jitter, (ix) shimmer, or (x) rhythm.

31. 20. The system of claim 18, wherein the plurality of speech classifications includes at least one of: (i) muttering (low and slurred), (ii) lisp, (iii) paralytic dysarthria, (iv) stuttering, or (v) intelligible.

32. The one or more processors: identifying a first video sample of a first non-verbal communication from the user at least partially contemporaneously with the first verbal communication; using the first video sample to identify a first plurality of non-verbal features of the first non-verbal communication, the first plurality of non-verbal features including at least one of gestures or eye contact of the user; 20. The system of claim 18, further configured to determine the first speech classification based on the first plurality of non-linguistic features.

33. 20. The system of claim 18, wherein the user suffers from one of a speech or language disorder and is receiving speech therapy at least partially contemporaneously with the providing of the instructions.

34. 20. The system of claim 18, wherein the user suffers from a disorder associated with the speech disorder and is receiving medication for the disorder at least partially contemporaneously with the providing of the instructions.

35. 1. A method for providing speech instructions based on characteristics of verbal communication from a user, the method comprising: identifying, by the one or more processors, a first audio sample of a first language communication from a user; generating, by the one or more processors, a plurality of first speech features of the first language communication using the first audio samples; selecting, by the one or more processors, an action from a plurality of actions for modifying one or more of the plurality of speech features defining the user's speech; and providing, by the one or more processors, instructions to present a message prompting a user to make the utterance defined by the one action selected from the plurality of actions.

36. determining the first speech classification further comprises determining that the linguistic communication is unintelligible; 36. The method of claim 35, wherein selecting the one action further comprises the user selecting the one action to modify at least one of the first plurality of speech features in the utterance.

37. determining that the first speech classification further comprises determining that the linguistic communication is understandable; 36. The method of claim 35, wherein selecting the one action further comprises the user selecting the one action to maintain one or more of the first plurality of speech features in the utterance.

38. wherein identifying the first plurality of speech features further comprises generating a score indicative of a severity of at least one of the first plurality of speech features; 36. The method of claim 35, wherein providing the instructions further comprises providing the instructions including the message identifying the score for presentation to the user.

39. 36. The method of claim 35, wherein the first plurality of speech features further comprises a corresponding plurality of scores, each of the plurality of scores defined along a scale for the respective speech feature.

40. 36. The method of claim 35, wherein identifying the first action further comprises identifying the first utterance classification based on at least one of: (i) an average of the plurality of scores; (ii) a weighted combination of the plurality of scores; (iii) a comparison to a dataset of a plurality of second scores; (iv) a neural network model; or (v) a generative transformer model.

41. 20. The system of claim 18, wherein identifying the first action further comprises applying a machine learning (ML) model to the first plurality of speech features, the machine learning (ML) model configured using a training dataset including a plurality of examples, each of the plurality of examples identifying (i) a respective second audio sample of a second language communication and (ii) a respective second classification from the plurality of speech classifications.

42. 36. The method of claim 35, wherein providing the instructions further comprises providing the instructions including the message identifying at least one of: (i) one or more of the first plurality of speech features; and (ii) the one action to modify the utterance.

43. identifying, by the one or more processors, a factor from a plurality of factors based on at least one of the first plurality of speech features; 43. The method of claim 42, wherein providing the indication further comprises providing the message identifying the one factor.

44. 36. The method of claim 35, further comprising generating, by the one or more processors, a second audio sample for playback to the user by modifying the first audio sample in accordance with the one action.

45. 36. The method of claim 35, wherein generating the second audio sample further comprises applying a speech synthesis model to the first audio sample and the one action to generate the second audio sample.

46. identifying, by the one or more processors, a second speech classification of the user's second language communication based on a plurality of second speech features, the plurality of second speech features being generated from a second audio sample identified at a time after providing the instruction; 36. The method of claim 35, further comprising determining, by the one or more processors, a progress metric based on a comparison of the first speech classification prior to the instruction and the second speech classification after submission of the instruction.

47. 36. The method of claim 35, wherein the first plurality of speech features further comprises at least one of: (i) respiration, (ii) voicing, (iii) articulation, (iv) resonance, (v) prosody, (vii) pitch, (viii) jitter, (ix) shimmer, (x) rhythm, (xi) pacing, or (xii) pauses.

48. identifying, by the one or more processors, a first video sample of a first non-verbal communication from the user at least partially contemporaneously with the first verbal communication; and identifying, by the one or more processors, a first plurality of non-verbal features of the first non-verbal communication using the first video sample, the first plurality of non-verbal features including at least one of a gesture, a facial expression, and an eye contact of the user; 36. The method of claim 35, wherein identifying the one action further comprises identifying the one action based on the first plurality of non-verbal features.

49. 36. The method of claim 35, wherein the user suffers from one of a speech or language disorder and is receiving speech therapy at least partially contemporaneously with the providing of the instructions.

50. 36. The method of claim 35, wherein the user is suffering from a disorder associated with the speech disorder and is receiving medication for the disorder at least partially contemporaneously with the providing of the instructions.

51. 1. A system for providing speech prompts based on characteristics of verbal communication from a user, the system comprising: One or more processors coupled to a memory, the processors comprising: Identifying a first audio sample of a first language communication from the user; generating a first plurality of speech features of the first language communication using the first audio samples; selecting an action from a plurality of actions for modifying one or more of the first plurality of speech features defining an utterance of the user; The system is configured to provide instructions to present a message prompting the user to make the utterance defined by the one action selected from the plurality of actions.

52. The one or more processors: determining that the verbal communication is unintelligible; 52. The system of claim 51, further configured for the user to select the one action to modify at least one of the first plurality of speech features in the utterance.

53. The one or more processors: determining that the verbal communication is understandable; 52. The system of claim 51, further configured for the user to select the one action to maintain one or more of the first plurality of speech features in the utterance.

54. The one or more processors: generating a score indicative of a severity of at least one of the first plurality of speech features; 20. The system of claim 18, further configured to provide the instructions, the instructions including the message identifying the score for presentation to the user.

55. 52. The system of claim 51, wherein the first plurality of speech features includes a corresponding plurality of scores, each of the plurality of scores defined along a scale for the respective speech feature.

56. 52. The system of claim 51, wherein the one or more processors are further configured to determine the first utterance classification based on at least one of: (i) an average of the plurality of scores; (ii) a weighted combination of the plurality of scores; (iii) a comparison to a dataset of a plurality of second scores; (iv) a neural network model; or (v) a generative transformer model.

57. 52. The system of claim 51 , wherein the one or more processors are further configured to apply a machine learning (ML) model to the first plurality of speech features, the ML model being configured using a training dataset including a plurality of examples, each of the plurality of examples identifying (i) a respective second audio sample of a second language communication and (ii) a respective second classification from the plurality of speech classifications.

58. 52. The system of claim 51, wherein the one or more processors are further configured to provide the instructions including the message identifying at least one of: (i) one or more of the first plurality of speech features; and (ii) the one action to modify the utterance.

59. 52. The system of claim 51, wherein the one or more processors are further configured to: identify a factor from a plurality of factors based on at least one of the first plurality of speech features; and provide the message identifying the one factor.

60. 52. The system of claim 51, wherein the one or more processors are further configured to generate a second audio sample for playback to the user by modifying the first audio sample according to the one action.

61. 61. The system of claim 60, wherein the one or more processors are further configured to apply a speech synthesis model to the first audio sample and the one action to generate the second audio sample.

62. The one or more processors: further configured to determine a second speech classification of the user's second language communication based on a plurality of second speech features generated from a second audio sample determined at a time subsequent to providing the instruction; 52. The system of claim 51, wherein the one or more processors are further configured to determine a progress metric based on a comparison of the first speech classification prior to the instruction and the second speech classification after submission of the instruction.

63. 52. The system of claim 51, wherein the first plurality of speech features further comprises at least one of: (i) respiration, (ii) voicing, (iii) articulation, (iv) resonance, (v) prosody, (vii) pitch, (viii) jitter, (ix) shimmer, (x) rhythm, (xi) pacing, or (xii) pauses.

64. The one or more processors: identifying a first video sample of a first non-verbal communication from the user at least partially contemporaneously with the first verbal communication; using the first video sample to identify a first plurality of non-verbal features of the first non-verbal communication, the first plurality of non-verbal features including at least one of a gesture, a facial expression, and an eye contact of the user; 52. The system of claim 51, further configured to identify the one action based on the first plurality of non-verbal features.

65. 52. The system of claim 51, wherein the user suffers from one of a speech or language disorder and is receiving speech therapy at least partially contemporaneously with the providing of the instructions.

66. 52. The system of claim 51, wherein the user suffers from a disorder associated with the speech disorder and is receiving medication for the disorder at least partially contemporaneously with the providing of the instructions.

67. 1. A method of providing needed relief to a speech expression deficiency in a user, comprising: obtaining, by one or more processors, a first metric associated with the user prior to completion of at least one of the plurality of sessions; repeating, by the one or more processors, providing the plurality of sessions to the user, each session of the plurality of sessions comprising: Identifying a first audio sample of a first language communication from the user; generating a first plurality of speech features of the first language communication using the first audio samples; identifying an action from a plurality of actions for modifying one or more of the first plurality of speech features defining an utterance of the user; repeating the providing step, including providing instructions to present a message prompting the user to perform the utterance defined by the one action selected from the plurality of actions; obtaining, by the one or more processors, a second metric associated with the user after completion of at least one of the plurality of sessions; wherein remission of the deficit in expressive speech is brought about for the user when the second metric (i) decreases from the first metric by a first predetermined margin, or (ii) increases from the first metric by a second predetermined margin.

68. 68. The method of claim 67, wherein the user is diagnosed with a condition comprising at least one of a speech disorder, an autism spectrum disorder (ASD), multiple sclerosis, a neurodegenerative disease, dementia, Parkinson's disease, Alzheimer's disease, an affective disorder, or schizophrenia.

69. 69. The method of claim 68, wherein the user is receiving a therapy at least partially concurrently with the at least one of the plurality of sessions, the therapy including at least one of a psychosocial intervention or a medication to address the condition.

70. 69. The method of claim 68, wherein said deficit in expressive speech is caused by said pathological condition.

71. 68. The method of claim 67, wherein the user is an adult at least 18 years of age.

72. 68. The method of claim 67, wherein the multiple sessions are provided over a period ranging from 3 days to 6 months.

73. the first language communication of the first audio sample includes an utterance of one or more words by the user; 68. The method of claim 67, wherein the first plurality of speech features further comprises at least one of: (i) respiration, (ii) voicing, (iii) articulation, (iv) resonance, (v) prosody, (vii) pitch, (viii) jitter, (ix) shimmer, (x) rhythm, (xi) pacing, or (xii) pauses.

74. At least one of the plurality of sessions: identifying a first speech classification of the first language communication based on the plurality of first speech features from a plurality of speech classifications; selecting an action from a plurality of actions based on the first speech classification, the action including modifying one or more of the speech features defining the user's speech; 68. The method of claim 67, wherein the plurality of speech classifications comprises at least one of: (i) mumbling, (ii) tongue-slurring, (iii) paralytic dysarthria, (iv) stuttering, or (v) intelligible.

75. The remission of the deficit in expressive speech of the user with a speech disorder is effected when the second metric decreases from the first metric by the first predetermined margin or when the second metric increases from the first metric by the second predetermined margin, the first metric and the second metric being: a Goldman-Fristoe Articulation Test-3 (GFTA-3) value, an Arizona Scale of Articulation Ability (Arizona-3) value, a Speech Intelligibility Index (SII) value, a percentage of intelligible words (PIW) value, a percentage of intelligible speech (PIU) value, a percentage of intelligible syllables (PIS) value, a percentage of correct consonants (PCC) value, a percentage of correct vowels (PVC) value, a percentage of correct vowels and diphthongs (PVC-R) value, a Stuttering Severity Inventory-4 (SSI-4) value, a Comprehensive Overview of the Experience of a Person Who Stutters 68. The method of claim 67, wherein the at least one of the following is a score for the OASES score, a maximum phonation time (MPT) value, a GRBAS scale value, a vocal range profile (VRP) value, a voice disorder index (VHI) value, a voice-related quality of life (V-RQOL) value, an auditory perceptual assessment of speech (CAPE-V) value, a deafening deafness score (DDK) value, a prosodic voice screening profile (PVSP) value, a Bzoch hypernasality scale value, a resonance severity index value, a nasalization rate value, a Western Comprehensive Aphasia Assessment (WAB) value, a Boston Diagnostic Aphasia Examination (BDAE) value, a Communication Effectiveness Index (CETI) value, an Adult Aphasia Examination (ABA-2) value, a DDK rate, a percentage of correct consonants revised (PCC-R) value, a Frenchay Dysarthria Assessment (FDA-2) value, or a Dysarthria Impact Profile (DIP) value.

76. 68. The method of claim 67, wherein the remission of the expressive speech of the user with ASD occurs when the second metric decreases from the first metric by the first predetermined margin or when the second metric increases from the first metric by the second predetermined margin, and the first metric and the second metric are at least one of an Autism Spectrum Disorders Observation Schedule (ADOS) value, a Test of Pragmatic Language (TOPL-2) value, a CETI value, a Interpersonal Responsiveness Scale-2nd Edition (SRS-2) value, a Comprehensive Assessment of Oral Language (CASL-2) value, and a Functional Communication Profile-R (FCP-R) value.

77. 68. The method of claim 67, wherein the remission of the expressive speech of the user with multiple sclerosis occurs when the second metric decreases from the first metric by the first predetermined margin or when the second metric changes from the first metric by the second predetermined margin, and the first metric and the second metric are at least one of a WAB value, a BDAE value, a CETI value, an ABA-2 value, a DDK velocity value, a PCC-R value, an FDA-2 value, or a DIP value.

78. 68. The method of claim 67, wherein the remission of the expressive speech of the user with an affective disorder occurs when the second metric decreases from the first metric by the first predetermined margin or when the second metric changes from the first metric by the second predetermined margin, and the first metric and the second metric are at least one of Hamilton Rating Scale for Depression (Ham-D) values.

79. 68. The method of claim 67, wherein the remission of the expressive speech of the user with schizophrenia occurs when the second metric decreases from the first metric by the first predetermined margin or when the second metric changes from the first metric by the second predetermined margin, and the first metric and the second metric are at least one of a Motivation and Pleasure Scale-Self-Report (MAP-SR) value, a Social Effort and Conscientiousness Scale (SEACS) social effort value, and a SEACS social conscientiousness value.

80. 68. The method of claim 67, wherein the first metric is determined based on a corresponding speech classification of a plurality of speech classifications in a first session of the plurality of sessions, and the second metric is determined based on the corresponding first speech classification in a second session of the plurality of sessions.

Citation Information

Cited By

  • Phonetic Transcription Model Using An Artificial Neural Network, And Method And System For Evaluating Articulation Accuracy Using The Same

    KR103006000B1