Systems and methods for assessing brain health
A computer-implemented method using machine learning models to analyze speech and text for brain health assessment addresses the inefficiencies of traditional methods, providing objective and accessible evaluations of brain health and neurological disorders.
Patent Information
- Application Number
- JP2025536997
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-22
- Filing Date
- 2023-12-22
- Publication Date
- 2025-12-25
AI Technical Summary
Traditional methods for assessing brain health through speech analysis are time-consuming, expensive, and inaccessible, lacking objective and practical measurements for changes due to fatigue and stress.
A computer-implemented method that analyzes speech and text to generate a speech score indicative of brain health, using machine learning models to process audio recordings, identify speech quality and content indicators, and compare them to control scores for accurate brain health assessment.
Provides an efficient and accessible means to assess brain health, offering objective measurements of speech quality and content, enabling identification of neurological disorders and their progression.
Smart Images

Figure 2025542401000001_ABST
Abstract
Description
[Technical Field]
[0001] The described embodiments relate to computing systems and computer-implemented methods for assessing brain health. In some embodiments, these computing systems and computer-implemented methods relate to analyzing speech and language to assess brain health. [Background technology]
[0002] Analysis of speech to determine or assess changes in brain health is a specialized medical and psychological field that requires advanced training by practitioners or clinicians. Traditional solutions, one-on-one sessions with trained professionals, are time-consuming, expensive, and / or inaccessible to many people with brain health issues that require diagnosis or ongoing treatment and management. Accurately identifying and assessing changes in brain health due to fatigue and stress is difficult. Objective measurements of these performance markers are not readily available or practical.
[0003] It would be desirable to overcome or ameliorate some of the drawbacks associated with such conventional methods and systems, or at least provide a useful alternative thereto.
[0004] Throughout this specification, the word "comprise" or variations such as "comprises" or "comprising" should be understood to imply the inclusion of the stated element(s), component(s) or step(s), but not the exclusion of other element(s), component(s) or step(s).
[0005] Any discussion of documents, acts, materials, devices, articles or the like which has been described in this specification should not be construed as an admission that any or all of such matters form part of the prior art or were common general knowledge in the art to which this disclosure pertains, as if in existence prior to the priority date of each of the appended claims. Summary of the Invention
[0006] The present disclosure is directed to a computer-implemented method comprising receiving an audio recording of sounds made by a subject, identifying from the audio recording a text representation of at least one sound made by the subject, providing the audio recording to an acoustic analysis model, receiving from the acoustic analysis model a speech quality dataset comprising one or more speech quality indicator(s), providing the text representation to a text analysis model, receiving from the text analysis model a speech content dataset comprising one or more speech content indicator(s), and identifying a speech score associated with the subject from the speech quality dataset and the speech content dataset, wherein the speech score is indicative of brain health of the subject.
[0007] In some embodiments, the speech score may include at least a naturalness score, which may include measures of at least one of fundamental frequency variability, intensity variability, formant transitions, speaking rate and rhythm, phonation measures, spectral measures, temporal measures, articulation, coarticulation, resonance, and prosody. In some embodiments, the speech score may include at least an intelligibility score.
[0008] In some embodiments, the speech content dataset includes at least a discourse complexity index, which may include measures of at least one of lexical diversity, syntactic complexity, allusive coherence, thematic development, argumentative structure, implicit and explicit information comparison, interactivity, intertextuality, modality and modulation, and pragmatic factors.
[0009] In some embodiments, the speech score comprises one or more of a communication effectiveness score, a dysarthria score, a disease severity score, a social communication score, a voice quality score, an intelligibility score, and / or a naturalness score.
[0010] In some embodiments, the speech quality dataset comprises one or more of a timing indicator, an articulation indicator, a resonance indicator, a prosody indicator and / or a voice quality indicator.
[0011] In some embodiments, the speech content dataset comprises one or more of a semantic complexity index, an idea density index, a verbal fluency index, a lexical diversity index, an information content index, a dialogue structure index and / or a grammatical complexity index.
[0012] In some embodiments, the method further comprises comparing the speech score associated with the subject to a control speech score to determine brain health of the subject, hi some embodiments, the control speech score is experimentally or statistically determined.
[0013] In some embodiments, the sounds produced by the subject are speech and / or sounds, and the speech and / or sounds are produced in response to a speech task presented to the subject, and / or the speech and / or sounds are recorded during a period in which the subject was continuously observed without instructions.
[0014] In some embodiments, the disclosed methods may include performing quality assurance of the data.
[0015] In some embodiments, quality assurance of the data includes one or more of: (a) determining if speech is present; (b) determining the recording time; (c) determining if the recording time is within predetermined recording limits; (d) removing abnormal noise; (e) removing silence; (f) determining the number of speakers recorded; and (g) separating speakers.
[0016] In some embodiments, a notification is sent to the mobile computing device if the voice recording fails a data quality assurance step.
[0017] In some embodiments, the determination of the utterance content score and the utterance quality score occurs in parallel.
[0018] In some embodiments, the disclosed method further includes receiving one or more refinement parameters, in some embodiments, the one or more refinement parameters include one or more brain health states, one or more brain health attributes, one or more brain health indicators, one or more neurological disorders, one or more subjects, and / or a medical history of the subject.
[0019] In some embodiments, in response to receiving the one or more refinement parameters, the acoustic analysis model and / or the text analysis model are adjusted such that the speech quality data set and / or the speech content data set are associated with the one or more refinement parameters.
[0020] In some embodiments, one or more of the text representation, the speech quality dataset, the speech content dataset, and the speech score are determined by one or more machine learning model(s). In some embodiments, at least one of the one or more machine learning model(s) is a neural network.
[0021] In some embodiments, the disclosed method further comprises converting the text representation into word embeddings.
[0022] In some embodiments, the method further comprises converting the audio recording into one or more of a series of filter bank spectra and / or one or more sound wave spectrograms before providing the audio recording to the acoustic analysis model.
[0023] In some embodiments, the textual representation is determined using natural language processing.
[0024] In some embodiments, the method further comprises determining that the textual representation cannot be identified, and if determining that the textual representation cannot be identified, omitting identifying the textual representation. In some embodiments, determining that the textual representation cannot be identified comprises determining that the audio recording does not include at least one morpheme, phoneme, or sound that can be represented in text.
[0025] In some embodiments, after determining that the audio recording does not contain at least one morpheme or sound that can be expressed in text, an indication that the audio recording does not contain at least one morpheme, phoneme, or sound that can be expressed in text is identified, and the indication is used as input for determining a speech score.
[0026] The present disclosure is also directed to non-transitory machine-readable media storing instructions that, when executed by one or more processors, cause a computing device to perform any of the methods of the present disclosure.
[0027] The present disclosure is also directed to a system for assessing brain health including a speech analysis module configured to identify a speech quality dataset from the audio recording, a speech text module configured to identify a text representation of the audio recording, a speech content module configured to identify a speech content dataset from the audio recording, and a brain health determination module configured to identify a speech score using the speech quality dataset and the speech content dataset, wherein the speech score is indicative of the brain health of one or more subjects captured in the audio recording.
[0028] In some embodiments, the system further comprises a screen configured to present instructions to one or more subjects.
[0029] In some embodiments, the system further comprises an audio recording device configured to record one or more subjects.
[0030] In some embodiments, the speech analysis module includes one or more machine learning models and / or information sifting. In some embodiments, the speech to text module includes a machine learning model. In some embodiments, the speech content module includes a machine learning model. In some embodiments, the machine learning model is a recurrent neural network.
[0031] The present disclosure is also directed to a computer-implemented method that includes receiving an audio recording of sounds made by a subject, providing the audio recording to an acoustic analysis model, receiving a speech quality dataset from the acoustic analysis model that includes one or more speech quality indicator(s), and identifying a speech score associated with the subject from the speech quality dataset, wherein the speech score is indicative of brain health of the subject. [Brief explanation of the drawings]
[0032] [Figure 1]FIG. 1 is a schematic diagram showing an overview of a method for assessing brain health, according to some embodiments. [Figure 2] FIG. 1 is a block diagram of a system for assessing brain health, according to some embodiments. [Figure 3] FIG. 1 is a block diagram of an alternative system for assessing brain health, according to some embodiments. [Figure 4] FIG. 1 is a process flow diagram of a method for assessing brain health, according to some embodiments. [Figure 5A] 1 is a screenshot example of a user interface for assessing brain health, according to some embodiments. [Figure 5B] 1 is a screenshot example of a user interface for assessing brain health, according to some embodiments. [Figure 5C] 1 is a screenshot example of a user interface for assessing brain health, according to some embodiments. [Figure 6] FIG. 1 is a process flow diagram of a method for performing a brain health assessment, according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0033] The described embodiments relate to computing systems and computer-implemented methods for assessing brain health. Some embodiments relate to the use of speech and / or text analysis to identify a speaker's neurological condition and / or disease. Some embodiments may include one or more machine learning models for performing the speech and / or text analysis.
[0034] In the context of the present disclosure, brain health may include or otherwise be defined as the presence and / or state of a subject's brain integrity and mental and cognitive function at a given age, regardless of the presence or absence of overt brain disease affecting normal brain function. Brain health may be and / or may be measured as the ability or inability to perform all, most, some, or all of the mental processes of cognition, including the ability to learn and make judgments, use language, remember, etc. In some embodiments, brain health may be defined statistically, i.e., by studying a patient population. In some embodiments, brain health may be a relative measurement compared to previous measurements of a single patient. Brain health may or may not be associated with the presence of a diagnosed or undiagnosed neurological disease.
[0035] FIG. 1 is a schematic diagram showing an overview of a method 100 performed by the system of the present disclosure.
[0036] In step 1 shown in FIG. 1 , according to described embodiments, sounds made by the subject are recorded. The sounds made by the subject may be made by the subject in response to some form of stimulus. In some embodiments, the subject may be presented with one or more speech and / or recording tasks that require the subject to speak specific phrases, carry on a conversation, make specific sounds, or record sounds for a predetermined or indefinite period of time. The speech task may be presented to the subject via a mobile computing device such as a smartphone, tablet, laptop, or computer kiosk. In some embodiments, the subject may not be presented with a specific speech task, but may instead be presented with a recording task. The recording task may be continuous over an extended period of time to record sounds and / or vocalizations that the subject may make naturally and / or spontaneously. Spontaneous sounds and / or vocalizations may include coughing, breathing, vomiting, wheezing, and / or nonsensical mumbling.
[0037] In step 1A, the recorded sounds may undergo one or more quality control processes to improve and / or remove audio recordings that may not be suitable for determining brain health.
[0038] In step 2, the recorded sounds are converted to text. In some embodiments, speech, spoken words, and / or non-lexical vocal / conversational sounds such as "uh-huh," "uh-huh," and / or throat clearing contained in the recorded sounds may be converted to text / text representations. The text representations created in step 2 and the recorded sounds created in step 1 may be provided to a text analysis module and an acoustic analysis module, respectively.
[0039] In step 3A, a text analysis module performs text analysis on the text representation created in step 2. In step 3B, an acoustic analysis module performs acoustic analysis on the audio recording recorded in step 1. The text analysis module and the acoustic analysis module can utilize one or more machine learning models to output a speech content dataset and a speech quality dataset, respectively. These datasets can describe various characteristics of the audio recording.
[0040] In step 4, the speech content dataset and speech quality dataset are analyzed for brain health using another machine learning model.
[0041] In step 5, the brain health analysis results in a determination of the subject's brain health. The determination of the subject's brain health may be indicative of whether the subject has a neurological disorder or may be indicative of the progression of an existing neurological disorder. The determination of the brain health may also be indicative of the subject's brain health level. The subject's brain health level may be on a scale from healthy to abnormal or unhealthy, or may be a binary determination associated with a brain health threshold.
[0042] 2 is a block diagram of a system 200 for assessing brain health, according to some embodiments. System 200 may include a mobile computing device 210, a database 245, and a speech analysis server 250 communicating over a network 240. Network 240 may include at least a portion of one or more networks having one or more nodes that, for example, originate, receive, forward, generate, buffer, store, route, switch, process, or any combination thereof, one or more messages, packets, signals, any combination thereof, etc. Network 240 may include, for example, one or more of a wireless network, a wired network, the Internet, an intranet, a public network, a packet-switched network, a circuit-switched network, an ad-hoc network, an infrastructure network, a public switched telephone network (PSTN), a cable network, a cellular network, a satellite network, a fiber optic network, combinations thereof, etc.
[0043] Mobile computing device 210 may be a mobile or handheld computing device such as a smartphone or tablet, a laptop, a smartwatch, a passive room-based microphone, or a PC, and in some embodiments may comprise multiple computing devices. Mobile computing device 210 may include one or more processors 215 and memory 220 that stores instructions (e.g., program code) executable by processor(s) 215. Processor(s) 215 may include one or more microprocessors, central processing units (CPUs), application specific instruction set processors (ASIPs), application specific integrated circuits (ASICs), or other processors capable of reading and executing instruction code.
[0044] In some embodiments, mobile computing device 210 and speech analysis server 250 may have a client-server architecture. Mobile computing device 210 may be a thin client tasked with presentation logic and input / output management for presenting speaking and / or recording tasks and collecting voice recordings, with all or most logic and data storage operations related to processing the voice recordings being handled by speech analysis server 250. For example, the thin client may be configured to receive data packets from speech analysis server 250 over network 240 that, when received by communications module 230 and executed by processor(s) 215, cause the thin client to display text, images, and / or interactive elements on the screen(s) of the mobile computing device and / or accept input from one or more users in the form of button operations, screen operations, video recordings, and / or audio recordings.
[0045] In some embodiments, the mobile computing device 210 may be a thick client, capable of collecting, storing, and / or pre-processing voice recordings before they are communicated to the voice analysis server 250. For example, the thick client may include one or more modules of executable program code that, when executed by the processor(s) 215, cause the thick client to pre-process the voice recordings to perform quality control of the voice recordings, as described below. The thick client may also be configured to prompt the user(s) / subject to re-record the speaking task or restart extended recording if the voice recording fails one or more quality control metrics.
[0046] The memory 220 may comprise one or more volatile or non-volatile memory types. For example, the memory 220 may comprise one or more of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The memory 220 is configured to store program code accessible by the processor(s) 215. The program code includes executable program code modules. In other words, the memory 220 is configured to store executable code modules configured to be executable by the processor(s) 215. The executable code modules, when executed by the processor(s) 215, may cause the processor(s) 215 to perform various methods, as described in more detail below. The memory 220 may comprise an evaluation application 225.
[0047] The assessment application 225 may be configured to present and manage the assessment process. The assessment application 225 may be a piece of software or pieces that are downloadable from a website, an application store, and / or an external data storage device. In some embodiments, the assessment application 220 may be pre-installed.
[0048] The mobile computing device 210 may include a communications module 230. The communications module 230 facilitates communication with components of the system 200, such as the speech analysis server 250 and / or the database 245, over a network 240.
[0049] The mobile computing device 210 may include a device I / O 235 configured to record or accept as input sounds, such as speech. In some embodiments, the mobile computing device may be configured to record and store any input sounds and communicate them to a speech analysis server 250, for example, via a network 240. In some embodiments, the mobile computing device 210 may be configured to stream any input sounds directly to the speech analysis server 250. In some embodiments, the mobile computing device may be configured to communicate speech input to a database 245 for storage and eventual processing by the speech analysis server 250. The device I / O 235 may also comprise one or more additional user input or user output peripherals, such as, for example, one or more of a display screen, a touchscreen display, a camera, an accelerometer, a mouse, a keyboard, or a joystick.
[0050] Network 240 may include a combination of network interface hardware and network interface software suitable for establishing, maintaining, and facilitating communications over associated communication channels. Network 240 may include at least a portion of one or more networks having one or more nodes that, for example, originate, receive, forward, generate, buffer, store, route, switch, process, or any combination thereof, one or more messages, packets, signals, any combination thereof, etc. Network 240 may include, for example, one or more of a wireless network, a wired network, the Internet, an intranet, a public network, a packet-switched network, a circuit-switched network, an ad-hoc network, an infrastructure network, a public switched telephone network (PSTN), a cable network, a cellular network, a satellite network, a fiber optic network, combinations thereof, etc.
[0051] Database 245 may form part of system 200, may be local to system 200, or may be remotely located from system 200 and accessible via network 240, for example. Database 245 may be configured to store data associated with system 200. Database 245 may be a centralized database. Database 245 may be a mutable data structure. Database 245 may be a shared data structure. Database 245 may be a data structure supported by a database system such as one or more of PostgreSQL, MongoDB, and / or ElasticSearch. Database 245 may be configured to store current states or current values of information associated with various attributes (e.g., "current knowledge").
[0052] In some embodiments, database 245 may be a SQL database that includes a table with a row entry for each user of system 200. For example, the row entries may include the subject's name, the subject's password, the subject's brain health determination, the subject's speech score, and / or other entries related to the user's brain health.
[0053] Speech analysis server 250 may include one or more processors 255 and memory 260 that stores instructions (e.g., program code) that, when executed by processor(s) 255, cause speech analysis server 250 to function according to the described methods. In some embodiments, speech analysis server 250 may operate in conjunction with one or more mobile computing devices 210 to facilitate the process of determining brain health, and in some embodiments may provide the determination to a user after the input has been appropriately analyzed.
[0054] The processor(s) 255 may comprise one or more microprocessors, central processing units (CPUs), application specific instruction set processors (ASIPs), application specific integrated circuits (ASICs), or other processors capable of reading and executing instruction code.
[0055] The memory 260 comprises one or more volatile or non-volatile memory types. For example, the memory 260 may comprise one or more of a random access memory (RAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The memory 260 is configured to store program code accessible by the processor(s) 255. The program code includes executable program code modules. In other words, the memory 260 is configured to store executable code modules configured to be executable by the processor(s) 255. The 255 executable code modules, when executed by the processor(s) 255, cause the speech analysis server 250 to perform specific functions, as described in more detail below. For example, the memory 260 may comprise a data processing module 262, a quality control module 264, a speech-to-text module 266, an acoustic analysis module 268, a text analysis module 270, a brain health status determination module 272, and / or a status determination module 274.
[0056] The speech analysis server 250 may also include a communications module 280 that facilitates communication with components of the system 200 over the network 240. The communications module 280 may include a combination of network interface hardware and software suitable for establishing, maintaining, and facilitating communications over the associated communications channels.
[0057] Data processing module 262 is configured to receive and / or request data from other components within system 200, such as mobile computing device 210 or database 245. In some embodiments, data processing module may be configured to receive and / or request data from devices external to system 200, such as, for example, an external database belonging to a medical record service provider or an external control testing system.
[0058] The data processing module 262 may receive data in the form of audio recordings. In some embodiments, the audio recordings are of one or more people speaking or making biological sounds. The one or more people speaking may be medical test subjects, medical practitioners, nurses, neurologists, family members, personal caregivers, or other types of people capable of administering and / or participating in the speaking task, recording task, and / or brain health assessment process.
[0059] In some embodiments, the data processing module may be configured to receive, via network 240, an audio recording of one or more people speaking. The audio recording, in some embodiments, may be a .WAV file, a .PCM file, an .AIFF file, an .MP3 file, or a .WMA file. The data processing module, in some embodiments, may be configured to temporarily store the received audio recording for subsequent communication to other modules in memory 260.
[0060] The data processing module 262, in cooperation with the communications module 280, may be configured to monitor receipt of individual data packets received over the network 240. In some embodiments, the voice recording may be transmitted using a lossless data transmission protocol, such as the Transmission Control Protocol (TCP), to ensure the integrity of the audio file. The data processing module 262 may request retransmission of particular TCP packet(s) if they are not received. The mobile computing device 210, the database 245, and / or other external devices may be configured to retransmit lost packets upon receiving an indication from the data handling module 262 that the packets were not received.
[0061] In some embodiments, where network connectivity and / or bandwidth may be limited, or where data transfer speed takes priority over data integrity, the audio recording may be transmitted using a lossy data transfer protocol, such as User Datagram Protocol (UDP).
[0062] The quality control module 264 may be configured to receive one or more audio recordings from the data processing module 262. The quality control module 264 is configured to determine whether the audio recordings are suitable for use in brain health assessment. The quality control process may determine whether the audio recordings are suitable for use in brain health assessment by performing one or more quality tests on the audio recordings. The quality tests may include, but are not limited to, verifying whether the audio file is empty, detecting whether other speakers are recorded in addition to the primary speaker, verifying whether the generated stimuli match the target stimuli according to one or more selected tasks and associated minimum requirements, and / or verifying the signal-to-noise ratio of the recording.
[0063] In some embodiments, one or more quality tests may each have its own pass threshold indicating an acceptable level of quality for that test. Some quality tests may be binary pass / fail tests, such as whether an audio file is empty. For some quality tests, the pass threshold may be the percentage of a recording that contains a particular unwanted quality; for example, a voice recording may pass the quality test if its signal-to-noise ratio is less than 50%. Other quality tests may have a pass threshold related to the strength of a particular feature of the voice recording; for example, if a particular quality test evaluates the presence of two or more speakers, the voice recording may fail the particular quality test if the additional speakers' voices do not reach a minimum decibel rating.
[0064] In some embodiments, the quality control module 264 may also perform one or more quality improvement processes, including, but not limited to, filtering background noise such as clicks, static artifacts, and / or audio artifacts; equalizing speech intensity to aid in analysis such as syllable emphasis where speech intensity is relevant; and / or trimming long periods of silence from the audio file, such as long pauses or silences at the beginning and end of a recording.
[0065] In some embodiments, if the quality control module 264 determines that an audio recording is not of sufficient quality, for example, by the audio recording not passing a certain number of quality checks (e.g., not meeting a pass threshold), the quality control module may be configured to mark the particular audio recording as inadequate. In some embodiments, if the audio recording is determined to be inadequate, the quality control module 264 may communicate a notification / indication to the mobile computing device 210 indicating that the particular audio recording is of unacceptable quality. In some embodiments, this notification / indication may include a prompt to re-record the particular audio recording by redoing the associated speaking task or to resume / restart extended recording of the audio.
[0066] The speech-to-text module 266 may be configured to receive as input an audio recording of sounds made by the subject and to output a text representation of the sounds made by the subject. The representation of the sounds made by the subject may be a direct written transcription of the spoken words. In some embodiments, the written words may not be a direct transcription but may be a representation of the spoken words. The representation may have repeated words, stutters, and loud pauses removed. If the speech is difficult to understand, the speech-to-text module 266 may, for example, do one or more of the following: not include the difficult to understand speech; replace the difficult to understand word with an indication that the written word cannot be determined; include a best guess of the speech; and / or create a list of possible words based on the quality of the difficult to understand speech or the context of the sentence to which the difficult to understand word belongs.
[0067] The acoustic analysis module 268 is configured to identify, from the audio recording, a speech quality dataset comprising one or more speech quality metrics. The one or more speech quality metrics may be indicative of the quality of the subject's speaking style. The subject's speech quality may be independent of the content of the subject's speech and may include metrics such as speaking rate, inter-word pauses, stuttering, pitch, and / or articulation.
[0068] The text analysis module 270 is configured to identify, from the textual representation of the sounds produced by the subject, a speech content dataset including one or more speech content indicators. The speech content indicators may be indicative of the content of the subject's speech. The content of the subject's speech may be unrelated to the quality of the subject's speech and may include indicators such as lexical complexity, syllabic complexity, correct / incorrect pronunciation, repetition, information efficiency and / or grammatical complexity, semantic complexity indicators, idea density indicators, verbal fluency indicators, lexical diversity indicators, information content indicators, and / or dialogue structure indicators. One or more of these indicators may be used to form a discourse complexity indicator. In some embodiments, the speech content dataset may include a discourse complexity indicator.
[0069] Discourse complexity may also be referred to as conversational intricacy, discourse depth, dialogic complexity, speech sophistication, linguistic complexity, narrative intricacy, rhetorical depth, discourse linguistic complexity, argumentative complexity, communication sophistication, and / or communication effectiveness. Discourse complexity measures refer to the multidimensional and intricate nature of the text or spoken language used in communication. In some embodiments, discourse complexity measures may include the structural, conceptual, and / or relational components of utterances that interact within a communication event.
[0070] In some embodiments, a discourse complexity index may be composed of one or more aspects that contribute to the index. For example, a lexical diversity index may be considered a discourse complexity index. In some embodiments, a discourse complexity index may include multiple different indices that are used in combination to provide a discourse complexity index. Multiple indices may be weighted differently to contribute to the discourse complexity index. In other embodiments, multiple indices, or a subset of multiple indices, may be weighted substantially equally to contribute to the discourse complexity index. For example, a weighting function may be used that appropriately weights individual indices in the discourse complexity index based on the context of the evaluation.
[0071] In some embodiments, discourse complexity indicators may include lexical diversity, which takes into account the range and variety of words used. In some embodiments, discourse complexity indicators may include syntactic complexity, which takes into account the use of diverse and sophisticated sentence structures. In some embodiments, discourse complexity indicators may include coherence of reference, which is how consistently and clearly subjects, objects, and concepts are connected and referenced throughout the discourse. In some embodiments, discourse complexity indicators may include thematic development, which is the depth and complexity with which a topic or theme is explored. In some embodiments, discourse complexity indicators may include argument structure, which is the presentation and support of claims, counterarguments, and solutions.
[0072] In some embodiments, discourse complexity indicators may include a comparison of implicit and explicit information, i.e., the balance between what is directly stated and what is implied or left unsaid. In some embodiments, discourse complexity indicators may include interactivity, which is a measurement of the level of engagement and interaction between the speaker and listener or writer and reader. In some embodiments, discourse complexity indicators may include intertextuality, which refers to references to other texts or discourses within a given discourse. In some embodiments, discourse complexity indicators may include modality and modulation, which are the use of linguistic resources to express possibility, necessity, obligation, or evaluation. In some embodiments, discourse complexity indicators may include pragmatic factors such as consideration of context, speaker intention, and listener / reader interpretation.
[0073] The brain health determination module 272 is configured to receive as input the speech quality dataset and the speech content dataset and determine a speech score indicative of the subject's speech in relation to the presence and / or progression of a neurological disease and / or condition and / or a change in the subject's brain health or condition. The speech score may include one or more composite scores, such as a communication effectiveness score, a dysarthria score, a disease severity score, a social communication score, a voice quality score, an intelligibility score, and / or a naturalness score. In some embodiments, the speech score may be compared to a statistically or experimentally determined scale to indicate the severity of a known neurological disorder and / or the subject's brain health. In some embodiments, the speech score may indicate the progression of an existing neurological disorder or known brain health issues / attributes specific to a particular subject. In some embodiments, the speech score may be determined or partially determined using previous speech scores. The previous speech scores may be from one or more subjects.
[0074] In some embodiments, the intelligibility score may be formed from speech intelligibility metrics. The intelligibility score may provide an indication related to the intelligibility of the speech, i.e., how clearly a speaker speaks so that the speech is understandable to a listener. In some embodiments, the intelligibility score may be formed from intrusive intelligibility metrics and non-intrusive intelligibility metrics. In some embodiments, the intelligibility score may take into account intelligibility metrics including an articulation index, a sound-speech transfer index, and coherence-based intelligibility.
[0075] In some embodiments, the naturalness score may also be referred to as smoothness, healthyness, or strangeness. In the context of speech, naturalness relates to the degree to which spoken words sound smooth, effortless, and typical of a human speaker. Naturalness refers to the absence of unnaturalness, artificiality, or awkwardness. When evaluating speech synthesis systems, such as text-to-speech engines, naturalness is an important criterion, reflecting how closely the synthesized speech resembles authentic human speech in intonation, rhythm, stress, and other prosodic features. In some embodiments, the naturalness score is calculated by assigning weights that differentiate different components of speech across speech subsystems. The naturalness score can be achieved through the use of machine learning or regression models. In some embodiments, the machine learning model may be supervised or unsupervised. In some embodiments, the regression models used include, but are not limited to, linear regression, logistic regression, polynomial regression, stepwise regression, Bayesian linear regression, quantile regression, principal component regression, elastic net regression, ridge regression, and lasso regression. The naturalness score formed from these components is sometimes referred to as a naturalness composite. In some embodiments, a supervised statistical machine learning framework is utilized to maximize the transparency of the features included. These different components that form the naturalness score include, but are not limited to, breathiness, pronunciation (voice quality), articulation, resonance, and prosody. In some embodiments, the naturalness composite may include a prosody measure that examines patterns of rhythm, stress, and intonation. The distribution and pattern of stressed syllables and the overall prosodic contour of an utterance contribute significantly to the perception of naturalness.
[0076] Natural speech has pitch variability, and the naturalness score may include a measure of fundamental frequency (F0) variability. This measure measures the mean, standard deviation, and contour of the fundamental frequency and evaluates its variability and pattern. Natural speech also has dynamic intensity, and the naturalness score may include a measure of intensity (or loudness) variability. This measure measures average intensity, its variability, and pattern over time. In some embodiments, the naturalness score may also include a measure of formant transitions, where rapid movements in formant frequencies (particularly F1 and F2) indicate fluid transitions between vowels and consonants. Speech rate and rhythm may also be part of the naturalness score, where the number of syllables or words per unit time is calculated. Additionally, examining variability in the duration of vowels, consonants, and pauses can provide insight into the rhythm of speech and can be used as a measure of speech rate and rhythm.
[0077] In some embodiments, these measures may be limited by upper or lower thresholds. Some thresholds may be dictated by underlying anatomy; for example, during typical conversation, adult voices rarely exceed the 50 Hz to 500 Hz range. Deviations of the measures beyond a predetermined threshold may be used to change the weighting of the measures in determining the naturalness score. In some embodiments, deviations from the threshold may result in the measures being weighted more or less in determining the naturalness score. In some embodiments, deviations may be used to determine potential errors or poor quality in the sound and / or text being analyzed.
[0078] In further embodiments, the naturalness score may include pronunciation or voice quality measurements that can evaluate parameters such as frequency perturbation (also known as jitter) and amplitude perturbation (also known as shimmer). High values of frequency perturbation and amplitude perturbation may indicate developmental disorders that result in unnaturalness (or a low naturalness score). Additionally, spectral measurements may also be included in the naturalness score. Spectral measurements may use the harmonic noise ratio (HNR) to measure the amount of noise in a speech signal. A lower HNR may indicate breathiness or hoarseness, which may result in less natural speech (resulting in a lower naturalness score). The naturalness score may also include a measure of resonance.
[0079] In some embodiments, the naturalness score may further include temporal measurements of the duration of examined and analyzed segments, pauses, and speech rate. Articulation and coarticulation measurements, including examining how sounds interact with each other, may also be included in the assessment of naturalness. Natural speech involves overlapping and blending of adjacent sounds. In some embodiments, the naturalness score, or naturalness composite, is formed, created, or generated from one or more metrics, including fundamental frequency variation, intensity variation, formant transitions, speech rate and rhythm, onset measurements, spectral measurements, temporal measurements, articulation, coarticulation, resonance, and prosody.
[0080] The status determination module 274 is configured to receive the speech score and determine the presence, progression, or regression of a neurological disorder and / or brain health condition. This determination may include a list of potential neurological disorders that may require further testing, an ordered list containing the probability / likelihood that the subject has one or more neurological disorders, and / or a yes / no determination regarding the presence of a particular neurological disorder or brain health condition. In some embodiments, the subject may already have been diagnosed with a neurological disorder and / or have a known brain health condition, and the system 200 may know this from previous speech and / or recording tasks or from medical data entered as part of the participation and / or enrollment process. If the subject has an existing disorder, the status determination module 274 may be configured to compare the previous speech score with the currently determined speech score to determine the progression of the existing disorder.
[0081] The speech analysis server may include or otherwise utilize an AI model incorporating a deep learning-based computational structure, including an artificial neural network (ANN). An ANN is a computational structure inspired by biological neural networks and includes one or more layers of artificial neurons configured or trained to process information. Each artificial neuron includes one or more inputs and an activation function for processing the received inputs and generating one or more outputs. The output of each neuron layer is connected to a subsequent neuron layer using a link. Each link may have a defined numerical weight that determines the strength of the link as information passes through several layers of the ANN. During the training phase, various weights and other parameters defining the ANN are optimized to obtain a trained ANN using inputs and known outputs for those inputs. This optimization may be performed through various optimization processes, including backpropagation. ANNs incorporating deep learning techniques include multiple hidden layers of neurons between an initial input layer and a final output layer. Multiple hidden layers of neurons enable the ANN to model complex information processing tasks, including the task of determining standard and non-standard user behaviors, performed by the system 200.
[0082] In some embodiments, the ML model may incorporate one or more variants of a convolutional neural network (CNN), a type of deep neural network, to perform various processing operations for determining brain health. A CNN includes various hidden layer neurons between an input layer and an output layer, and convolves an input through the various hidden layer neurons to generate an output.
[0083] FIG. 3 is a block diagram of a mobile computing device 310 for assessing brain health, according to some embodiments. The mobile computing device 310 may include the same or similar components as the audio analysis server 250 necessary to perform the disclosed methods. The client device 310 may be an “all-in-one” system, including the assessment application 225, data processing module 262, quality control module 264, speech-to-text module 266, acoustic analysis module 268, text analysis module 270, brain health assessment module 272, and / or status assessment module 274, as described above with reference to FIG. 2. The client device 310 may also include a data storage device 320, which may be used to store unprocessed audio recordings and / or brain health assessments and data related to brain health assessments performed by the mobile computing device 310. The client device 310 may communicate with a database 245 via the communication module 230 and / or the network 240. In some embodiments, client device 310 may be configured to perform all of the processing steps for performing the disclosed uses.
[0084] 4 is a process flow diagram of a method 400 for determining brain health, according to some embodiments. Method 400 may be performed by an embodiment of system 200 or an embodiment of system 300, or any other configuration of hardware and / or software installed and / or executed on one or more servers, systems, and / or client devices. The steps of method 400 depicted in FIG. 4 and described below may correspond more or less to steps of method 100, described above and in FIG. 1. In some embodiments, one or more steps of method 400 may correspond to the same steps of method 100.
[0085] At 410, the data processing module 262 receives an audio recording of a user or subject speaking. In some embodiments, step 410 may correspond to step 1 of FIG. 1. The audio recording may be in a .WAV file, a .PCM file, an .AIFF file, an .MP3 file, or a .WMA format in some embodiments. The audio recording may be lossy or lossless. The audio recording may be a pre-recorded audio recording stored in and transferred from the database 245, or the audio recording may be recorded by and then communicated from the mobile computing device 210. In some embodiments, the audio recording may be recorded and communicated (i.e., streamed) in real time from the mobile computing device 210 to the audio analysis server 250.
[0086] In some embodiments, data handling module 262 may receive two or more audio recordings at a time. In this example, data handling module 262 may be configured to temporarily store two or more audio recordings and transmit each recording to other modules in audio analysis server 250 as needed. For example, if a subject completes three speech tasks and / or augmented audio recordings as part of an assessment, data handling module 262 may receive all three recordings consecutively. Data handling module 262 may communicate any one of the three audio recordings to one or more modules in memory 260 and store the remaining two audio recordings. In some embodiments, data handling module 262 may communicate one of the two stored audio recordings after a predetermined period of time, or may transmit one of the two stored audio recordings after receiving an indication from one or more modules in memory 260 indicating that one or more modules are ready to receive another audio recording.
[0087] In optional step 412, the quality control module 264 performs a quality check on the data to determine whether the audio recording is suitable for use in brain health assessment. In some embodiments, optional step 412 may correspond to step 1A of FIG. 1. The quality control process can determine whether the audio recording is suitable for use in brain health assessment by performing one or more quality tests on the audio recording. The quality tests may include, but are not limited to, verifying whether the audio file is empty, detecting whether other speakers are recorded in addition to the primary speaker, verifying whether the generated stimuli match the target stimuli according to one or more selected tasks and associated minimum requirements, and / or verifying the signal-to-noise ratio of the recording. The quality checks may also include one or more, or a combination of two or more, of determining whether a voice is present, determining the duration of the recording, determining whether the recording duration is within a predetermined recording time limit, removing abnormal noise, removing silence, determining the number of speakers recorded, and separating speakers.
[0088] In some embodiments, in optional step 412, one or more quality tests may each have its own pass threshold indicating an acceptable level of quality for that test. Some quality tests may be binary pass / fail tests, such as whether an audio file is empty. For some quality tests, the pass threshold may be the percentage of the recording that contains a particular unwanted quality; for example, a voice recording may pass the quality test if its signal-to-noise ratio is less than 50%. Other quality tests may have a pass threshold related to the strength of a particular feature of the voice recording; for example, if a particular quality test evaluates the presence of more than one speaker, the voice recording may fail the particular quality test if the additional speakers' voices do not reach a minimum decibel rating.
[0089] In some embodiments, the quality control module 264 may optionally perform one or more quality improvement processes in step 412, including, but not limited to, filtering background noise such as clicks, static, and / or audio artifacts; equalizing speech intensity to aid in analysis such as syllable emphasis where speech intensity is relevant; and / or trimming long periods of silence from the audio file, such as large pauses in speech and silence at the beginning and end of a recording.
[0090] In some embodiments, if the quality control module 264 determines that an audio recording is not of sufficient quality, for example, by the audio recording not passing a certain number of quality checks (e.g., not meeting a pass threshold), the quality control module may be configured to mark the particular audio recording as inadequate. In some embodiments, if the audio recording is determined to be inadequate, the quality control module 264 may communicate a notification / indication to the mobile computing device 210 indicating that the particular audio recording is of unacceptable quality. In some embodiments, this notification / indication may include a prompt to re-record the particular audio recording by redoing the associated speaking task or to resume / restart extended recording of the audio.
[0091] At 415, the speech-to-text module 266 receives the voice recording from either the data processing module 262 or the quality control module 264. In some embodiments, step 415 may correspond to step 2 of FIG. 1. The speech-to-text module 315 is configured to analyze the voice recording and generate a text representation of the spoken words captured by the voice recording. For example, the data processing module may receive as input an audio file of the recorded voice and output a file containing text, such as a .docx, .PDF, and / or .txt file. The speech defined by the voice recording may be converted to text through the use of a machine learning model trained to accept the voice recording as input.
[0092] In some embodiments, the speech-to-text module 266 comprises a speech-to-text model configured to receive a representation of an audio recording, such as a sequence of filterbank spectral features or a sound wave spectrogram, and to output strings of text indicating the words spoken in the audio recording.
[0093] In some embodiments, the speech-to-text module 266 may first convert the received audio file into a sequence of filterbank spectral features. A filterbank is an array of bandpass filters that separates the input signal into multiple frequency components, each responsible for a single frequency subband of the original signal. The set of filterbank spectral features represents the frequencies present in the audio file.
[0094] The speech-to-text module 266 may include a speech-to-text model trained and configured to receive a sequence of filterbank spectral features and output a textural representation of the sounds produced by the subject.
[0095] The speech-to-text model may include one or more artificial neural networks (ANNs) to accomplish the task of converting spoken language into text. ANNs are computational structures inspired by biological neural networks and include one or more layers of artificial neurons configured or trained to process information. Each artificial neuron includes one or more inputs and an activation function for processing the received inputs and generating one or more outputs. The outputs of each neuron layer are connected to subsequent neuron layers using links. Each link may have a defined numerical weight that determines the strength of the link as information passes through several layers of the ANN. During the training phase, various weights and other parameters defining the ANN are optimized to obtain a trained ANN using inputs and known outputs for those inputs. This optimization may be performed through various optimization processes, including backpropagation. ANNs incorporating deep learning techniques include multiple hidden layers of neurons between an initial input layer and a final output layer. Multiple hidden layers of neurons enable ANNs to model complex information processing tasks, including the speech analysis and text generation tasks performed by system 200.
[0096] In some embodiments, the speech-to-text model may include an encoder recurrent neural network (RNN) configured to listen and a decoder RNN configured to spell. An RNN is a type of ANN in which connections between neurons form a directed or undirected graph that follows a temporal sequence. RNNs can exhibit temporal dynamic behavior; that is, they can process inputs using time as a contributing factor to the final decision on the input. This makes RNNs suitable for speech recognition. The encoder and decoder RNNs may be jointly trained simultaneously using the same training dataset.
[0097] The listener RNN accepts as input a set of spectral features from the filter bank, which are transformed through the layers of the RNN into higher-level features. The higher-level features may be shorter phonetic sequences, such as individual phonemes and / or various versions of individual phonemes depending on their position within a word and / or surrounding phonemes. A phoneme is the smallest unit of speech that distinguishes one word (or word element) from another. Once the listener RNN processes the input and determines the sequence of high-level features, it can provide the determined high-level features to the decoder RNN.
[0098] The decoder RNN transcribes spoken speech into written speech one character at a time. To determine the first or next word of a subject's speech, the decoder RNN can generate a probability distribution conditioned on all previously seen characters. The probability distribution for determining the next character is a function of the decoder RNN's current state and the current context. The decoder RNN's current state is a function of the previous state, previously emitted characters, and the context. The context is represented by a context vector generated by the attention mechanism.
[0099] At each time step in the speech-to-text decision process, the attention mechanism generates a context vector. This context vector encapsulates the information contained in the acoustic signal (i.e., the sounds made by the subject) necessary to determine the next character. Specifically, at each time step, the attention context function computes a scalar energy for each time step. The scalar energy is converted into a probability distribution of the next character over the time step (or attention) using a normalization function such as softmax.
[0100] The final output of the speech-to-text model can be determined using a left-to-right beam search algorithm. At each time step in the decision process, each sub-hypothesis in the beam is expanded with all possible characters, and only the most likely beam is retained. When a sentence-end token is encountered, it is removed from the beam and added to the set of complete hypotheses, generating a textural representation of the sounds produced by the subject in a step-by-step manner.
[0101] In some embodiments, the speech-to-text model may determine that it is unable to determine a textual representation for the audio recording. The speech-to-text model may determine that the audio recording does not contain any textually representable sounds. For example, the speech recording may not contain textually representable morphemes such as "er," "ly," "ish," or "ic." In some embodiments, the audio recording may contain bodily sounds that the speech-to-text model is unable to represent in text, such as borborygmus, coughing, breathing, gasping, and / or snorting. In some embodiments, the speech-to-text model may represent bodily sound occurrences by labeling them, e.g., to indicate that a subject coughed, it may generate text such as "[cough]," "*cough*," or "{cough}," and may indicate bodily sound occurrences using parentheses, asterisks, and / or any other symbols. In some embodiments, the speech-to-text model may generate a description indicating one or more sound(s) made by one or more recorded bodies.
[0102] In some embodiments, if the speech-to-text model determines that a text representation cannot and / or should not be determined, the step of generating a text representation may be skipped. In some embodiments, determining that a text representation should not be generated may include determining that the audio recording does not include at least one morpheme, phoneme, or sound for which a text representation can be created. After determining that a text representation cannot and / or should not be determined, the speech-to-text module 266 may generate an indication that a text representation was not created. In some embodiments, the indication that a text representation was not created may be sent or otherwise communicated to the text analysis module 270 for use in method 400.
[0103] In some embodiments, method 400 may not include step 415 of converting the recording of the subject's sound into a text representation and / or determining whether a text representation can be generated and / or needs to be generated and / or can be determined.
[0104] At 420, the acoustic analysis module 268 receives the audio recording from the data processing module 262 or the quality control module 264. In some embodiments, step 420 may correspond to step 3B of Figure 1. The acoustic analysis module may process the audio recording to identify a data set of acoustic analysis metrics.
[0105] The acoustic analysis module 268 can process the audio recordings using machine learning models. The machine learning models can be trained using datasets of fully labeled data, partially labeled data, or unlabeled data. These data can be audio recordings of speech. In some embodiments, the audio recordings can be of a subject completing one or more speech tasks and / or one or more augmented audio recordings of the subject. The ML model can be iteratively trained on a training dataset and tested using a validation dataset. In some embodiments, the training dataset and validation dataset can be subsets of a larger dataset or can be randomly sampled to generate multiple test and validation datasets.
[0106] In some embodiments, the acoustic analysis module 268 may include an acoustic analysis model configured to perform acoustic analysis and classification of inputs. In some embodiments, the acoustic analysis model may be a trained RNN and may be trained or otherwise configured to receive audio files or representations of audio files, such as sound spectrograms, waveform representations, filter bank spectral features, and / or numerical encodings. In some embodiments, each audio file may be converted into a time-series waveform representation, which may then be decomposed into a set of specific frequencies and / or frequency bands. In some embodiments, the time-series waveform representation and / or the set of specific frequencies and / or frequency bands may be an image file, for example, a .JPEG, .BMP, or .PDF. This set of frequencies and / or frequency bands may then be provided as multivariate inputs to one or more ML models.
[0107] The acoustic analysis model may process the input and return a classification or other determination regarding one or more qualities of the input audio file or audio file representation, which may include, for example, prosody and timing, articulation, resonance, and / or quality.
[0108] In some embodiments, to perform the acoustic analysis, the acoustic analysis module 268 may use Bayesian statistical models in conjunction with deep learning and recurrent neural networks (RNNs) to incorporate the time course nature of progression or change. For example, long short-term memory networks are a special class of RNNs that outperform many traditional machine learning and deep learning architectures in tracking time course data. In some embodiments, measures of variable importance can be obtained to provide analytical and interpretable models, allowing for inference of the effects of intervention / progression. Output from this analysis stage is interpreted in light of clinical data (e.g., disease severity, cognition, fatigue).
[0109] In some embodiments, acoustic analysis module 268 may include multiple ML models, each trained on a particular dataset, each dataset configured to represent a particular brain health attribute, neurological disorder, and / or brain health indicator. In some embodiments, one of the multiple models may be trained or otherwise configured to classify two or more brain health attributes, neurological disorders, and / or brain health indicators when the two or more brain health attributes, neurological disorders, and / or brain health indicators often co-occur and / or share similar acoustic signatures and / or attributes.
[0110] In some embodiments, acoustic analysis module 268 may be configured with multiple speech quality ML models, each trained to classify a particular speech quality. In some embodiments, acoustic analysis module 268 may include an ML model that can accept as input the decisions of multiple ML models configured to classify a particular speech quality. For example, the multiple ML models may return two or more clear indications of the prosody and timing, articulation, resonance, and / or quality of an input speech expression. The clear indication may be a numerical representation on a predetermined scale or a classification of a particular brain health condition or neurological disease. The predetermined scale may be determined empirically or based on previous test results associated with the current subject.
[0111] The clear instructions may be provided to a summary ML model, which may provide a summary classification of the speech input. The summary classification may be a numerical representation of the subject's brain health. For example, the summary classification may be a numerical value on a predetermined scale, determined experimentally or based on previous test results associated with the subject.
[0112] In some embodiments, training one or more acoustic analysis models may include a sliding window data selection approach to control for one or more brain health attributes, neurological disorders, and / or brain health indicators. The sliding window may evaluate each input during the training process to determine its suitability, appropriateness, and / or relevance to previous and / or subsequent inputs. For example, the sliding window may determine that an audio input is indicative of a particular brain health attribute, neurological disorder, and / or brain health indicator. The sliding window may then curate subsequent inputs that are indicative of the same or related brain health attributes, neurological disorders, and / or brain health indicators.
[0113] In some embodiments, the training process may include a semi-supervised training approach. A semi-supervised training approach may include using datasets of both labeled and unlabeled data. For example, a training dataset may include a small number of labeled data and a large number of unlabeled data, e.g., a relatively small number of abnormal brain health records labeled as indicative of brain health changes. A training dataset may also include a large number of unlabeled data, which may include both normal and abnormal brain health data but without associated tags / labels.
[0114] In some embodiments, the semi-supervised training approach may be a self-training approach, in which an initial ML model is trained on a small collection of labeled data to create a first classifier, i.e., a base model. The first classifier may then label one or more larger unlabeled datasets to create a collection of pseudo-labels for the unlabeled dataset. The labeled dataset is then combined with a selection of the most confident pseudo-labels from the pseudo-labeled dataset to create a new, fully labeled dataset. The most confident pseudo-labels may be selected manually or determined by the ML model. This new, fully labeled dataset is then used to train a second classifier. This second classifier, by virtue of having a larger labeled training dataset, may exhibit improved classification performance compared to the first model. The above process may be repeated any number of times, generally resulting in a better performing classifier.
[0115] In some embodiments, the semi-supervised training approach may be a joint training approach, in which two first classifiers are initially trained simultaneously on two different labeled datasets or "views," each containing different features of the same instances. For example, one dataset may contain frequencies below a certain threshold, while another dataset may contain frequencies above a certain threshold. In this approach, each set of features is sufficient for each classifier to reliably determine the class of each instance.
[0116] After initial training of the two first classifiers, a larger pool of unlabeled data may be separated into two different views and presented to the first classifier to receive pseudo-labels. The classifiers then co-train each other using the pseudo-labels with the highest confidence. If the first classifier confidently predicts the true label of a data sample and the other classifier makes a prediction error, the data with the pseudo-label confidently assigned by the first classifier updates the second classifier, and vice versa. Finally, the predictions from the updated classifiers are combined to obtain a single classification result. As with the self-learning approach, this process can be iteratively repeated to improve classification performance.
[0117] In some embodiments, training of the ML model may use a deep generative model to correct for imbalances between normal and abnormal brain health data. Generative models treat semi-supervised learning problems as specialized missing data imputation tasks for classification problems, effectively treating data imbalances as a classification problem rather than an input problem. Generative models utilize probability distributions that can determine the probability of observable features given a desired outcome. Generative models have the ability to generate new data instances based on previous data instances, helping train models with better performance on label-limited datasets.
[0118] The use of semi-supervised learning models may be used to account for the lack of certain types of data associated with particular brain health conditions. For example, speech data from subjects with brain tumors may be scarce, and the semi-supervised techniques described above may help create a more reliable model. In another example, two or more brain health conditions may manifest very similarly in terms of speech, making it correspondingly difficult for a model to reliably distinguish between them. Semi-supervised learning is used to improve the accuracy of one or more models and reliably classify one condition from another.
[0119] In some embodiments, the acoustic analysis model may include an information sieving process. The information sieving process may be configured to perform a hierarchical decomposition of the information provided to it. The information sieving process may include multiple, increasingly finer sieves. Each level of sieving may be configured to retrieve a single latent feature that is maximally informative about multivariate dependencies in the data. In some embodiments, a representation of a subject's voice recording may be provided to the sieving process. The sieving process may extract specific qualities associated with the representation that are determined to be indicative of specific brain health attributes, neurological disorders, and / or brain health indicators. In some embodiments, each layer of the information sieving process may be configured to extract specific indicators and / or features of the voice representation that are associated with specific speech characteristics. The sieving process may return each latent feature in the latent feature dataset. In some embodiments, the information sieving process may be configured to evaluate the latent features and return a determination regarding the specific speech characteristic being evaluated.
[0120] Some embodiments may configure two or more information sieves, each configured to assess a particular brain health and / or speech characteristic. Each determination received from each information sieve may be used as an input for determining the subject's brain health. In some embodiments, the acoustic analysis model may use each determination from each information sieve to determine a final speech determination.
[0121] In some embodiments, the training data used to train one or more ML or artificial intelligence models may include one or more data curation steps and / or one or more data preparation steps. One or more data curation steps and / or one or more data preparation steps may include removing low-quality or inaccurate data. It may also include correcting inaccurate or low-quality data. One or more of these processes may also include relabeling incorrectly labeled data. One or more data curation steps and / or one or more data preparation steps may further include labeling unlabeled data. It may also include removing unnecessary portions of the data while leaving the remaining useful data intact.
[0122] At 425, the acoustic analysis model outputs a determination of the quality and / or characteristics of the recorded sound. In some embodiments, step 425 may correspond to step 3B of FIG. 1. In some embodiments, the determination may be a speech quality dataset including one or more speech quality metrics. The speech quality metrics may be one or more of prosody and timing, articulation, resonance, and / or quality.
[0123] Prosodic features include measurements of speech timing, stress, rhythm, intonation, and stress changes at the individual sound and repetition level. These may include measurements and variations in amplitude, frequency, and energy dynamics. Prosodic features may be determined at word / intraword intervals, speech formation / breathing, across speaker interactions, stuttering episodes, blocks / prolongations, syllable rate, articulatory rate, or over continuous recording periods. Articulation can be determined using the frequency distribution of the speech signal. These may include power spectral density (PSD) and Mel-frequency cepstral features (MFCCs), formant slope, voicing onset time, and / or vowel articulatory scores. Resonance may be determined using power and energy distribution features such as MFCCs, octave ratios, frequency / intensity thresholds, and / or formants. Quality may be determined using source features designed to measure VF oscillation pattern, periodicity, and / or aerodynamics.
[0124] At 430, the text representation of the sounds made by the subject is provided to the text analysis module 270. In some embodiments, step 430 may correspond to step 3A of FIG. 1. In some embodiments, the text analysis module 270 may configure a text analysis model. The text analysis model may be a machine learning model configured to perform natural language processing (NLP). In some embodiments, an ANN may be trained using a dataset of text. The text analysis model 170 may utilize word embeddings to represent relationships between words or sentences in n-dimensional space. Word embeddings may be in the form of real-valued vectors that encode the meaning of words, such that words that are closer in vector space are expected to be similar in meaning, use, content, and / or interpretation. The text analysis module may be configured to convert the text representation into one or more word embeddings, which may then be input to the text analysis model, which may output one or more quality determinations of the text.
[0125] In some embodiments, the text analysis module 270 may receive an indication from the speech-to-text module 266 that a text representation was not created.
[0126] At 435, the text analytics model returns an utterance content dataset including one or more utterance content indicators. In some embodiments, step 435 may correspond to step 3A of Figure 1. The utterance content indicators may indicate one or more of lexico-semantic performance and verbal fluency, grammatical and morphosyntactic performance, and / or discourse-level performance (writing production and comprehension).
[0127] Lexical-semantic performance can be assessed by examining word recall, verbal knowledge, nonverbal semantic knowledge and comprehension, and verbal fluency. Grammatical and morphosyntactic performance can be assessed by measuring sentence comprehension and grammar in task tasks and continuous speech tasks / long recordings. Verbal fluency can be assessed by correctly spoken words / sentences, incorrectly spoken words / sentences, and abruptions and repetitions. Lexical density can be assessed by lexical range and / or the ratio of types to tokens. Information content can be assessed by speech efficiency, semantic versus conceptual content, speech schema, and / or cohesion. Grammatical complexity can be assessed by morphological complexity, word types, semantic complexity, and subject-verb-object (SVO) order.
[0128] In some embodiments, the speech content indices may be numeric values indicating a level of ability, performance, execution, or other measurement on a scale associated with each indice. For example, one of the indices may indicate verbal fluency on a scale from 0 to 100, with 0 associated with no verbal fluency and 100 associated with excellent verbal fluency. For example, if a subject's spoken words are presented in text, the verbal fluency scale may be scored as 50 / 100. This may indicate that the subject has average verbal fluency compared to similar subjects.
[0129] In some embodiments, the scale can be empirically determined from collecting data from multiple subjects. The scale can also be an individualized scale indicating a subject's previous performance. In some embodiments, the index can indicate whether the subject performed better or worse than one or more previous records.
[0130] In some embodiments, text analysis module 270 may output a single indicator of the subject's speech content, determined via a textual representation of every word spoken by the subject during the recording. This single indicator may indicate, for example, one or more of the subject's speech content in general compared to other similar subjects, whether the subject's speech content is particularly indicative of one or more particular brain health conditions, neurological diseases, and / or disorders, and / or an indication of improvement or decline in the subject's speech content.
[0131] In some embodiments, after receiving an indication that a text representation was not created, the text analysis module 270 may determine and / or generate a speech content dataset that includes one or more speech content indicators that indicate the fact that the subject did not utter any sounds that are deemed capable of and / or need to be converted from speech to a text representation.
[0132] In some embodiments, step 415 is not performed (i.e., not generating a text representation and / or not determining whether a text representation can be generated and / or will be generated and / or needs to be determined), followed by step 430 of providing the text representation to text analysis module 270 and step 435 of determining a text content dataset. In some embodiments, a text representation may still be determined by utterance to text module 266, but the text representation may not be provided to text analysis module 270, and text analysis module 270 may not determine a content dataset for the utterance. The text analysis module may receive a text representation or an indication that a text representation could not be determined, but in some embodiments, may not determine a context dataset for the utterance.
[0133] 1 , steps 415, 420, 425, 430, and / or 435 may occur sequentially, in parallel, synchronously (i.e., performed or executed one after the other or in parallel as part of a single series of steps and / or actions), and / or synchronously (i.e., performed separately at substantially different times, which may result in data that may be stored for later and / or subsequent use). In some embodiments, the steps of converting any spoken words and / or utterances to text 415, providing the text to a text analysis model 430, and receiving from the text analysis model a speech content dataset including one or more speech content indicators 435 may occur sequentially but in parallel with the steps of providing the audio recording to an acoustic analysis model 420 and receiving from the acoustic analysis model a speech quality dataset including one or more speech quality indicators 425.
[0134] In some embodiments, steps 415, 420, 425, 430 and / or 435 may be performed sequentially in the order depicted in FIG. 4, or in any other order that would result in performing a method of determining brain health according to the present disclosure.
[0135] In some embodiments, if the speech-to-text model has not determined a textual representation, the speech-to-text model may generate a speech content dataset indicating that no recorded sounds are available for textual representation, hi some embodiments, the speech content dataset may include indicating that bodily sounds have been recorded but are not available for textual representation.
[0136] At 440, the brain health determination module 272 may receive the speech quality dataset and the speech content dataset and determine a speech score. In some embodiments, step 440 may correspond to step 4 of FIG. 1. The speech score may be indicative of the subject's brain health. The speech score may indicate the presence of a neurological disease in the subject, or the subject's brain health if no neurological disease was previously known (i.e., recent onset of neurological disease / deteriorating brain health). In some embodiments, the speech score may indicate progression of a known neurological disease and / or a change in the subject's brain health. The progression of a known neurological disease and / or a change in brain health may be determined by comparing the newly determined speech score to one or more previously determined speech scores associated with the subject.
[0137] The speech score may include one or more composite scores indicative of various speech characteristics, which may be one or more of a communication effectiveness score, a dysarthria score, a disease severity score, a social communication score, a voice quality score, an intelligibility score, and / or a naturalness score.
[0138] In some embodiments, at 445, the brain health determination module 272 may include one or more brain health models. In some embodiments, the brain health model may be a type or variation of a NN trained or otherwise configured to receive as input scores, metrics, or other determinations of speech content and / or speech quality generated, classified, calculated, or otherwise determined by one or more of the acoustic analysis module 268 and / or text analysis module 270, and to output a brain health determination. The brain health determination may be a set of metrics indicative of brain health and / or the progression / presence of one or more neurological diseases. In some embodiments, the determination may be a classification regarding the presence and / or progression of a known or suspected brain health condition and / or neurological disorder. The brain health NN may be composed of and / or otherwise include a combination of differently weighted features indicative of brain health and / or aspects of brain health obtained from / during and / or configured during a training process.
[0139] In some embodiments, the brain health model may derive one or more indicators and / or determinations indicative of and / or related to brain health via a combination of attributes and / or features that map to states, symptoms, disorders, and / or signs of brain health and / or neurological disorders. In some embodiments, the signs of brain health may include listener ratings, speaker ratings, patient-reported outcomes, and / or performance on another test.
[0140] In some embodiments, the brain health model may return a set of values, indices, or decisions indicative of one or more multivariate subject brain health attributes, each of which is associated with one or more aspects of speech content and / or speech quality as defined by one or more outputs of the speech quality model and / or speech content model, as described above.
[0141] In some embodiments, each of the one or more brain health indicators may be a numerical value on a continuous scale. The continuous scale may be empirically determined from data collected from multiple subjects with or without brain health or changes in brain health. In some embodiments, the scale may be determined based on data associated with a particular subject, such as the subject's medical history, pre-testing of the subject before initiating and / or continuing use of system 100, and / or previous determinations by system 100.
[0142] In some embodiments, the brain health NN may be trained using a ground truth or other training dataset that includes experimentally derived or determined data, which may be determined via disease severity scores determined through assessment by trained experts, paper-and-pencil language tests designed by trained experts and administered to subjects with or without known brain health status, clinical assessments or impressions collected as part of standard clinical practice and / or academic research, and / or patient-reported outcomes recorded, for example, via patient visits, surveys, clinical studies, and / or routine physician visits.
[0143] In some embodiments, the brain health NN may utilize, include, incorporate, and / or otherwise utilize any one or more of the training and / or data labeling, generation, and / or curation techniques / methods as discussed and / or referenced herein.
[0144] In some embodiments, aspects of the recorded sounds produced by the subject and / or their textual representations that are particularly predictive and / or exhibit known or unknown correlations with indicators, symptoms, and / or brain health may be considered and incorporated into the disclosed methods. In some embodiments, a trade-off between efficiency and accuracy may be made to assess the overall usefulness of one or more particular indicators and / or types of data, or aspects of types of data, to ensure adequate processing / performance time.
[0145] In some embodiments, the speech score and / or brain health status determination made by brain health status determination module 272 may be used by status determination module 274 to determine one or more brain health statuses and / or neurological conditions of the subject, and / or the progression of one or more brain health statuses and / or neurological conditions. This may be done by comparing the newly determined speech score to one or more previous speech scores associated with the subject. In some embodiments, the condition may be determined by comparing one or more speech scores, including the newly determined speech score.
[0146] In some embodiments, the audio recording may be generated based on task prompts for the speaker to perform the speech task. The speaker may be presented with a speech task that is performed and recorded by an audio recording interface, such as device I / O 235. The mobile computing device 210 may automatically present task prompts instructing the user or patient to speak aloud a specific phrase, sound, or sequence of sounds. Each task may have minimum requirements for an attempt to be considered successful. In some embodiments, the assessment application 225 may perform analysis on the audio recording to determine whether the minimum requirements have been met. In some embodiments, the mobile computing device may communicate the audio recording to the audio analysis server 250 without performing any analysis.
[0147] In some embodiments, during setup and / or calibration of the system and / or at any point in method 400 prior to determining a speech score (at 440), or prior to determining progression of one or more brain health conditions and / or neurological disorders (at 445), one or more modules of speech analysis server 250 may be configured to receive one or more refinement parameters. The refinement parameters may include, for example, one or more brain health conditions, one or more brain health attributes, one or more brain health indicators, one or more neurological disorders, and / or the subject's medical history. Upon receiving the one or more refinement parameters, one or more models of system 100 may be configured to adjust their decision processes to address, incorporate, and / or otherwise take into account the one or more refinement parameters.
[0148] For example, the text analysis module 270 may receive a refinement parameter that indicates that the subject whose sounds are being recorded stutters. Accordingly, upon receiving this refinement parameter, the text analysis module 270 may be configured to remove repeated utterances from the text representation during and / or after conversion of the voice recording to text.
[0149] In another example, brain health determination module 272 may receive refinement parameters indicating that the patient has one or more brain health diagnoses and / or diseases. In some embodiments, brain health determination module 272 may be configured to tailor its determination to return a determination that is indicative of the severity and / or progression of a brain health diagnosis and / or disease, rather than indicating the presence of a particular brain health diagnosis and / or disease.
[0150] In yet another exemplary embodiment, the acoustic analysis module 268 may receive the refinement parameter that the subject is neurodiverse. The analysis module 268 may then be configured to adjust its determination to account for incorrect sounds and / or mispronounced words, thereby returning a speech quality dataset and associated speech quality metrics indicative of the particular subject's normal speaking habits.
[0151] It should be noted that despite the fact that the tasks are referred to as speech tasks, the subject may not, for example, form intelligible words, phrases, syllables, and / or sentences while performing the task(s). However, this does not mean that the subject is not performing or otherwise participating in the speech task. In other words, a speech task simply prompts the subject to make possible sounds, attempt to speak, and / or otherwise produce some sound using the vocal cords. In some embodiments, the speech task may explicitly require the production of intelligible, otherwise intelligible speech, or at least portions of intelligible, otherwise intelligible speech, such as vowel sounds (e.g., / ah / , / i / , / o / , / u / , / ae / , / e / ), other morphemes and / or phonemes.
[0152] The speech task may involve reading set text provided to the subject via a mobile computing device, such as a smartphone, laptop computer, desktop computer, and / or kiosk. In some embodiments, the speech task may be provided by one or more people assisting the subject, or may be displayed on a separate mobile computing device or printed sheet. The minimum requirement for reading set text may include the subject making a three-second audio recording containing the speech.
[0153] Speech tasks may also include unscripted, simultaneously generated speech on random or designated topics. In some embodiments where a topic is designated, the topic may be provided by a mobile computing device recording the sounds made by the subject, such as via a screen / monitor and / or recorded audio cues. In some embodiments, one or more people assisting the subject may provide the designated topic via another mobile computing device, verbal instructions, and / or a printed sheet. Minimum requirements may include a three-second audio recording containing the speech.
[0154] Speech tasks may also include engaging in conversation with other speakers on random or designated topics. In some embodiments where a topic is designated, the topic may be provided by a mobile computing device that is recording sounds made by the subject, such as via a screen / monitor and / or recorded audio cues. In some embodiments, one or more people assisting the subject may provide the designated topic via another mobile computing device, verbal instructions, and / or a printed sheet. Minimum requirements may include one interaction (i.e., one utterance by each subject).
[0155] The speech task may also include the production of a series of single vowels (e.g., / a / , / i / , / o / , / u / , / ae / , / e / ).Minimum requirements may include a 0.01-second audio recording containing the sounds produced by the subject.
[0156] In some embodiments, the subject may be recorded using a continuous recording protocol, which may be task-independent and may also record biological sounds (e.g., breathing, coughing, vomiting, etc.), or speech and / or one or more speech tasks performed by the subject.
[0157] The speech task may also include pronouncing a single vowel (e.g., / a / , / i / , / o / , / u / , / ae / , / e / ) continuously in one breath for as long as possible. The subject may select the vowel themselves or may be prompted to pronounce a particular vowel, for example, by a mobile computing device recording the speech, a separate mobile recording device, or one or more people assisting the subject. The minimum requirement may include 0.01 seconds of audio recording containing the speech.
[0158] Speech tasks may also include repeatedly producing alternating or consecutive syllable strings (e.g., pataka, pa-ta-pata, pa-pa-pa, ta-ta-ta-ta, kaka-kaka). In some embodiments, the subject may select the alternating or consecutive syllable string themselves, or the syllable string may be provided to the subject by a mobile computing device that is recording the sounds produced by the subject, such as via a screen / monitor and / or recorded audio cues. In some embodiments, one or more people assisting the subject may provide a particular syllable string or selection of syllable strings via another mobile computing device, verbal instructions, and / or a printed sheet. Minimum requirements may include one second of recorded audio containing the utterance.
[0159] Speech tasks may include saying the days of the week, saying them in a particular order, saying them in no particular order, starting with a particular day of the week, or starting with a random day of the week.Minimum requirements may include a one-second audio recording containing sounds made by the subject.
[0160] The speech task may also include counting to a predetermined number or to the highest possible number. In some embodiments, the subject may choose which number to start counting from, which number to count to, and / or a specific interval between the start and end numbers. In some embodiments, the subject may receive instructions regarding the predetermined start number, end number, and / or interval. A minimum requirement may include saying at least one digit.
[0161] The speech task may include saying words that follow a consonant-vowel consonant structure (e.g., hard, heat, hurt, hoot, hit, hot, hat, head, hub, hoard). In some embodiments, the subject may freely choose and pronounce specific words that fit the consonant-vowel consonant structure, or the subject may be provided with a list and / or order of words to recite. The minimum requirements may include an example of a consonant-vowel consonant structure.
[0162] Speech tasks may also include repeating written or spoken phrases or sentences of varying lengths. In some embodiments, the written phrases or sentences may be presented via the screen / monitor of a mobile computing device or, for example, on written paper. A minimum requirement may include the subject pronouncing at least one word.
[0163] The speech task may also include repeating multi-syllable words (e.g., computer, computer, computer). In some embodiments, the subject may, for example, select a multi-syllabic word, be instructed to say a particular multi-syllabic word, or select a multi-syllabic word from a predetermined list. Minimum requirements may include pronunciation of at least two word repetitions.
[0164] The speech task may include producing words with vowel changes resulting in two or more vowels (e.g., bay, pay, dye, tie, goat, coat). In some embodiments, the subject may, for example, choose a word themselves, choose from a list of words, or be given a series of words to pronounce. A minimum requirement may include pronouncing at least one example.
[0165] The speech task may also include repeating or reading related words of increasing length or complexity (e.g., profit, profitable, profitability; thick, thicker, thickening). In some embodiments, the subject may, for example, choose words themselves, be presented with specific words, or be given a list of words to choose from. A minimum requirement may include one attempt to pronounce each word in the string.
[0166] The speech task may also include the subject verbally describing or writing down what they see in the image or what they see around them. In some embodiments, the image may be a single image or a series of images. In some embodiments, the subject may use one or more input devices, such as a keyboard, touchpad, mouse, dial, toggle, digital keyboard, on-screen keyboard, and / or buttons, to enter a description configured to interact with a mobile computing device, such as a mobile computing device making a recording of sounds made by the subject, a mobile computing device determining the subject's brain health, or another mobile computing device. In some embodiments, the subject may handwrite the description, which may be manually entered, scanned, or otherwise subsequently transferred to a digital form. Minimum requirements may include at least three seconds of audio recording containing sounds made by the subject, at least one identifiable word pronounced by the subject, or at least one word written by hand and / or entered into a mobile computing device.
[0167] Speech tasks may also include listening to or reading a story and then having the subject retell what they heard or read. The subject may listen to the story via a pre-recorded audio recording played by a mobile computing device, such as at least one of the devices making the brain health assessment. In some embodiments, the subject may be read aloud by one or more assisting person(s). In some embodiments, the subject may read the story from a computing device that includes a screen / monitor. Minimum requirements may include at least three seconds of audio recording containing speech or one word spoken by the subject.
[0168] The speech task may also include generating a list of words in a particular category (e.g., words beginning with F, A, or S, or words that fit into semantic categories such as food, or animals, or furniture). In some embodiments, the subject may be presented with a particular category, or the subject may choose their own category or choose from a list of categories. A minimum requirement may include pronouncing one word.
[0169] The speech task may include repeating fictitious words (e.g., yeecked or throofed). In some embodiments, subjects may make up fictitious words themselves or may be presented with a list of fictitious words to recite. A minimum requirement may include one attempt by the subject per word.
[0170] The speech task may include listening to and repeating a few words or sentences aloud (e.g., a single word or sentence spoken in noise, such as cocktail or white noise). In some embodiments, the mobile computing device may play a recording containing, for example, a single word / sentence spoken in noise. A minimum requirement may include one attempt to pronounce each word in the string.
[0171] The speech task may include describing what the viewer sees while looking at a picture of a single item (e.g., the subject looks at a picture of a dog and says "dog"). In some embodiments, the subject may be presented with one or more photographs via a mobile computing device including a screen / monitor, or the subject may be presented with a printed photograph featuring a single item from one or more person(s) assisting the subject. Minimum requirements may include only one attempt to identify and pronounce the single item.
[0172] In some embodiments, the minimum requirements may be configured to ensure that sufficient data exists to enable system 200 to perform any of the methods for determining brain health according to the present disclosure. The minimum requirements may not be indicative of what a subject must do to fulfill or otherwise be considered to pass a speech task. In some embodiments, a subject may be able to provide the minimum required responses, inputs, and / or recordings and still be considered to have failed to satisfactorily complete the speech task. In other words, the minimum requirements for a subject's recorded voice may not be indicative of the subject's particular performance when undertaking one or more particular speech tasks. In some embodiments, if a subject fails to fulfill the specific minimum requirements for one or more speech tasks, this may indicate that the subject has failed one or more speech tasks.
[0173] FIG. 5A is a screenshot 500 of a quality assurance check that may be performed, according to some embodiments. The quality assurance check may be a microphone check configured to determine whether the microphone being used to record sounds made by the subject is functioning properly. The quality assurance check screenshot 500 may include instructions 510 instructing the subject and / or a user assisting the subject on the steps that must be taken to complete the quality assurance check. The screenshot 500 may also include a quality assurance progress indicator 520 configured to indicate how far the quality assurance check has progressed and / or whether the quality assurance check was successful. In some embodiments, the quality assurance check may be one or more of a microphone check to determine whether the microphone is responsive to input, a level check to determine whether the input is within an acceptable decibel range, a network connectivity check that is an available memory and / or storage check, and / or any other type of quality assurance check configured to ensure that recording and / or processing proceeds as required.
[0174] FIG. 5B is a screenshot of a speech task screen 530 according to some embodiments. The speech task screen 530 is an example of a speech task screen corresponding to a speech task that has not yet been started and / or completed. In some embodiments, the speech task screen 530 may be presented to the subject and / or a user assisting the subject upon launching, starting, logging in, or otherwise accessing the assessment application 225. The speech task screen 530 may prompt the subject and / or a user assisting the subject to complete a particular task. The speech task screen 530 may include a task prompt 540 that instructs the subject to speak a particular phrase, make a particular sound, engage in a cognitive speech task (e.g., say as many words beginning with "H" as possible), and / or perform any other type of speech task. In some embodiments, the task prompt 540 may instruct the subject and / or a user to begin an extended recording. Speech task screen 530 may also include a task count 535 that indicates the subject's and / or user's progress through the multiple tasks involved in assessing brain health. In some embodiments, speech task screen 530 may also include a button 545 that, when activated by the subject and / or user, causes the client device to record sounds by the subject and / or user. In some embodiments, button 545 may be animated, and the animation may indicate the duration of the recording, how much time is left to record a particular speech task and / or recording session, or may indicate that recording has started.
[0175] 5C is a screenshot of a speech task screen 550, according to some embodiments. The speech task screen 550 is an example of a speech and / or recording task currently being recorded, as indicated by button 555. In some embodiments, button 555 may be interacted with when the speech and / or recording task is completed and / or when the subject and / or user wishes to stop and / or pause recording. Button 555 may also be animated, in some embodiments, where the animation may indicate the duration of the recording, how much time is left to record a particular speech task and / or recording session, or may indicate that recording has begun. The speech task screen 550 may also include a recording timer 560 configured to indicate the progress and / or total recording time of the speech or recording task.
[0176] 6 is a process flow diagram of a method 600 of determining brain health performed by the system 200, according to some embodiments. Optionally, at 610, the subject and / or a user assisting the subject may initiate, launch, and / or otherwise access the assessment application 225 via the mobile computing device 210. The assessment application 225 may be downloaded and installed onto the mobile computing device 210 via the network 240. The assessment application 225 may be a web application executed via a web browser (not shown) installed on the client device 210. Upon initiating the assessment application 225, the subject and / or user is presented with a login screen where they can log in with an existing account or create a new account. In some embodiments, each subject may have their own account, or medical institutions and / or practitioners may have their own accounts, and each subject / patient may have a user profile associated with the account.
[0177] In some embodiments, at 615, system 100 may receive an indication of one or more brain health attributes to be assessed by system 100. In some embodiments, the subject and / or user may be presented with a data entry field and / or list of brain health attributes / conditions to be assessed. The selection of a particular brain health attribute may determine which speech and / or recording tasks are presented to the subject and / or user and / or which speech and / or recording tasks can be presented to the subject and / or user. In some embodiments, the selection of a particular brain health attribute may determine one or more machine learning model(s) used to process the recorded speech tasks and / or sounds. In some embodiments, the selection of a particular brain health attribute may be used to configure one or more data filtering processes, such as a sliding data selection filter. The data filtering processes may select or filter aspects of the recorded sounds and / or textual representations of the recorded sounds to more accurately assess the particular brain health attribute.
[0178] At 620, the subject and / or user is presented with a speaking task and / or a recording task. The speaking task may include speaking specific words, making specific sounds, engaging in a conversation, and / or saying a word(s) according to specific criteria (e.g., as many rhyming words as the subject can think of). The recording task may include recording the sounds made by the subject over an extended period of time.
[0179] At 625, the speaking or recording task is recorded by system 100. In some embodiments, the task is recorded by client device 210. The client device may include device I / O 235 configured to record audio, such as a built-in microphone. In some embodiments, client device 210 may be a "kiosk" that includes a computing device connected to one or more microphones, connected via audio connections such as USB, XLR, and / or TRS. At 630, once the speaking or recording task is completed, system 100 may present the subject and / or user with the next or subsequent speaking or recording task to complete.
[0180] At 635, the one or more recordings are processed according to method 400 to generate a brain health determination. In some embodiments, the recordings are communicated via network 240 to audio analysis server 250 for processing. In some embodiments, client device 210 may be configured to process the recordings to determine the brain health. In some embodiments, the audio recordings and / or the brain health determination may be communicated to and stored on database 245.
[0181] At 640, a determination of the subject's brain health is returned. The determination may be a diagnosis of a particular brain health condition and / or neurological disorder. The diagnosis may be a binary determination of whether the subject is likely to have a particular condition. In some embodiments, the determination may be a numerical value. The numerical value may be a measure of the severity of the condition or the likelihood that the subject has a particular condition. In some embodiments, the determination may include multiple numerical values, each associated with a particular aspect of brain health. The numerical values may represent values on a scale, where the scale is associated with the progression, loss, and / or progression of brain health and / or neurological disorders.
[0182] Those skilled in the art will appreciate that numerous variations and / or modifications may be made to the above-described embodiments without departing from the broad general scope of the present disclosure, and the present embodiments are, therefore, to be considered in all respects as illustrative and not restrictive.
[0183] Those skilled in the art will understand that any particular method and / or technique described in connection with or used in combination with any one or more individual element(s) in the above-described embodiments may be equally applied to any other applicable elements described above without departing from the broad general scope of the present disclosure.
Claims
1. 1. A computer-implemented method comprising: receiving an audio recording of the sounds made by the subject; determining a text representation of at least one sound uttered by the subject from the voice recording; providing the audio recording to an acoustic analysis model; receiving a speech quality dataset from the acoustic analysis model, the speech quality dataset including one or more speech quality indicator(s); providing the text representation to a text analysis model; receiving an utterance content dataset from the text analysis model, the utterance content dataset including one or more utterance content indicator(s); determining a speech score associated with the subject from the speech quality dataset and the speech content dataset; Including, The speech score is an index indicating the brain health state of the subject. method.
2. The computer-implemented method of claim 1 , wherein the speech scores include at least a naturalness score.
3. 3. The computer-implemented method of claim 2, wherein the naturalness score includes at least one measure of fundamental frequency variability, intensity variability, formant transitions, speaking rate and rhythm, phonation measures, spectral measures, temporal measures, articulation, coarticulation, resonance, and prosody.
4. The computer-implemented method of any one of claims 1 to 3, wherein the speech score comprises at least an intelligibility score.
5. 5. The computer-implemented method of claim 1, wherein the speech scores include one or more of a communication effectiveness score, a dysarthria score, a disease severity score, a social communication score, and / or a voice quality score.
6. 6. The computer-implemented method of claim 1, wherein the speech quality dataset comprises one or more of a timing index, an articulation index, a resonance index, a prosody index, and / or a voice quality index.
7. A computer-implemented method according to any one of claims 1 to 6, wherein the speech content dataset comprises at least a discourse complexity measure.
8. 8. The computer-implemented method of claim 7, wherein the discourse complexity indicators include at least one of lexical diversity, syntactic complexity, referential coherence, thematic development, argument structure, implicit and explicit information comparison, interactivity, intertextuality, modality and modulation, and pragmatic factors.
9. 9. The computer-implemented method of any one of claims 1 to 8, wherein the speech content dataset comprises one or more of a semantic complexity index, an idea density index, a verbal fluency index, a lexical diversity index, an information content index, a dialogue structure index, and / or a grammatical complexity index.
10. 10. The computer-implemented method of claim 1, further comprising comparing the speech score associated with the subject to one or more previous speech scores associated with the subject to identify changes in the brain health of the subject.
11. 10. The computer-implemented method of claim 1, further comprising comparing the speech score associated with the subject with a control speech score to identify the brain health of the subject.
12. 12. The computer-implemented method of claim 11, wherein the control speech score is an experimentally determined score or a statistically determined score.
13. the sounds produced by the subject are spoken words and / or sounds; the spoken words and / or sounds are in response to a speech task provided to the subject, and / or the spoken words and / or sounds are recorded during a period of continuous observation of the subject without prompting; 10. A computer-implemented method according to any one of the preceding claims.
14. 10. A computer-implemented method according to any one of the preceding claims, further comprising performing quality assurance of the data.
15. The quality assurance of the data is (a) determining whether a voice is present; (b) determining the recording duration; (c) determining that the recording time is within a predetermined recording limit; (d) removing abnormal noise; (e) removing silence; (f) determining the number of speakers recorded; (g) separating speakers; 15. The computer-implemented method of claim 14, comprising one or more of:
16. 16. The computer-implemented method of claim 14 or claim 15, sending a notification to a mobile computing device if the audio recording fails the data quality assurance step.
17. 10. A computer-implemented method according to any one of the preceding claims, wherein determining the utterance content score and the utterance quality score is performed in parallel.
18. further comprising receiving one or more refinement parameters.
10. A computer-implemented method according to any one of the preceding claims.
19. The one or more refinement parameters are: one or more brain health conditions; one or more brain health attributes; one or more brain health indicators; one or more neurological disorders; The subject's medical history; 20. The computer-implemented method of claim 18, comprising:
20. 20. The computer-implemented method of claim 18 or 19, wherein in response to receiving one or more refinement parameters, an acoustic analysis model and / or a text analysis model are adjusted such that the speech quality dataset and / or the speech content dataset are associated with the one or more refinement parameters.
21. 21. The computer-implemented method of any one of claims 1 to 20, wherein one or more of the text representation, the speech quality dataset, the speech content dataset, and the speech score are determined by one or more machine learning model(s).
22. 22. The computer-implemented method of claim 21, wherein at least one of the one or more machine learning model(s) is a neural network.
23. 23. The computer-implemented method of any one of claims 1 to 22, further comprising converting the text representation into word embeddings.
24. and converting the audio recording into one or more of a series of filter bank spectra and / or one or more sound spectrograms prior to providing the audio recording to the acoustic analysis model. A computer-implemented method according to any one of claims 1 to 23.
25. 25. The computer-implemented method of any one of claims 1 to 24, wherein the textual representation is determined using natural language processing.
26. The method further comprises: determining that the textual representation cannot be determined; omitting the step of determining a text representation if it is determined that the text representation cannot be determined; 26. A computer-implemented method according to any preceding claim, comprising:
27. Determining that the text representation cannot be determined is determining that the voice recording does not contain at least a single morpheme, phoneme, or sound that can be represented in text; 27. The computer-implemented method of claim 26, comprising:
28. after determining that the voice recording does not contain at least a single morpheme, phoneme, or sound that can be represented in text; determining an indication that the voice recording does not contain at least a single morpheme, phoneme, or sound that can be expressed in text; the instructions are used as input for determining the speech score.
28. A computer-implemented method according to claim 26 or claim 27.
29. A non-transitory machine-readable medium storing instructions that, when executed by one or more processors, cause a computing device to perform the method of any one of claims 1 to 28.
30. 1. A system for assessing brain health, comprising: a speech analysis module configured to determine a speech quality dataset from the audio recording; a speech-to-text module configured to determine a text representation of the audio recording; a speech content module configured to determine a speech content dataset from the audio recording; a brain health determination module configured to determine a speech score using the speech quality dataset and the speech content dataset; Equipped with the speech score being indicative of the brain health of one or more subjects captured in the audio recording; system.
31. 31. The system of claim 30, further comprising a screen configured to present instructions to the one or more subjects.
32. 32. The system of claim 30 or claim 31, further comprising an audio recorder configured to record audio of the one or more subjects.
33. The system of any one of claims 30 to 32, wherein the speech analysis module comprises one or more machine learning models and / or information sifting.
34. 34. The system of claim 30, wherein the speech-to-text module comprises a machine learning model.
35. The system of any one of claims 30 to 34, wherein the speech content module comprises a machine learning model.
36. 36. The system of any one of claims 30 to 35, wherein the machine learning model is a recurrent neural network.
37. 1. A computer-implemented method comprising: receiving an audio recording of the sounds made by the subject; providing the audio recording to an acoustic analysis model; receiving a speech quality dataset from the acoustic analysis model, the speech quality dataset including one or more speech quality indicator(s); determining a speech score associated with the subject from the speech quality dataset; and Including, the speech score being indicative of the brain health of the subject; method.