Techniques for speech language model training and application
By employing speech language models to analyze speech samples, the system effectively addresses the inefficiencies and inaccuracies of traditional methods for assessing nerve damage and diseases, particularly for mild cognitive impairment.
Patent Information
- Application Number
- JP2024187994
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-26
- Filing Date
- 2024-10-25
- Publication Date
- 2025-05-13
AI Technical Summary
Current methods for assessing nerve damage and diseases are often manual, inaccurate, and inconsistent, relying on handwritten forms and lacking efficiency in diagnosis.
The development of a system that uses speech language models to analyze speech samples and determine characteristics indicative of mild cognitive impairment (MCI), employing machine learning models trained in one language to be applied in another.
This approach enables accurate and consistent assessment of MCI by analyzing speech patterns, providing a more efficient and reliable method compared to traditional manual assessments.
Smart Images

Figure 2025074056000001_ABST
Abstract
Description
[Technical field]
[0001] This invention relates to speech analysis, and more particularly to the automated assessment and diagnosis of one or more medical conditions based on collected speech samples. [Background technology]
[0002]
[0002] Assessment of nerve injuries and diseases and other medical conditions is often performed manually by medical personnel and may be based on handwritten pencil-and-paper forms. However, manual assessments can be inaccurate and / or inconsistent. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] U.S. Pat. No. 10,152,988 [Patent Document 2] U.S. Patent No. 10,311,980 [Patent Document 3] US Patent Application Publication No. 2023 / 072242 Summary of the Invention [Means for solving the problem]
[0004] An apparatus for speech language model training and application techniques is presented. In one embodiment, the apparatus includes a processor and a memory coupled to the processor. In one embodiment, the memory stores code executable by the processor for training a first speech model in a first language, the first speech model being used to determine one or more characteristics of speech data indicative of MCI, training a second speech model for use in a second language using at least a portion of the first speech model trained in the first language, applying the second speech model trained for use in the second language to the user's speech data captured in the second language, and determining an assessment of MCI for the user based on an output from the second speech model trained for use in the second language.
[0005] In another embodiment, the apparatus includes means for training a first speech model in a first language, the first speech model being used to determine one or more characteristics of speech data indicative of MCI. In one embodiment, the apparatus includes means for training a second speech model for use in a second language using at least a portion of the first speech model trained in the first language. In some embodiments, the apparatus includes means for applying the second speech model trained for use in the second language to the user's speech data captured in the second language. In one embodiment, the apparatus includes means for determining an assessment of MCI for the user based on output from the second speech model trained for use in the second language.
[0006] A method for speech language model training and application techniques is presented. In one embodiment, the method includes training a first speech model in a first language, where the first speech model is used to determine one or more characteristics of speech data indicative of MCI, training a second speech model for use in a second language using at least a portion of the first speech model trained in the first language, applying the second speech model trained for use in the second language to the user's speech data captured in the second language, and determining an assessment of MCI for the user based on an output from the second speech model trained for use in the second language.
[0007] A computer program product is presented that includes a computer-readable storage medium. In an embodiment, the computer-readable storage medium stores computer usable program code executable to perform operations for techniques of training and applying a spoken language model. In some embodiments, one or more of the operations may be substantially similar to one or more steps described above with respect to the disclosed apparatus, system, and / or method.
[0008] In order that the advantages of the invention may be readily understood, a more particular description of the invention briefly described above will now be presented by reference to specific embodiments illustrated in the accompanying drawings. With the understanding that these drawings merely depict exemplary embodiments of the invention and therefore should not be considered as limiting its scope, the invention will be described and explained with additional particularity and detail through the use of the accompanying drawings, in which: [Brief description of the drawings]
[0009] [Figure 1A] FIG. 1 is a schematic block diagram illustrating one embodiment of a system for the technique of training and applying a spoken language model. [Figure 1B]FIG. 1 is a schematic block diagram illustrating a further embodiment of a system for the technique of training and applying a spoken language model. [Diagram 2]
[0010] FIG. 1 is a schematic block diagram illustrating one embodiment of a system for processing speech data with a mathematical model to perform a medical diagnosis. [Diagram 3]
[0011] FIG. 1 is a schematic block diagram of one embodiment of a training corpus of utterance data. [Figure 4]
[0012] FIG. 1 is a schematic block diagram illustrating one embodiment of a list of prompts for use in diagnosing a medical condition. [Diagram 5]
[0013] FIG. 1 is a schematic block diagram illustrating one embodiment of a system for selecting features for training a mathematical model for diagnosing a medical condition. [Figure 6A]
[0014] FIG. 2 is a schematic block diagram illustrating one embodiment of a graph of feature value and diagnostic value pairs. [Figure 6B]
[0015] FIG. 13 is a schematic block diagram illustrating a further embodiment of a graph of pairs of feature values and diagnostic values. [Figure 7]
[0016] FIG. 1 is a schematic flow chart diagram illustrating one embodiment of a method for selecting features to train a mathematical model for diagnosing a medical condition. [Figure 8]
[0017] FIG. 1 is a schematic flow chart diagram illustrating one embodiment of a method for selecting a prompt for use with a mathematical model to diagnose a medical condition. [Figure 9]
[0018] FIG. 1 is a schematic flow chart diagram illustrating one embodiment of a method for training a mathematical model for diagnosing a medical condition that is adapted to a set of selected prompts. [Figure 10]
[0019] FIG. 1 is a schematic block diagram illustrating one embodiment of a computing device that may be used to train and develop mathematical models for diagnosing medical conditions. [Figure 11]
[0020] FIG. 1 is a schematic block diagram illustrating one embodiment of a diagnostic device. [Figure 12]
[0021] FIG. 1 is a schematic flow chart diagram illustrating one embodiment of a method for the technique of training and applying a spoken language model. [Figure 13]
[0022] FIG. 13 is a schematic flow chart diagram illustrating a further embodiment of a method for speech language model training and application technique. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0010]
[0001] In general, the subject matter of this specification aims to assess the risk or likelihood of a user having mild cognitive impairment (MCI) based on speech or voice data analysis using a trained machine learning model. As used herein, MCI refers to a small but measurable degree of cognitive decline during the prodromal phase, e.g., between the expected decline due to normal aging and the more serious decline due to dementia. MCI is usually not detectable by everyday conversation or observation, and it is presumed that the majority of people with MCI are unaware that they have cognitive problems. People with MCI struggle to remember recent conversations and events, keep track of schedules and appointments, or use new guidelines for tasks. Differentiating between normal aging and MCI is a difficult task even for the most attentive primary care physician, and doing so is important for timely intervention and optimal treatment outcomes.
[0011]
[0002] As discussed in more detail below, the subject matter described herein aims to use artificial intelligence, particularly machine learning, to determine a user's assessment of MCI using speech data. In particular, a first machine learning model may be trained for speech data in a first language, e.g., English, and a second machine learning model may be trained for use in a second language, e.g., Japanese, using at least a portion of the first machine learning model. In this manner, as described in more detail below, a speech model may be created to analyze speech data in one language using a speech model trained in a different language. The solutions proposed herein may utilize or otherwise relate to the solutions described in U.S. Pat. No. 10,152,988, U.S. Pat. No. 10,311,980, and U.S. Patent Publication No. 2023 / 072242, which are incorporated herein by reference in their entirety.
[0012]
[0003] As used herein, artificial intelligence (AI) is broadly defined as the field of computer science that deals with the automation of intelligent behavior. AI systems can be designed to emulate and simulate human intelligence and corresponding behavior using machines. This can take many forms, including symbolic AI or symbol-manipulation AI. AI can deal with analyzing abstract symbols and / or human-readable symbols. AI can form abstract connections between data or other information or stimuli. AI can form logical conclusions. AI is the intelligence exhibited by a machine, program, or software. AI is defined as the study and design of intelligent agents, which are systems that perceive their environment and take actions that maximize their probability of success.
[0013]
[0004] AI may have various attributes such as deduction, reasoning, and problem solving. AI may include knowledge representation or learning. AI systems may perform natural language processing, perception, motion detection, and information manipulation. At higher levels of abstraction, it may result in social intelligence, creativity, and general intelligence. A variety of approaches are employed, including but not limited to integrating cybernetics and brain simulation, symbolic, sub-symbolic, and statistics.
[0014]
[0005] Various AI tools may be employed, alone or in combination, including search and optimization, logic, probabilistic methods for uncertain reasoning, classifiers and statistical learning methods, neural networks, deep forward propagation neural networks, deep recurrent neural networks, deep learning, control theory, and languages.
[0015]
[0006] Machine learning (ML) plays an important role in a wide range of important applications using large amounts of data, such as data mining, natural language processing, image recognition, speech recognition, and many other intelligent systems. There are some basic common threads regarding the definition of ML. As used herein, ML is defined as a field of study that gives computers the ability to learn without being explicitly programmed. For example, it is possible to run a machine learning algorithm / model using data on past or historical traffic patterns, e.g., to train the machine learning algorithm / model to predict traffic patterns at a busy intersection. If the program has learned / trained correctly from past patterns, it may correctly predict future traffic patterns.
[0016]
[0007] There are various ways that an algorithm can model a problem based on its interaction with experience, environment, or input data. Machine learning algorithms can be categorized to help think about the role of input data and the model preparation process that leads to the correct selection of the most appropriate category for a problem to get the best results. Known categories are supervised learning, unsupervised learning, semi-supervised learning, and reinforcement learning.
[0017] (a) In the category of supervised learning, the input data, called training data, has known labels or outcomes. A model is prepared through a training process that requires predictions to be made and corrected when they are wrong. The training process continues until the model achieves a desired level of accuracy on the training data. Example problems are classification and regression.
[0018] (b) In the category of unsupervised learning, the input data is unlabeled and has no known outcome. A model is prepared by deducing structures present in the input data. Example problems are association rule learning and clustering. An example algorithm is k-means clustering.
[0019] (c) Semi-supervised learning lies somewhere between unsupervised learning (no labeled training data) and supervised learning (having fully labeled training data). Researchers have found that unlabeled data, when used in conjunction with a small amount of labeled data, can provide significant improvements in learning accuracy.
[0020] (d) Reinforcement learning is another category that differs from standard supervised learning in that the correct input / output pairs are never presented. Furthermore, there is an emphasis on online performance, which involves finding a balance between exploring new knowledge and exploiting current knowledge that has already been discovered.
[0021]
[0008] Certain machine learning techniques are widely used, including decision tree learning, association rule learning, artificial neural networks, inductive logic programming, support vector machines, clustering, Bayesian networks, reinforcement learning, representation learning, and genetic algorithms. In some embodiments, multiple machine learning algorithms may be applied using ensemble learning. As used herein, ensemble learning may refer to a machine learning technique that combines multiple algorithms to create a single predictive model.
[0022]
[0009] The learning process in machine learning algorithms is generalization from past experience. After going through a training dataset, the generalization process is the ability of a machine learning algorithm to perform accurately on new examples and tasks. The learner needs to build a generalized model of the problem space to enable the machine learning algorithm to produce sufficiently accurate predictions in future cases. The training examples may come from some generally unknown probability distribution.
[0023] In theoretical computer science, computational learning theory performs computational analysis of machine learning algorithms and their performance. Training datasets are limited in size and may not capture all forms of distribution in future datasets. Performance is represented by probabilistic bounds. Generalization error is quantified by bias-variance decomposition. The time complexity and feasibility of learning in computational learning theory is said to be feasible if the computation is performed in polynomial time. Positive results are determined and classified when a class of features can be learned in polynomial time, and negative results are determined and classified when they cannot be learned in polynomial time.
[0024]
[0023] Figure 1A illustrates one embodiment of a system 100 for speech language model training and application techniques. In one embodiment, the system 100 includes one or more hardware devices 102, one or more diagnostic devices 104 (e.g., one or more diagnostic devices 104a, one or more back-end diagnostic devices 104b, etc., located on one or more hardware devices 102), one or more data networks 106 or other communication channels, and / or one or more back-end servers 108. In an embodiment, a particular number of hardware devices 102, diagnostic devices 104, data networks 106, and / or back-end servers 108 are illustrated in Figure 1, but one skilled in the art will recognize in light of this disclosure that any number of hardware devices 102, diagnostic devices 104, data networks 106, and / or back-end servers 108 may be included in the system 100 for speech collection and / or speech language model training and application techniques.
[0025]
[0024] Generally, the diagnostic device 104 is configured to train a first speech model in a first language, the first speech model is used to determine one or more characteristics of speech data indicative of MCI, train a second speech model for use in a second language using at least a portion of the first speech model trained in the first language, apply the second speech model trained for use in the second language to the user's speech data captured in the second language, and determine an assessment of MCI for the user based on output from the second speech model trained for use in the second language. The diagnostic device 104, including its various sub-modules, may be located on one or more information handling devices 102, one or more servers 108, one or more network devices, etc. in the system 100. The diagnostic device 104 is described in more detail below.
[0026] In one embodiment, the system 100 includes one or more hardware devices 102. The hardware devices 102 and / or one or more backend servers 108 (e.g., computing devices, information handling devices, etc.) may include one or more of a desktop computer, a laptop computer, a mobile device, a tablet computer, a smartphone, a set-top box, a gaming console, a smart TV, a smart watch, a fitness band, an optical head mounted display (e.g., a virtual reality headset, smart glasses, etc.), an HDMI or other electronic display dongle, a personal digital assistant, and / or another computing device including a processor (e.g., a central processing unit (CPU), a processor core, a field programmable gate array (FPGA) or other programmable logic, an application specific integrated circuit (ASIC), a controller, a microcontroller, and / or another semiconductor integrated circuit device), a volatile memory, and / or a non-volatile storage medium. In an embodiment, the hardware devices 102 communicate with the one or more backend servers 108 via a data network 106, which will be described below. In further embodiments, the hardware devices 102 are capable of executing various programs, program codes, applications, instructions, functions, etc.
[0027] In various embodiments, the diagnostic device 104 may be embodied as hardware, software, or some combination of hardware and software. In one embodiment, the diagnostic device 104 may include executable program code stored on a non-transitory computer-readable storage medium for execution on a processor of the hardware device 102, the back-end server 108, or the like. For example, the diagnostic device 104 may be embodied as executable program code executing on one or more of the hardware device 102, the back-end server 108, one or more combinations thereof, or the like. In such an embodiment, various modules that perform the operations of the diagnostic device 104, as described below, may be located on the hardware device 102, the back-end server 108, a combination of the two, or the like.
[0028]
[0027] In various embodiments, the diagnostic device 104 may be embodied as a hardware appliance that may be installed or deployed on the backend server 108, on the user's hardware device 102 (e.g., a dongle, a protective case for the phone 102 or tablet 102 that includes one or more semiconductor integrated circuit devices within the case that communicates with the phone 102 or tablet 102 wirelessly and / or via a data port such as a USB or dedicated communications port, or another peripheral device), or elsewhere on the data network 106, and / or co-located with the user's hardware device 102. In an embodiment, the diagnostic device 104 may include a hardware device such as a secure hardware dongle or other hardware appliance device (e.g., a set-top box, a network appliance, etc.) that attaches to another hardware device 102, such as a laptop computer, a server, a tablet computer, a smartphone, by either a wired connection (e.g., a USB connection) or a wireless connection (e.g., Bluetooth, Wi-Fi, Near Field Communication (NFC), etc.), attaches to an electronic display device (e.g., a television or monitor using an HDMI port, a DisplayPort port, a MiniDisplayPort port, a VGA port, a DVI port, etc.), operates substantially independently on a data network 106, etc. The hardware appliance of the diagnostic device 104 may include a power interface, a wired and / or wireless network interface, a graphical interface that outputs to a display device (e.g., a graphics card and / or GPU with one or more display ports), and / or a semiconductor integrated circuit device as described below, configured to perform the functions described herein with respect to the diagnostic device 104.
[0029] In such an embodiment, the diagnostic device 104 may include a semiconductor integrated circuit device (e.g., one or more chips, dies, or other discrete logic hardware), such as a field programmable gate array (FPGA) or other programmable logic, firmware for the FPGA or other programmable logic, microcode for executing on a microcontroller, an application specific integrated circuit (ASIC), a processor, a processor core, etc. In one embodiment, the diagnostic device 104 may be mounted on a printed circuit board having one or more electrical lines or connections (e.g., to volatile memory, non-volatile storage media, network interfaces, peripheral devices, graphical / display interfaces). The hardware appliance may include one or more pins, pads, or other electrical connections configured to transmit and receive data (e.g., in communication with one or more electrical lines of a printed circuit board, etc.), as well as one or more hardware circuits and / or other electrical circuits configured to perform various functions of the diagnostic device 104.
[0030]
[0029] In one embodiment, the semiconductor integrated circuit device or other hardware appliance of the diagnostic apparatus 104 includes and / or is communicatively coupled to one or more volatile memory media, which may include, but are not limited to, random access memory (RAM), dynamic RAM (DRAM), cache, etc. In one embodiment, the semiconductor integrated circuit device or other hardware appliance of the diagnostic apparatus 104 includes and / or is communicatively coupled to one or more non-volatile memory media, which may include, but are not limited to, NAND flash memory, NOR flash memory, nano random access memory (nanoRAM or NRAM), nanocrystalline wire-based memory, silicon oxide based sub-10 nanometer process memory, graphene memory, silicon-oxide-nitride-oxide-silicon (SONOS), resistive RAM (RRAM), programmable metallization cell (PMC), conductive bridge RAM (CBRAM), magnetoresistive RAM (MRAM), dynamic RAM (DRAM), phase change RAM (PRAM or PCM), magnetic storage media (e.g., hard disk, tape), optical storage media, and the like.
[0031] In one embodiment, the data network 106 includes a digital communication network that transmits digital communications. The data network 106 may include wireless networks such as wireless cellular networks, local wireless networks such as Wi-Fi networks, Bluetooth networks, near field communication (NFC) networks, ad-hoc networks, etc. The data network 106 may include a wide area network (WAN), a storage area network (SAN), a local area network (LAN), an optical fiber network, the Internet, or other digital communication networks. The data network 106 may include two or more networks. The data network 106 may include one or more servers, routers, switches, and / or other networking equipment. The data network 106 may also include one or more computer-readable storage media, such as hard disk drives, optical drives, non-volatile memory, RAM, etc.
[0032] In one embodiment, the one or more backend servers 108 may include one or more network-accessible computing systems, such as one or more web servers hosting one or more websites, enterprise intranet systems, application servers, application programming interface (API) servers, authentication servers, etc. The backend servers 108 may include one or more servers located remotely from the hardware device 102. The backend servers 108 may include at least a portion of the diagnostic device 104, may include the hardware of the diagnostic device 104, may store executable program code of the diagnostic device 104 on one or more non-transitory computer-readable storage media, and / or may otherwise perform one or more of the various operations of the diagnostic device 104 described herein for shared content tracking and attribution.
[0033]
[0032] Figure 1B is an exemplary system 109 for diagnosing a medical condition using a person's speech. Figure 1B includes a medical condition diagnosis service 140 that may receive the person's speech data and process the speech data to determine whether the person has a medical condition. For example, the medical condition diagnosis service 140 may process the speech data to compute a yes or no decision as to whether the person has a medical condition, or may compute a score indicative of the possibility or likelihood that the person has the medical condition and / or the severity of the condition.
[0034]
[0033] As used herein, diagnosis refers to any determination as to whether a person may have a medical condition, or any determination as to the possible severity of a medical condition. A diagnosis may include any form of assessment, conclusion, opinion, or judgment regarding a medical condition. In some cases, a diagnosis may not be accurate and a person diagnosed with a medical condition may not actually have the medical condition.
[0035]
[0034] The medical condition diagnostic service 140 may receive the person's speech data using any suitable technique. For example, the person may speak to the mobile device 110, which may record the speech and transmit the recorded speech data to the medical condition diagnostic service 140 over the network 130. Any suitable technique and any suitable network may be used for the mobile device 110 to transmit the recorded speech data to the medical condition diagnostic service 140. For example, an application or "app" may be installed on the mobile device 110 that uses REST (Representational State Transfer) API (Application Programming Interface) calls to transmit the speech data over the Internet or a mobile telephone network. In another example, a medical provider may have a medical provider computer 120 that is used to record the person's speech and transmit the speech data to the medical condition diagnostic service 140.
[0036] In some implementations, the medical condition diagnostic service 140 may be installed on the mobile device 110 or the healthcare provider computer 120 so as to avoid the need to transmit the speech data over a network. The example of FIG. 1B is not limiting and any suitable technique may be used to transmit the speech data for processing by the mathematical model.
[0037]
[0036] The output of the medical condition diagnostic service 140 may then be used for any suitable purpose, for example, information may be presented to the person who provided the speech data or to medical personnel treating the person.
[0038]
[0037] Figure 2 is an exemplary system 200 for processing speech data with a mathematical model to perform medical diagnosis. In processing the speech data, features may be calculated from the speech data and then processed by the mathematical model. Any suitable type of features may be used.
[0039]
[0038] The features may include acoustic features, which are any features computed from the speech data that do not involve or depend on performing speech recognition on the speech data (e.g., the acoustic features do not use information about the spoken words in the speech data). For example, the acoustic features may include mel-frequency cepstral coefficients, perceptual linear prediction features, jitter, or shimmer.
[0040]
[0039] The features may include linguistic features, which are calculated using the results of speech recognition. For example, the linguistic features may include speaking rate (e.g., number of vowels or syllables per second), number of pause fillers (e.g., "ums" and "ahs"), difficulty of the word (e.g., less common words), or part of speech of the word after the pause filler.
[0041] 2, speech data is processed by an acoustic feature computation component 210 and a speech recognition component 220. The acoustic feature computation component 210 may compute acoustic features from the speech data, such as any of the acoustic features described herein. The speech recognition component 220 may perform automatic speech recognition on the speech data using any suitable technique (e.g., Gaussian mixture models, acoustic modeling, language modeling, and neural networks).
[0042]
[0041] Because the speech recognition component 220 may use acoustic features in performing speech recognition, some processing of these two components may overlap, and therefore other configurations are possible. For example, the acoustic feature computation component 210 may compute the acoustic features required by the speech recognition component 220, and the speech recognition component 220 may thereby not need to compute any acoustic features.
[0043]
[0042] The linguistic feature computation component 230 may receive the speech recognition results from the speech recognition component 220 and process the speech recognition results to determine linguistic features, such as any of the linguistic features described herein. The speech recognition results may be in any suitable format and may include any suitable information. For example, the speech recognition results may include a word lattice that includes multiple possible sequences of words, information about pause fillers, and the timing of words, syllables, vowels, pause fillers, or any other unit of speech.
[0044]
[0043] The medical condition classifier 240 may process the acoustic and linguistic features with a mathematical model to output one or more diagnostic scores indicative of whether the person has a medical condition, such as a score indicative of the probability or likelihood that the person has the medical condition and / or a score indicative of the severity of the medical condition. The medical condition classifier 240 may use any suitable technique, such as a support vector machine, or a classifier implemented in a neural network, such as a multi-layer perceptron, e.g., a fully connected dense network, a convolutional neural network, etc.
[0045]
[0044] The performance of the disease condition classifier 240 may depend on the features computed by the acoustic feature computation component 210 and the linguistic feature computation component 230. Furthermore, a set of features that works well for one disease condition may not work well for another disease condition. For example, word difficulty may be an important feature for diagnosing Alzheimer's disease, but may not be useful for determining whether a person has a concussion. As another example, features related to the pronunciation of vowels, syllables, or words may be important for Parkinson's disease, but less important for other disease conditions. Thus, a technique is needed to determine a first set of features that works well for a first disease condition, and this process may need to be repeated to determine a second set of features that works well for a second disease condition.
[0046] In some implementations, medical condition classifier 240 may use other features, which may be referred to as non-speech features, in addition to acoustic and linguistic features. For example, features may be obtained or calculated from a person's demographic information (e.g., gender, age, or place of residence), information from a medical history (e.g., weight, recent blood pressure readings, or previous diagnoses), or any other suitable information.
[0047]
[0046] The selection of features for diagnosing a medical condition may be more important in situations where the amount of training data for training a mathematical model is relatively small. For example, to train a mathematical model for diagnosing a concussion, the training data required may include speech data of a number of individuals immediately after experiencing a concussion. Such data may exist in small amounts, and obtaining additional examples of such data may take a significant period of time.
[0048]
[0047] Training a mathematical model using a small amount of training data can result in overfitting, where the mathematical model fits the particular training data, but the model may not work well on new data due to the small amount of training data. For example, the model may be able to detect all of the concussions in the training data, but may have a high error rate when processing production data of people with possible concussions.
[0049]
[0048] One technique for preventing overfitting when training a mathematical model is to reduce the number of features used to train the mathematical model. The amount of training data required to train the model without overfitting increases as the number of features increases. Therefore, using a smaller number of features allows the model to be built with a smaller amount of training data.
[0050]
[0049] When a model needs to be trained with a smaller number of features, it becomes more important to select features that allow the model to perform well. For example, when a large amount of training data is available, hundreds of features may be used to train the model, and the likelihood that the appropriate features are used is higher. Conversely, when a small amount of training data is available, only 10 or so features may be used to train the model, and it is more important to select the most important features to diagnose the disease state.
[0051]
[0050] Examples of features that can be used to diagnose a medical condition are now presented.
[0051] Acoustic features may be computed using short-time segment features. When processing speech data, the duration of the speech data may vary. For example, some utterances may be one or two seconds long, while some utterances may be several minutes or longer. For consistency in processing the speech data, the speech data may be processed in short-time segments (sometimes called frames). For example, each short-time segment may be 25 milliseconds long, and the segments may proceed in 10 millisecond increments, such that there is a 15 millisecond overlap across two consecutive segments.
[0052]
[0052] Spectral features (such as mel-frequency cepstral coefficients or perceptual linear prediction), prosodic features (such as pitch, energy, or probability of utterance), voice quality features (such as jitter, jitter of jitter, shimmer, or harmonic to noise ratio), and entropy (e.g., to capture how accurately an utterance is pronounced, entropy can be calculated from the posterior distribution of an acoustic model trained on natural speech data) are non-limiting examples of short-time segment features.
[0053]
[0053] Short-time segment features may be combined to compute acoustic features for speech. For example, a two-second speech sample may produce 200 short-time segment features for pitch, which may be combined to compute one or more acoustic features for pitch.
[0054]
[0054] The short-time segment features may be combined to compute acoustic features for the speech sample using any suitable technique. In some implementations, the acoustic features may be calculated using statistics of the short-time segment features (e.g., the arithmetic mean, the standard deviation, the skewness, the kurtosis, the 1st quartile, the 2nd quartile, the 3rd quartile minus the 1st quartile, the 3rd quartile minus the 2nd quartile, the 0.01 percentile, the 0.99th percentile, the 0.99th percentile minus the 0.01 percentile, the percentage of the short-time segment whose value is above a threshold (e.g., if the threshold is 75% of the range plus a minimum), the percentage of the segment whose value is above a threshold (e.g., if the threshold is 90% of the range plus a minimum), the slope of a linear approximation of the value, the offset of a linear approximation of the value, a linear error calculated as the difference between the linear approximation and the actual value, or a quadratic error calculated as the difference between the linear approximation and the actual value). In some implementations, the acoustic features may be computed as utterance embeddings to represent partial or complete speech. The utterance embeddings may include short-term segment features based on self-supervised pre-trained models such as wav2vec or Trillson and identity vectors such as i-vectors or x-vectors of the utterance representation. The identity vectors may be computed using any suitable technique, such as performing a matrix-vector transformation using factor analysis techniques and a Gaussian mixture model for the i-vector or a neural network model for the x-vector.
[0055]
[0055] The following are non-limiting examples of linguistic features: Speech rate, such as by calculating the duration of all spoken words divided by the number of vowels or any other suitable measure of speech rate. Number of pause fillers, which may indicate hesitation in speech, such as (1) the number of pause fillers divided by the duration of spoken words, or (2) the number of pause fillers divided by the number of spoken words. Measures of word difficulty or less common use of words. For example, word difficulty may be calculated using statistics of one-gram probability of spoken words, such as by classifying words according to their frequency percentiles (e.g., 5%, 10%, 15%, 20%, 30%, or 40%). Parts of speech of words after pause fillers, such as (1) the count of each part of speech class divided by the number of spoken words, or (2) the count of each part of speech class divided by the sum of all part of speech counts.
[0056]
[0056] In some implementations, the linguistic features may include a determination of whether the person answered a question correctly. For example, a person may be asked what year it is or who is the president of the United States. The person's speech may be processed to determine what the person said in response to the question and to determine whether the person answered the question correctly. Additionally, in some implementations, the linguistic features may include a determination of whether the person read correctly, e.g., whether they read a presented passage correctly. In such implementations, a word error rate is calculated, e.g., using automatic speech recognition (ASR) results, and compared to an expected reading script. In some implementations, when the question prompt is intended to assess a verbal fluency test, e.g., by asking the user to list words in a category such as animals, an evaluation is made to determine whether the user's response actually belongs to the expected category by checking or calculating the distance between the word vectors.
[0057]
[0057] To train a model for diagnosing a medical condition, a corpus of training data may be collected. The training corpus may include examples of utterances from which a person's diagnosis is known. For example, it may be known that a person does not have a concussion, or that they have a mild, moderate, or severe concussion.
[0058]
[0058] Figure 3 shows an example of a training corpus including speech data for training a model to diagnose concussion. For example, the rows of the table in Figure 3 may correspond to entries in a database. In this example, each entry includes an identifier for a person, a known diagnosis for that person (e.g., no concussion or mild, moderate, or severe concussion), an identifier for a prompt or question presented to the person (e.g., "How are you feeling today?"), and a filename for a file containing the speech data. The training data may be stored in any suitable format using any suitable storage technology.
[0059]
[0059] The training corpus may store representations of human speech using any suitable format. For example, the speech data items of the training corpus may include digital samples of an audio signal received at a microphone, or may include processed versions of the audio signal, such as Mel-frequency cepstral coefficients.
[0060]
[0060] A single training corpus may contain speech data related to multiple medical conditions, or a separate training corpus may be used for each medical condition (e.g., a first training corpus for concussion and a second training corpus for Alzheimer's disease). A separate training corpus may be used to store speech data for people with no known or diagnosed medical conditions, when this training corpus may be used to train models for multiple medical conditions.
[0061]
[0061] Figure 4 shows examples of stored prompts that may be used to diagnose a medical condition. Each prompt may be presented to a person, by a human (e.g., a medical professional) or a computer, to obtain the person's utterance in response to the prompt. Each prompt may have a prompt identifier so that it may be cross-referenced with the prompt identifier in the training corpus. The prompts of Figure 4 may be stored using any suitable storage technique, such as a database.
[0062] 5 is an exemplary system 500 that may be used to select features for training a mathematical model for diagnosing a medical condition, and then to train a mathematical model using the selected features. The system 500 may be used multiple times to select features for different medical conditions. For example, a first use of the system 500 may select features for diagnosing a concussion, and a second use of the system 500 may select features for diagnosing Alzheimer's disease.
[0063] 5 includes a training corpus 510 of speech data items for training a mathematical model to diagnose a medical condition. The training corpus 510 may include any suitable information, such as speech data of a plurality of people with and without a medical condition, a label indicating whether a person has a medical condition, and any other information described herein.
[0064]
[0064] The acoustic feature computation component 210, the speech recognition component 220, and the linguistic feature computation component 230 may be implemented as described above to compute acoustic and linguistic features for the speech data in the training corpus. The acoustic feature computation component 210 and the linguistic feature computation component 230 may compute multiple features so that the best performing features may be determined. This may be in contrast to FIG. 2, where these components are used in a production system and thus may compute only previously selected features.
[0065]
[0065] The feature selection score calculation component 520 may calculate a selection score for each feature (which may be an acoustic feature, a linguistic feature, or any other feature described herein). To calculate a selection score for a feature, a pair of numbers may be created for each utterance data item in the training corpus. The first number of the pair is the value of the feature, and the second number of the pair is an indicator of a medical condition diagnosis. The value of the indicator of a medical condition diagnosis may have two values (e.g., 0 if the person does not have the medical condition and 1 if the person has the medical condition) or may have a larger number of values (e.g., a real number between 0 and 1, or multiple integers indicating the likelihood or severity of the medical condition).
[0066]
[0066] Thus, for each feature, a number pair may be obtained for each speech data item of the training corpus. Figures 6A and 6B show two conceptual plots of number pairs for a first feature and a second feature. In the case of Figure 6A, there does not appear to be a pattern or correlation between the values of the first feature and the corresponding diagnostic values, whereas in the case of Figure 6B, there appears to be a pattern or correlation between the values of the second feature and the diagnostic values. Thus, it may be concluded that the second feature is likely to be a useful feature for determining whether a person has a medical condition, whereas the first feature is not.
[0067]
[0067] The feature selection score calculation component 520 may calculate a selection score for the feature using the feature value and diagnostic value pairs. The feature selection score calculation component 520 may calculate any suitable score that indicates a pattern or correlation between the feature value and the diagnostic value. For example, the feature selection score calculation component 520 may calculate a Rand index, an adjusted Rand index, mutual information, adjusted mutual information, Pearson correlation, absolute Pearson correlation, Spearman correlation, or absolute Spearman correlation.
[0068]
[0068] The selection score may indicate the usefulness of the feature in detecting a disease state. For example, a high selection score may indicate that the feature should be used in training the mathematical model, and a low selection score may indicate that the feature should not be used in training the mathematical model.
[0069]
[0069] The feature stability determination component 530 may determine whether a feature (which may be an acoustic feature, a linguistic feature, or any other feature described herein) is stable or unstable. To make the stability determination, the speech data items may be divided into multiple groups, which may be referred to as folds. For example, the speech data items may be divided into five folds. In some implementations, the speech data items may be divided into folds such that each fold has an approximately equal number of speech data items for different gender and age groups.
[0070] The statistics of each fold can be compared to the statistics of other folds. For example, for a first fold, the median (or mean or any other statistic related to the center or middle of a distribution) feature value (M 1 Statistics may also be calculated for other fold combinations. For example, the median of the feature values (M 0 ) and statistics measuring the variability of feature values, such as the interquartile range, variance, or standard deviation (V 0 A feature may be determined to be unstable if the median of the first fold is too different from the median of the second fold. For example, a feature may be determined to be unstable if the median of the first fold is too different from the median of the second fold.
[0071]
number
[0072] It may be determined to be unstable if C is a scaling factor. The process may then be repeated for each of the other folds. For example, the median of the second fold may be compared to the medians and variabilities of the other folds as described above.
[0073]
[0071] In some implementations, after comparing each fold to the other folds, a feature may be determined to be stable if the median of each fold is not too far from the medians of the other folds. Conversely, a feature may be determined to be unstable if the median of any fold is too far from the medians of the other folds.
[0074] In some implementations, the feature stability determination component 530 may output a Boolean value for each feature to indicate whether the feature is stable or not. In some implementations, the feature stability determination component 530 may output a stability score for each feature. For example, the stability score may be calculated as the maximum distance (e.g., Mahalanobis distance) between the median of a fold and the median of another fold.
[0075]
[0073] The feature selection component 540 may receive the selection scores from the feature selection score calculation component 520 and the stability determination from the feature stability determination component 530 and select a subset of features to be used to train the mathematical model. The feature selection component 540 may select a number of features with the highest selection scores that are also sufficiently stable.
[0076]
[0074] In some implementations, the number of features selected (or the maximum number of features selected) may be preset. For example, the number N may be determined based on the amount of training data, and the N features may be selected. The selected features may be determined by removing unstable features (e.g., features determined to be unstable or features having a stability score below a threshold) and then selecting the N features with the highest selection scores.
[0077]
[0075] In some implementations, the number of features selected may be based on the selection score and a stability determination. For example, the selected features may be determined by removing unstable features and then selecting all features having a selection score above a threshold.
[0078]
[0076] In some implementations, the selection score and the stability score may be combined when selecting features. For example, for each feature, a combined score may be calculated (such as by adding or multiplying the selection score and the stability score for the feature), and features may be selected using the combined score.
[0079]
[0077] The model training component 550 may then train a mathematical model using the selected features. For example, the model training component 550 may iterate over the utterance data items of the training corpus to obtain selected features for the utterance data items, and then train the mathematical model using the selected features. In some implementations, a dimensionality reduction technique, such as principal component analysis or linear discriminant analysis, may be applied to the selected features as part of the model training. Any suitable mathematical model may be trained, such as any of the mathematical models described herein.
[0080] In some implementations, other techniques such as wrapper methods may be used for feature selection or may be used in combination with the feature selection techniques presented above. Wrapper methods may select a set of features, train a mathematical model using the selected set of features, and then use the trained model to evaluate the performance of the set of features. If the number of possible features is relatively small and / or the training time is relatively short, all possible sets of features may be evaluated and the best performing set may be selected. If the number of possible features is relatively large and / or the training time is a significant factor, optimization techniques may be used to iteratively find a set of features that perform well. In some implementations, a set of features may be selected using system 500, and then a subset of these features may be selected as the final set of features using wrapper methods.
[0081]
[0079] Figure 7 is a flow chart of an exemplary embodiment of feature selection for training a mathematical model for diagnosing a medical condition. In Figure 7 and other flow charts herein, the order of steps is exemplary, other orders are possible, not all steps are required, steps may be combined (in whole or in part) or sub-divided, and in some embodiments some steps may be omitted or other steps may be added. The method described by any flow chart described herein may be implemented, for example, by any of the computers or systems described herein.
[0082]
[0080] At step 710, a training corpus of speech data items is obtained. The training corpus may include representations of audio signals of speech of a person, medical diagnostic indications of the person from whom the utterance was obtained, and any other suitable information, such as any of the information described herein.
[0083]
[0081] In step 720, speech recognition results are obtained for each speech data item of the training corpus. The speech recognition results may be pre-computed, stored with the training corpus, or stored elsewhere. The speech recognition results may include any suitable information, such as a phonetic transcription, a list of the highest scoring phonetic transcriptions (e.g., a best N list), a lattice of possible phonetic transcriptions, and timing information, such as start and end times of words, pause fillers, or other speech units.
[0084]
[0082] In step 730, acoustic features are computed for each utterance data item of the training corpus. The acoustic features may include any features that are computed without using speech recognition results of the utterance data item, such as any of the acoustic features described herein. The acoustic features may include or be computed from data used in the speech recognition process (e.g., mel-frequency cepstral coefficients or perceptual linear prediction), but the acoustic features do not use speech recognition results, such as information about words or pause fillers present in the utterance data item.
[0085]
[0083] In step 740, linguistic features are computed for each utterance data item of the training corpus. The linguistic features may include any features computed using the speech recognition results, such as any of the linguistic features described herein.
[0086]
[0084] In step 750, a feature selection score is calculated for each acoustic feature and for each linguistic feature. To calculate the feature selection score for a feature, the value of the feature for each utterance data item in the training corpus may be used along with other information, such as known diagnostic values corresponding to the utterance data item. The feature selection score may be calculated using any of the techniques described herein, such as by calculating absolute Pearson correlation. In some implementations, feature selection scores may be calculated for other features as well, such as features related to a person's demographic information.
[0087]
[0085] In step 760, a number of features are selected using the feature selection scores. For example, a number of features having the highest selection scores may be selected. In some implementations, a stability determination may be calculated for each feature, and a number of features may be selected using both the feature selection scores and the stability determination, such as by using any of the techniques described herein.
[0088] In step 770, a mathematical model is trained using the selected features. Any suitable mathematical model may be trained, such as a neural network or a support vector machine. After the mathematical model is trained, it may be deployed to a production system, such as diagnostic device 104, system 109 of FIG. 1B, to perform diagnosis of a medical condition.
[0089]
[0087] The steps of Figure 7 may be performed in a variety of ways. For example, in some implementations, steps 730 and 740 may be performed in a loop that loops over each of the utterance data items in the training corpus. In a first iteration, acoustic and linguistic features may be calculated for a first utterance data item, in a second iteration, acoustic and linguistic features may be calculated for a second utterance data item, and so on.
[0090]
[0088] When using the developed model to diagnose a medical condition, the person being diagnosed may be presented with a series of prompts or questions to obtain utterances from the person. Any suitable prompts may be used, such as any of the prompts in Figure 4. After a feature is selected, as described above, a prompt may be selected such that the selected prompt provides useful information regarding the selected feature.
[0091]
[0089] For example, assume that the selected feature is pitch. While pitch has been determined to be a useful feature for diagnosing a medical condition, some prompts may be better than others for obtaining useful pitch features. Very short utterances (e.g., yes / no answers) may not provide enough data to accurately calculate pitch, and therefore prompts that result in longer responses may be more useful for obtaining information about pitch.
[0092]
[0090] As another example, assume that the selected feature is word difficulty. While word difficulty has been determined to be a useful feature for diagnosing a medical condition, some prompts may be better than others at obtaining useful word difficulty features. A prompt that asks a user to read a presented passage will generally result in the utterance of the words in the passage, and thus the word difficulty features will have the same value each time the prompt is presented, and thus the prompt is not useful for obtaining information about word difficulty. In contrast, an open-ended question such as "What was your day like today?" may result in a greater variation in the vocabulary of responses, and thus may yield more useful information regarding word difficulty.
[0093]
[0091] Selecting a set of prompts may also improve the performance of the system for diagnosing a medical condition and provide a better experience for the person being evaluated. By using the same set of prompts for each person being evaluated, the system for diagnosing a medical condition may provide more accurate results because the data collected from multiple people may be more similar than if different prompts were used for each person. Furthermore, using a set of predefined prompts makes the evaluation of a person more predictable and allows for a desired period of evaluation appropriate for the evaluation of a medical condition. For example, to evaluate whether a person has Alzheimer's disease, it may be acceptable to collect a large amount of data using more prompts, but to evaluate whether a person has a concussion during a sporting event, it may be necessary to use a smaller number of prompts to obtain results more quickly.
[0094]
[0092] In some implementations, a prompt may be selected by calculating a prompt selection score. The training corpus may have multiple or even more utterance data items for a single prompt. For example, the training corpus may include examples of prompts used with different people, or the same prompt may be used multiple times with the same person.
[0095] FIG. 8 is a flow chart of an exemplary implementation of prompt selection for use with a developed model for diagnosing a medical condition.
[0094] Steps 810-840 may be performed for each prompt (or a subset of prompts) in the training corpus to calculate a prompt selection score for each prompt.
[0096] In step 810, a prompt is obtained, and in step 820, a speech data item corresponding to the prompt is obtained from the training corpus.
[0096] In step 830, a medical diagnostic score is calculated for each utterance data item corresponding to a prompt. For example, the medical diagnostic score for an utterance data item may be a number output by a mathematical model (e.g., the mathematical model trained in FIG. 7) that indicates the likelihood that a person has a medical condition and / or the severity of the medical condition.
[0097]
[0097] In step 840, a prompt selection score is calculated for the prompt using the calculated medical diagnostic score. The calculation of the prompt selection score may be similar to the calculation of the feature selection score, as described above. For each utterance data item corresponding to the prompt, a pair of numbers may be obtained. For each pair, the first number of the pair may be the calculated medical diagnostic score calculated from the utterance data item, and the second number of the pair may be the person's known medical condition diagnosis (e.g., the person is known to have that medical condition or severity of the medical condition). Plotting these number pairs may result in a plot similar to Figure 6A or Figure 6B. Depending on the prompt, there may or may not be a pattern or correlation in the number pairs.
[0098]
[0098] The prompt selection score for a prompt may include any score that indicates a pattern or correlation between the calculated medical diagnosis score and a known medical condition diagnosis. For example, the prompt selection score may include a Rand index, an adjusted Rand index, mutual information, adjusted mutual information, Pearson correlation, absolute Pearson correlation, Spearman correlation, or absolute Spearman correlation.
[0099]
[0099] In step 850, it is determined whether other prompts remain to be processed. If prompts remain to be processed, processing may proceed to step 810 to process the additional prompts. If all prompts have been processed, processing may proceed to step 860.
[0100]
[0100] In step 860, a number of prompts are selected using the prompt selection scores. For example, a number of prompts with the highest prompt selection scores may be selected. In some implementations, a stability determination may be calculated for each prompt, and a number of prompts may be selected using both the prompt selection scores and the prompt stability determination, such as by using any of the techniques described herein.
[0101]
[0101] In step 870, the selected prompts are used in the deployed medical condition diagnosis service. For example, when diagnosing a person, the selected prompts may be presented to the person to obtain the person's utterances in response to each of the prompts.
[0102]
[0102] In some implementations, other techniques, such as wrapper methods, may be used for prompt selection or may be used in combination with the prompt selection techniques presented above. In some implementations, a set of prompts may be selected using the process of Figure 8, and then a subset of these prompts may be selected as the final set of features using wrapper methods.
[0103]
[0103] In some implementations, a person involved in creating the medical condition diagnosis service may assist in the selection of the prompts. The person may use their knowledge or experience to select a prompt based on the selected feature. For example, if the selected feature is word difficulty, the person may review the prompts and select the prompts that are more likely to provide useful information regarding word difficulty. The person may select one or more prompts that are more likely to provide useful information for each of the selected features.
[0104]
[0104] In some implementations, the person may review the prompts selected by the process of Figure 8 and add or remove prompts to improve the performance of the medical condition diagnosis system. For example, two prompts may each provide useful information about the difficulty of a word, but much of the information provided by the two prompts may be redundant, and using both prompts may not provide a significant benefit over using only one of them.
[0105]
[0105] In some implementations, a second mathematical model that is adapted to the selected prompt may be trained after prompt selection. The mathematical model trained in FIG. 7 may process a single utterance (in response to a prompt) to generate a medical diagnostic score. If the process of making a diagnosis involves processing multiple utterances corresponding to multiple prompts, each of the utterances may be processed by the mathematical model of FIG. 7 to generate multiple medical diagnostic scores. To determine an overall medical diagnosis, multiple medical diagnostic scores may need to be combined in some way. Thus, the mathematical model trained in FIG. 7 may not be adapted to the set of selected prompts.
[0106]
[0106] When the selected prompts are used in a session to diagnose a person, each of the prompts may be presented to the person to obtain an utterance corresponding to each of the prompts. Instead of processing the utterances separately, the utterances may be processed simultaneously by the model to generate a medical diagnosis score. Thus, the model may be adapted to the selected prompts because it is trained to simultaneously process utterances corresponding to each of the selected prompts.
[0107]
[0107] Figure 9 is a flow chart of an exemplary implementation of training a mathematical model to be fitted to a set of selected prompts. In step 910, a first mathematical model is obtained, such as by using the process of Figure 7. In step 920, a number of prompts are selected using the first mathematical model, such as by the process of Figure 8.
[0108]
[0108] In step 930, a second mathematical model is trained to simultaneously process the speech data items corresponding to the selected prompts to generate a medical diagnostic score. When training the second mathematical model, a training corpus including sessions having speech data items corresponding to each of the selected prompts may be used. When training the mathematical model, inputs to the mathematical model may be fixed to the speech data items from the sessions corresponding to each of the selected prompts. Outputs of the mathematical model may be fixed to the known medical diagnosis. Parameters of the model may then be trained to optimally simultaneously process the speech data items to generate a medical diagnostic score. Any suitable training technique may be used, such as stochastic gradient descent.
[0109]
[0109] The second mathematical model may then be deployed as part of a medical condition diagnostic service, such as diagnostic device 104, the service of FIG. 1, etc. The second mathematical model may provide better performance than the first mathematical model because it has been trained to process utterances simultaneously rather than individually, thereby allowing the training to be better able to combine information from all utterances to generate a medical condition diagnostic score.
[0110]
[0110] Figure 10 illustrates components of one embodiment of a computing device 1000 for implementing any of the techniques described above. Although the components are illustrated in Figure 10 as being on a single computing device, the components may be distributed among multiple computing devices, such as, for example, a system of computing devices including end-user computing devices (e.g., smartphones or tablets) and / or server computing devices (e.g., cloud computing).
[0111]
[0111] The computing device 1000 may include any components typical of a computing device, such as volatile or non-volatile memory 1010, one or more processors 1011, and one or more network interfaces 1012. The computing device 1000 may also include any input and output components, such as a display, a keyboard, and a touch screen. The computing device 1000 may also include various components or modules that provide specific functions, and these components or modules may be implemented in software, hardware, or a combination thereof. Below, some examples of components are described for one exemplary implementation, and other implementations may include additional components or may exclude some of the components described below.
[0112]
[0112] The computing device 1000 may have an acoustic feature computation component 1021 that may compute acoustic features for the utterance data items as described above. The computing device 1000 may have a linguistic feature computation component 1022 that may compute linguistic features for the utterance data items as described above. The computing device 1000 may have a speech recognition component 1023 that may generate speech recognition results for the utterance data items as described above. The computing device 1000 may have a feature selection score computation component 1031 that may compute selection scores for the features as described above. The computing device 1000 may have a feature stability score computation component 1032 that may make a stability determination or compute a stability score as described above. The computing device 1000 may have a feature selection component 1033 that may select features using the selection scores and / or the stability determination as described above. The computing device 1000 may have a prompt selection score computation component 1041 that may compute selection scores for prompts as described above. The computing device 1000 may have a prompt stability score calculation component 1042 that may perform a stability determination or calculate a stability score as described above. The computing device 1000 may have a prompt selection component 1043 that may select a prompt using the selection score and / or the stability determination as described above. The computing device 1000 may have a model training component 1050 that may train a mathematical model as described above. The computing device 1000 may have a medical condition diagnosis component 1060 that may process the utterance data items to determine a medical diagnosis score as described above.
[0113]
[0113] Computing device 1000 may include or have access to various data stores, such as training corpus data store 1070. The data stores may use any known storage technology, such as files, relational or non-relational databases, or any non-transitory computer-readable medium.
[0114]
[0114] Figure 11 illustrates one embodiment of a diagnostic device 104 for the speech language model training and application technique. In an embodiment, the diagnostic device 104 may be substantially similar to one or more of the device diagnostic device 104a and / or back-end diagnostic device 104b, as described above with respect to Figure 1A. In the illustrated embodiment, the diagnostic device 104 includes a training module 1102, an ML module 1104, an assessment module 1106, and an audio module 1108, which are described in more detail below.
[0115]
[0115] In one embodiment, the training module 1102 is configured to train a first speech model in the first language. A speech model, as used herein, may refer to a machine learning model trained to analyze, predict, forecast, process, etc., speech data, e.g., audio data of a user's voice when speaking. An example of a speech model may be an x-vector embedding speech model. X-vector models may be fixed-length representations of variable-length speech segments. They are embeddings extracted from deep neural networks (DNNs) that use vectors as inputs. X-vectors are known to capture speaker characteristics even when the speaker is not seen during DNN training. However, x-vectors are just one example of a speech model that may be used. One skilled in the art will recognize other speech models that may be utilized.
[0116]
[0116] In one embodiment, the first speech model is used to determine one or more characteristics of speech data indicative of MCI or another medical condition. For example, the speech data may include voice or audio data from a user or multiple users. In the case of multiple users, at least a subset of the users has MCI or another medical condition. The training module 1102 may train the first speech model on speech data captured in or associated with a first language, e.g., English, French, German, Japanese, etc.
[0117] In such an embodiment, the first speech model may be trained to analyze sublingual characteristics of the speech data. For example, the sublingual characteristics may include acoustic features, which include any features computed from the speech data that do not involve or depend on performing speech recognition on the speech data (e.g., the acoustic features do not use information about the spoken words in the speech data). For example, the acoustic features may include mel-frequency cepstral coefficients, perceptual linear prediction features, jitter, shimmer, speech tone, speech rate, speech patterns, etc.
[0118]
[0118] The characteristics may include linguistic features, which are computed using the results of speech recognition. For example, the linguistic features may include speaking rate (e.g., number of vowels or syllables per second), number of pause fillers (e.g., "ums" and "ahs"), difficulty of the word (e.g., less common words), or part of speech of the word after the pause filler.
[0119]
[0119] In one embodiment, the training model 1102 may access a training corpus, described above with reference to Figures 3 and 4, that includes speech data for training the first speech model. For example, the training corpus may include database entries, where each entry includes an identifier for a person, a known diagnosis for the person (e.g., not MCI, has MCI, etc.), an identifier for a prompt or question presented to the person (e.g., "How are you feeling today?"), a filename for a file containing the speech data, and / or a link to a file containing the speech data.
[0120] In one embodiment, the training module 1102 is configured to further train the first speech model on the non-speech data for assessing MCI. In such an embodiment, when the assessment is performed, the non-speech data of the user is provided to the first speech model to determine the assessment of MCI.
[0121]
[0121] Non-speech data, as used herein, may refer to other data that may be captured and indicative of one or more symptoms of MCI. For example, non-speech data may include data captured from one or more sensors associated with the user, images or videos of the user, descriptive information about the user, etc. The training module 1102 may receive non-speech training data from a data store or database, a remote location, a website, etc.
[0122] In one embodiment, the non-speech data includes demographic information such as gender, age, weight, height, place of birth or residence, etc. The ML module 1104 described below may receive demographic information from a user (e.g., in response to a prompt), public records (e.g., publicly accessible records or data available online), user profiles, social media, etc.
[0123] In one embodiment, the non-speech data includes motion data, such as gait data. As used herein, gait data may refer to data that describes how a user walks, jogs, runs, etc. In one embodiment, the ML module 1104 may receive the user's gait data from motion sensors, such as accelerometers, gyroscopes, etc., e.g., data captured from the user's device.
[0124]
[0124] In one embodiment, the non-speech data includes activity data. Activity data may refer to data describing different activities a user performs throughout the day, such as brushing teeth, breakfast, exercise, eating, sleeping, etc. Activity data may include sensor data as described above, but may also include audio, image, or video data captured using audiovisual cameras, for example located in or around the user's home. The ML module 1104 may interact or communicate with those audiovisual cameras to receive the user's activity data.
[0125]
[0125] In one embodiment, the non-speech data includes driving-related data related to the user's driving history. Driving-related data may refer to data indicating that the user is authorized to drive, data describing how the user drives, how often the user drives, where the user drives, how long (distance and / or time) the user drives, driving accident information, etc. In such an embodiment, the ML module 1104 may receive driving-related data from the user (e.g., self-reported information) and may access public records such as DMV databases (e.g., via API, screen scraping, etc.) to determine whether the user has a valid driver's license, determine accident information about the user, etc. Other driving data may be captured by sensors associated with the user, such as a GPS sensor or an accelerometer, to determine the user's driving location, speed, distance, etc.
[0126]
[0126] In one embodiment, the non-speech data includes medication information. The medication information may include the type of medication the user takes, how often the user takes the medication, side effects of the medication, dosage of the medication, whether the medication requires a prescription or is available over the counter, the last time the user took the medication, etc. In one embodiment, the ML module 1104 may receive medication information from the user (e.g., self-reported), from tracked or monitored activity data (e.g., camera data or video data showing the user taking the medication).
[0127]
[0127] In one embodiment, the non-speech data includes motor function data describing one or more motor functions of the user. Motor function data, as used herein, may refer to data describing the user's ability to perform a task that requires some movement of the user's body, such as data describing the types of tasks the user can perform (e.g., changing a light bulb, changing batteries in a device, opening a jar, etc.), the amount of time it takes to complete a task, etc. The ML module 1104 may receive the user's motor function data from the user, a device associated with the user, camera content or video content, etc.
[0128]
[0128] In one embodiment, the training module 1102 is configured to train a second speech model for use in a second language (different from the first language) using at least a portion of a first speech model trained in a first language. For example, if the first speech model is trained on speech data related to English, the first speech model may be used to further train a second speech model for a different language, such as Japanese. In such an embodiment, at least a portion of the first speech model may be used to train the second speech model, may be used as a base model for the second speech model (which may be further refined or trained using speech data from the second language, for example), and so on. Even if the languages being analyzed are different, the first speech model may serve as a base model or training model for the second speech model due to similarities in sublingual characteristics between the different languages.
[0129]
[0129] In one embodiment, the training module 1102 may employ transfer learning to train a second speech model using at least a portion of the first speech model. Transfer learning, as used herein, may refer to applying knowledge of an already trained machine learning model to a different machine learning model, e.g., for a different task. For example, applying knowledge obtained from a first speech model trained in a first language to detect MCI to a second speech model trained in a different language to detect MCI. In this manner, using transfer learning allows for the efficiency of speech model training, especially when training data is unavailable or is insufficient, missing, or incomplete.
[0130] In one embodiment, the ML module 1104 is configured to apply a second speech model trained for use in the second language to the user's speech data captured in the second language. In such an embodiment, the ML module 1104 may receive speech data in the second language and provide the received speech data as input to the second speech model (trained at least in part using the first speech model for the different language).
[0131]
[0131] The ML module 1104 may receive speech data from a user, a database or other data store, etc. For example, the ML module 1104 may be in communication with a database or data store and may access a user's speech data from the database based on one or more parameters, variables, etc. (e.g., based on data ranges, type of data (e.g., speech data and / or non-speech data), etc.).
[0132] As described above with respect to the acoustic feature computation component 210, the speech recognition component 220, the linguistic feature computation component 230, and / or the medical condition classifier 240, in an embodiment, the ML module 1104 may extract one or more speech features (e.g., acoustic features and / or linguistic features) from the speech recordings (e.g., the baseline response data and / or the test case response data) and input the one or more extracted speech features to a second speech model. The second speech model may output information that the assessment module 1106 may use to assess the likelihood that the user has a medical condition, such as MCI.
[0133] In further embodiments, in addition to inputting the extracted speech features into the second speech model to diagnose a medical condition, the ML module 1104 may input other supplemental data associated with the user, such as the non-speech data described above, into the model, which the assessment module 1106 may use to diagnose a medical condition based on the results. For example, the ML module 1104 may input sensor data (e.g., along with the extracted speech features or other audio data), camera data, video data, driving-related data, medication data, activity data, demographic data, motor function data, etc., from the user's computing device 102 into the model to determine an assessment or other diagnosis of a medical condition, such as MCI, for the user.
[0134] In one embodiment, the ML module 1104 may extract one or more image features from image data (e.g., one or more images of a user, the user's face, another body part of the user associated with a medical condition, a video, etc.) from an image sensor such as a camera of the computing device 102 and may input the one or more image features into a model (e.g., along with extracted audio features, etc.). In a further embodiment, the ML module 1104 may make an assessment or other diagnosis based at least in part on touch input received from the user on a touchscreen, touchpad, etc. of the computing device 102.
[0135]
[0135] In one embodiment, the ML module 1104 may be configured to provide output using the second speech model to assess a medical condition or make another diagnosis based on one or more acoustic features of the received speech data, regardless of one or more linguistic features of the received speech data, for example, based on partial linguistic characteristics of the speech, without any linguistic features, using only one or more predefined linguistic features, without any automatic speech recognition, etc.
[0136] In one embodiment, the assessment module 1106 is configured to determine an assessment of MCI for the user based on output from a second speech model trained for use in the second language. The assessment may include a probability or likelihood that the user has MCI, a prognosis for the user, a recommended treatment, etc.
[0137]
[0137] Thus, in some embodiments, the assessment and / or diagnosis of the output of the second speech model may be independent of the language and / or dialect of the received verbal response, so that the assessment module 1106 may use acoustic features of the received speech and / or non-speech data to provide an assessment and / or diagnosis for the user in a different language. In other words, a speech model trained using data in a first language, such as English, may be used to train or refine a different speech model for a different language, such as Japanese, and to provide accurate results that may be used to assess the likelihood that a user has MCI or another medical condition.
[0138]
[0138] In one embodiment, the audio module 1108 is configured to capture and process speech data received from a user. In such an embodiment, the audio module 1108 may receive a plurality of different recorded audio clips of the user's speech and generate a single recorded audio clip of the user's speech from a plurality of short audio clips of the user's verbal response to at least one inquiry by combining the plurality of short audio clips into a single audio clip having a length that meets a threshold length, e.g., 30 seconds, 1 minute, etc. In an embodiment, the audio module 1108 may record the user's verbal response to an unprompted dialogue (or monologue), such as, for example, a general discussion between two or more parties where the dialogue is not obviously prompting speech (which may be captured using sensors on the user's device, IoT device, etc.).
[0139] In such embodiments, the audio module 1108 questions and / or queries the user with multiple questions, prompts, requests, etc. to elicit multiple different responses. In an embodiment, the audio module 1108 may audibly and / or verbally ask the user questions (e.g., using a speaker of the computing device 102, such as an integrated speaker, headphones, Bluetooth speakers, or headphones). For example, due to certain underlying medical conditions such as MCI, the user may have difficulty reading the questions and / or prompts, and asking the user audibly may simplify and / or speed up the diagnosis.
[0140] In further embodiments, audio module 1108 may display one or more questions and / or other prompts to the user (e.g., on an electronic display screen of computing device 102), another user (e.g., a caretaker, parent, medical professional, administrator, etc.) may read the one or more questions and / or other prompts to the user, etc. In various embodiments, the one or more questions or prompts may be selected, such as described above with respect to prompt selection component 1043, to facilitate diagnosis of one or more medical conditions.
[0141] In an embodiment, multiple audio modules 1108 located on multiple different computing devices 102 may query and / or ask questions of multiple different users. For example, multiple distributed audio modules 1108 may collect voice samples for clinical trials to train speech models to diagnose medical conditions, to collect test data to facilitate prompt selection, etc.
[0142] In one embodiment, the audio module 1108 queries and / or otherwise queries the user at predefined health states, such as known healthy states, predefined stages of a medical condition, etc., to collect one or more baseline audio recordings, training data, or other data. In an embodiment, the audio module 1108 queries and / or otherwise queries the user in response to a potential medical event or other trigger.
[0143]
[0143] The audio module 1108 may query the user to collect an audio recording or other test case data for a test case in response to the user requesting a medical evaluation based on data from a sensor of the computing device 102, such as a wearable device or mobile device, and / or based on receiving another trigger indicating that an injury may have occurred, that one or more symptoms of an illness have been detected, etc.
[0144]
[0144] In one embodiment, the audio module 1108 may ask and / or prompt the user regarding general information, tasks the user is performing, activities in which the user has been involved, the user's daily activities, etc., to elicit responses from the user that can be recorded and analyzed.
[0145]
[0145] For example, the audio module 1108 may audibly and / or textually ask the user questions such as "What did you do today?", "What did you eat today?", "Did you drive today?", "Did you take your medicine today?", "What medications are you taking?", "Are you having trouble with daily tasks?", "What day is it today?", "What day of the week is it today?", "What year is it?", "What time is it?", audibly list words and / or numbers to the user and ask the user to repeat them, display a series of images to the user and ask the user to repeat a description of the series of images, etc.
[0146] In one embodiment, audio module 1108 is configured to receive response data (e.g., voice data of verbal responses) in response to one or more questions and / or other prompts from audio module 1108. For example, in an embodiment, audio module 1108 may use a microphone of computing device 102, such as a smartphone, to record a user's verbal responses (e.g., answers) to one or more questions or other prompts from audio module 1108.
[0147]
[0147] In one embodiment, the audio module 1108 may present a prompt to the user to elicit a longer, free-form response. For example, the prompt may be "Tell me about your day," in which case the user is expected to give a longer response than a one-word answer or a yes / no answer. In one embodiment, when the audio module 1108 detects that the user has finished making a response, the audio module 1108 may present one or more follow-up prompts or questions to continue the conversation. In this manner, the audio module 1108 may receive conversational speech data from the user that includes or indicates various linguistic and sub-linguistic characteristics. The audio module 1108 may process and analyze the speech data to select snippets from the different responses and combine the snippets into a single audio clip.
[0148] In one embodiment, the audio module 1108 may store the received response data, such as an audio recording, on a computer-readable storage medium of the computing device 102, 110 such that the detection module 1106 may access and / or process the received response data to provide data to a speech model for diagnosing and / or assessing a medical condition, for training a speech model for diagnosing and / or assessing a medical condition, etc. In another embodiment, the audio module 1108 may provide the received response data directly (e.g., without otherwise storing the data, without temporarily storing and / or caching the data, etc.) to the detection module 1106 for diagnosing and / or assessing a medical condition.
[0149]
[0149] In one embodiment, the audio module 1108 is configured to remove audio clips from the user that are shorter than a predefined or threshold length. For example, the audio module 1108 may remove responses that are one word, shorter than a threshold length (e.g., 10 words), etc. In further embodiments, the audio module 1108 may remove responses that are shorter than a predefined or threshold duration, e.g., shorter than 3 seconds, etc. In this manner, the audio module 1108 may filter out outliers that may not contain substantial data that contributes to training a speech model or that are otherwise useful for providing an assessment of a medical condition, e.g., MCI.
[0150]
[0150] In one embodiment, the audio module 1108 is configured to remove audio clips from the plurality of short audio clips that do not exhibit predefined speech characteristics. In one embodiment, the predefined speech characteristics may include number of words, number of syllables, pronunciation, tone, acoustic speech features, prosodic speech features, linguistic speech features, signal-to-noise ratio, length of speech, etc. In such an embodiment, the audio module 1108 flags certain audio clips or snippets for removal from the set of short audio clips, such that the flagged audio clips are not included in the combined audio clip if they do not exhibit certain speech features, e.g., sublingual features, that are useful for assessing whether a user has a medical condition, such as MCI.
[0151]
[0151] In one exemplary embodiment, the audio module 1108 may capture verbal or spoken responses from a user during a telephone conversation, telephone survey, etc. During a telephone survey, a caller may use a predefined script that includes questionnaires and example questions to elicit free speech from the participant. The free speech questions may include aspects of daily life, shopping, housework, work, etc.
[0152]
[0152] In one embodiment, the audio module 1108 captures audio clips or files in a format, such as uncompressed WAV format, to eliminate the potential adverse effects of compression on the model's final predictions. Because the audio is captured by conversation over a telephone, including a potentially bandwidth-limited landline, Bluetooth, or recording device, the audio module 1108 resamples the recording to ensure a consistent sampling rate, for example, resampling to 8 kHz.
[0153]
[0153] In one embodiment, the audio module 1108 performs initial segmentation by checking pauses and speech duration. Additionally, the audio module 1108 filters out audio segments that are not indicative of a medical condition, e.g., MCI. In one embodiment, the audio module 1108 measures the duration of voice activity by excluding pauses between words and summing the duration of speech segments as determined by an automatic speech recognition (ASR) system. The audio module 1108 may remove short audio samples based on the duration of voice activity. The audio module 1108 concatenates the resulting audio clips / files into one audio sample per call.
[0154] In one embodiment, the training module 1102 and / or the ML module 1104 convert the variable length audio input signal into a fixed size representation, e.g., a feature vector embodied as an x-vector. Based on that representation, in one embodiment, the training module 1102 builds models at two different levels, a first level that is a per-speaker model, and a second level that is a per-segment model. For the per-speaker model, given a segment of (concatenated) audio for each speaker, a single feature vector is generated. In one embodiment, a fully connected DNN architecture may be applied. For the per-segment model, the audio input for each speaker is separated into a set of fixed-length segments. Based on the findings of preliminary experiments, the segment length is fixed at 5 seconds and is created every 2.5 seconds. Each segment is treated as an independent sample and the model is trained to return the corresponding label.
[0155]
[0155] In one embodiment, for each model type, the final layer is a probability layer that returns two probability values, one for the normal class indicator and one for the MCI class indicator (both totaling 1). To determine the final prediction for a given speaker in the per-segment model, in one embodiment, the assessment module 1106 averages the output prediction probabilities of the MCI class for all segments that are higher than the 5th percentile and lower than the 95th percentile for the speaker. If the average is greater than or equal to 0.51, the final prediction is determined to be MCI. In the per-speaker model, the final predicted label for a given speaker is determined by averaging the predicted outputs of each segment.
[0156] 12 illustrates one embodiment of a method 1200 for a technique for training and applying a spoken language model. In one embodiment, the method 1200 is performed by a computing device 102, a diagnostic device 104, a device diagnostic device 104a, a back-end diagnostic device 104b, a training module 1102, an ML module 1104, an assessment module 1106, an audio module 1108, a mobile computing device 102, a back-end server computing device 108, or the like.
[0157] In one embodiment, a method 1200 begins by training 1202 a first speech model in a first language, where the first speech model is used to determine one or more characteristics of speech data indicative of MCI. In one embodiment, the method 1200 trains 1204 a second speech model for use in a second language using at least a portion of the first speech model trained in the first language.
[0158] In one embodiment, the method 1200 applies 1206 a second speech model trained for use in the second language to the user's speech data captured in the second language. In one embodiment, the method 1200 determines 1208 an assessment of MCI for the user based on output from the second speech model trained for use in the second language, and the method 1200 ends.
[0159] 13 illustrates one embodiment of a method 1300 for a technique for training and applying a spoken language model. In one embodiment, the method 1300 is performed by a computing device 102, a diagnostic device 104, a device diagnostic device 104a, a back-end diagnostic device 104b, a training module 1102, an ML module 1104, an assessment module 1106, an audio module 1108, a mobile computing device 102, a back-end server computing device 108, or the like.
[0160] In one embodiment, a method 1300 begins by training 1302 a first speech model in a first language, where the first speech model is used to determine one or more characteristics of speech data indicative of MCI. In one embodiment, the method 1300 also trains 1304 a second speech model for use in a second language using at least a portion of the first speech model trained in the first language, e.g., by transfer learning.
[0161] In one embodiment, the method 1300 prompts the user for a verbal response in the second language (1306), captures multiple audio responses from the user, e.g., multiple audio clips (1308), and filters the audio responses to create a single audio clip (1310). In such an embodiment, the method 1300 may filter out audio clips that are not of a threshold direction or length, do not contain desired linguistic characteristics, etc.
[0162] In one embodiment, the method 1300 applies 1312 a second speech model trained for use in the second language to the user's speech data captured in the second language. In one embodiment, the method 1300 determines 1314 an assessment of MCI for the user based on output from the second speech model trained for use in the second language, and the method 1300 ends.
[0163]
[0163] The means for training a first speech model in a first language used to determine one or more characteristics of speech data indicative of MCI may, in various embodiments, include a diagnostic device 104, a device diagnostic device 104a, a back-end diagnostic device 104b, a training module 1102, a mobile computing device 102, a back-end server computing device 108, electronic speakers of the computing devices 102, 108, headphones, electronic display screens of the computing devices 102, 108, user interface devices, network interfaces, mobile applications, processors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), programmable logic, other logic hardware, and / or other executable program code stored on a non-transitory computer-readable storage medium. Other embodiments may include substantially similar or equivalent means for training a first speech model in a first language.
[0164]
[0164] The means for training a second speech model for use in a second language using at least a portion of a first speech model trained in a first language may, in various embodiments, include the diagnostic device 104, the device diagnostic device 104a, the back-end diagnostic device 104b, the training module 1102, the mobile computing device 102, the back-end server computing device 108, electronic speakers of the computing devices 102, 108, headphones, electronic display screens of the computing devices 102, 108, user interface devices, network interfaces, mobile applications, processors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), programmable logic, other logic hardware, and / or other executable program code stored on a non-transitory computer-readable storage medium. Other embodiments may include substantially similar or equivalent means for training a second speech model in a second language.
[0165]
[0165] The means for applying the second speech model trained for use in the second language to the user's speech data captured in the second language may, in various embodiments, include the diagnostic device 104, the device diagnostic device 104a, the back-end diagnostic device 104b, the ML module 1104, the mobile computing device 102, the back-end server computing device 108, a mobile application, machine learning, artificial intelligence, an acoustic feature computation component 210, a speech recognition component 220, a Gaussian mixture model, an acoustic model, a language model, a neural network, a deep neural network, a medical condition classifier 240, a classifier, a support vector machine, a multi-layer perceptron, a processor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), programmable logic, other logic hardware, and / or other executable program code stored on a non-transitory computer-readable storage medium. Other embodiments may include substantially similar or equivalent means for applying the second speech model.
[0166]
[0166] The means for determining an assessment of MCI for a user based on output from a second speech model trained for use in a second language may, in various embodiments, include a diagnostic device 104, a device diagnostic device 104a, a back-end diagnostic device 104b, an assessment module 1106, a mobile computing device 102, a back-end server computing device 108, a mobile application, a processor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), programmable logic, other logic hardware, and / or other executable program code stored on a non-transitory computer-readable storage medium. Other embodiments may include substantially similar or equivalent means for determining an assessment of MCI.
[0167]
[0167] The means for creating a recorded audio clip from a plurality of short audio clips may, in various embodiments, include the diagnostic device 104, the device diagnostic device 104a, the back-end diagnostic device 104b, the audio module 1108, the mobile computing device 102, the back-end server computing device 108, a mobile application, a processor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), programmable logic, other logic hardware, and / or other executable program code stored on a non-transitory computer-readable storage medium. Other embodiments may include substantially similar or equivalent means for creating a recorded audio clip from a plurality of short audio clips.
[0168]
[0168] The methods and systems described herein may be deployed in part or in whole through a machine that executes computer software, program code, and / or instructions on a processor. As used herein, "processor" is meant to include at least one processor, and the plural and singular forms should be understood to be interchangeable unless the context clearly indicates otherwise. Any aspect of the present disclosure may be implemented as a method on a machine, as part of a machine, as a system or apparatus related to a machine, or as a computer program product embodied in a computer-readable medium that executes on one or more of the machines. The processor may be part of a server, a client, a network infrastructure, a mobile computing platform, a stationary computing platform, or other computing platform.
[0169]
[0169] A processor may be any type of computing or processing device capable of executing program instructions, codes, binary instructions, and the like. A processor may be or include a signal processor, digital processor, embedded processor, microprocessor, or any variation such as a coprocessor (mathematical coprocessor, graphic coprocessor, communication coprocessor, and the like) that may directly or indirectly facilitate the execution of program codes or program instructions stored thereon. In addition, a processor may enable the execution of multiple programs, threads, and codes. Threads may be executed simultaneously to improve the performance of the processor and to facilitate the simultaneous operation of applications. In embodiments, the methods, program codes, program instructions, and the like described herein may be implemented in one or more threads. A thread may give rise to other threads that may be assigned priorities associated with the threads. The processor may execute these threads based on priorities or any other order based on instructions provided in the program code. The processor may include a memory that stores the methods, codes, instructions, and programs as described herein and elsewhere. The processor may access a storage medium through an interface that may store the methods, codes, and instructions as described herein and elsewhere. A storage medium associated with a processor for storing methods, programs, code, program instructions, or other types of instructions executable by a computing or processing device may include, but may not be limited to, one or more of a CD-ROM, a DVD, memory, a hard disk, a flash drive, RAM, ROM, cache, etc.
[0170]
[0170] A processor may include one or more cores, which may increase the speed and performance of a multiprocessor. In embodiments, a processor may be a dual-core processor, a quad-core processor, other chip-level multiprocessor that combines two or more independent cores (called a die), etc.
[0171]
[0171] The methods and systems described herein may be deployed in part or in whole through machines executing computer software on servers, clients, firewalls, gateways, hubs, routers, or other such computers and / or networking hardware. The software programs may be associated with servers, which may include file servers, print servers, domain servers, Internet servers, intranet servers, and other variations such as secondary servers, host servers, distributed servers, etc. The servers may include one or more of memory, processors, computer-readable media, storage media, ports (physical and virtual), communication devices, and interfaces accessible through wired or wireless media, etc. to other servers, clients, machines, and devices. The methods, programs, or codes as described herein and elsewhere may be executed by the servers. In addition, other devices necessary for the execution of the methods described herein may be considered part of the infrastructure associated with the servers.
[0172]
[0172] A server may provide an interface to other devices, including but not limited to clients, other servers, printers, database servers, print servers, file servers, communication servers, distributed servers, and the like. Additionally, this coupling and / or connection may facilitate remote execution of a program across a network. The networking of some or all of these devices may facilitate parallel processing of a program or method in one or more locations without departing from the scope of this disclosure. In addition, any of the devices attached to the server through an interface may include at least one storage medium capable of storing methods, programs, code, and / or instructions. A central repository may provide program instructions to be executed on different devices. In this embodiment, the remote repository may function as a storage medium for program code, instructions, and programs.
[0173]
[0173] A software program may be associated with a client, which may include a file client, a print client, a domain client, an internet client, an intranet client, and other variations such as a secondary client, a host client, a distributed client, etc. A client may include one or more of a memory, a processor, a computer-readable medium, a storage medium, a port (physical and virtual), a communication device, and an interface accessible to other clients, servers, machines, and devices through a wired or wireless medium, etc. Methods, programs, or codes as described herein and elsewhere may be executed by a client. In addition, other devices necessary for the execution of the methods described herein may be considered part of the infrastructure associated with the client.
[0174]
[0174] A client may provide an interface to other devices, including but not limited to servers, other clients, printers, database servers, print servers, file servers, communication servers, distributed servers, and the like. Additionally, this coupling and / or connection may facilitate remote execution of a program across a network. The networking of some or all of these devices may facilitate parallel processing of a program or method in one or more locations without departing from the scope of this disclosure. In addition, any of the devices attached to a client through an interface may include at least one storage medium capable of storing methods, programs, applications, code, and / or instructions. A central repository may provide program instructions to be executed on different devices. In this embodiment, the remote repository may function as a storage medium for program code, instructions, and programs.
[0175]
[0175] The methods and systems described herein may be deployed in part or in whole through a network infrastructure. The network infrastructure may include elements such as computing devices, servers, routers, hubs, firewalls, clients, personal computers, communication devices, routing devices, and other active and passive devices, modules, and / or components known in the art. The computing devices and / or non-computing devices associated with the network infrastructure may include storage media such as flash memory, buffers, stacks, RAM, ROM, etc., apart from other components. The processes, methods, program codes, instructions described herein and elsewhere may be executed by one or more of the network infrastructure elements.
[0176]
[0176] The methods, program codes, and instructions described herein and elsewhere may be implemented on a cellular network having multiple cells. The cellular network may be either a Frequency Division Multiple Access (FDMA) network or a Code Division Multiple Access (CDMA) network. The cellular network may include mobile devices, cell sites, base stations, repeaters, antennas, towers, etc. The cell network may be a GSM, GPRS, 3G, EVDO, mesh, or other network type.
[0177]
[0177] The methods, program codes, and instructions described herein and elsewhere may be implemented on or through a mobile device. The mobile device may include a navigation device, a cell phone, a mobile personal digital assistant, a laptop, a palmtop, a netbook, a pager, an e-book reader, a music player, and the like. These devices may include storage media, such as flash memory, a buffer, a RAM, a ROM, and one or more computing devices, apart from other components. The computing device associated with the mobile device may be enabled to execute the program code, the method, and the instructions stored thereon. Alternatively, the mobile device may be configured to execute instructions in cooperation with other devices. The mobile device may communicate with a base station interfaced with a server and configured to execute the program code. The mobile device may communicate over a peer-to-peer network, a mesh network, or other communication network. The program code may be stored on a storage medium associated with the server and executed by a computing device embedded within the server. The base station may include a computing device and a storage medium. The storage device may store the program code and instructions executed by the computing device associated with the base station.
[0178]
[0178] The computer software, program code, and / or instructions may be stored and / or accessed on machine-readable media, which may include computer components, devices, and recording media that hold for a period of time digital data used in computations, semiconductor storage known as random access memory (RAM), mass storage for further permanent storage typically in the form of magnetic storage such as optical disks, hard disks, tapes, drums, cards, and other types, processor registers, cache memory, volatile memory, non-volatile memory, optical storage such as CDs, DVDs, flash memory (e.g., USB sticks or keys), removable media such as floppy disks, magnetic tape, paper tape, punch cards, standalone RAM disks, Zip drives, removable mass storage, offline, dynamic memory, static memory, read / write storage, mutable storage, read only, random access, sequential access, addressable location, addressable file, addressable content, network attached storage, storage area networks, bar codes, magnetic ink, and other computer memory.
[0179]
[0179] The methods and systems described herein may transform physical and / or intangible items from one state to another. The methods and systems described herein may also transform data representing physical and / or intangible items from one state to another.
[0180]
[0180] References throughout this specification to "one embodiment," "an embodiment," or similar words mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment. Thus, appearances of the phrases "in one embodiment," "in an embodiment," and similar words throughout this specification may all refer to the same embodiment, but not necessarily, and may mean "one or more, but not all, embodiments," unless expressly specified otherwise. The terms "including," "comprising," and "having," and variations thereof, mean "including but not limited to," unless expressly specified otherwise. An enumerated list of items does not imply that any or all of the items are mutually exclusive and / or mutually inclusive, unless expressly specified otherwise. The terms "a," "an," and "the" also refer to "one or more," unless expressly specified otherwise.
[0181]
[0181] Furthermore, the described features, advantages, and characteristics of the embodiments may be combined in any suitable manner. Those skilled in the art will recognize that an embodiment may be practiced without one or more of the specific features or advantages of a particular embodiment. In other instances, additional features and advantages may be recognized in certain embodiments that may not be present in all embodiments.
[0182]
[0182] These features and advantages of the embodiments will become more fully apparent from the following description and appended claims, or may be learned by practice of the embodiments as set forth hereinafter. As will be appreciated by one of ordinary skill in the art, aspects of the present invention may be embodied as a system, method, and / or computer program product. Thus, aspects of the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, which may all be generally referred to herein as a "circuit," "module," or "system." Furthermore, aspects of the present invention may take the form of a computer program product embodied in one or more computer-readable medium(s) having program code embodied thereon.
[0183]
[0183] Many of the functional units described herein have been labeled as modules (or components) to more particularly emphasize their implementation independence. For example, a module may be implemented as a hardware circuit comprising custom VLSI circuits or gate arrays, off-the-shelf semiconductors such as logic chips, transistors, or other discrete components. A module may also be implemented in programmable hardware devices such as field programmable gate arrays, programmable array logic, programmable logic devices, etc.
[0184]
[0184] Modules may also be implemented in software for execution by various types of processors. An identified module of program code may comprise one or more physical or logical blocks of computer instructions, which may be organized as, for example, an object, procedure, or function. Nevertheless, an executable file of an identified module may comprise disparate instructions stored in different locations that may not be physically located together, but which, when logically joined, comprise the module and achieve the stated purpose of the module.
[0185]
[0185] In practice, a module of program code may be a single instruction or many instructions, and may be distributed across several different code segments, among different programs, and across several memory devices. Similarly, operational data may be identified and illustrated herein within a module, and may be embodied in any suitable form and organized within any suitable type of data structure. The operational data may be collected as a single data set, or may be distributed across different locations, including across different storage devices, or may simply exist at least in part as an electronic signal over a system or network. If a module or part of a module is implemented in software, the program code may be stored in and / or propagated across one or more computer-readable media.
[0186]
[0186] A computer program product may include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present invention.
[0187]
[0187] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memory ("RAM"), read-only memory ("ROM"), erasable programmable read-only memory ("EPROM" or flash memory), static random access memory ("SRAM"), portable compact disk read-only memory ("CD-ROM"), digital versatile disk ("DVD"), memory sticks, floppy disks, punch cards or mechanically encoded devices such as ridge structures in grooves that allow instructions to be recorded thereon, and any suitable combination of the foregoing. A computer-readable storage medium, as used herein, should not be construed as a transitory signal per se, such as electric waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through a fiber optic cable), or electrical signals transmitted through electrical wires.
[0188]
[0188] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to the respective computing / processing device or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical transmission fiber, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in a computer-readable storage medium in the respective computing / processing device.
[0189]
[0189] The computer readable program instructions for carrying out the operations of the present invention may be either assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state setting data, or source or object code written in any combination of one or more programming languages, including object oriented programming languages such as Smalltalk, C++, and traditional procedural programming languages such as the "C" programming language or similar programming languages. The computer readable program instructions may execute completely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or completely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (e.g., through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute computer readable program instructions by utilizing state information of the computer readable program instructions to individualize the electronic circuitry to perform aspects of the present invention.
[0190] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.
[0191]
[0191] These computer-readable program instructions may be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing apparatus create means for performing the functions / operations specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions may also be stored on a computer-readable storage medium that may direct a computer, programmable data processing apparatus, and / or other device to function in a particular manner, such that the computer-readable storage medium on which the instructions are stored comprises an article of manufacture including instructions that perform aspects of the functions / operations specified in one or more blocks of the flowcharts and / or block diagrams.
[0192]
[0192] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device that causes the computer, other programmable apparatus, or other device to perform a series of operational steps to create a computer-implemented process, such that the instructions executing on the computer, other programmable apparatus, or other device perform the functions / operations specified in one or more blocks of the flowcharts and / or block diagrams.
[0193]
[0193] The schematic flow chart diagrams and / or schematic block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of apparatus, systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the schematic flow chart diagrams and / or schematic block diagrams may represent a module, segment, or portion of code that includes one or more executable instructions of a program code that implements a specified logical function(s).
[0194]
[0194] It should also be noted that in some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may in fact be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved. Other steps and methods may be conceived that are equivalent in function, logic, or effect to one or more blocks, or portions thereof, of the illustrated figures.
[0195]
[0195] Although various arrow types and line types may be employed in the flowcharts and / or block diagrams, it is understood that they are not intended to limit the scope of the corresponding embodiments. In fact, some arrows or other connectors may be used to indicate only the logical flow of the illustrated embodiment. For example, the arrows may indicate waiting or monitoring periods of unspecified duration between the recited steps of the illustrated embodiment. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or a combination of dedicated hardware and program code.
[0196]
[0196] The present invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described embodiments are to be considered in all respects only as illustrative and not restrictive. The scope of the invention is, therefore, indicated by the appended claims rather than by the foregoing description. All changes that come within the meaning and range of equivalency of the claims are intended to be embraced within their scope. [Explanation of symbols]
[0197] 100 Systems 102 Hardware devices, computing devices 104 Diagnostic Equipment 104a Device diagnostic equipment 104b Back-end diagnostic equipment 106 Data Network 108 Backend Server 109 System 110 Mobile Devices 120 Healthcare Provider Computers 130 Network 140 Medical diagnosis services 200 Systems 210 Acoustic Feature Calculation Component 220 Speech Recognition Component 230 Linguistic Feature Computation Component 240 Disease Classifier 500 Systems 510 Training Corpus 520 Feature Selection Score Calculation Component 530 Feature Stability Judgment Component 540 Feature Selection Component 1000 Computing Devices 1010 Memory 1011 Processor 1012 Network Interface 1021 Acoustic Feature Calculation Component 1022 Linguistic Feature Computation Component 1023 Speech Recognition Component 1031 Feature Selection Score Calculation Component 1032 Feature Stability Score Calculation Component 1033 Feature Selection Component 1041 Prompt Selection Score Calculation Component 1042 Prompt Stability Score Calculation Component 1043 Prompt Selection Component 1050 Model Training Components 1060 Medical Condition Diagnostic Component 1070 Training Corpus Data Store 1102 Training Module 1104 ML module 1106 Assessment Module 1108 Audio Module
Claims
1. A processor; The apparatus, coupled to the processor, training a first speech model in a first language, the first speech model being used to determine one or more characteristics of speech data indicative of mild cognitive impairment (MCI); training a second speech model for use in a second language using at least a portion of the first speech model trained in the first language; applying the second speech model trained for use in the second language to user speech data captured in the second language; and determining an assessment of MCI for the user based on output from the second speech model trained for use in the second language; a memory storing code executable by the processor to cause the processor to An apparatus comprising:
2. The apparatus of claim 1 , wherein the first speech model trained in the first language is configured to analyze sub-linguistic characteristics of the speech data.
3. The apparatus of claim 2 , wherein the sub-language characteristics include at least one of speech tone, speech rate, speech pattern, acoustic speech characteristics, prosodic speech characteristics, and linguistic speech features.
4. 2. The apparatus of claim 1, wherein the first speech model trained in the first language is further trained to analyze non-speech data, and the non-speech data of the user is further provided to the first speech model trained in the first language to determine the assessment of MCI for the user.
5. The apparatus of claim 4 , wherein the non-speech data includes demographic information about the user.
6. The apparatus of claim 4 , wherein the non-speech data includes gait data, the gait data describing a manner in which the user walks.
7. The device of claim 4 , wherein the non-speech data includes activity data, the activity data captured from one or more sensors associated with the user.
8. The apparatus of claim 4 , wherein the non-speech data comprises driving-related data relating to the user's driving history.
9. The apparatus of claim 4 , wherein the non-speech data includes medication information for the user.
10. The apparatus of claim 4 , wherein the non-speech data includes data describing one or more motor functions for the user.
11. The apparatus of claim 1 , wherein the speech data comprises recorded audio clips of verbal responses by the user in response to at least one of at least one query and an unprompted interaction.
12. 12. The apparatus of claim 11, wherein the code is further executable by the processor to create the recorded audio clip from a plurality of short audio clips of the user's verbal response to the at least one query by combining the plurality of short audio clips into a single audio clip having a length that meets a threshold length.
13. The apparatus of claim 12 , wherein the code is further executable by the processor to present a plurality of prompts to the user to elicit the plurality of short audio clips.
14. The apparatus of claim 12 , wherein the plurality of audio clips comprises snippets taken from a longer conversation exhibiting various predefined linguistic characteristics.
15. The apparatus of claim 12 , wherein the code is further executable by the processor to remove audio clips of the plurality of short audio clips that are shorter than a threshold length.
16. The apparatus of claim 12 , wherein the code is further executable by the processor to remove audio clips from the plurality of short audio clips that do not exhibit predefined speech characteristics.
17. 17. The apparatus of claim 16, wherein the predefined speech characteristics include at least one of a number of words, a number of syllables, pronunciation, tone, signal-to-noise ratio, and length of speech.
18. The apparatus of claim 1 , wherein the first speech model, the second speech model, or a combination of the first speech model and the second speech model comprises an x-vector embedded speech model.
19. training a first speech model in a first language, the first speech model being used to determine one or more characteristics of speech data indicative of mild cognitive impairment (MCI); training a second speech model for use in a second language using at least a portion of the first speech model trained in the first language; applying the second speech model trained for use in the second language to user speech data captured in the second language; determining an assessment of MCI for the user based on output from the second speech model trained for use in the second language; A method comprising:
20. means for training a first speech model in a first language, the first speech model being used to determine one or more characteristics of speech data indicative of mild cognitive impairment (MCI); means for training a second speech model for use in a second language using at least a portion of the first speech model trained in the first language; means for applying the second speech model trained for use in the second language to user speech data captured in the second language; means for determining an assessment of MCI for the user based on output from the second speech model trained for use in the second language; An apparatus comprising:
Citation Information
Patent Citations
Emotion-identifying apparatus, emotion-identifying method, and emotion-identifying program
JP2018097292A
Determination method, program, and determination system
JP2023096356A
System and method for cross-language speech impairment detection
US20230147895A1
Selecting speech features for building models for detecting medical conditions
US10152988B2
Medical assessment based on voice
US10311980B2