Acoustic and natural language processing models for velocity-based screening and monitoring of behavioral health

An improved acoustic model with an encoder, decoder, and classifier system, trained on transcribed speech data, addresses the challenge of accurately predicting behavioral and mental health conditions by integrating segment fusion and natural language processing, achieving enhanced prediction accuracy and reduced data requirements.

JP2026136266APending Publication Date: 2026-08-25ELLIPSIS HEALTH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026087566
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-05-19
Filing Date
2026-05-25
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing technologies face challenges in accurately diagnosing behavioral and mental health conditions, particularly due to the high cost and the possibility of undiagnosed individuals, and existing acoustic models have limited effectiveness in predicting these conditions.

Method used

An improved acoustic model using an encoder and decoder system trained on transcribed speech data unrelated to behavioral or mental health, combined with a classifier, reduces the need for labeled training data and enhances prediction accuracy by employing segment fusion and natural language processing models to integrate linguistic and acoustic features.

Benefits of technology

The proposed model achieves improved prediction accuracy for behavioral and mental health conditions, with area under the curve (AUC) ranging from 0.75 to 0.79, specificity and sensitivity of 0.68, and the ability to process single audio segments effectively, reducing the reliance on extensive labeled data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026136266000001_ABST
    Figure 2026136266000001_ABST
Patent Text Reader

Abstract

This invention provides an acoustic natural language processing model for predicting whether a person has a behavioral or mental health condition based on input speech. [Solution] A method for detecting behavioral or mental health status using an acoustic model including an encoder and a classifier, comprising: (a) acquiring a speech sample comprising a plurality of speech segments; (b) processing the speech sample with an encoder to generate an abstract feature representation, wherein the encoder is pre-trained to perform a first task other than detecting behavioral or mental health status; and (c) processing the abstract feature representation with a classifier to generate an output indicating whether or not the person has a behavioral or mental health status, wherein the classifier is trained on a training dataset comprising a plurality of speech samples from a plurality of speakers, and the speech samples are labeled as originating from or not originating from a speaker with a health status.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross-reference This application claims priority to U.S. Provisional Patent Application No. 62 / 926,245, filed Oct. 25, 2019; U.S. Provisional Patent Application No. 63 / 021,617, filed May 7, 2020; U.S. Provisional Patent Application No. 63 / 021,625, filed May 7, 2020; U.S. Provisional Patent Application No. 63 / 027,238, filed May 19, 2020; and U.S. Provisional Patent Application No. 63 / 027,240, filed May 19, 2020, the entire contents of each of which are incorporated herein by reference.

Background Art

[0002] Behavioral and mental health conditions are common in populations and can be costly to society. While there are treatments available for such conditions, there is a possibility that multiple individuals may not be diagnosed.

Summary of the Invention

Means for Solving the Problems

[0003] The present disclosure provides an improved acoustic model for use in predicting whether a subject has a behavioral or mental health condition of interest. The present disclosure also provides a method for training such a model. The acoustic model described herein may have an automatic speech recognition (“ASR”) system. The ASR system may have an encoder and a decoder. The encoder and decoder may be trained on transcribed speech data that is unrelated to behavioral or mental health. ​​​​​​​​​​​​​​​​​​​​​​​​The acoustic model may also have a classifier. After the ASR system is trained, the decoder The classifier may be discarded, and the classifier will determine that the subject has a behavioral or mental health condition. Audio data labeled as originating from or not originating from the subject The encoder may be trained on the classifier, or frozen. This training scheme is necessary to train this type of acoustic model. This can reduce the amount of training data related to behavioral or mental health. Furthermore, the end-to-end acoustic model described herein is more effective than existing acoustic models for target audiences. It is possible to more accurately predict whether a person has a behavioral or mental health condition of interest. In particular, when predicting whether a patient has depression, the end products described herein The end acoustic model has an area under the curve of 0.75 to 0.79 ("area-under-t"). It has a he-curve (AUC), a specificity of 0.68, and a sensitivity of 0.68. This has been demonstrated. Existing i-vector and convolutional neural networks ("convo The CNN (Central Neural Network) model uses AUC, a specific type of neural network. Sex and sensitivity are 0.60, 0.58, and 0.58, and 0.64, 0.60 respectively. , and only 0.60.

[0004] In addition to encoders and classifiers, acoustic models also include segment fusion models. The system can process a single audio segment at a time. Segment fusion is By combining information from segment-level output, session-level acoustic modeling is achieved. It can output a Delscore. The acoustic system generates each segment by the classifier. The average of all predictions for the ment can be calculated. A more complex version is a classifier. Using several representations of segments generated by the module, then these Other machine learning methods can be used to compute the final prediction from the input. The methods include LSTM, RCNN, and multi-layer perceptron. This may include rceptrons (MLPs), random forests, and other models. More complex combination methods for meaning yield larger gains (e.g., AUC 0.79 vs This can result in 0.75). Even using more complex methods, the underlying segment can be modified. The ring cannot be changed, and the gain is purer due to better fusion of the segment outputs. It can be obtained.

[0005] This disclosure also relates to natural language processing. Using the "ing:NLP" model, we determine whether the subject has a behavioral or mental health condition. This document provides a system and method for predicting the outcome. The NLP model described herein is It may have an encoder, a language model, and one or more classifiers. The encoder is used for the target audience. Receive an audio sample transcribed from and encode the audio sample (e.g., real numbers) It can generate vectors. The language model and classifier can generate encoded speech samples. The sample is processed to generate a prediction indicating whether the subject has a behavioral or mental health condition. It is possible. Language models are initially not necessarily related to behavioral or mental health conditions. It may be trained on encoding-based representations that do not use this method. For example, a language model can be trained on Wikipedia. It may be trained on a corpus of edia articles. The language model can then be behavioral or mental. can be fine-tuned with encoded text related to the health status. Thereafter, one or more classifiers can be trained to predict whether a subject has a behavioral or mental health condition . The training data for the classifier can include multiple transcripts and encoded audio samples from multiple subjects. Each audio sample can be associated with a label indicating whether the subject who provided the audio sample has a behavioral or mental health condition.

[0006] The above training process can provide several improvements in the technical field of automated mental health detection . The use of general and domain-specific text corpora to pre-train and determine the language model can reduce the number of labeled audio samples required to train end-to-end NLP. Furthermore, the pre-trained and fine-tuned language model can be used in different end-to-end NLP models that detect different behavioral or mental health conditions . The reuse of such language models for multiple tasks can further reduce the training time.

[0007] The above-described acoustic model and NLP model can be fused with each other to generate a more robust composite model .

[0008] In one aspect, the present disclosure provides a method for detecting a behavioral or mental health condition in a subject using an acoustic model including an encoder and a classifier, the method comprising: (a) obtaining an audio sample including a plurality of audio segments from the subject, (b) processing the audio sample with an encoder to generate an abstract feature representation of the audio sample, wherein: (a) obtaining an audio sample including a plurality of audio segments from the subject; (b) processing the audio sample with an encoder to generate an abstract feature representation of the audio sample; wherein: (a) obtaining an audio sample including a plurality of audio segments from the subject; (b) processing the audio sample with an encoder to generate an abstract feature representation of the audio sample; wherein: (a) obtaining an audio sample including a plurality of audio segments from the subject; (b) processing the audio sample with an encoder to generate an abstract feature representation of the audio sample; In addition, the encoder does not detect behavioral or mental health status in the subject. (c) an abstract feature representation, pre-trained to perform the first task, steps The data is processed by a classifier to produce an output indicating whether the subject has a behavioral or mental health condition. This is a step in which the classifier performs training data that includes multiple voice samples from multiple speakers. The dataset is trained, and the voice samples of multiple voice samples are behavioral or Labeled as originating from or not originating from a speaker with a mental health condition. The method includes the step of being erased. In some embodiments, the method includes the sound before (b). The step of converting a voice sample into a filter bank or Mel-frequency cepstrum coefficients is further It includes. In some embodiments, the classifier is a binary classifier, and the output is that the subject is behavioral Alternatively, it is a binary output indicating whether or not the person has a mental health condition. In some embodiments, The classifier is a multi-class classifier, and its output is a multi-class classification of the behavioral or mental health status of the subjects. Includes probability distributions across numerical levels or severity levels. In some embodiments, the output is relevant to the target. Includes segment output for each segment of multiple segments of the voice sample from the person. The method involves fusing segment outputs to detect the behavioral or mental health status of the subject. This further includes: In some embodiments, the first task is automatic speech recognition, speaker recognition, This is an emotion classification or a sound classification. In some embodiments, (a) is a telemedicine session This includes obtaining an audio sample. In some embodiments, (a) the subject (b) and (c) include obtaining audio samples from a mobile device. It is executed at least partially on the device. In some embodiments, (b) and ( c) is executed at least partially on a remote server. In some embodiments, this The law includes non-speech models, such as laughter models, breathing models, or pause models, for voice samples. The method further includes the step of processing the following. In some embodiments, the method is performed before (b), The process further includes a step of determining whether the audio sample meets a quality threshold.

[0009] In another aspect, the Disclosure relates to the use of acoustic sensors to detect the behavioral or mental health status of a subject. This provides a method for training Dell, and the acoustic model includes an encoder and a classifier. The law stipulates that (a) the behavioral or mental health status of subjects may be detected on the first training dataset. (b) The step of training the encoder to perform a first task other than (b) (a) Following this, indicate whether the subject has a behavioral or mental health condition of interest. To generate the output, a second training dataset different from the first training dataset is used. The next step is to train the encoder and classifier, wherein the second training dataset is multiple Includes multiple audio samples from several speakers, and the audio samples of the multiple audio samples are of interest. Derived from or not derived from a speaker having a certain behavioral or mental health condition. Steps, which are labeled as such, include the first task. In some embodiments, the first task The functions include automatic speech recognition, speaker recognition, emotion classification, or sound classification. In some embodiments, (b) processes an abstract feature representation of the audio sample from the encoder to generate output. This includes training the classifier to do so. In some embodiments, during (b), encoding The part is fixed. In some embodiments, the encoder is not fixed during (b). In some embodiments, (a) and (b) are supervised learning processes. In this implementation, the classifier is a binary classifier, and the output is the behavioral or mental health status of the subject. The output is a binary value indicating whether or not it has something. In some embodiments, the classifier performs multi-class classification. It is a device, and its output is a measure of multiple levels or severity of the behavioral or mental health status of the subject. Includes a probability distribution over a certain range. In some embodiments, the output is an audio sample from the subject. The method includes segment output for each segment of multiple segments, and the segment This further includes fusing the output to detect the behavioral or mental health status of the subject.

[0010] In another aspect, the Disclosure relates to detecting behavioral or mental health conditions in subjects. This provides a method for training an acoustic model, and this method (a) a first training data set In this context, an automated speech recognition (ASR) system is trained to transcribe audio samples. A step, wherein the ASR system comprises an encoder and a decoder, (b) a step of discarding the decoder, and (c) a second training dataset different from the first training dataset. On the training dataset, audio samples from participants are processed to determine whether the participants are behavioral or mental. To train an encoder and classifier to generate an output indicating whether or not the person has a specific health condition, The step is to have a second training dataset that has behavioral or mental health conditions. Multiple labels, labeled as originating from the speaker or not. The method includes a step that includes a deleted audio sample. In some embodiments, this method includes Before (a), filter multiple unlabeled audio samples into a filter bank or mel-circle. The process further includes the step of converting the wavenumber to cepstrum coefficients. In some embodiments, this method The law requires that, prior to (c), multiple labeled audio samples be filtered into a filter bank or mel frequency cable. The process further includes the step of converting to a Pstrum coefficient. In some embodiments, (a) is The steps involve training an encoder to generate an abstract feature representation of an audio sample, and sound The abstract feature representation of the voice sample is processed to generate a transcribed voice sample. The process includes the step of training the coder. In some embodiments, (c) generates an output. To do this, the classifier is designed to process the abstract feature representation of the audio sample from the encoder. This includes training. In some embodiments, the encoder is fixed during (c). In some embodiments, the encoder is not fixed during (c). Then, (a) and (c) are supervised learning processes. In some embodiments, this The law involves multiple labeled audio samples and multiple methods that generated multiple labeled audio samples. Steps to train a classifier on a third training dataset that includes metadata about the speaker. Further includes the following. In some embodiments, the metadata includes the age of each of the multiple speakers. Race, ethnicity, sex, gender, income, education, place of origin, or medical history Includes one or more of the above. In some embodiments, the encoder is a convolutional neural Network (convolutional neural network: CNN) and long short-term memory network It includes twork (LSTM). In some embodiments, the CNN is a visual geo Visual Geometry Group (VGG) Network In some embodiments, the classifier is a recurrent convolutional neural network. (recurrent convolutional neural network) k: RCNN), attention-enabled LSTM, self-attention network, and converters from the group Includes a selected model. In some embodiments, the classifier is a binary classifier, and the output is The output is a binary value indicating whether or not the subject has a behavioral or mental health condition. In the embodiment, the classifier is a multi-class classifier, and the output is behavioral or precise in the subject. Includes probability distributions across multiple levels or degrees of divine health. In some embodiments, The output is a sample of multiple segments from the subject, for each segment. The method includes segment outputs and merges segment outputs to assess the behavioral or mental health of the subjects. This further includes detecting the state.

[0011] In another aspect, the disclosure describes how a natural language processing (NLP) model is used to describe the actions of a subject. The method provides a way to detect dynamic or mental health states, and the NLP model is a language model and one or The method includes multiple classifiers, and the method includes (a) voice segments including multiple voice segments from the subject (b) the step of obtaining a sample, and (b) processing the speech sample or its derivative with a language model. A step of generating language model output, wherein the language model uses a first dataset and It is trained on two datasets, the first of which is used to study behavioral or mental health status. The second dataset includes unrelated text, and the second dataset is related to behavioral or mental health status. Including text, the first dataset is substantially larger than the second dataset, (c) Generate output indicating whether the subject has a behavioral or mental health condition. To that end, the process includes the step of processing the language model output using one or more classifiers. In some embodiments, the method generates a transcribed audio sample before (b). The steps involve transcribing an audio sample and then using an encoder to transfer the transcribed audio sample. The further step includes generating a pull embedding. In some embodiments, the language model Dell includes a long-term short-term memory (LSTM) network or transducer. Several embodiments Then, one or more classifiers include binary classifiers, and (c) the subject is behavioral or mental A binary classification is created to indicate whether a person has a normal health condition or not, or whether they have a behavioral or mental health condition. This includes performing. In some embodiments, one or more classifiers include regression classifiers. (c) confirms the behavioral or mental health status of the subject across multiple levels or severity levels. This includes generating a rate distribution. In some embodiments, the method generates an output. This further includes the step of fusing binary classification and probability distribution. In some embodiments, the Dataset 1 includes publicly available text corpora.

[0012] In another aspect, the Disclosure relates to natural language processing to detect behavioral or mental health conditions. We provide a method for training Dell, and the natural language processing model is (i) a language model and (i i) A classifier comprising a method comprising (a) training a language model on a first encoded text. The first encoded text is a step, and is unrelated to behavioral or mental health status. (b) a step including text, and (b) a second encoded text, and optionally meta This is a step to fine-tune the language model with data information, where the second encoded text is, Steps and (c) multiple target audiences, including text related to behavioral or mental health conditions. To detect behavioral or mental states on multiple encoded audio samples from A step in training a classifier, comprising encoding multiple encoded audio samples. The encoded audio samples were provided by individuals who were active or sensitive. A label indicating whether or not the individual has a divine health status, and optional metadata information, are associated with this information. This includes steps and. In some embodiments, the language model includes long-term short-term memory (L Includes an STM network. In some embodiments, training (a) is non-monotonic stochastic. This includes a gradient descent process. In some embodiments, training (a) involves dropout or This includes the DropConnect operation. In some embodiments, the language model uses a converter. Includes. In some embodiments, the second encoded text includes additional behavioral or mental (b) includes text related to health status, and the fine-tuning of (b) includes multitasking learning. In one embodiment, this method involves multiple encoded audio samples from multiple subjects. The step involves training additional classifiers to detect additional behavioral or mental states on the system. Furthermore, the encoded audio samples of multiple encoded audio samples include, The subjects who provided the encoded audio samples had additional behavioral or mental health conditions. It is associated with a label indicating whether or not it does. In some embodiments, it is behavioral or mental. The primary health condition is anxiety disorder, and the additional behavioral or mental health condition is depression. In this embodiment, the fine-tuning in (b) includes discriminative fine-tuning of different layers of the language model. In one embodiment, the fine-tuning of (b) is done to train layers of the language model using a sloping triangle. This includes using learning rates. In some embodiments, the classifier is a binary classifier and a regression classifier. Including similar instruments, the training in (c) (i) whether the test subject has a behavioral or mental health condition (ii) training a binary classifier to predict whether or not, and (ii) the behavior or mental state of the subject This includes training a regression classifier to predict a numerical score indicating the severity of a specific health condition. In some embodiments, the output of the natural language processing model is the output of the binary classifier and the regression. It is based at least partially on the output of a similar device. In some embodiments, the method is (c) Next, (d) the step of obtaining voice samples from the subjects, and (e) the step of creating a natural language processing model. The system processes audio samples and determines whether the test subject has a behavioral or mental health condition. The process further includes the step of predicting whether the audio sample is multiple (e) includes multiple responses to a number query, and to process the audio sample multiple times This involves using a natural language processing model, where multiple responses are distributed in a different order each time. It is placed. In some embodiments, the natural language processing model receives multiple responses from multiple subjects. Includes an automatic speech recognition model for transcribing audio samples. In some embodiments, A natural language processing model is an encoder for encoding multiple transcribed speech samples. It includes an n-gram model, a skip gram model. In some embodiments, the encoder is an n-gram model, a skip gram model. The selection is made from a group consisting of Dell, neural networks, and byte-pair encoders. In some embodiments, the labels are the results of a standardized mental health questionnaire.

[0013] In another aspect, this disclosure may determine whether the subject has or will have a behavioral or mental health condition. A method for determining whether there is a high probability, comprising: (a) obtaining audio data from the subject. Step (b) Process the audio data by computer to obtain at least one of the audio data (c) a step of identifying the linguistic features and at least one acoustic feature The linguistic features and at least one acoustic feature are processed by computer to obtain one or more scores. It generates a score and uses one or more scores to determine if the subject has a behavioral or mental health condition. The steps of generating a determination of whether or not there is a high probability of having, and the steps of generating in (d)(c) A step of outputting an electronic report containing instructions for the judgment made, wherein (b) to (d) are It runs in less than 5 minutes, and the determination generated in (c) has an area under the curve of at least approximately 0.70 (A The present invention provides a method comprising a step having UC. In some embodiments, A The UC is at least about 0.75. In some embodiments, the AUC is less Both are approximately 0.80. In some embodiments, the electronic report is used to determine that the subject is performing the action. If it indicates that the person has or is likely to have a dynamic or mental health condition, behavior or Includes psychoeducational materials on mental health.

[0014] In another aspect, this disclosure may indicate whether the subject has or may have a behavioral or mental health condition. A method for determining that there is a high probability of this, comprising the steps of (a) obtaining voice data from the subject and (b) Computer processing of the audio data to determine at least one audio feature in the audio data and the step of identifying at least one acoustic feature, and (c) at least one acoustic feature and computer processing of at least one acoustic feature to determine the behavioral or mental health of the subject. (d)(c) A step of outputting an electronic report showing the judgment provided in (b) or (c) The computer processing includes at least one of the criteria for the determination provided in (c) This provides a method, including steps, for optimizing performance metrics.

[0015] In another aspect, this disclosure may determine whether the subject has or will have a behavioral or mental health condition. This method provides a method for determining whether the probability is high, and this method (a) subjects and medical During a telemedicine session using a telemedicine application between the provider and the recipient, the recipient's audio (b) Acquisition of stream and video stream, and (b) Acoustic model, natural language processing Steps to obtain one or more models, including a Neural Programming (NLP) model and a video model. And one or more models determine whether the subject has a behavioral or mental health condition. (c) A step in which the system is trained to determine whether it is likely to have or have Process audio or video streams using one or more models to reach the target audience. This indicates whether or not a person has, or is likely to have, a behavioral or mental health condition. (d) While the telemedicine session is in progress, the healthcare provider The user interface of the health application running on the user's device The method includes the step of sending a decision. In some embodiments, the method is used for natural language processing. The model is used to determine one or more topics or words within the audio stream. The step of sending one or more topics or words to the user interface further Includes. In some embodiments, the determination includes a confidence interval for the determination. This method involves repeating steps (a) to (d) sequentially during a telemedicine session. (b) includes demographic or medical history information relating to the subject. In some embodiments, (b) includes demographic or medical history information relating to the subject. This includes selecting one or more models based at least partially on the report.

[0016] Another aspect of this disclosure, when executed by one or more computer processors, Includes machine executable code that implements the systems described above or elsewhere in this specification. To provide a non-temporary computer-readable medium.

[0017] Another aspect of this disclosure relates to one or more computer processors and coupled thereto The system is provided, which includes computer memory. The computer memory may be one or more When executed by a computer processor, the methods described above or elsewhere in this specification Includes machine-executable code that performs one of the following actions.

[0018] In another aspect, the Disclosure relates to one or more computer processors and one or more When executed by a computer processor, one or more computer processors are used for Based at least partially on the input audio containing multiple segments from the subject, the subject is involved Acoustic motors configured to predict whether a person has a mental, behavioral, or emotionally healthy state. A memory containing machine-executable instructions for implementing Dell, wherein the acoustic model extracts input sound An encoder configured to generate symbolic representations, wherein the encoder is of interest to the subject. Perform tasks other than predicting whether a person has a certain behavioral or mental health condition. To do this, a transfer learning framework is used to pre-train an encoder, including the note The system processes the abstract representation of the input voice to determine the behavioral or mental health status of the subject. A classifier configured to produce an output indicating whether or not it has And at least one classifier is based on the speaker having a behavioral or mental health condition of interest. Instructions regarding audio samples labeled as originating from or not originating from The system provides a refined system comprising at least one classifier. In this embodiment, the encoder is a Visual Geometry Group ("VGG") net. Includes a stack of work and long-term short-term memory ("LSTM") networks. In this configuration, at least one classifier is a recurrent convolutional neural network. From the group consisting of "RCNN", attention-enabled LSTM, self-attention network, or converter Includes a selected model. In some embodiments, at least one classifier outputs It is further configured to process metadata about the subject in order to generate several In some embodiments, the metadata includes the subject's age or gender. In some embodiments, The encoder is trained on the audio sample transcribed by the decoder, and the decoder is sys It is not part of the system. In some embodiments, the task is automatic speech recognition, speaker recognition, and sensing. This is information classification or sound classification. In some embodiments, the segment output is averaged. In some embodiments, the segment outputs are fused using a machine learning algorithm. In some embodiments, the encoder is pre-trained with the decoder, and the encoder and decoder The coder includes an automatic speech recognition (ASR) system. In some embodiments, the decoder This is one of the attention unit, the long-term short-term memory network, and the beam search unit. This includes multiple classes. In some embodiments, at least one classifier includes a binary classifier. In some embodiments, at least one classifier includes a multi-class classifier, and the output This represents the probability of multiple severity levels of behavioral or mental health conditions of interest in the subjects. Includes cloth. In some embodiments, the output is each segment of multiple segments of the input audio. This is a segment output about the state, and the system uses it to obtain the predicted mental state. It is configured to fuse the learned representations of the segment outputs of at least one classifier. It also features a segment fusion module.

[0019] Further aspects and advantages of this disclosure are shown and described only in exemplary embodiments of this disclosure. As will be readily apparent to those skilled in the art from the following detailed explanation, Other different embodiments are possible, and some of their details deviate from this disclosure. Without doing so, modifications are possible in various obvious respects. Therefore, the drawings and descriptions are essential. It should be considered an example, not an limitation.

[0020] Embedding by reference All publications, patents, and patent applications referred to herein are treated as if each individual publication were a separate publication. It is specifically and individually indicated that a product, patent, or patent application is incorporated by reference. To the same extent, publications incorporated herein by reference and To the extent that a patent or patent application is inconsistent with the disclosures contained herein, this specification shall not be subject to such restrictions. It is intended to replace and / or take precedence over the material used for shielding.

[0021] Novel features of the present invention are described in detail in the appended claims. Features of the present invention A better understanding of the advantages and benefits is provided below, which describes exemplary embodiments in which the principles of the present invention are utilized. See the detailed description and attached drawings (hereinafter also referred to as "Figures" and "Drawings"). This will be obtained by... [Brief explanation of the drawing]

[0022] [Figure 1] This diagram schematically illustrates a system configured to predict whether a subject has a behavioral or mental health condition of interest, based at least partially on the voice input from that subject. [Figure 2] This diagram schematically shows the metadata section used to predict metadata about the subject from the input audio. [Figure 3] This diagram schematically shows an i-vector estimator configured to estimate an i-vector from input audio. [Figure 4] This is a flowchart of an exemplary process for training the system shown in Figure 1. [Figure 5] This is a flowchart of another exemplary process for training the system shown in Figure 1. [Figure 6] This diagram schematically illustrates a system configured to assess, screen, predict, or monitor the behavioral or mental health status of subjects using audio data, video data, and / or metadata related to those subjects. [Figure 7] This is a schematic diagram of a segment fusion module. [Figure 8] This figure shows a computer system programmed or otherwise configured to carry out the methods provided herein. [Figure 9] This diagram schematically illustrates a system that uses a natural language processing ("NLP") model to predict whether a subject has a behavioral or mental health condition. [Figure 10] Figure 9 is a flowchart of an exemplary process for training the NLP model. [Figure 11] This chart shows the distribution of Patient Health Questionnaire-8 (PHQ-8) and Generalized Anxiety Disorder-7 (GAD-7) scores in the dataset. [Figure 12] This is the score matrix in Figure 11. [Figure 13] This chart shows the accuracy of the trained model in predicting raw PHQ-8 and GAD-7 scores. [Figure 14] This chart shows the ROC (Recovery over Time) of various training models. [Figure 15] This is a bar graph showing the age distribution in two corpora of data. [Figure 16] This chart shows the distribution of PHQ-8 scores for two corpora of data from Figure 15. [Figure 17] This chart shows the binary classification results of a trained model. [Figure 18] This chart shows the data count for each age bucket and the ROC for each age bucket. [Figure 19] This is a schematic diagram illustrating a telemedicine system. [Figure 20] This figure shows the performance data for acoustic models and NLP models. [Modes for carrying out the invention]

[0023] Various embodiments of the present invention have been shown and described herein, but such embodiments are just a few examples. It will be obvious to those skilled in the art that it is provided only in this way. Those skilled in the art will not deviate from the present invention. Multiple transformations, modifications, and substitutions can be made without doing so. Please understand that various alternative forms may be used for the embodiment described above.

[0024] The terms "at least," "greater than," or "greater than or equal to" refer to a set of two or more numbers. Whenever it precedes the first number, it means "at least," "greater than," or "greater than or equal to." The term applies to each number in that series of numbers. For example, 1, 2, or 3 or more are: Equivalent to 1 or more, 2 or more, or 3 or more.

[0025] The terms "less than or equal to," "less than," or "less than or equal to" are used when referring to the first number in a series of two or more numbers. Whenever preceding, the terms “less than,” “less than,” or “less than or equal to” are part of the same series. This applies to each individual number. For example, 3, 2, or 1 or less is 3 or less, 2 or less, or 1 or less. It is equivalent to the one below.

[0026] Acoustic model Figure 1 shows the row of the row of interest to the subject, based at least partially on the voice input from the subject. A system 100 configured to predict whether or not a person has a dynamic or mental health condition. This is a general overview. Behavioral or mental health conditions include fatigue, loneliness, low motivation, stress, and Menstrual cramps, anxiety disorders, drug or alcohol addiction, post-traumatic stress disorder ("post-traumatic stress disorder") PTSD (Post-Traumatic Stress Disorder), Schizophrenia, Bipolar Disorder These may include harm, dementia, suicidal ideation, etc. Behavioral or mental health conditions may be used in the diagnosis of mental disorders. This may be related to, coexist with, or defined in relation to the statistical manual.

[0027] System 100 is on an Internet-connected device or connected to the Internet. Microphone connected to a device (for example, via Bluetooth® connection) The device can acquire input audio via a microphone or microphone array. Wearable devices (e.g., smartwatches), mobile phones, tablets, laptops Top computers, desktop computers, smart speakers, home assistance devices (For example, Amazon Alexa® devices or Google Home) (Registered trademark) Device may also be used. The device is used for mental health applications. It may be used. Mental health applications involve the subject's work and home life, sleep, mood, The subject can be visually or audibly prompted to answer questions about their medical history, etc. The subject's response to the prompt may be used as input voice. System 100 This can be implemented on a mobile application and on the target user's mobile device. Local input audio can be processed. Alternatively or additionally, mobile devices The audio can be transmitted to a remote location for processing. In some cases, the processing is partial. It may run on a local device and partially on a remote server.

[0028] Alternatively or additionally, the input voice may be obtained through a clinical encounter with a medical professional. That's good. For example, a recording device can capture audio from a patient while they are on a doctor's appointment. A doctor's appointment can be made in person or through telemedicine, which is conducted remotely.

[0029] System 100 consists of an encoder subsystem 110 and a decoder subsystem 120. , and may have a classification subsystem 130. System 100 and its subsystems are 1 It can be implemented on one or more computers in one or more locations.

[0030] The encoder subsystem 110 and the decoder subsystem 120 work together to process the input audio. An automatic speech recognition ("ASR") system may be formed to generate a transcription of the text. Generally, The encoder subsystem 110 can generate high-level acoustic features from the input audio. Yes, it is possible. The decoder subsystem 120 consumes high-level acoustic features to convert the string into... A probability distribution can be generated. The system samples from the probability distribution and inputs It can generate transcriptions of force sounds.

[0031] The encoder subsystem 110 first performs functions other than predicting behavioral or mental health status. It can be trained on tasks. For example, an encoder can be used for automatic speech recognition and sentiment classification. It can be trained with a decoder for tasks such as sound classification. This training is complete. It is not necessary. Even partial training of the encoder is not possible if the encoder is not pre-trained. It is possible to improve performance compared to combining. After training the encoder, the first task is deco The encoder can be discarded and used to predict behavioral or mental health conditions of interest. It can be used for tasks intended to be performed. This is known as transfer learning. ru.

[0032] The encoder subsystem 110 is a convolutional neural network ("CNN"). CNN112 may have a convolutional layer and a fully connected layer. 112 has at least approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more convolutional layers. They may have. CNNs have up to approximately 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 It may have convolutional layers. CNN112 has at least about 1, 2, 3, 4, or more. It may have fully connected layers. CNN112 has up to approximately 4, 3, 2, or 1 fully connected layers. The input to CNN112 may also be a spectrogram feature. The features include a 25-millisecond window and 10-millisecond intervals across 5-second segments of the input audio. It may have a frame rate of seconds. In other cases, the input is other front-end features. That's also good. CNN112 is part of the Visual Geometry Group ("VGG") network. It can be done. The VGG network improves the representation of high-level acoustic features. It is possible.

[0033] The LSTM network 114 has at least approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, It may have 10, 15, 20, or more LSTM layers. LSTM Network 114 It may have at least approximately 1, 2, 3, 4, 5, 6, or more fully connected layers. The input to TM network 114 may also be the output of CNN 112. MFCC is, It has a 25-millisecond window and a 10-millisecond frame rate throughout the entire input audio. Obtain. In some cases, the LSTM network 114 is a bidirectional LSTM (bi-vire It may also be a ctional LSTM (BLSTM).

[0034] The decoder subsystem 120 receives high-level audio signals from the encoder subsystem 110. It may have a feature receiving attention unit 122 and an LSTM network 124. Knit 122 is an LSTM network 124 that outputs high-level acoustic features at each output step. This may allow for focusing on (or "paying attention to") a subset. Attention unit 122 and The LSTM network 124 can generate probability distributions across strings. The intention unit 122 and the LSTM network 124 perform connectionist time series classification (" connectionist temporal classification: CT It can be trained using the C function. The decoder subsystem 120 is LSTM The network 124 receives a probability distribution on a character sequence, and the probability distribution gives rise to the possibility Beam search unit 1 traverses the transcriptions and selects the best transcription according to specific criteria. It may have 26 more.

[0035] In some cases, the decoder subsystem 120 is used only during training of system 100. It may be used. That is, the decoder subsystem 120 is deactivated during inference. Alternatively, it may be discarded. Training of System 100 is described in more detail with reference to the following diagrams. The classifier network 132 can generate a decision for a single segment. To generate a decision about the entire session (which consists of multiple segments), The system connects the segment fusion module 140 to the internal layer of the classifier network 132. We can supply one of them (usually the last layer). Then, segment fusion motor Joule 140 can generate a single prediction for the entire session. In the implementation, the segment fusion module uses MLP, LSTM, RCNN, for prediction. Random forests or similar methods can be used. Segment fusion module. Lu140 also includes different modalities, including models with different modalities (acoustics, NLP, image processing, etc.). These basic models may be used to combine with each other.

[0036] The classification subsystem 130, like the decoder subsystem 120, also includes an encoder subsystem. High-level acoustic features can be received from the system 110. Classification subsystem 1 30 processes high-level acoustic features to identify the behavioral or mental health status of the subject of interest. It is possible to predict whether or not it has a classification ("segment output"). More specifically, classification ("segment output") The output of subsystem 130 has the subject's condition (e.g., depression or bipolar disorder). The classification subsystem 130 is the posterior probability of the voice segment from the subject. It may have a classifier network 132. The classifier network 132 is a recurrent CN N ("recurrent CNN: RCNN"), attention-enabled LSTM, self-attention network It may be a workpiece or a converter. The classifier network 132 performs regression, order prediction, and bi It can perform value classification, multi-class classification, and more. In the case of binary classification, a classifier network is used. Work 132 performs binary prediction regarding whether the subject has a behavioral disorder or a mental disorder. It is possible to do so. In the case of multi-class classification, the classifier network 132 will classify the subject (for example) If the subject's PHQ-9 score or GAD-7 score indicates behavioral or mental health disorder It is possible to predict the severity or level of the condition.

[0037] System 100 includes metadata and / or an identity vector ("identity vec") Using the tor:i vector, the subject has a behavioral or mental health condition of interest. It becomes possible to predict more accurately whether or not this will happen. Metadata is data about the subject. For example, the subject's age, race, ethnicity, sex, gender, income, education, location, medical history, etc. This is also acceptable. Such metadata may indicate the behavioral or mental health status of the subject. Tem100 can retrieve metadata from a database, or input from the target person. Metadata can be predicted from the voice. Figure 2 shows how such predictions are made. The resulting metadata unit 200 is schematically shown. Each metadata unit 200 is Multiple different nuclei are configured to predict different types of metadata about the subject. It may have a multi-network classifier. For example, the metadata unit 200 may have a metadata unit 200 for the subject One neural network classifier is trained to predict age, and the other is located at the subject's location. It may have a neural network classifier that is trained to predict the same thing. Generally The metadata unit 200 predicts demographic data, past medical history, time, location, etc. The acoustic model can use known or inferred metadata to analyze the patient's condition. It is possible to predict the state of physical or mental health more accurately.

[0038] In some cases, the above metadata can be used in mental health applications. The experience of the person can be adapted or personalized. For example, if the patient is elderly, The font size of the mental health application can be increased. As another example, The wording of the question may use a specific regional dialect or be in a specific context (for example, Stem uses the word "roommate" when asking students about their home life. (It is possible to adjust for specific demographic groups, etc.) .

[0039] On the other hand, the i-vector may be a low-dimensional feature extracted from the input speech. Figure 3 shows i A schematic representation of the i-vector estimator 300, configured to estimate a vector, is shown. The estimator 300 uses a Gaussian mixture model to estimate such i vectors. It is possible.

[0040] In some cases, metadata and / or i-vectors are used to classify high-level acoustic features into subsets. High-level acoustic characteristics from encoder subsystem 110 before being passed to stem 130 It may be added to. In some other cases, metadata and / or i vectors may be used as substitutes. It is then added to the output of the classifier network 132 and passes through network 134. This is possible. Network 134 is, for example, a deep neural network ("d EEP Neural Network (DNN), Random Forest Classifier, or Support Vector Machine (SV) It can be "M").

[0041] Alternatively, the system can use an end-to-end model with transfer learning. The first few layers of this model (CNN and LSTM) help with the ASR task. It can be borrowed and initialized. By doing so, the system can create a new network. This can be done and trained with the transcribed audio data. The first layer of the model is pre-trained. After training, the system will freeze them during training for classification or prediction tasks. These weights can be continuously updated. Pre-training of CNNs and LSTMs is a system This allows the neural network to learn a more restrictive representation than when training all layers from scratch. To teach.

[0042] The end-to-end model can generate output from individual audio segments. In the case of an audio session containing multiple of these segments, the system will... By averaging predictions from all segments, including the overall mental health Predictions can be generated. In other embodiments, the system can generate additional neural networks. Individual segments can be merged using the workpiece. Segments are classified into subs. The final hidden layer of the stem output may be represented by a vector. (Per session) The sequence of elements is projected onto a single vector by maximum pooling, and then followed It can then be fed into an additional network (e.g., an MLP network). Next, a classification task is performed. Alternatively, the model can be trained for either a regression task.

[0043] The system uses an automatic speech recognition (ASR) task to identify the first few people in the network. That layer can be pre-trained. The pre-training step ensures that the network is functioning correctly. It is possible to start with a feature representation. To pre-train the first few layers Even when using a "weak" (high character error rate) model, significant performance improvements can be achieved. can.

[0044] The final layer (not shown) of the classification subsystem 130 has multiple output classes, for example, behavioral or This is a softmax layer configured to generate a probability distribution across mental health states. That's fine.

[0045] The acoustic models described above are at least approximately 60%, 65%, 70%, 80%, 85%, and 90%. , it may have a specificity of 95% or more. The acoustic model has at least about 60%, 65%. It may have a sensitivity of %, 70%, 80%, 85%, 90%, 95%, or higher. Acoustic Mo To increase Dell's specificity, sensitivity must be reduced, and vice versa. Acoustic model at least approximately 60%, 65%, 70%, 80%, 85%, 90%, 95%, or more. The above curve area ("AUC") may be present. The acoustic model is less than that of conventional systems. At least approximately 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, or more. This can provide improvements in the relative performance (e.g., sensitivity, specificity, or AUC).

[0046] Figure 7 schematically shows the segment fusion module 140. Unit 140 can receive inputs, which are outputs from the classification subsystem. The inputs are individual The classification results from each audio segment can be reflected. The process is patient and Multiple such segments, including audio sessions, can be collected. The stem arranges the segments and then uses maximum pooling to make them a single vector. They are projected onto a multilayer perceptron ("MLP") network and then onto a digital array. It can be supplied to the Advanced Learning Network. The model then provides a series of... To generate a prediction of the overall session output from machine learning analysis performed on the session. This is possible. The segment fusion module 140 can process all of the audio in a given session. The learned representations (classifier outputs) of each segment are merged across the segments. Then, you can get an overall prediction for that session. In its simplest form, The integration module 140 simply calculates the average forecast across all segments. That's fine. In more complex versions, the module would then input data for each audio segment within the session. It receives corresponding pre-trained representations and combines these representations using a machine learning model. (They can be merged). The learned representations correspond to the inner layers of the speech segment classifier. It is possible. Combinatorial or fusion models include MLP, LSTM, RCNN, and may include other similar models. Furthermore, the segment fusion module 140 is multi It may also be used to combine the results of modal inputs, for example, all modalities Used to combine acoustic segments, NLP, and visuals for the final decision, including the element. It may also be used.

[0047] The subsystem in Figure 1 is implemented on one or more computing devices. This is also acceptable. Computing devices include servers, desktops, or laptop computers. This may include computers, electronic tablets, mobile devices, etc. • Devices can be located in one or more locations. Computing devices A general-purpose processor is a graphics processing unit. GPU (graphics processing unit), application-specific integrated circuit (application-sp) Technical Integrated Circuit (ASIC), FieldPro Field-programmable gate array ys: FPGA), may have machine learning accelerators, etc. Computing device The chair is, for example, a dynamic random access memory or a static random access memory. Memory such as read-only memory, flash memory, and hard drives It may have. Memory is stored in the computing device at runtime. To train them, or to determine whether the subject has a behavioral or mental health condition of interest. It can be configured to store instructions that allow prediction. Computing devices are network The network communication device may further include a computer Connecting devices communicate with each other and with any number of user devices via the network. It can make it possible to trust. For example, a network communication device is a system Computing devices that implement 100 will predict the behavior or mental state of the subject. It will be possible to communicate with medical professionals' mobile devices regarding the patient's health status. The network may be a wired or wireless network. For example, The company operates fiber optic networks, Ethernet® networks, and satellite networks. Bluetooth, cellular network, Wi-Fi® network, Bluetooth It may also be an h(registered trademark) network, etc. In other implementation forms, computation The device is a distributed computer accessible via the internet. It may also be a computing device. Such computing devices are cloud It can be thought of as a computing device.

[0048] Training acoustic model Figure 4 is a flowchart of an exemplary process 400 for training system 100. Yes. Process 400 is a system of one or more computers located in one or more locations. This can be executed by a system. In Figure 4, such a computer is referred to as a "training system." They are collectively referred to as "Mu".

[0049] In operation 410, the training system uses an encoder sub to transcribe the audio data. The system 110 and the decoder subsystem 120 can be trained. Training data This may include raw audio data and the corresponding transcript of that raw audio data. This may be unrelated to the topic of behavioral or mental health. The raw audio data is unlabeled. It may be possible to do so. In other words, raw audio data is a conversation in which the mental or behavioral state is unknown. It may originate from a person. In some cases, training data may be from a public audio corpus. It may be derived from. Action 410 may be a supervised learning action.

[0050] In sub-operation 412 of operation 410, the training system filters the raw audio data. The mel-frequency cepstral coefficient ("mel-frequency cepstral coefficient") It can be converted to m coefficients:MFCC. Operation 410 In operation 414, the training system filters the encoder subsystem 110. Banks or MFCCs can be mapped to robust abstract feature representations. In suboperation 416 of operation 410, the training system is connected to the decoder subsystem 120. This allows the processing of abstract feature representations to generate output. Sub-operation 410 In operation 418, the training system receives the output of the decoder subsystem 120 as audio data. Compared with known transcriptions of the encoder subsystem 110 and decoder subsystem The 120 weights and biases can be updated to account for the differences. More specifically, training The system uses a cost function to calculate the difference between the generated output and the known transcription. This is possible. Regarding the weights and biases of encoder and decoder subsystems. By calculating the derivative of the cost function, the training system minimizes the cost function. Therefore, weights and biases can be iteratively adjusted over multiple cycles. If the output satisfies convergence conditions such as the magnitude of the calculated cost being small, then the training is complete. It can be completed.

[0051] In operation 420, the training system ignores or discards the decoder subsystem 120. It is possible. In other words, the decoder subsystem 120 can perform the remaining training operations or It does not need to be used in the inference.

[0052] In operation 430, the training system uses the classification subsystem 130 to collect labeled audio data. It can be trained on about the subject. Labeled audio data can be used for behavioral or behavioral activities of interest. Whether or not it originates from a person who has been determined to have a mental health condition. This may also be audio data labeled as: Interesting behavioral or mental health status. This may be any such condition as described herein. Labels are for clinical diagnosis, standardization. This could be a score from a mental health questionnaire (e.g., PHQ-9). Therefore, the classification subsystem 130 uses a standardized mental health questionnaire (e.g., PHQ). - Use the answers to a specific subset of questions from questions 1 and 2 of question 9 to perform an action. It can be trained to predict subclasses of physical or mental health conditions. (Action 410) Similarly, operation 430 may be a supervised learning operation. Sub-operation 432 of operation 430 In this process, the training system converts raw audio data into a filter bank or MFCC. This is possible. In suboperation 434, the training subsystem uses a previously trained encoder. The sub-system 110 can be made to generate an abstract feature representation of the audio data. In suboperation 436 of operation 430, the training subsystem interacts with the classification subsystem 130. From abstract characteristic representations, the behavior or mental health status of the subject who is the source of the audio data can be determined. It is possible to generate the output shown. In suboperation 438 of operation 430, the training system The system compares the output to the subject's known behavioral or mental health status and explains the differences. The weights and biases within the classification subsystem 130 can be updated. The training system For multiple audio samples, until the output of the classification subsystem 130 satisfies the convergence condition This process can be repeated.

[0053] In operation 430, the encoder subsystem 110 may be fixed. The weights and biases do not need to be updated. Alternatively, encoder subsystem 11 Weights and biases of 0 are particularly relevant when multiple labeled audio data are available. This may be adjusted in coordination with the weights and biases of the classification subsystem 130. This can result in a more robust system.

[0054] The system uses metadata and / or i-vectors to assess the behavioral or mental health of the subject. When predicting a state, the training system uses metadata and / Alternatively, the i vector can be initialized to 0. In operation 440, the training system, Metadata and / or before the classification subsystem 130 or after the classifier network 132 You can add an i-vector and continue training. Metadata and / or i-vector The system 100 is configured to be added to the output of the encoder subsystem 110. If present, the training system will process such output as well as attached metadata and / or iB The entire classification subsystem 130 can be continuously trained on the model. Alternatively, The data and / or i vector are added to the output of the classifier network 132. If stem 100 is configured, the training system will train only network 134. It is possible.

[0055] Process 400, in operation 410, the encoder uses the first training dataset. It is trained to perform one task (i.e., automatic speech recognition) and operation 430 In this process, the encoder and classifier use the second training dataset to perform the second task ( In other words, they are trained to perform actions such as predicting the mental or behavioral state of a subject. This is a transfer learning process. The encoder is pre-trained to perform the first task. This means having a robust second training dataset with a sufficient amount of clinically labeled data. This may be beneficial because obtaining accurate audio data can be difficult. The first task is automatic speech recognition. However, in other embodiments, the first task is This could also be a classification of emotions, sounds, etc.

[0056] Figure 5 is a flowchart of an exemplary process 500 for training system 100. Yes. Process 500 may be a substitute for process 400. Process 500 is 1 It can be performed by one or more computer systems located in one or more locations. Yes, it is possible. In Figure 5, such computers are collectively referred to as "training systems."

[0057] In operation 510, the training system encodes the transcribed audio data. The bus system 110 and decoder subsystem 120 can be trained. Operation 51 0 may be the same as or similar to operation 410 in Figure 4. In operation 520, the training system This refers to audio data labeled with the speaker's behavioral or mental health status from which the audio data originates. While training the classification subsystem 130 on the data, the encoder subsystem 110 and The decoder subsystem 120 can continue to be trained. During operation 520, encoding Classification subsystem 130 of the cost function for the contributions of the decoder and decoder subsystems The contribution of the encoder can increase. Therefore, in operation 530, the training system uses the encoder. By fixing the system 110 and ignoring or discarding the decoder subsystem 120, This allows for fine-tuning of the classification subsystem 130.

[0058] Metadata and / or i-vectors are added to the output of the classifier network 132. When system 100 is configured, the training system will include metadata and / or i-vectors. The system can perform the operation 540 which is added in this way, and the training system can (i) classify Either freeze the device network 132 and train only network 134, or (ii) ) Continue training the classifier network 132 while training network 134. Operation 5 At step 50, the training system trains a model for segment fusion. Prior to training, project the sequence of segment outputs of various segments into a single vector. You may do so.

[0059] Acoustic Model Example 1 In one example, the acoustic model shown in Figure 1 was used to predict anxiety disorders and depression in a group of subjects. The acoustic model's classifier was trained to perform binary classification. The acoustic model's encoder is shown in Figure It was pre-trained to perform the automated speech recognition task as described in 4. The model in which only the encoder weights are updated ("the first model") and, A model in which both encoder and decoder weights have been updated ("second model") is used beforehand. The participants were trained to complete a participant health questionnaire that served as a label for depression. 8 (i.e., the PHQ-9 with the suicidal ideation question removed) and its role as an anxiety disorder label He had been diagnosed with generalized anxiety disorder-7. The first model had a specificity of 0.71, 0.71 The first model predicted depression with a sensitivity of 0.79, an AUC of 0.79, and an F1 of 0.54. The second model predicted depression with 0 Depression was predicted with a specificity of 0.72, a sensitivity of 0.72, and an AUC of 0.79. The second mode... The LU has a specificity of 0.68, a sensitivity of 0.69, an AUC of 0.75, and an F1 of 0.49. The disease was predicted.

[0060] Using transfer learning, depression classification can be improved compared to acoustic models trained without transfer learning. The performance of the acoustic model for this improved by 27%, from an AUC of 0.62 to an AUC of 0.79. It was done.

[0061] Natural Language Processing Models This disclosure also uses natural language processing models ("NLP") to understand how subjects behave or think The system and method are provided for predicting whether or not a person has a divine state of health. This allows us to obtain audio samples from the subjects. The subjects may have work or family The system can provide voice samples in response to prompts related to daily life. To predict whether a subject has a behavioral or mental health condition, an NLP model is used. It can be used to process audio samples. The NLP model is a general text, Different combinations of main-specific text and audio samples from multiple subjects. Training may be conducted at this stage. The audio sample is provided by the person who provided the audio sample, and This can be associated with clinical labels indicating whether or not one has a mental health condition. Bell uses standardized health questionnaires, such as the Patient Health Questionnaire 9 ("PHQ-9"). The results may be used as a basis. In some cases, clinical labels predict a subclass of depression. The PHQ-9 can be used for this purpose (for example, the answers to questions 1 and 2 in the PHQ-9). This could be an answer to a subset of questions from (mi). Alternatively, the clinical label is a clinician. It may also be based on a diagnosis from [source].

[0062] Figure 9 shows whether the subject has a behavioral or mental health condition using an NLP model. The system 900 for predicting symptoms is outlined below. Symptoms are described in the diagnostic and statistical manual for mental disorders. Diagnostic and Statistical Manual f Mental Disorders (DSM) or other similar reliable sources This may include clinically defined symptoms, or symptoms related to those defined in the DSM. The symptoms may be present or coexisting. For example, the condition may include fatigue, loneliness, low motivation, etc. Trauma, depression, anxiety, drug or alcohol addiction, post-traumatic stress disorder (PTS) D) This could include schizophrenia, bipolar disorder, dementia, suicidal ideation, etc.

[0063] System 900 includes an automatic speech recognition ("ASR") subsystem 905, an encoder, and a sub-system 905. The system includes a subsystem 910, a language model subsystem 915, and a classification subsystem 925. But that's fine.

[0064] The ASR subsystem 905 can generate a transcript of the input voice from the subject. In some cases, the ASR subsystem 905 may be a third-party ASR, such as Google ASR. This may include third-party ASR models. Third-party ASR may also be a best-case hypothesis ASR. Furthermore, word uncertainty may be considered, and information regarding word confusion may be included. In other cases, The ASR subsystem 905 may include custom ASR models.

[0065] System 900 can acquire input audio in several different ways. The Mu900 acquires input audio by sending one or more queries to the target. It is possible. System 900 supports audio formats, visual formats, Alternatively, queries can be sent in audiovisual format. For example, sys The Tem900 is an electronic display and speaker for the user's computing device. Queries can be sent via this method. Queries can be sent about the subject's mood, sleep, appetite, and energy. It may be related to ghee levels, relationships, work, medical history, medication, etc. In some cases, the query is Standardized questionnaires for mental health, such as the PHQ-9 or General Anxiety Questionnaire 7 ("G The questions may be from or based on AD-7. In some cases, System 900 can send queries to the subject as part of a dynamic conversation with the subject. It can. That is, each query can be based on previous queries and the responses of the subjects to such previous queries. In other cases, the queries and their order can be predefined. Additionally or alternatively, the system 900 can obtain input speech by passively listening to the subject. The system 900 can, for example, passively listen to the voice of the subject during normal daily activities or during a conversation with a healthcare provider. The response of the subject to the query can function as input speech to the ASR subsystem 905.

[0066] The encoder subsystem 910 can convert the transcribed speech from the ASR subsystem 905 into a vector of real numbers (i.e., embeddings) in a vector space. The vector can represent individual words. Vectors that are close to each other in the vector space can represent words that are semantically similar in that such words are often displayed together in text or are otherwise associated with each other. The encoder subsystem 910 can use several different models or techniques to convert the transcribed speech into a vector For example, the encoder subsystem 910 can use an n - gram or skip - gram model, a feed - forward or recurrent neural network, matrix factorization, byte - pair encoding, sub - word regularization, or any combination of such models and techniques. These models and techniques are described in more detail in the following papers, which are incorporated herein by reference: T. Mikolov et al., Distributed Representations of Words an 、Distributed Representations of Words an d Phrases and their Compositionality,201 3,https: / / arxiv.org / pdf / 1310.4546.pdf;J. Pennington et al., GloVe: Global Vectors for Word Representation, 2014, https: / / nlp.stanford .edu / pubs / glove.pdf; R. Sennrich et al., Neural Machine Translation of Rare Words with Subword Units, 2015, https: / / arxiv.org / pdf / 1508.07909.pdf; T. Kudo, Subword Regulariz ation: Improving Neural Network Translation on Models with Multiple Subword Candidat es, 2018, https: / / arxiv.org / pdf / 1804.10959 .pdf. The encoder subsystem 910 can convert words, syllables, phonemes, or characters from the transcribed speech into vectors according to the specific model or technology used by the encoder subsystem 910.

[0067] The language model subsystem 915 can process the vectors generated by the encoder subsystem 910 and additional metadata information, such as metadata regarding the person who provided the utterance (e.g., age, gender, sex, ethnicity, location, income, medical history, etc.), or metadata regarding the query and the person's response to those queries (e.g., order of questions, type of questions, etc.). The language model It may have a ("LSTM") network 916. The LSTM network is recurrent Recurrent neural network RNN is a type of RNN. RNNs are used in relation to time series data, such as audio data. This is a neural network with cyclic connections that can encode [something]. RN N may include an input layer configured to receive a sequence of time-series inputs. RNN is , and may further include one or more hidden recurrent layers that maintain the state at each time step In this configuration, each hidden recurrent layer can calculate the output of that layer and the next state. The next state may depend on the previous state and the current input. The state is maintained over time steps. The dependencies within the input sequence may be obtained.

[0068] An LSTM network can be composed of LSTM units. An LSTM unit is a... The cell may include an input gate, an output gate, and a forget gate. It can play a role in tracking the dependencies between elements. The input gate receives a new value in the cell. The amount of data flowing in can be controlled, and the forget gate controls the extent to which values ​​remain in the cell. The output gate can then calculate the output activation of the LSTM unit based on the value in the cell. The extent to which it is used can be controlled. The activation function of the LSTM gate is logis A tick function would also work.

[0069] Alternatively, the language model subsystem 915 may have a converter 917. Model 17 may also be a model without iterative connections. Instead, it may rely on an attention mechanism. The attention mechanism either focuses on a specific input area while ignoring others, or "corresponds" This is possible because certain input regions may not be very relevant to the model. Performance can be improved. In each time step, the attention unit, among other things, The inner product of the context vector and the input at the time step can be calculated. The output of the unit determines where the most relevant information is located within the input sequence. It can be interpreted. The converter is A. Vaswani et al., Attentio n is All You Need,2017,https: / / arxiv.org Further details can be found in / pdf / 1706.03762.pdf, which is a reference. This is incorporated herein and reproduced in Appendix A. The converter 917 is used for any input region. When deciding whether to respond, non-verbal metadata information may be relied upon.

[0070] The classification subsystem 925 includes a binary classifier 926, a regression classifier 927, and an inverse binary classifier. It may have 928. Each of the three classifiers may be trained for a different purpose. Binary The classifier 926 classifies subjects as having behavioral or mental health conditions, or behavior It can be trained to classify individuals as those lacking a specific health condition. (Regression Classifier 92) 7. Behavioral or mental health according to a certain scale, for example, the PHQ-9 scale for depression. It can be trained to predict health conditions. The output layer of the regression classifier 927 uses the softmax function. Applying this to the possible scores, for example, 28 possible scores from 0 to 27 for PHQ-9 It can generate a probability distribution. The inverse binary classifier 928 is similar to the binary classifier 926. to train to classify the subject as having a behavioral or mental state or not having a behavioral health state, but with words reversed (e.g., from "My name is Michael Jordan" to "Jordan Michael is my name") it is possible to train on the transcribed speech. This approach can enable the system 900 to capture word dependencies not captured by the binary classifier 9 26. Inference may be repeated up to 10 times for the subject. In each iteration, the system 9 00 can concatenate the subject's responses in different orders. This creates various permutations of the same session driven by the rearrangement of the responses. The classifiers 926, 927, and 9

[0071] 28 can return slightly different outputs in each iteration. Then, the system 90 0 can optimize the results by averaging the outputs or performing other statistical analyses. Finally, the system 900 can combine the outputs of the three classifiers to generate a final prediction. The system 900 can make more accurate predictions for subjects participating in multiple sessions. The above NLP model can have a specificity of at least about 60%, 65%, 70%, 80%, 85%, 90 %, 95%, or more. The NLP model can have a sensitivity of at least about 60%, 65%, 70%, 80%, 85%, 90%, 95%, or more. To increase the specificity of the acoustic model, it is necessary to decrease the sensitivity, and vice versa. The NLP model can have a specificity of at least about 60%, 65%, 70%, 80%, 85%, 90%, 95%, or more. The system 900 can make more accurate predictions for subjects participating in multiple sessions.

[0072] The above NLP model can have a specificity of at least about 60%, 65%, 70%, 80%, 85%, 90 %, 95%, or more. The NLP model can have a sensitivity of at least about 60%, 65%, 70%, 80%, 85%, 90%, 95%, or more. To increase the specificity of the acoustic model, it is necessary to decrease the sensitivity, and vice versa. The NLP model can have a specificity of at least about 60%, 65%, 70%, 80%, 85%, 90%, 95%, or more. It can have an AUC of at least 1. %, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25%, or higher relativity. This can provide improvements in performance (e.g., sensitivity, specificity, or AUC).

[0073] The subsystems and their components shown in Figure 9 are one or more computing devices. It may be mounted on a device. Computing devices include servers and desktops. Alternatively, it may be a laptop computer, electronic tablet, mobile device, etc. Computing devices can be located in one or more locations. A computing device is a general-purpose processor, a graphics processing unit (GPU), or a special type of computing device. Application-oriented integrated circuits (ASICs), field-programmable gate arrays (FPGAs) Computing devices may have features such as dynamic random action. Seth memory or static random access memory, read-only memory, flash memory It may have additional memory such as memory and hard drives. Memory is used at runtime. The routing device is configured to store instructions that cause the subsystem to perform its functions. It may be. The computing device may further have a network communication device. It is also acceptable. Network communication devices are computing devices that connect to the network. It can enable communication with each other and with any number of user devices via this. The network may be a wired or wireless network. For example, a network This includes fiber optic networks, Ethernet® networks, and satellite networks. Cellular network, Wi-Fi® network, Bluetooth (Registered trademark) Network may also be used. Other implementation forms may include computing • The device is accessible via the internet through several distributed computing resources. It may also be a cloud computing device. Such computing devices are cloud computing devices. It can be thought of as a computing device.

[0074] NLP training model Figure 10 shows an exemplary process 1000 for training a model in system 900. This is a diagram. Process 1000 is one or more locations. It can be executed by a computer system.

[0075] The system uses an LSTM network 9 on a publicly available data corpus (1005). 16 or the converter 917 can be trained. The publicly available data corpus is text A text corpus is also acceptable. Text corpora do not necessarily relate to behavioral or mental health. They don't have to be related. Instead, the text corpus is a general-purpose text corpus. This is also fine. Text corpora can be large and capture the general characteristics of the text language. This is acceptable. For example, a text corpus could include Wikipedia articles. Operation The training task in 1005 is language modeling, for example, an LSTM network 916 or This may involve training the converter 917 to predict the next word in a word sequence. The output of the LSTM network 916 or the converter 917 is a probability distribution across multiple words. That's fine.

[0076] The training in operation 1005 includes dropout and DropConnect operations. Dropout occurs when a random subset of nodes in a neural network is obtained. This is a process that is eliminated during training. Different subsets can be eliminated for each training example. DropConnect removes a random subset of weights during training. That is, it is a process (which is set to 0). Similar to dropout, it differs for each training example. A subset can be removed. Dropout and DropConnect are This can help prevent over-fitting.

[0077] Training in operation 1005 is performed using non-monotonic mean stochastic gradient descent (stochastic g The process may further include radiant descent (SGD). SGD is a scalar descent process. Training loss is achieved by iteratively adjusting the model weights with gradient steps. It is a process of reducing errors. Training deep networks is a non-convex optimization problem.

number

number

[0078] Following the training in action 1005, the system performs the target task, namely, the action and skill. Adjust the LSTM network 916 or the transducer 917 for detecting the sanitary condition. This can be done (1010). Operation 1010 is an LSTM network on a domain-specific data corpus. This may include training the twerk 916 or the transducer 917. For example, texts relating to behavioral and mental health, such behavioral and mental Transcripted voice data from patients being tested for the condition, as well as domain-specific data This may include additional non-verbal metadata information about the corpus (e.g., its source). Domain-specific corpora contain texts about specific conditions for single-task learning. It may include text on multiple different conditions for multitasking learning. obtain.

[0079] Training in action 1010 may include discriminative fine-tuning. LSTM network 916 Alternatively, different layers of the converter 917 can take in different types of information, so different layers They can benefit from having different learning speeds. Generally, deeper layers are You can benefit from a higher learning rate. The learning rate for a particular layer is also adjusted over time. This is also possible. For example, the system linearly increases the learning rate until the condition is met. Next, the speed is linearly reduced. This method is called the "slant triangle learning rate". It can be called "ed triangular learning rates (STLR)". This process is described by J. Howard et al., Universal Languages. uage Model Finetuning for Text Classific ation,2018,https: / / arxiv.org / pdf / 1801.06 Further details are provided in 146.pdf, which is incorporated herein by reference. This is reproduced in Appendix A.

[0080] Training in operation 1010 further involves gradually unraveling the language model, and longer words. Backpropagation over time to process word dependencies, and in the LSTM network 916 This may include multiple pooling operations.

[0081] Following the fine-tuning in operation 1010, the system will perform each task. Classifiers 926, 927, and 928 can be trained (1015). Operation 101 Training in step 5 involves the ASR model, encoder model, LSTM network 916, etc. or a converter 917 and / or one of classifiers 926, 927, or 928, etc. It can be an end-to-end process including. However, classifier 926, Units 927 and 928 may or may not be trained independently of each other.

[0082] The training data includes several metadata, such as metadata about the subjects who provided the voice recordings. In addition to data information, it may also be a labeled audio sample that is transcribed and encoded. The audio sample was created using the method described with reference to Figure 9, that is, by sending a series of queries to the target audience. It can be collected by sending. The system uses a random order for training. This allows us to concatenate the responses of specific subjects. Classifiers 926, 927, and 928 are sequential. The order may differ. This technique may help alleviate the shortage of audio samples. By administering PHQ-9 to subjects, it is possible to label audio samples.

[0083] NLP Example 1 In the first example, the inventors conducted approximately 11,000 sessions over approximately 16,000 sessions. Audio recordings were collected from specific subjects. Some subjects participated in multiple sessions. The ages of the participants ranged from 18 to over 65, with an average age of approximately 30. The participants were software engineers. I provided an audio sample in response to a prompt presented via the application. The sessions are related to topics such as "work" and "family life." Each session consists of 4 It included ~6 prompts, with an average of 4.52 prompts, and the resulting sessions were Each lasted an average of about 5 minutes.

[0084] In addition to answering prompts, each participant completed the PHQ-9 (excluding the suicidal ideation question). We completed the "PHQ-8" and GAD-7. The results of these standardized questionnaires are sound The voice samples proved useful as labels for depression and anxiety, respectively. PHQ-8 and For both GAD-7, a score above 10 is mapped to the presence of symptoms, and a score of 10 is mapped to the presence of symptoms. Lower scores were mapped to the absence of symptoms. Table 1 shows the training data and test data mentioned above. This provides statistics for both conditions; "-" indicates the absence of a condition, and "+" indicates the presence of a condition.

[0085] [Table 1]

[0086] Table 2 shows the co-occurrence of depression and anxiety in both training and test data. Statistics are provided along with the training data in bold text. The statistics are based on approximately 16,000 training data. 18.5% of the sessions resulted in a positive label for both depression and anxiety. However, the study showed that 14% of the trial data sessions resulted in positive labels for both. Approximately 15% of the training data sessions were labeled "mismatched," i.e., depression or anxiety. This resulted in a label indicating a positive result for the disease but not for both.

[0087] [Table 2]

[0088] Figure 11 shows the raw PHQ-8 scores in the training and test datasets. The percentage distribution of GAD-7 scores is shown. The largest difference is between the PHQ-8 and GAD-7 scores. This is the case where the value is 0, and there is a 5% discrepancy. PHQ-8 and GAD-7 after normalization. The overall correlation between Q-8 and GAD-7 is 0.80.

[0089] Figure 12 shows rows of PHQ-8 and GAD-7 scores from training and test data sessions. These are columns. Please note the difference in score ranges. Each question has four possible scores (i.e., It has values ​​of 0, 1, 2, and 3. Therefore, the GAD-7 score ranges from 0 to 21. The PHQ-8 score ranges from 0 to 24. Within each scale, a higher value indicates a higher score. This indicates the severity of the condition. As shown in Figure 12, the majority of sessions occur near the diagonal, and two This is consistent with the high correlation to mental health status. Also, for each GAD-7 label, the PHQ-8 label There are many variations in the row, not the other way around. That is, the row in Figure 12 shows more variation than the column. This is significant. This may reflect the fact that anxiety disorders tend to be a prerequisite for depression. .

[0090] The language model was trained as described with reference to operations 1005 and 1010 in Figure 10, and then micro After adjustment, the inventors use the above training data to train the classifier according to operation 1015. A classifier was used. One group of classifiers was trained to detect anxiety disorders, and the other group was trained to detect depression. It was trained to detect [something]. Next, the trained model was tested using test data. Ta.

[0091] Figure 13 shows the trained model predicting raw PHQ-8 and GAD-7 scores. This chart shows the accuracy. The model is most accurate when predicting low and high scores. Yes, and it is not the most accurate when predicting a score between 8 and 12. This range is for healthy individuals. This is expected, as it represents the natural boundary between the body and the individual diagnosed with a positive test result.

[0092] Table 3 shows the properties of the binary classifier, including specificity, sensitivity, and area under the ROC curve ("AUC"). This provides statistics on cognitive function. The model is 0.828 for depression and 0 for anxiety. An AUC of 0.792 was achieved.

[0093] [Table 3]

[0094] The model's performance depends on whether the speaker has both anxiety and depression, or neither. It is best in both cases. In either case, it can be called a "consistent" session. The AUCs for PHQ-8 and GAD-7 were 0.861 and 0.84, respectively. It increases to 1. The prior distribution of the matching-only data is approximately 0.20 for the positive class. It changes to 0.16. This is not true after data rebalancing. Improved result The results remained the same, increasing even after rebalancing, with PHQ-8 and GAD-7 each at 0. The results were 0.863 and 0.849. This finding indicates that class identification is based on the individual values ​​of either state. This suggests that collaborative modeling of depression and anxiety disorders is better than individual modeling.

[0095] The trained model uses signals to distinguish between positive and negative cases for each condition. By using specific word sequences and their dependencies, it predicts depression more accurately than anxiety disorders. It is possible. To investigate, the inventors have found that during a given time in a test session, To estimate the amount of usable predictive information, the word sequence was gated in the forward direction. For example For example, in an 800-word session, you start with the first word and write one word at a time. By adding this, 800 cumulative gate samples were generated. 3078 test sessions Regarding the prediction, the inventors generated approximately 2.4 million predictions. Based on these predictions, A value called "Intra-session model variation" was calculated. This process is performed separately for each condition. We performed the following steps. In both cases, the model was optimized for AUC in the test set. Therefore, the test set is identical for both models.

[0096] Table 4 provides results for this measure of variability in the depression model within a session. The variability is highest when +, + (i.e., both conditions are present), and when -, - (i.e., The lowest value is when none of the conditions exist, and the lowest value is somewhere in between in the case of a mixture. This is binary depression. A model adjusted to match the maximum AUC for disease classification shows higher values ​​for this measure within a session. This suggests that a word sequence queue related to variability is being used.

[0097] [Table 4]

[0098] Table 5 provides results for this measure of variability in the anxiety disorder model. However, Here, (1) the overall variability is lower than that of depression, and (2) the variation in the cases of - and - The dynamism is much lower than expected considering the other three values. The same test data is in both tables. Because the same NLP model and method are used, this is a word sequence for anxiety disorders. This suggests that the cue may be weaker or less common than those for depression. They are instigating it.

[0099] [Table 5]

[0100] Figure 14 shows the complete test dataset, consistent sessions only (i.e., PHQ-8). (and sessions where GAD-7 sessions were consistent), and sessions where the data was rebalanced. This shows various AUCs of the model, including the AUC for a consistent session.

[0101] NLP Example 2 In the second example, the same approximately 16,000 sessions of audio, the age of each speaker, and the corresponding P The HQ-8 depression label was used. Table 6 shows statistics for both training and test data. The data is shown in italics. "GP" indicates the general population corpus, and "SP" indicates the corpus. The near-population corpus is shown. "Depression + / " refers to PHQ-8 (i.e., separate sessions) (By scoring both above 10 and below 10) it consistently responded This indicates individuals who have participated in two or more sessions.

[0102] [Table 6]

[0103] The main difference between the GP corpus and the SP corpus is the age distribution. The distribution is shown in Figure 15. The ages of the subjects in the GP corpus and the SP corpus do not overlap. 67% of the P Corpus participants are 60 years of age or older. Further differences exist between the two corpora. There is a possibility that when the subjects of the SP corpus give a short answer, they will be asked additional questions. On the other hand, the target group for the GP corpus was limited to 4-6 questions. The time allotted for collecting responses from participants was limited to 5 minutes, after which the session ended. Most participants It was also anticipated that the process would be repeated five times at a frequency of once a week. On the other hand, multiple sessions Participants in the GP corpus who have completed the training will wait at least 3 months between sessions and will be single Within the session, they were not part of the structured schedule.

[0104] The SP corpus was collected in Southern California. The sessions within the SP corpus were: Sessions in the GP corpus are shorter on average than those in the SP corpus. An average of 450 words per session, and an average of 800 words per session for the GP corpus. The average number of responses per session in the SP corpus is also the same as in the GP corpus. The average number of responses was higher than in the (6.1) corpus. In the example, it is used only for test data. The gender distribution between the GP corpus and the SP corpus is the same. It appears that 62% of the subjects in the SP corpus were female, and 58% of the subjects in the GP corpus were female. She is a woman.

[0105] Figure 16 is a chart showing the distribution of PHQ-8 scores for two corpora. The same applies to the cloth, especially for higher PHQ-8 scores. The prevalence of depression is higher in SP Co. The percentage is 30% in the Pass Corpus and 26.7% in the GP Corpus.

[0106] In this example, the classifier is trained on the GP training corpus only according to operation 1015 in Figure 10. Table 7 shows the models described herein and F. Ringeval et al., AVEC 201 9 Workshop and Challenge:State-of-Mind,D Etecting Depression with AI, and Cross-Cult ural Affect Recognition,2019,https: / / arx AVEC 2 as described in iv.org / pdf / 1907.11510.pdf Performance statistics for the 019 model are provided, which are incorporated herein by reference and reproduced in Appendix A. RMSE is an error metric that is inversely correlated with performance, while CCC is an error metric that is inversely correlated with performance. This is a correlation metric that correlates with [the specified value]. The models described herein have been tested on the GP corpus. In this case, it had both a lower RMSE and a higher CCC than the AVEC model.

[0107] [Table 7]

[0108] Figure 17 shows the model described herein for both the GP test corpus and the SP test corpus. This chart shows Dell's binary classification results. The AUC of the GP corpus was 0.828. However, the AUC of the SP corpus was 0.761. The corpus includes differences in the primary age distribution. Considering the differences, the trained model was unexpectedly portable. The patients participated in the longitudinal trial as described above. The classification performance of the GP training model was evaluated across multiple sessions. This strongly depends on the consistency of patient self-reported PHQ-8 scores across the collection of data. Of the 161 unique patients, 119 consistently experienced depression over multiple sessions. -or had a PHQ-8 score that consistently indicated depression ("SP consistent"). The remaining 42 One patient had inconsistent PHQ-8 results ("SP mismatch") across multiple sessions. They had. Overall, patients who reported consistently were more concise than those who reported inconsistently. There was a tendency for fewer responses. Figure 17 shows that sessions were executed one at a time, and the target Even if a user is unaware of their own score, it will manifest as a function of user consistency in model performance. This indicates a significant difference. The AUC of the SP corpus model is 0 for consistent patients. The ratio is 0.82, and 0.61 in inconsistent patients. Age and in two corpora Despite significant discrepancies in other factors, the model is similar to the SP Corpus in the case of the GP Corpus. - This was also performed on consistent users of the path. This data is particularly relevant for consistent patients. It exhibits good portability.

[0109] Table 8 provides statistics on model performance by age group. The number of participants under a certain age is small by design. The performance of the model on the GP test corpus is based on GP training. The performance is strongly correlated with the age distribution of the corpus. Very low data sample size is robust to the results. This affects gender, and the same applies to the SP test corpus.

[0110] [Table 8]

[0111] For the SP test corpus, performance at actual age was also investigated. Each age threshold (e.g., 30, Regarding 35, 40, 45, etc., the inventors have identified all subjects below that threshold and All subjects who exceeded the threshold were combined. Figure 18 shows the results for each age group. This chart shows the data count (solid line) and the AUC for each age bucket. Figure 18 As the age threshold increases, that is, as more and more older subjects are added to the bucket... This shows that model performance decreases as more data is added. Model performance also decreases with younger data. It decreases slightly as the person is removed from the bucket.

[0112] Table 9 provides statistics on model performance by age group.

[0113] [Table 9]

[0114] Table 10 provides statistics on model performance by ethnicity. The model is compared with other groups. And it didn't work very well for Hispanic subjects. This is because of the group By assigning higher weights to the samples during training, this particular model can be applied to this population. This is the case if we can train it. For several subgroups, though not all, 1 The two sizes fit all models and work well. In some subgroups, The same invention is used, but in training, data from that group is primarily weighted or included Using this, more attention is paid to creating a model tailored to that subgroup. You may need to pay.

[0115] [Table 10]

[0116] Additional data Figure 20 and Table 11 show the results of training with the same audio data used in NLP Examples 1 and 2. Addition of both acoustic and NLP models when performing binary depression prediction during testing. The performance data is shown below.

[0117] [Table 11]

[0118] Table 11 shows the AU values ​​for both the acoustic and NLP models that are close to or exceed 0.80. This demonstrates that C is achieved. Model fusion provides an additional 2-3% improvement in AUC performance. These systems do not use any information other than the audio sample itself; in other words, metadata. Patient history or other information (such as visual information) will not be used for acoustic and NLP results. The LP model performs better overall than the sound system, but both systems are shown in Figure 2. As shown in 0, it is in line with, or better than, the reference study of primary care providers (PCPs). The results are positive and strong. However, comparisons with the PCP trial are indirect due to differences in settings and data. It is.

[0119] Composite model Figure 6 uses audio data, video data, and / or metadata about the subjects. to assess, screen, predict, or monitor the behavioral or mental health status of the subjects. The system 600 configured as shown is schematically represented. System 100 in Figure 1 is a part of system 600. It may be a component of the same system. For example, system 100 is the acoustic model 6 of system 600. It may also be used as 17. The system in Figure 9 is also a component part of system 600. It's acceptable. For example, System 900 is NLP Model 616 of System 600. It may be used.

[0120] System 600 is a signal that can preprocess audio and video data from the subject. It may have a preprocessor 605. For example, the signal preprocessor 605 may have a preprocessor 605 within the audio data. The noise is segmented and reduced, or beamforming, acoustic echo cancellation, echo It can perform suppression, reverberation removal, or even noise injection. Signal preprocessor 60 5 can also generate audio and video quality confidence values. Audio quality reliability values ​​are, for example, the quality of each audio and video signal and audio The length of the audio and video samples can be taken into consideration.

[0121] Furthermore, the signal preprocessor 605 adds metadata to the audio and video data. This data can be preprocessed in such a way for consumption by Model 615. It may also be supplied to bus 610 in the form of a third-party or custom ASR system 62 It may be multiplied by 0. The ASR system 620 provides machine-readable transcription and transcription reliability of input speech. It can generate degrees. Similar to the signal preprocessor 605, the ASR system 620 It can then supply its output to bus 610 for later consumption by other components. .

[0122] Model leader 622 accesses model 615 from model repository 623. Model 615 can do this. Model 615 is a natural language processing model 616, an acoustic model 617, and a video model. This may include model 618 and metadata model 619. Natural language processing model 616 is the target The vocabulary content of the input speech from the user can be considered. The acoustic model 617 considers the input speech Non-lexical content may also be considered. The acoustic model 617 is, for example, system 100 in Figure 1. It is also possible. Video model 618 may, for example, consider footage of the subject's facial expressions. Furthermore, metadata model 619 includes the age, race, ethnicity, gender, sex, income, education of the subject. Other factors related to the subject, such as location and medical history, can be taken into consideration. Model 615 is a Consume pre-processed input data from S610 to determine the behavioral or mental health status of the subject. It can be evaluated, screened, predicted, or monitored. Each model produces a separate output. It can be generated. However, the models may be interdependent. That is, One model can generate its own output by consuming the output of another model.

[0123] The output of each model is provided for calibration, confidence, and the desired descriptor module 625. This module 625 can calibrate the model output and scale the scaled score. Module 625 can generate A and generate a reliability scale for the score. Readable labels can be assigned to scores. Module 625 modifies its output The weight and fusion engine 630 can be provided. The engine 630 has a model output This is incorporated into an integrated classification of the behavior or mental health status of the subject from which the input data originated. It can be matched. Engine 630 can apply static weights to Model 615. It is possible. Alternatively, the weights may be dynamic. For example, the weights of a given model output may be... In several embodiments, the model can be modified based on the confidence level of the classification. For example, NLP Model 616 classifies an individual as not being pushed down with a confidence level of 0.56. However, acoustic model 617 renders a classification that is pushed down to a confidence level of 0.97. In this case, engine 630 can be given a greater weight than acoustic model 617.

[0124] In some cases, the weights of a given model scale linearly according to its reliability level. The model's base weights may be multiplied by the time. In some cases, the model's output weights may be multiplied by time. It may also be the base. For example, engine 630 is generally when the subject is speaking. While NLP model 6161 can be assigned larger weights, the subject is speaking... When not available, a larger weight can be assigned to the video model 618. If acoustic model 617 and video model 618 suggest that the subject is not telling the truth ( For example, due to frequent eye movements, pitch modulation, or increased speech rate, engine 63 0 allows for the application of lower weights to the NLP model 616.

[0125] Engine 630 provides its fused and weighted output to the multiplex output module 635. The multiplex output module can provide fused and weighted outputs combined with other information. In combination, the final result, for example, a prediction of the subject's behavioral or mental health status, is generated. It is possible.

[0126] Fusion considers not only the model inputs but also the range of information that has different effects on the model. This is possible. Examples of information that have different effects on the model include the spread of states and the distribution of label values. (Data skew patterns), metadata, sample length, sample data quality, etc. are included. Born.

[0127] System 600 queries or It can be used with an automated query module that presents a sequence of queries to the target audience. The automated query module is partially based on one or more target mental states to be evaluated. Based on this, queries can be presented and / or formulated. Queries are based on at least the following from the target audience. It can also be configured to elicit one response. The automated query module must have at least one To elicit a response, query the target audience in audio, visual, or text format. It can be sent to. The automated query module will receive at least one response from the target. It is possible to receive data including audio and video data from the subject. The system 600 obtains, for each of multiple different sessions, for a single session. Therefore, or upon completion of one or more of several different sessions Using audio data, video data, and metadata about the subjects, we can identify the subjects and related information. It can generate one or more assessments of the assigned mental state.

[0128] Patient health questionnaire 9 designed to screen patients for depression Compared to conventional screening tools such as ("PHQ-9"), System 600 is It can be more attractive and lead to a higher level of adoption. System 600 (for example, The composite acoustic and NLP models are at least approximately 60%, 65%, 70%, 80%, and 85%. , it may have a specificity of 90%, 95%, or more. System 600 has at least approximately Having a sensitivity of 60%, 65%, 70%, 80%, 85%, 90%, 95%, or higher. Obtain. System 600 is at least about 60%, 65%, 70%, 80%, 85%, 90%. The system may have an area under the curve (AUC) of 95% or more. At least 1%, 2%, 3%, 4%, 5%, 10%, 15%, 20%, 25% more than Tem. To provide an improvement in relative performance (e.g., sensitivity, specificity, or AUC) of % or more. It is possible.

[0129] System 600 also compares with written questionnaires like PHQ-9 to the target audience. This can lead to a more faithful and complete response from the system. Similar systems can refer to this statement. It is described in PCT / US2019 / 037953, which is fully incorporated into the book.

[0130] Longitudinal modeling The system described herein can be used to track the progression of a patient over time. This can be called longitudinal analysis. In long-term analysis, input from the current session... Generating predictions by supplementing force voice with input voice from one or more past sessions. This is possible. Current and past audio data can be represented as vectors in the response matrix. The model can generate predictions for each vector in the matrix. The longitudinal handler is It is possible to search for any correlation between past audio data and current data. Longitudinal analysis can help return more accurate predictions for current data. Factors that may affect prior values ​​regarding physical health status, such as time of day, day of the week, month, and weather at location. This can be taken into consideration. The model can be trained with this information for better predictive performance. It is possible.

[0131] System Output The system 600 in Figure 6 determines whether the patient is at risk of a mental or physiological condition. It can output an electronic report that identifies the user's electronic device. It can be configured to be displayed in the graphical user interface of the user. This can be the patient themselves or the patient's healthcare provider. Electronic reports are mental or This may include a quantification of the risk of a physiological condition, such as a normalized score. The score is The data can be normalized to the entire population or to a subpopulation of the target group. (Electronic Report) It may also include the confidence level of the normalized score. The confidence level is normalized This can indicate the reliability of the score (i.e., the degree to which the normalized score is reliable).

[0132] Electronic reports may include visual and graphic elements. For example, a patient may have several different conditions. Having multiple scores from multiple screening or monitoring sessions that occurred during that time. In this case, the visual graphic element could be a graph showing the progression of the patient's score over time. .

[0133] System 600 is used by patients or patient-related contact persons, healthcare providers, healthcare payers, and It can output electronic reports to another third party. System 600 is a screening Even while monitoring or diagnostics are in progress, electronic reports can be generated in virtually real time. This is possible. Normalized scores or confidence levels in the screening, monitoring, or diagnostic process. In response to changes in reliability, electronic reports are updated and resent to users in virtually real time. It is possible.

[0134] In some cases, the electronic report may include one or more descriptors relating to the patient's mental state. It is possible. The descriptor can be used as a qualitative measure of the patient's mental state (e.g., "mild depression"). This is possible. Alternatively or additionally, the descriptor may include topics mentioned by the patient during screening. It could be a word cloud. The descriptor can be displayed graphically, for example, in a word cloud. .

[0135] The models described herein are for a specific purpose or for receiving the output of a system. It can be optimized based on entities that can perform this task. For example, the model can be optimized based on whether the patient has a high success rate. It may be optimized for sensitivity when estimating whether or not a person has a divine state. (e.g., insurance companies) Healthcare payers are minimizing the number of insurance payments made to patients with false positive diagnoses. In some cases, such a model is preferred so that it can be obtained. In other cases, the model This may be optimized for specificity in estimating whether a patient has a mental state. Healthcare providers may prefer such a model. The system is a relationship in which output is sent. The system can select an appropriate model based on the user. After processing, the system outputs the results to the relevant parties. It can be sent.

[0136] Alternatively, the models described herein may be used by clinicians, healthcare providers, insurance companies, or government regulations. Voice and It can be adjusted or configured to process other data. Alternatively or additionally, the model can be refined. Degree, recall, F1, equivalent error rate ("equal error rate: EER") Positive predictive value (PPV), Negative predictive value (NPV), positive sexual likelihood ratio (``LR+''), negative likelihood ratio (``likelihood ratio nega tive:LR-"), coincident correlation coefficient ("concordance correlati") Pearson correlation coefficient (CCC), Pearson c orrelation coefficient:PCC”), root mean square error (“ Root mean squared error (RMSE), mean absolute error (m EA Absolute Error (MAE), or any other relevant performance metrics Rick can be adjusted, configured, or trained to optimize it.

[0137] The electronic report is a "word cloud" extracted from the text transcript of the patient's voice or It may include a "topic cloud". Word clouds can be displayed in larger font sizes, with different font sizes. Using different colors, different fonts, different typefaces, or any combination thereof most frequently This may also be a visual representation of individual words or phrases using the words and phrases specified. Describing the frequency of a word or phrase in this way generally indicates that depressed patients use it more often than non-depressed patients. It can be useful because certain words or phrases are said with high frequency. For example, depressed people often say dark words. They may use words or phrases that indicate a dark or morbid mood. They felt worthless. They talk about feeling like they've failed, or they say "always," "never," or " They may use absolute words such as "completely." People with depression also tend to be more sensitive to the general population. Compared to this, the first-person pronouns with higher frequency (e.g., "I", "me") and the second-person pronouns with lower frequency It can use personal or third-person pronouns. The system is trained with a machine learning algorithm to drop We performed a semantic analysis of the word sets of people who are depressed and those who are not, and based on the word sets, we analyzed the meaning of the words. It can be classified as either depressed or not depressed. Word cloud analysis is, It can also be done using unsupervised learning. For example, the system is labeled Analyze non-existent word sets, search for patterns, and separate people into groups based on their mental state. It is possible. The generated words may indicate a decrease or increase in the risk of depression (that is, (This is associated with an increased or decreased risk of depression.)

[0138] Similarly, electronic reports may include predicted personality traits of the patient. Personality traits (e.g., Introversion or extroversion can be inferred from the length of the utterance.

[0139] The electronic report may also include evidence-based psychoeducational materials and support strategies. The support strategy can be adjusted to the patient's score. Materials and support strategies are available. It may be provided directly to the patient in the form of images, texts, and assignments, or as materials and support. The abbreviation may be provided to the patient's healthcare provider who can guide the psychoeducational process.

[0140] Usage example The acoustic and NLP models described herein were used to monitor teenagers for depression. It can be used for this purpose. The model independently classifies teenagers as being at risk of depression. To determine voice-based biomarkers that can be used, a group of teenagers were subjected to a machine. Machine learning analysis can be performed. Teenage depression may have different causes than adult depression. Hormonal changes also introduce teenage behaviors that would otherwise be atypical for adults. It is possible. Systems for screening or monitoring teenagers are necessary for these unique circumstances. A model that is tuned to recognize motion is needed. For example, when depressed or Teenagers who are upset are more likely to express anger than adults who are more likely to withdraw when upset. They may also be more easily irritated. Therefore, the questions from the assessment should be different from those for adults. Voice-based biomarkers can be induced in teenagers. When conducting tests, or when studying the mental state of teenagers, screening adults or Screening or monitoring methods different from those used for monitoring It can be used. Clinicians can use voice-based biomeris specific to teenage depression. The evaluation can be modified to specifically induce certain behaviors. The system uses these evaluations. They are trained to determine teenage-specific models for predicting mental states. Teenagers also have families (foster care, adoptive parents(s), two biological parents) (Care provided by one biological parent, guardian / relative, etc.), medical history, sex, age, and socioeconomic status. The model may be segmented by state, and these segments are incorporated into the model's predictions. That's fine.

[0141] The models described herein can also be used to monitor older adults for depression and dementia. It can be used. Older adults also possess certain voice-based biomarkers that younger adults do not have. It is possible. For example, older people may have a tense or thin voice due to aging. Elderly individuals may exhibit aphasia or dysarthria, and survey questionnaires, follow-up, or meetings may be necessary. They have difficulty understanding spoken language and may use repetitive language. Clinicians consider elderly patients to be particularly susceptible. Develop research to extract specific voice-based biomarkers from or algorithms It can be developed using Zoom. Specifically, to predict the mental state of elderly patients. By categorizing patients by age, it is possible to develop machine learning algorithms. Different worlds may have different views on gender roles, morals, and cultural norms. Differences may exist among elderly patients in their 20s. The model considers age bracket, gender, race, and socioeconomic status. It can be trained to incorporate the patient's condition, physical medical status, and family involvement.

[0142] The system could be used to test airline pilots for their mental health. Airline pilots have a tough job and experience a lot of stress from long flights. And fatigue may be experienced. Clinicians or algorithms may be used to diagnose these conditions. A screening or monitoring method can be developed for this purpose. For example, a system This is the Minnesota Multiphasic Personality Questionnaire. c Personality Inventory (MMPI) and MMPI-2 are used for testing. It may also be based on an evaluation of a similar query.

[0143] The system could also be used to screen military personnel for mental health reasons. For example, the system is used to examine PTSD, and to diagnose primary care post-traumatic stress disorder. Similar subject matter as asked in the Statistical Manual of Statistics (DSM)-5 (PC-PTSD-5). An evaluation can be performed using queries that have the following characteristics: In addition to PTSD, the system can also detect depression, Screening military personnel for panic disorder, phobic disorder, anxiety disorder, and hostility. The system can do this. The system uses different methods to screen soldiers before and after deployment. The system can be used by the military by segmenting by occupation. People can be segmented by branch, officer or subcontractor, gender, age, ethnicity, travel / dispatch Military personnel can be segmented by the number of times they have been deployed, their marital status, their medical history, and other factors. Cut.

[0144] The system, for example, performs background checks to determine the likelihood of success. It can be used to evaluate gun buyers. The evaluation assesses their mental readiness to own small arms. To evaluate promising buyers, clinicians or algorithmically designed The investigation uses questions and follow-up questions to find out if a prospective gun buyer has ever been in court or To determine whether it could be deemed a danger to him or others by other authorities. It may have the following requirements.

[0145] Scoring The models described herein generate scores at various stages of mental or behavioral health assessment. It is possible to do this. The generated score can be a scaled score or a binary score. It is also possible. Scaled scores can have multiple values, and binary scores are, It can be one of two discrete values. The model is also used to monitor different mental states. Throughout the evaluation process, specific mental states are assessed using specific binary scores and specific scales. To update the graded score, binary scores and scaled scores are used at various stages of evaluation. You can exchange scores.

[0146] The scores generated by the binary or scaled system are used for each component in the evaluation. It may be generated after each response to Eli, or formulated based in part on the previous query. It may be modified. In the latter case, each limit score fine-tunes the prediction of depression or another mental state. and acts to make predictions more robust. Peripheral predictions are (certain intermediate mental states) (correlated with) After a certain number of queries and responses, thus for predicting mental state This can increase the level of trust.

[0147] In the case of scaled scores, an improvement in the score means that clinicians can tell the patient that the patient is experiencing This may make it possible to determine the severity of one or more mental states with greater accuracy. For example, an improvement in the scaled score is observed when observing multiple intermediate depressive states. This allows clinicians to determine whether a patient has mild, moderate, or severe depression. This is possible. Performing multiple scoring iterations also adds redundancy and increases robustness. By adding this, clinicians and administrators can help eliminate false negatives. For example, predicting the initial mental state is difficult because there are relatively few audio segments available for analysis, and NL The P algorithm is sufficient to determine the semantic context of the patient's recorded speech. Because it may not have information, it may be noisy. A single peripheral prediction itself is noisy. Even with estimates containing many variances, predictions can be refined by adding more measurements. This reduces the overall variance of the system, resulting in more accurate predictions. The survey is not simply conducted because people may have motives lying about their condition. This may be more practical than the predictions that can be obtained by doing so. When the investigation is conducted, multiple false positives Sexual and false negative results are obtained, allowing patients who need treatment to slip through the cracks. Furthermore, trained clinicians can recognize voice and face-based biomarkers. It is possible, but it is not possible to analyze the large amount of data that the models disclosed herein can analyze. It may not be possible.

[0148] Scaled scores can be used to describe the severity of mental state. The scored number may be, for example, between 1 and 5, or between 0 and 100. Often, a larger number indicates a more severe or acute form of the mental state experienced by the patient. The scaled score may include integers, percentages, or decimals. Symptoms for which the score can indicate severity include depression, anxiety, stress, PTSD, and fear. This includes, but is not limited to, phobic disorders, schizophrenia, and panic disorder. For example, a score of 0 on the depression-related aspects of the assessment may indicate the absence of depression, while a score of 50 may indicate the absence of depression. A score of 100 may indicate moderate depression, while a score of 100 may indicate severe depression. The score may be a combination of multiple scores. The mental state is a composition of mental substates and It may also be expressed as follows: the patient's complex mental state is expressed as individual scores from mental substates. A weighted average may also be used. For example, the composition score for depression could be anger, sadness, self-image, self This could be a weighted average of individual scores for values, stress, loneliness, isolation, and anxiety.

[0149] The scaled score is generated using a model that employs a multi-label classifier. This classifier could be, for example, a decision tree classifier, a k-nearest neighbor classifier, or a neural network classifier. It may also be a classifier based on a mark. The classifier identifies specific patients at the intermediate or final stage of evaluation. Multiple labels can be generated for a person, and the labels indicate the severity or degree of a particular mental state. It indicates the degree. For example, a multi-label classifier uses a softmax layer to normalize the probabilities. It can output multiple possible numbers. The label with the highest probability represents the mental state the patient experienced. It can indicate the severity of the condition.

[0150] The scaled score may also be determined using a regression model. The fit can be determined from training examples, which are represented as a sum of weighted variables. It can be used to extrapolate scores from patients with known weight. The weights are audiovisual. Partially derived from signals (e.g., voice-based biomarkers), such as patient demographics. It may be partially based on features that can be partially derived from patient information. Final score or intermediate score The weights used to predict the score can be obtained from previous intermediate scores. .

[0151] The scaled score may be scaled based on a confidence scale. The degree refers to the recording quality, the type of model used to analyze the patient's voice from the recording (e.g., For example, audio, visual, and semantic, which model had the most users during a specific period? Time analysis related to usage, and specific audio-based biometrics within audiovisual samples. The decision can be made based on the marker's time point. Multiple confidence levels are used to determine the intermediate score. A scale can be adopted. The reliability scale being evaluated is a specific scaled score. It may be averaged to determine the weighting for each factor.

[0152] A binary score can reflect the binary result from the system. For example, the system The system can classify whether a user is depressed or not. This can be done using classification algorithms such as neural networks or ensemble methods. It is possible to do so. A binary classifier can output a number between 0 and 1. If A exceeds the threshold (for example, 0.5), the patient may be classified as having "depression." If the score falls below the threshold, the patient may be classified as “not depressed.” The system, The system can generate multiple binary scores for multiple intermediate states of the evaluation. To generate an overall binary score for evaluation, binary scores from the intermediate states of the evaluation are used. It can be weighted and summed up.

[0153] The output of the models described herein is a calibrated score, for example, a score with a unit range. It can be converted to a clinical model. The output of the models described herein may be used additionally or alternatively in clinical applications. It can be converted into a score with clinical value. A score with clinical value is used for qualitative diagnosis. (For example, a high risk of severe depression.) Or, a score with clinical value may be , a normalized qualitative score that is normalized for the general population or a specific subgroup of patients It is possible. Normalized qualitative scores are risk percentages for the general population or subpopulation. It may show a page.

[0154] The system described herein is more effective than standardized mental health questionnaires or assessment tools. With a small error (e.g., less than 10%) or higher accuracy (e.g., 10% or more), the target group It may be possible to identify mental states (e.g., mental disorders or behavioral disorders). Error rate Alternatively, accuracy is used to identify or assess one or more medical conditions, including mental states. It can be established against usable benchmark criteria by the entity. This could be a clinician, healthcare provider, insurance company, or government regulatory body. (Benchmark) The criteria may be independently validated clinical diagnoses.

[0155] Confidence scale The models described herein can be used with confidence scales. Confidence scales are used for depression. A score generated by a machine learning algorithm to accurately predict which mental state This can be a measure of how effective it may be. The confidence scale is based on the conditions under which the score was obtained. It may depend on [something]. Confidence measures can be expressed as integers, decimals, or percentages. The conditions include the type of recording device, the surrounding space where the signal was acquired, background noise, the patient's speech habits, and the speaker. The patient's verbal fluency, the length of their responses, the assessed truthfulness of their responses, and the incomprehensible units of speech. This may include the frequency of words and phrases. The quality of the signal or sound makes it more difficult to analyze the sound. Under certain conditions, the confidence scale may have a smaller value. In some embodiments, the calculated By weighting binary or scaled scores with confidence levels, the confidence level can be calculated. A can be added to the calculation. In other embodiments, the reliability measure may be provided separately. For example, the system can determine that a patient has a depression score of 0.93 with 75% confidence. This can be communicated to clinicians.

[0156] The reliability level is also the training data used to train the model that analyzes patient speech. It may also be based on the quality of the label. For example, if the label is not a formal clinical diagnosis, the patient If based on a survey or questionnaire completed by [company name], the label quality is judged to be lower. Therefore, the reliability level of the score may be lower. In some cases, the survey or The survey may be determined to contain a certain level of fraud. In such cases, the label The quality may be judged as lower, and therefore the reliability level of the score may be lower.

[0157] In particular, if the reliability scale is affected by the environment in which it is evaluated, it is necessary to improve the reliability scale. To that end, various scales may be used depending on the system. For example, the system may use one or It uses multiple signal processing algorithms to remove background noise or the impulse response Measurements are used to determine what is caused by objects and features in the environment in which the audio sample was recorded. The system can determine how to eliminate the effects of reverberation. Use the context to determine the identity of missing or incomprehensible words. You can find clues.

[0158] Furthermore, the system uses user profiles to gather information on behavior, ethnic background, gender, and age. People can be grouped based on similar groups or other categories. Since these people may have similar voice-based biomarkers, similar voice-based biomarkers Since people who exhibit symptoms may exhibit depression in similar ways, the system can diagnose depression with greater reliability. It may be possible to predict the onset of the disease.

[0159] For example, people with depression from different backgrounds may speak slowly, with a monotonous pitch or Low pitch fluctuations, excessive pauses, vocal tone (rough or noisy voice), Disorders may manifest as inconsistent speech, distraction or loss of concentration, silent responses, and stream-of-consciousness narratives. These can be classified into several categories. These voice-based biomarkers are one or more of the patients analyzed. It may belong to the number segment.

[0160] Clinical scenarios The models described herein analyze voices from primary care and health interactions. It is possible to use the system to enable trained healthcare providers to perform detailed assessments of patients. It can enhance inferences regarding divine health. The system can also perform preliminary screening or This includes monitoring calls (for example, setting up medical appointments with trained mental health professionals). For the purpose of this, mental health is assessed from calls made to healthcare provider organizations by promising patients. It can be used to evaluate. For primary screening, medical professionals assess the patient's mental state. To determine the need for health treatment, patients can be asked specific questions in a specific order. The recording device records promising patient responses to one or more of these questions. This can be done. Before this is done, consent can be obtained from promising patients. The model described can process voice snippets collected from promising patients.

[0161] The system uses standard clinical encounters to train voice biomarker models. The system can collect records of clinical encounters regarding physical complaints. The complaint may relate to injury, illness, or a chronic condition. The system will assess the patient's condition. With permission, conversations between the patient and healthcare provider during an appointment can be recorded. Complaints can indicate feelings about a patient's health. In some cases, physical complaints may indicate that the patient is experiencing physical discomfort. It causes significant distress, affects the patient's overall character, and in some cases leads to depression. It could cause that.

[0162] Audio-based biomarkers can be associated with experimental or physiological measurements. Biomarkers of this condition can be associated with mental health-related measurements. For example, they can be linked to: The effectiveness of psychiatric treatment, or compared with logs collected by medical professionals such as therapists. They determine whether voice-based analysis aligns with assessments commonly performed in the field. For verification, it can be compared with the answers to the survey questions.

[0163] Voice-based biomarkers can be associated with physical health-related measurements, for example, disease. These vocalization issues, among others, produce vocal sounds that need to be considered in order to generate feasible predictions. It can contribute to patients' recovery. Furthermore, it can contribute to the time scale of patients' recovery from illness or injury. Comparing depression predictions over time with patient health outcomes over that timescale, treatment can help patients It can be used to check whether depression or depression-related symptoms are improving. The biomarkers were collected at multiple time points to determine the clinical effectiveness of the system. This can be compared with data on brain activity.

[0164] Model training is performed continuously while audio data is being collected. Therefore, it may be continuous. Continuously add voice-based biomarkers to the system. The model can be used for training between multiple epochs. It can be updated using the 'Ta' command.

[0165] The system can use a reinforcement learning mechanism, and in this mechanism, reliability To induce voice-based biomarkers that lead to a high prediction of depression, survey questions were used. It can be modified accordingly. For example, a reinforcement learning mechanism selects a question from a group. It may also be possible. Based on a previous question or a sequence of previous questions, the reinforcement mechanism may This allows for the selection of questions that may yield highly reliable predictions of depression.

[0166] The system determines which questions or sequences of questions may elicit specific responses from the patient. The system can determine this. The system uses machine learning to, for example, generate probabilities. This allows for the prediction of specific triggers. The system also uses a softmax layer. Using this, multiple probabilities of triggering can be generated. The system can answer specific questions, as well as... When will these questions be asked, the time until the survey in which they were asked, the time the questions were asked, and the questions themselves. Points within the treatment course can be used as features.

[0167] The system uses voice-based biomarkers to dynamically influence the course of treatment. This may include methods to record user triggers over a certain period of time. From the induction, it is possible to determine whether the treatment was effective or not. For example, voice-based If Iomarkers show less depression over a long period of time, this indicates that the prescribed treatment has been completed. This could serve as evidence that the treatment is effective. On the other hand, voice-based biomarkers over a long period of time If the symptoms become more depressive, the system will prompt healthcare providers to change treatments. This may encourage them to do so, or to pursue the current treatment process more actively.

[0168] The system can spontaneously recommend changes in treatment. The system continuously collects data. In embodiments where processing and analysis are performed, the system processes depression (or another mental disorder or This allows for the detection of sudden increases in voice-based biomarkers indicating behavioral disorders. This can occur over a relatively short timeframe during the course of treatment. The system also involves a series of treatments. If it has been inactive for a specific period (e.g., 6 months, 1 year), we will voluntarily recommend making changes. It is possible.

[0169] The system may be able to track the probability of a specific response to a drug. For example. The system then uses voice-based biometric data collected before, during, and after a series of treatments. By tracking markers, it is possible to analyze changes in scores indicating mental or behavioral disorders.

[0170] The system, being trained on similar patients, can determine the appropriate medication for a particular patient. The system can track the probability of a response. The system uses this data to track similar populations. Based on patient responses from statistics, patient responses can be predicted. The total may include age, sex, weight, height, medical history, or a combination thereof.

[0171] Furthermore, the system analyzes the patient's biomarkers based on the questions asked. By analyzing, it is possible to tell whether or not the treatment is continuing. For example, the patient, Becoming defensive, stopping for long periods, cramming, or the patient lying down faithfully according to the treatment plan They can act in such a way. The patient also feels sad about not following the treatment plan. It can express shame, embarrassment, or sadness.

[0172] The system can predict whether a patient will follow a series of treatments or medications. Tem uses data from multiple patients to predict whether a patient will continue a series of treatments. The system can use training data from voice-based biomarkers. Specific voice-based biomarkers can be identified as predictors of defense. For example... If a patient has voice-based biomarkers indicating fraud, they are less likely to adhere to their treatment plan. It may be specified as low.

[0173] The system can establish a baseline profile for each individual patient. Each patient may have a specific speech style and specific voice-based biomarkers. It expresses emotions such as happiness, sadness, anger, and sorrow. For example, some people Some people laugh when they feel stressed, or shout when they are happy. To speak loudly or softly, to speak clearly or whisper, to have a broad or narrow vocabulary. They may speak freely or with more hesitation. Some people may have an extroverted personality. However, other people may be more introverted.

[0174] Some people may hesitate to speak out more than others. Some people feel Some people may become more cautious about expressing their emotions. Sometimes, some people may be denying their own feelings.

[0175] A person's baseline mood or mental state, and therefore a person's voice-based biomarker, This can change over time. The model may be continuously trained to account for this. The model also doesn't need to predict depression very often. The model's predictions over time are precise. These results can be recorded by medical professionals. These results indicate a progression from the patient's depressive state. It can be used for that purpose.

[0176] The system creates a certain number of profiles to take into account various types of individuals. In some cases, these profiles may include, for example, an individual's gender, age, ethnicity, This may be related to the language used and the occupation.

[0177] Certain profiles may have similar voice-based biomarkers. For example, the elderly. They may have a thinner, more breathy voice than younger people. These weak voices are often heard through a microphone. This could make it more difficult for the von to pick up certain biomarkers, and they are young people. They may speak more slowly than others. Furthermore, older adults may contaminate behavioral therapy, and However, younger people may be less likely to share multiple pieces of information.

[0178] Men and women may express themselves differently, and this is due to different biomarkers. This can lead to, for example, men being able to express more positive or more intensely negative emotions. Women are better able to articulate their emotions.

[0179] In addition, people from different cultures have different ways of dealing with or expressing emotions. Sometimes, or when expressing negative emotions, one may feel guilt and shame. To make the system more effective in acquiring different voice-based biomarkers, culture It is sometimes necessary to segment people based on their background.

[0180] The system segments and clusters by personality type, It is possible to consider people with different personality types. This allows clinicians to consider personality types. They may be familiar with how those types of people might express feelings of depression. Because it is possible, it can be done manually. Clinicians can analyze these segmented groups of people. Developing specific survey questions to extract specific voice-based biomarkers from various sources. It is possible.

[0181] Voice-based biomarkers can detect if a person is concealing information or attempting to circumvent testing methods. Even if that is the case, it should not be used to determine whether that person is depressed. This is possible because many voice-based biomarkers can be involuntary speech. For example, a patient may be vague, or their voice may tremble.

[0182] Certain voice-based biomarkers may correlate with specific causes of depression. For example, depression To find specific words, phrases, or sequences thereof that indicate a disease, meaning for multiple patients Analysis is performed. The system also determines the effectiveness of the user by treating the user. The effectiveness of the treatment options can be tracked. Finally, the system is better available. Reinforcement learning can be used to determine treatment methods.

[0183] Additional usage examples The systems disclosed herein enhance the care provided by healthcare providers. It can be used for, for example, one or more of the disclosed systems, patient care providers It can be used to facilitate the transfer to a specific mental state threshold after the system has evaluated it. If the system generates a score exceeding the value, it will dedicate the patient to further investigation and analysis. They can be referred to a specialist. For example, if the patient is receiving treatment through a telemedicine system. Alternatively, if a specialist is in the same location as the patient, the patient may be referred before the assessment is complete. Yes. For example, a patient may receive treatment at a clinic with one or more specialists.

[0184] The disclosed system can direct the patient's clinical process after scoring. For example, if a patient was being assessed using a client device, the patient would be assessed After completion, cognitive behavioral therapy You can refer to y:CBT services. They are also called healthcare providers. You may have appointments with healthcare providers, or appointments made by the system. The system may suggest one or more medications. The system may also suggest specific dietary therapies. Alternatively, further exercise therapy may be suggested. The recommended exercise regimen is at least part In particular, the patient's demographics (e.g., age and sex), past medical history, or health data generated by the patient. It may also be based on data (e.g., weight, cardiovascular or lung health).

[0185] The systems and models described herein can be used for accurate case management. In the first surgery, the patient converses with the case manager. In the second operation, one or more entities Tee passively records the conversation with the patient's consent. The conversation may be face-to-face. i. In another embodiment, the case manager can conduct a remote conversation. For example The conversation may also be a conversation using a telemedicine platform. In the third action, The models described herein process recorded conversations and provide real-time results to the payer. It can be sent to [a specific location]. Real-time results may include a score corresponding to mental state. In the fourth step, the case manager updates the care plan based on real-time results. For example, a certain score exceeding a certain threshold can be used to communicate between caregivers and patients. This could affect future interactions and may allow providers to ask patients different questions. It has a tendency. The score triggers the system to suggest specific questions related to the score. It's even possible to do that. The conversation can be repeated in the updated care plan.

[0186] The systems and models described herein are for primary care screening or monitoring. It can be used in the ring. In the first surgery, the patient visits a primary care provider. In operation 2, the voice is captured by the primary care provider's tissue for electronic transcription. The system may also provide a copy for analysis. In the third step... Primary care providers can use the analysis to inform care pathways in real time. It can be received. This can facilitate a warm handover to a behavioral health professional. It can be used to direct primary care providers along a specific care pathway.

[0187] The systems and models described herein are enhanced employee support plans. eAssistance Plan (EAP) is used for navigation and triage. It is possible. In the first action, the patient can call the EAP line. The second action In the step, the system records audiovisual data and screens patients. Yes, it is possible. Real-time screening or monitoring results are available to the provider in real time. It can be distributed to. Based on the real-time results collected, the provider can determine high-risk Patients can be adaptively screened for specific topics. Real-time screening Leaning or monitoring data may also be provided to other entities. For example If real-time screening or monitoring data is provided to clinicians on call, It may be provided, used to schedule introductions, and used for educational purposes. It may be used for other purposes. The interaction between the patient and the EAP is direct. It may be in person or remotely. The person in charge of the EAP line should have a positive screen when the patient is present. It can help guide patients to the appropriate level of treatment by providing real-time warnings. EAP can also respond to the results of assessments performed on the patient, for example, the patient's mental state. You may be instructed to ask questions based on the score. The audio data described herein Data that may be collected and analyzed in real time, or recorded and analyzed later. That's fine.

[0188] Telemedicine In some cases, the models described herein may refer to patients and healthcare providers (health care providers). From one or more telemedicine sessions with a re provider (HCP) It can process audio and video. Figure 19 shows the telemedicine system 1900. The telemedicine system 1900 allows patients and HCPs to conduct telemedicine sessions regarding the patient's health. This makes it possible to perform the procedure. The telemedicine system 1900 is a patient device 1905, HCP device 1910, telemedicine server 1915, and telemedicine database This may include S1920, patient device 1905, HCP device 1910, and telemedicine. Server 1915 can communicate via network 1930. Patient device 1905 and HCP device 1905 are mobile devices (e.g., smartphones) This could be an electronic tablet, laptop, or desktop computer.

[0189] Patient device 1905 and HCP device 1910 are used with telemedicine application 19 It can run 25 instances. Telemedicine application 1925 is Tandoron's desktop application, web application, mobile application This could also be a reference, etc. Each instance of telemedicine application 1925 This means that a user of that instance (e.g., a patient) interacts with another user (e.g., a healthcare provider). It may have a user interface that enables the establishment of a secure communication link. The user interface allows the user to access the user's device (for example, patient device 190). 5) Use the camera and microphone above to record audio and video, and other users Recorded by other users using the other device (e.g., HCP device 1905) It can be made possible to consume the audio and video that have been recorded. Two devices It continuously exchanges audio and video streams over a secure communication link. It can be exchanged, facilitating real-time video conferencing between two users. Telemedicine Each instance of Application 1925 is an audio stream and a video stream. It may have audio and video codecs for compressing and decompressing files. Depending on the circumstances, the user interface may display demographic or clinical information about the patient. Further information can be displayed on the CP. This information is provided by the telemedicine server 1915. You can also search the Telemedicine Database 1920.

[0190] The telemedicine system 1900 stores audio and video from video conferences into a telemedicine database. It can be stored in 1902. Subsequently, the acoustic, NLP, and video described herein The model processes audio and video, for example, one of the participants in a video conference. For example, it can be determined whether a patient has a behavioral or mental health disorder.

[0191] Additionally or alternatively, the telemedicine system 1900 may be used when a video conference is taking place. It can process audio and video from patients in real time. In addition, the Telemedicine Database 1920 is a collection of acoustic, NLP, and video models as described herein. It can store the data. The telemedicine server 1915 receives orders from the patient device 1905. Obtain audio and video streams and select the appropriate model from the telemedicine database 1920. The system retrieves the data and processes the audio and video streams using a model to identify the patient's behavioral impairment. It can determine whether or not a person has a physical or mental disorder. The telemedicine server 1915 can determine whether or not a person has a physical or mental disorder. The model output is provided in real time to the user interface of the HCP device 1905. It is possible to do this. The output includes qualitative or quantitative scores, confidence intervals, word clouds, etc. The output may include any of the outputs described herein, including the video conference with the patient. It can support HCP when providing guidance. The telemedicine server 1915 is based on the output. The patient's user interface can be further modified. For example, the output may change to the patient's interface. If it indicates that the patient is depressed, the telemedicine server 1915 will provide cognitive behavioral therapy options. This can be added to the user interface.

[0192] In the case of the real-time processing described above, the telemedicine server 1915 will find available information about the patient. By using demographic or clinical data, the telemedicine database 1920 You can select the appropriate model from among them. For example, the telemedicine server 1915 is a device that allows patients to If you are a young person, you should use a youth model (for example, one that primarily provides audio and video instruction from young people). A refined model can be selected. Additionally or alternatively, the telemedicine server 191 5. If such demographic information is not yet known, use an image recognition process Demographic information about the patient can be determined. For example, the telemedicine server 1915 Using an image recognition process, it is possible to determine the patient's gender, age, race, etc.

[0193] In some cases, the patient's speech may be modeled as described herein immediately before the telemedicine session. This can be analyzed, and as a result, during the session, the healthcare provider can predict the patient's Questions can be asked to assess the patient's condition. In other cases, immediately after the telemedicine session. It is possible to analyze the patient's speech.

[0194] In telemedicine or in-person clinical encounters, the patient's voice characteristics should be matched to those of the healthcare provider. It may be beneficial to do so. By doing so, it increases the possibility of achieving intimacy with the patient. It can be raised.

[0195] In some cases, the telemedicine system 1900 connects the patient to a "care buddy". This is possible. Caregivers can take into account at least the location, age, behavioral or mental state, personality traits, etc. Assignments may be made based on the number of people. Communication between the patient and their caregiver is handled by the telemedicine system. This may also be done via Tem1900. Caregiver buddies are provided with a communication template. This may include weekly check-in calls and questions to be asked of each other during the call. That's fine.

[0196] quality control There may be situations where the input voice provided by the patient is unacceptable. In such cases, The system described in the specification can flag input voice in real time. In the example, the corresponding user is unable to generate audio, or the quality is below optimal or It can generate sound in terms of quantity. The acoustic quality detector analyzes the collected sound and the sound It can generate a real-time warning if the quality (e.g., its volume) is too low. The system can also determine the total number of words in real time, and if the number of words is sufficiently high If not available, a new set of prompts can be supplied. The new prompts are It may be designed to elicit longer or more multiple responses. In another example, the user, One can try to make the system into a game (for example, to get an incentive, or (to avoid diagnosis). In such cases, the ASR model considers the voice to be "good". The audio can be processed to determine if it differs significantly from the user's voice. Next, the input from the test user is compared with this model in real time to identify word patterns. This method checks whether it is too far from what good users expect. A user who plays audio from another source instead of speaking live to Tem, or who is asked It is possible to capture users who talk about the question but do not attempt to talk. And the system The system can either display a warning to the user or tag the audio file.

[0197] Non-speech model In some cases, the systems described herein include a breathing model, a laughing model, and a pause. This may include non-speech models, including stop models. Modeling of breathing is useful in predicting anxiety or mania. It may be useful. Modeling laughter (or its absence) may be useful in predicting depression. Pauses may also indicate certain behavioral or mental health conditions. Output of non-speech models. This can be fused with the output of the acoustic model.

[0198] Neural Network This disclosure describes various types of neural networks. The system uses multiple layers of computation to produce one or more outputs, for example, to predict a subject's blood glucose level. It can be used. A neural network is one that is located between the input layer and the output layer. Alternatively, it may include multiple hidden layers. The output of each layer leads to another layer, for example, the next hidden layer or output layer. It can be used as input. Each layer of the neural network responds to the input to the layer. You can specify one or more transformation operations that should be performed. These can be called neurons. The output of a particular neuron is regulated by a bias and activated. Functions, for example, the rectified linear unit (R). This could be a weighted sum of inputs to the neuron, multiplied by eLU or the sigmoid function.

[0199] The step of training a neural network is to train it to generate predictive outputs. The steps involve providing input to a neural network that does not exist and comparing the predicted output with the predicted output. The steps involve adjusting the algorithm's weights to account for the difference between the predicted output and the predicted output. This may include steps to update the bias, and more specifically, using a cost function, The difference between the measured output and the predicted output can be calculated. Network weights and biases. By calculating the derivative of the cost function with respect to the cost function, the weights and biases are obtained from the cost function It can be iteratively adjusted over multiple cycles to minimize the effect. Training is The predicted output is determined by the convergence condition, for example, the cost function, and the calculated cost It can be completed when the condition of being small in size is met.

[0200] This disclosure describes a convolutional neural network (CNN). A CNN is a convolutional neural network. Neurons in several layers called the 3 layers process only a small portion of the input dataset (for example, speech). This is a neural network that receives input from short time segments of data. These small parts can be called the receptive fields of neurons. Each of these neurons within the convolutional layer Rons can have the same weight. In this way, the convolutional layer handles the input dataset. It is possible to detect specific features in the intended part. CNNs also use the nucleus of the convolutional layer. The output of the ¹¹ cluster and the conventional layers of a feedforward neural network are similar. It may have a pooling layer combined with a fully connected layer.

[0201] This disclosure describes recurrent neural networks (RNNs). N is a time-series data that can encode dependencies in audio data. It is a neural network with circular connections. RNN processes a sequence of time-series inputs. It may include an input layer configured to receive. RNNs also have one or more state-maintaining layers. It may contain a number of hidden recurrent layers. At each time step, each hidden recurrent layer is The output of that layer and the next state can be calculated. The next state is calculated from the previous state and the current input. It can depend on force. The state can be maintained over time steps, and input It can capture dependencies within a sequence.

[0202] An example of an RNN is an LSTM, which can be composed of LSTM units. The LSTM unit is... It can consist of a cell, an input gate, an output gate, and a forget gate. The cell is an input cell. It can take on the role of tracking the dependencies between elements within the system. The input gate is new The degree to which a value flows into a cell can be controlled, and the forget gate controls the degree to which a value remains in the cell. The output gate can be controlled, and the value in the cell activates the output of the LSTM unit. The degree to which the transformation is used to calculate the transformation can be controlled. LSTM gate activation relationship The number can be a logistic function. LSTM can also be bidirectional.

[0203] Computer system This disclosure provides a computer system programmed to carry out the methods described herein. Provided. Figure 8 shows the implementation of system 100 in Figure 1, or the training process in Figures 4 and 5. A computer system programmed to perform or otherwise configured to perform This indicates 801.

[0204] Computer system 801 has a single-core or multi-core processor, or parallel A central processing unit (central p) can be configured as multiple processors for column processing. Processing unit: CPU, referred to as "processor" and "computer" in this specification. The computer system includes the "processor" 805. The computer system 801 also includes memory or memory. Location 810 (e.g., random access memory, read-only memory, flash memory) ) and an electronic storage unit 815 (e.g., a hard disk) and one or more other systems A communication interface 820 (e.g., a network adapter) for communicating with the system and , cache, other memory, data storage, and / or electronic display adapter Includes any peripheral device 825, memory 810, storage unit 815, interface 8 20 and peripheral device 825 connect to CPU 805 via a communication bus (solid line) such as the motherboard. It communicates with the data storage unit (or It may also be a data repository. Computer system 801 has a communication interface. -820 operates the computer network ("Network") 830 with the help of 820. It can be connected. Network 830 is the Internet, Internet and / or extranets, or intranets and / or intranets that communicate with the Internet. Alternatively, it can be an extranet. Network 830 may, in some cases , telecommunications and / or data networks. Network 830 is a cloud con One or more distributed computing that can enable pute and other distributed computing This may include computer servers. Network 830 may, in some cases, include computers With the help of system 801, the devices connected to computer system 801 A peer-to-peer network that can be enabled to operate as a liaison or server. The work can be implemented.

[0205] The CPU 805 is a set of machine-readable instructions that can be incorporated into a program or software. It can be executed. Instructions can be stored in memory locations such as memory 810. The instruction may target CPU805, and CPU805 thereafter... The CPU can be programmed or configured to implement the method. Examples of actions performed include fetching, decoding, executing, and writing back.

[0206] CPU 805 may be part of a circuit such as an integrated circuit. The circuit may include several other components. In some cases, the circuit is a collection of components designed for a specific purpose. It is an integrated circuit (ASIC).

[0207] The storage unit 815 stores files such as drivers, libraries, and saved programs. It can store user data, for example, user player data. It can store reference and user programs. Computer System 801 In some cases, computer systems may be accessed via an intranet or the internet. Located on a remote server that communicates with 801, or otherwise outside the computer system 801 It may include one or more additional data storage devices.

[0208] Computer system 801 can connect to one or more remote computers via network 830. It can communicate with the computer system. For example, computer system 801 can communicate with the computer system. It can communicate with the remote computer system of the remote computer system. Examples include personal computers (e.g., portable PCs), slates, or tablets. Net PCs (for example, Apple® iPad®, Samsung® (Registered trademark) Galaxy tablets, phones, smartphones (e.g., Apple (registered trademark)) iPhone®, Android® compatible devices, Blackbe This includes rry(registered trademark), or a mobile information terminal. Users can use network 830. It is possible to access the computer system 801 via this.

[0209] The methods described herein include, for example, a memory 810 or an electronic storage unit 815. Machines (for example, computer processors) stored in the electronic memory location of the Pewter System 801 It can be implemented by executable code. Machine executable code or machine-readable code The code may be provided in software form. During use, the code is on processor 805 This can be executed by taking the code from the memory unit 815. This can be obtained and stored in memory 810 for easy access by processor 805. In some situations, the electronic memory unit 815 can be excluded, and machine executable instructions This is stored in memory 810.

[0210] The code is intended for use in machines that have a processor adapted to run the code. It can be pre-compiled and configured, or compiled during runtime. The code can be executed in a way that the code has been pre-compiled or compiled. It can be supplied in a programming language that can be selected to enable this. Cut.

[0211] The embodiments of the systems and methods provided herein, such as computer system 801, It can be realized in programming. Various aspects of this technology are typically mechanical (or Processor) Executable code and / or carried on some kind of machine-readable medium or The related data that is materialized can be considered as a "product" or "product" in form. The code is available in memory (e.g., read-only memory, random access memory, and cloud memory). It can be stored in electronic storage devices such as hard disks or memory. "This type of media includes tangible memory such as that found in computers and processors, or various semiconductor memory. Includes any or all of related modules such as tape drives and disk drives. It can also be used to provide non-temporary memory for software programming at any time. The software, in whole or in part, is not compatible with the Internet or various other telecommunications networks. Communication may occur via a computer or platform. Such communication may occur, for example, via a computer or platform. From the processor to another computer or processor, for example, a management server or host computer Software from the user to the application server computer platform It can be made possible to load it. Therefore, it can carry software elements. Another type of medium is wired and across the physical interface between local devices. Optical, used via terrestrial networks and various air links. This includes waves, electric waves, and electromagnetic waves. It also includes wired or wireless links, optical links, and other means of transporting such waves. The physical elements used for transmission can also be considered as a medium for carrying software. If used, unless limited to non-temporary and tangible “storage” media, a computer or machine Terms such as "readable medium" in the context of machines are involved in providing instructions to the processor for execution. It refers to any medium.

[0212] Therefore, machine-readable media such as computer executable code are tangible storage media, transport It can take multiple forms, including but not limited to wave media or physical transmission media. Non-volatile storage media are used, for example, to implement databases shown in drawings. Optical or magnetic disks, including any storage device such as a computer. Volatile storage media are such as the main memory of a computer platform. Includes dynamic memory. Tangible transmission media include coaxial cable; computer system Copper wires and optical fibers, including wires with internal buses. The carrier transmission medium is an electrical signal Or electromagnetic signals, or radio frequency (RF) and infrared Sound or light wave forms that are generated during infrared (IR) data communication. It can take this form. Therefore, a common form of computer-readable media is, for example, f Loppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic Airtight media, CD-ROM, DVD or DVD-ROM, Blu-ray (registered trademark), optional. Other optical media, punched card paper tape, any other physical storage media having a pattern of holes Body, RAM, ROM, PROM and EPROM, FLASH (registered trademark) - EPROM, Any other memory chip or cartridge, a carrier wave that carries data or instructions, such A cable or link that carries a carrier wave, or a computer programmed with code and / or any other medium from which the data can be read. Most computer-readable media require one or more sequences of one or more instructions to be executed. It can be involved in transporting data to the processor.

[0213] The computer system 801, for example, can elicit speech from the user, one or more User interface for providing the above query: Includes an electronic display 835 equipped with UI)840, or capable of communicating with it. Examples of UI include graphical user interfaces. User interfaces (GUI) and web-based user interfaces are included. These are rare, but not limited to them.

[0214] The methods and systems described herein can be implemented by one or more algorithms. The algorithm is executed by software during execution by the central processing unit 805. The algorithm may be, for example, the acoustic model, machine learning model, as described herein. Alternatively, it may be either a training process.

[0215] Preferred embodiments of the present invention have been shown and described herein, but such embodiments are not examples. It will be obvious to those skilled in the art that this invention is provided solely as such. The present invention is not intended to be limited by any particular example. As has been made clear, the descriptions and examples of embodiments in this specification should be interpreted in a limited sense. This does not mean that. Without departing from the present invention, several modifications, changes, and substitutions can be conceived by those skilled in the art. It will come to mind. Furthermore, all aspects of the present invention depend on various conditions and variables. Please understand that this is not limited to the specific descriptions, compositions, or relative proportions described in the specification. Various alternative embodiments of the present invention described herein may be used when carrying out the present invention. It should be understood that it can be used. Therefore, the present invention may be used in any such alternative form, modification This is considered to include the original form, the deformed form, or equivalents. The following claims define the present invention. The scope is defined, and the methods and structures within these claims, as well as their equivalents, are defined thereto. Therefore, it is intended to be included.

Claims

1. Using acoustic models including encoders and classifiers to study the behavioral or mental health of subjects A method for detecting a state, (a) The step of obtaining an audio sample containing multiple audio segments from the subject person. 、 (b) Processing the audio sample with the encoder to obtain the abstract characteristics of the audio sample A step of generating a sign expression, wherein the encoder performs the behavioral or pre-trained to perform a first task other than detecting mental health status There are steps, (c) Processing the abstract feature representation with the classifier to determine that the subject is the behavioral or personality A step of generating an output indicating whether or not the person has a divine health state, wherein the classifier, It is trained on a training dataset that includes multiple voice samples from multiple speakers, and the The audio samples of multiple audio samples are provided by speakers with the aforementioned behavioral or mental health conditions. Steps that are labeled as originating from or not originating from, Methods that include...

2. (b) Before that, the audio sample is put into a filter bank or Mel-frequency cepstrum coefficients The method according to claim 1, further comprising the step of converting.

3. The classifier is a binary classifier, and the output is that the subject is in the behavioral or mental health state. The method according to claim 1, wherein the output is a binary output indicating whether or not a state exists.

4. The classifier is a multi-class classifier, and the output is the behavioral or The method described in claim 1 includes a probability distribution across multiple levels or severity levels of mental health status. method.

5. The output is each segment of the plurality of segments of the voice sample from the subject. The method includes segment output for the target, and the method merges the segment output for the target Claim 1 further includes the step of detecting the behavioral or mental health status of the person. Method of description.

6. The claim states that the first task is automatic speech recognition, speaker recognition, emotion classification, or sound classification. The method described in 1.

7. (a) a claim including the step of obtaining the voice sample during a telemedicine session. The method described in item 1.

8. (a) The step of obtaining the voice sample from the subject's mobile device Including (b) and (c) being performed at least partially on the mobile device, The method described in item 1.

9. The claim 8 is as follows, wherein (b) and (c) are performed at least partially on a remote server. The method.

10. The aforementioned audio sample is processed by a non-speech model including a laughter model, a breathing model, or a pause model. The method according to claim 1, further comprising the step of processing with a ru.

11. (b) Before that, the step of determining whether the audio sample meets the quality threshold is performed. The method according to claim 1, further comprising:

12. When executed by one or more computer processors, the subject's behavior or mental state Non-temporal computer including machine executable instructions that implement a method for detecting divine health status A readable medium, wherein the method is (a) The step of obtaining an audio sample containing multiple audio segments from the subject person. 、 (b) Processing the audio sample with an encoder to obtain an abstract feature table of the audio sample A step of generating a reality, wherein the encoder controls the behavior or It is pre-trained to perform a first task other than detecting mental health status. , steps and, (c) Processing the abstract feature representations with a classifier to determine that the subject is the behavioral or mental A step of generating an output indicating whether or not the person has a health condition, wherein the classifier is a plurality It is trained on a training dataset that includes multiple voice samples from speakers, and the multiple The audio samples are derived from speakers who have the aforementioned behavioral or mental health condition. Steps that are labeled as either being or not derived from, Methods that include...

13. This method involves training an acoustic model to detect the behavioral or mental health status of a subject. The acoustic model includes an encoder and a classifier, and the method is (a) On the first training dataset, the encoder is used to perform the actions of the subject. The system is trained to perform a first task other than detecting physical or mental health conditions. Top and, (b) Following (a), a second training dataset different from the first training dataset is used. A step of training the encoder and the classifier on the platform, wherein the subject is Generate output indicating whether the user has a behavioral or mental health condition of interest, and the second The training dataset includes multiple voice samples from multiple speakers, and the multiple voice samples The audio samples for pull are derived from speakers with the aforementioned behavioral or mental health conditions of interest. Steps that are labeled as either being or not derived from, Methods that include...

14. The claim states that the first task is automatic speech recognition, speaker recognition, emotion classification, or sound classification. Method 13.

15. (b) an abstraction of the audio sample from the encoder in order to generate the output Claim 13 includes the step of training the classifier to process characteristic feature representations. The method.

16. The method according to claim 13, wherein the encoder is fixed during (b).

17. The method according to claim 13, wherein the encoder is not fixed during (b).

18. The method according to claim 13, wherein (a) and (b) are supervised learning processes.

19. The classifier is a binary classifier, and the output is that the subject is in the behavioral or mental health state. The method according to claim 13, wherein the output is a binary value indicating whether or not a state exists.

20. The classifier is a multi-class classifier, and the output is the behavioral or Claim 13 includes a probability distribution over multiple levels or severity levels of mental health status. The method.

21. The output is in each of the multiple segments of the voice sample from the subject. The method includes a segment output relating to the subject, and the method merges the segment output to determine the subject Claim 13 further includes the step of detecting the behavioral or mental health status in the above. Method of description.

22. When executed by one or more computer processors, the behavioral or mental health of the subject Non-functional instructions including machine executable instructions that implement a method for training an acoustic model to detect health conditions. A temporary computer-readable medium wherein the acoustic model includes an encoder and a classifier, The method described above is (a) On the first training dataset, the encoder is used to perform the actions of the subject. The system is trained to perform a first task other than detecting physical or mental health conditions. Top and, (b) Following (a), a second training dataset different from the first training dataset is used. A step of training the encoder and the classifier on the platform, wherein the subject is Generate output indicating whether the user has a behavioral or mental health condition of interest, and the second The training dataset includes multiple voice samples from multiple speakers, and the multiple voice samples The audio samples of the pull are derived from speakers who possess the aforementioned behavioral or mental health conditions of interest. Steps that are labeled as either being or not derived from, Methods that include...

23. This method involves training an acoustic model to detect the behavioral or mental health status of a subject. hand, (a) In order to transcribe the audio sample, automatic speech recognition is performed on the first training dataset. A step of training an ASR system, wherein the ASR system includes an encoder and Equipped with a decoder, step and (b) The step of discarding the decoder, (c) Processing the audio sample from the subject to determine whether the subject exhibits the behavioral or mental To generate an output indicating whether or not the person is in good health, the first training dataset and This involves training the encoder and classifier on a different second training dataset. The second training dataset is used for speakers with the aforementioned behavioral or mental health conditions. Multiple labeled items, labeled as originating from or not originating from Steps, including audio samples, Methods that include...

24. (a) Before that, the plurality of unlabeled audio samples are filtered into a filter bank or a mel frequency cable. The method according to claim 23, further comprising the step of converting to a Pstrum coefficient.

25. (c) Before that, the plurality of labeled audio samples are placed in a filter bank or a mel frequency cable. The method according to claim 23, further comprising the step of converting to a Pstrum coefficient.

26. (a) The encoder shall generate an abstract feature representation of the audio sample Train the decoder to process the abstract feature representation of the audio sample and transcribe it. The method according to claim 23, comprising the step of training to generate a given voice sample.

27. (c) an abstraction of the audio sample from the encoder in order to generate the output Claim 23 includes the step of training the classifier to process characteristic feature representations. The method.

28. The method according to claim 23, wherein the encoder is fixed during (c).

29. The method according to claim 23, wherein the encoder is not fixed during (c).

30. The method according to claim 23, wherein (a) and (c) are supervised learning processes.

31. Multiple labeled audio samples and multiple stories that generated the multiple labeled audio samples. Steps to train the classifier on a third training dataset that includes metadata about the person The method according to claim 23, further comprising P.

32. The metadata includes the age, race, ethnicity, gender, income, education, and location of each of the multiple speakers. The method according to claim 31, comprising one or more of the following: place of birth or medical history.

33. The encoder is a convolutional neural network (CNN) and a long-term short-term memory network. The method according to claim 23, including twerking (LSTM).

34. The aforementioned CNN is the Visual Geometry Group (VGG) network, The method described in item 23.

35. The aforementioned classifier is a recurrent convolutional neural network (RCNN), with caution L The model includes a selection from the group consisting of STM, self-awareness network, and transducer. The method described in item 23.

36. The classifier is a binary classifier, and the output is that the subject is in the behavioral or mental health state. The method according to claim 23, wherein the output is a binary value indicating whether or not a state exists.

37. The classifier is a multi-class classifier, and the output is the behavioral or Claim 23 includes a probability distribution over multiple levels or severity levels of mental health status. The method.

38. The output is in each of the multiple segments of the voice sample from the subject. The method includes a segment output relating to the subject, and the method merges the segment output to determine the subject Claim 23 further includes the step of detecting the behavioral or mental health status in the above. Method of description.

39. When executed by one or more computer processors, the behavioral or mental health of the subject A machine executable instruction that implements a method for training an acoustic model to detect a health state. A non-temporary computer-readable medium, wherein the method is (a) In order to transcribe the audio sample, automatic speech recognition is performed on the first training dataset. A step of training an ASR system, wherein the ASR system includes an encoder and Equipped with a decoder, step and (b) The step of discarding the decoder, (c) Processing the audio sample from the subject to determine whether the subject exhibits the behavioral or mental To generate an output indicating whether or not the person is in good health, the first training dataset and This involves training the encoder and classifier on a different second training dataset. The second training dataset is used for speakers with the aforementioned behavioral or mental health conditions. Multiple labeled items, labeled as originating from or not originating from Steps, including audio samples, Non-temporary computer-readable media, including [specific examples of such media].

40. It is a system, One or more computer processors, When execution is performed by the one or more computer processors, the one or more computer processors The computer processor receives input audio containing multiple segments from the subject, at least Based on partial assessment, whether the subject has a behavioral or mental health condition of interest. Memory containing machine-executable instructions that implement an acoustic model configured to predict The aforementioned acoustic model, An encoder configured to generate an abstract representation of the input sound, The coder determines whether the subject has the behavioral or mental health condition of interest. To perform tasks other than prediction, transfer learning frameworks are used for prior training. The encoder is refined, The abstract expression of the input voice is processed so that the subject can engage in the behavior or At least one configured to generate an output indicating whether or not the person has a mental health condition. A classifier wherein at least one classifier is the behavioral or mental health of interest. Labeled as originating from or not originating from a speaker in good health. A classifier trained on audio samples, Includes memory, A system that includes these features.

41. The encoder is a Visual Geometry Group ("VGG") network and The following is a stack of long-term short-term memory ("LSTM") networks as described in claim 40. Stem.

42. The aforementioned at least one classifier is a recurrent convolutional neural network ("RC"). Selected from the group consisting of NN), attention-enabled LSTM, self-attention network, or converter. The system according to claim 40, including a model that can be used.

43. The at least one classifier generates the output by metadata about the subject. The system according to claim 40, further configured to process data.

44. The system according to claim 43, wherein the metadata includes the age or gender of the subject.

45. The encoder is trained on the transcribed audio sample using a decoder. The system according to claim 40, wherein the decoder is not part of the system.

46. Claim 40, wherein the task is automatic speech recognition, speaker recognition, emotion classification, or sound classification. The system described.

47. The system according to claim 40, wherein the segment output is averaged.

48. The segment output is merged using a machine learning algorithm, as described in claim 40. The system.

49. The encoder is pre-trained by the decoder, and the encoder and decoder perform automatic voice recognition. The system according to claim 40, comprising an ASR (Autonomous System Recognition) system.

50. The decoder comprises an attention unit, a long-term short-term memory network, and a beam search unit. The system according to claim 49, comprising one or more of the above.

51. The system according to claim 40, wherein the at least one classifier includes a binary classifier.

52. The at least one classifier includes multiple classifiers, and the output is for the target group. Claim 40 includes a probability distribution over multiple severities of a behavioral or mental health condition. The system.

53. The output is a segment of the input audio for each of the plurality of segments. The output is, and the system obtains the predicted mental state, at least one of the A segment configured to merge the learned representation of the segment output of the classifier. The system according to claim 40, further comprising a fusion module.

54. Using natural language processing (NLP) models to detect the behavioral or mental health status of subjects. A method wherein the NLP model includes a language model and one or more classifiers, The method is (a) The step of obtaining an audio sample containing multiple audio segments from the subject person. 、 (b) To generate language model output, the speech sample or a derivative thereof is used in the language model A step processed by Dell, wherein the language model processes a first dataset and a second dataset. The dataset is trained on, and the first dataset is related to the behavioral or mental health status. The second dataset includes unrelated text and relates to the behavioral or mental health status. Including related text, the first dataset is more substantial than the second dataset. The processing steps are large in scale, (c) Processing the output of the language model with one or more classifiers, and determining that the subject is A step of generating an output indicating whether or not the person has a behavioral or mental health condition, Methods that include...

55. (b) Before that, the audio sample is transcribed in order to generate the transcribed audio sample. Steps include generating an embedded version of the transcribed audio sample using an encoder. The method according to claim 54, further comprising a step.

56. Claims that the language model includes a long-term short-term memory (LSTM) network or a converter. The method described in 54.

57. The one or more classifiers include a binary classifier, and (c) the subject is the behavioral or indicating that the person has a mental health condition or does not have the aforementioned behavioral or mental health condition. The method according to claim 54, comprising the step of generating a binary classification.

58. The one or more classifiers include a regression classifier, and (c) the behavior of the subject This includes the step of generating a probability distribution across multiple levels or severity levels of mental health status. The method according to claim 57.

59. The step of fusing the binary classification and the probability distribution to generate the output further The method according to claim 58, including the method described in claim 58.

60. The first dataset includes a publicly available text corpus, as described in claim 54. Method of loading.

61. When executed by one or more computer processors, natural language processing (NLP) ) A method for detecting behavioral or mental health status in subjects using a model. A non-temporary computer-readable medium containing machine-executable instructions to be executed, wherein the NLP model The method includes a language model and one or more classifiers, (a) The step of obtaining an audio sample containing multiple audio segments from the subject person. 、 (b) To generate language model output, the speech sample or a derivative thereof is used in the language model A step processed by Dell, wherein the language model processes a first dataset and a second dataset The dataset is trained on, and the first dataset is related to the behavioral or mental health state. The second dataset includes unrelated text and relates to the behavioral or mental health status. Including related text, the first dataset is more substantial than the second dataset. Large steps, (c) Processing the output of the language model with one or more classifiers, and determining that the subject is A step of generating an output indicating whether or not the person has a behavioral or mental health condition, Methods that include...

62. Methods for training natural language processing models to detect behavioral or mental health conditions The natural language processing model includes (i) a language model and (ii) a classifier, The notation method is (a) A step of training the language model with a first encoded text, the The first encoded text includes text unrelated to the behavioral or mental health status. Hmm, steps and, (b) The language model is converted into a second encoded text and optionally metadata information A step of making fine adjustments on the report, wherein the second encoded text is the active or Steps and texts related to mental health, (c) The behavior or precision on multiple encoded audio samples from multiple subjects A step of training the classifier to detect a divine state, wherein the plurality of encodings The encoded audio sample of the encoded audio sample A label indicating whether the person to whom the sample was provided has the aforementioned behavioral or mental health condition. and steps associated with optional metadata information, Methods that include...

63. The language model includes a long-term short-term memory (LSTM) network, as described in claim 62. The method.

64. The method according to claim 63, wherein the training in (a) includes a non-monotonic stochastic gradient descent process. 。

65. The training in (a) includes a dropout or DropConnect operation, according to the claim. The method described in 63.

66. The method according to claim 62, wherein the language model includes a converter.

67. The second encoded text is text related to additional behavioral or mental health conditions. The method according to claim 62, wherein the fine-tuning in (b) includes multitasking learning.

68. The additional actions on the multiple encoded audio samples from the multiple subjects A step of training additional classifiers to detect a target or mental state, wherein the plurality Among the encoded audio samples, the encoded audio sample is the encoding If the person who provided the recorded audio sample has the aforementioned additional behavioral or mental health condition Claim 67 further includes a step associated with a label indicating whether or not The method.

69. The aforementioned behavioral or mental health condition is an anxiety disorder, and the aforementioned further behavioral or mental health condition The method according to claim 68, wherein the patient is depressed.

70. (b) The fine-tuning in (b) includes discriminative fine-tuning of different layers in the language model. The method according to claim 62.

71. (b) The fine-tuning in the above is the sloped triangle learning rate for training the layers of the language model. The method according to claim 62, comprising the step of using

72. The classifier includes a binary classifier and a regression classifier, and the training of (c) is (i) test pair The binary classifier predicts whether the person has the aforementioned behavioral or mental health condition. (ii) training in the severity of the behavioral or mental health condition in the subject The claim includes the step of training the regression classifier to predict a numerical score indicating a degree. The method described in 62.

73. The output of the natural language processing model is obtained from the output of the binary classifier and the output of the regression classifier. The method according to claim 72, at least in part.

74. (c) followed, (d) A step of obtaining an audio sample from the subject, (e) Processing the speech sample using the natural language processing model and the test subjects A step of predicting whether the person has the aforementioned behavioral or mental health condition, The method according to claim 62, further comprising:

75. The aforementioned audio sample includes multiple responses to multiple queries, where (e) is the natural language The process includes the step of processing the audio samples multiple times using a word processing model, and the multiple The method according to claim 74, wherein the responses are arranged in a different order each time.

76. The natural language processing model transcribes the multiple voice samples from the multiple subjects. The method according to claim 62, comprising an automatic speech recognition model for doing so.

77. The natural language processing model encodes the plurality of transcribed speech samples. The method according to claim 76, comprising an encoder.

78. The encoder is an n-gram model, a skip-gram model, or a neural network. The method according to claim 77, selected from the group consisting of a byte pair encoder and a byte pair encoder.

79. The method according to claim 62, wherein the label is the result of a standardized mental health questionnaire. Law.

80. Determine whether the subject has or is likely to have a behavioral or mental health condition. A method for doing so, (a) The step of obtaining voice data from the subject, (b) at least one linguistic feature and at least one acoustic feature in the audio data Steps of computer processing to process the audio data for identification, (c) Compute the at least one linguistic feature and the at least one acoustic feature Process to generate one or more scores, and use the one or more scores, Whether the subject has or is likely to have the aforementioned behavioral or mental health condition Steps to generate a judgment, (d) The step of outputting an electronic report containing the judgment instructions generated in (c) Therefore, steps (b) to (d) are executed in less than 5 minutes, and the determination generated in (c) is at least A step having an area under the curve (AUC) of approximately 0.70, Methods that include...

81. The method according to claim 80, wherein the AUC is at least about 0.

75.

82. The method according to claim 81, wherein the AUC is at least about 0.

80.

83. The aforementioned electronic report indicates that the determination is that the subject has the aforementioned behavioral or mental health condition. If it indicates that there is or is likely to have, it relates to the behavioral or mental health condition. The method according to claim 80, including psychological education materials.

84. To determine whether the subject has or is likely to have a behavioral or mental health condition. This method, (a) The step of obtaining voice data from the subject, (b) at least one speech feature and at least one acoustic feature in the audio data Steps of computer processing to process the audio data for identification, (c) Compute the at least one audio feature and the at least one acoustic feature The data is processed to determine whether the subject has or is likely to have the aforementioned behavioral or mental health condition. A step that provides a determination of whether there is, (d) The step of outputting an electronic report showing the determination provided in (c), 、 (b) or (c) the computer processing of the determination provided in (c) or A method for optimizing at least one performance metric, including a specificity.

85. Determine whether the subject has or is likely to have a behavioral or mental health condition. A method for doing so, (a) Telemedicine session of the telemedicine application between the subject and the healthcare provider The process includes the steps of acquiring the audio stream and video stream of the subject, and 、 (b) one or more including an acoustic model, a natural language processing model (NLP), and a video model A step of acquiring multiple models, wherein the subject is in the behavioral or mental health state One or more mice determine whether or not they have, or are likely to have. The steps Dell takes to be trained and acquire, (c) The audio stream or the video stream is provided to the one or more modifiers The process is performed to determine whether the subject has the aforementioned behavioral or mental health condition. A step of generating a determination indicating whether there is a high probability of, (d) While the telemedicine session is in progress, the decision is made by the healthcare provider. - The user interface of the health application running on the device The steps to send, Methods that include...

86. Using the aforementioned natural language processing model, one or more topics in the audio stream Determine a topic or word, and use the user interface to access one or more topics or words. The method according to claim 85, further comprising the step of sending to a destination.

87. The method according to claim 85, wherein the determination includes a confidence interval for the determination.

88. The telemedicine session further includes the step of continuously repeating (a) to (d) during the telemedicine session. The method according to claim 85.

89. (b) based at least in part on demographic or medical history information relating to the subject The method according to claim 85, further comprising the step of selecting one or more of the aforementioned models.