A multimodal system for speech-based mental health assessment with emotional stimuli and its use
The multimodal system addresses emotional state and speech behavior in mental health assessments by integrating emotional stimuli and recognition techniques, improving accuracy and reliability in voice-based evaluations.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-01-24
- Publication Date
- 2026-03-05
AI Technical Summary
Existing voice-based mental health assessment technologies fail to adequately account for emotional states and emotional regulation during user responses, leading to biased and inaccurate assessments due to longer answers being inherently biased and ASR technologies missing meaningful speech behaviors.
A multimodal system that incorporates emotional stimuli and recognition, utilizing a Task Construction Module, Stimulus Output Module, Answer Receiving Module, Feature Construction Module, Feature Extraction Module, and Feature Fusion Module to capture acoustic, linguistic, and emotional features, and employs autoencoders for high-level feature fusion and classification.
Enhances the accuracy of mental health assessments by considering emotional states and speech behaviors, providing a more reliable and unbiased evaluation of mental health status.
Smart Images

Figure 0007824592000004 
Figure 0007824592000005 
Figure 0007824592000006
Abstract
Description
FIELD OF THE INVENTION
[0001] The invention relates to the fields of artificial intelligence, machine learning, computational networks, and computer engineering. In particular, the present invention relates to a multimodal system for voice-based mental health assessment utilizing induced emotion and emotion recognition techniques, and methods for using the same.
[0002] The global spread of COVID-19 (novel coronavirus) has led to an accelerating rise in the prevalence of mental illness, reportedly affecting 10-15% of the world's population. Given the current rate of increase, mental illness has a significant economic impact, even compared to cancer, cardiovascular disease, diabetes, and respiratory diseases. Suicide due to mental health issues is now the second leading cause of death among people aged 15-29, resulting in significant social disruption and loss of productivity. In response to these alarming trends, the World Health Organization (WHO) declared depression to be the leading cause of mental illness worldwide in 2016. However, in the traditional healthcare system, the number of patients suffering from mental illness continues to increase significantly, and the more severe the number of patients, the more difficult it becomes to provide appropriate treatment. However, early detection and early intervention techniques for mental illness can dramatically improve the cure rate, which will reduce the financial burden on patients and even improve their daily productivity and quality of life. Therefore, automated screening and continuous monitoring have proven to be effective alternatives that can smoothly realize screening. Voice-based and artificial intelligence (AI)-based screening technologies are gaining popularity due to their proven success in detecting various mental illnesses, including depression and anxiety, across a wide range of users. These technologies are recognized as inexpensive and scalable solutions due to their easy accessibility via digital devices such as mobile phones. By utilizing vocalizations as "biomarkers," these technologies can more robustly address misconduct and misbehavior during assessment or monitoring sessions than traditional outpatient screening mechanisms, such as the Patient Health Questionnaire-9 (PHQ-9) or Generalized Anxiety Disorder-7 (GAD-7). Voice-based artificial intelligence (AI) screening systems and their uses typically rely on scripted dialogue and collect voice-based responses (potentially along with other signals). Based on the collected responses, they may utilize models built on acoustic features or models built on acoustic and NLP features to generate a final classification, as is well documented in the prior art. To avoid accepting unhelpful responses as input, certain statistically based methods may be used to validate each response before proceeding to the next step. Simple solutions may measure the total time spent on the recorded response, while more complex solutions may utilize voice activation or ASR (automatic speech recognition) to ensure a sentence contains a sufficient number of words. If a preconfigured threshold is not met, the user may be presented with the same question (or a different question) and asked to answer it again. This type of approach with traditional technology has unintended consequences. First, longer answers are inherently biased and may not always be relevant. While shorter answers tend to be more candid about a user's mental health, longer answers tend to be longer and contain more words, which can lead to incorrect ratings if the user is dishonest. Second, the ASR (Automatic Speech Recognition) technology used in such systems and methods typically generates only recognizable sentences / words, and may ignore certain speech behaviors such as pauses, hesitations, trembling, wavering, etc. Filtering based on the output of such ASR technology may miss meaningful answers for the screening model. Another issue is that the emotional state of the user when answering the questions may have some effect on the results of the mental health assessment. However, prior art answer validation techniques are not a sufficient solution in this respect either, as they do not take into account the emotions involved in inputting the answers. Therefore, to ensure a relatively high level of input quality, better mechanisms are needed to improve model performance. In recent years, smartphone technology and wearable devices for noninvasive, continuous monitoring of physiological and psychological data have attracted significant research interest. Advances in acoustic and speech processing have opened up new areas of behavioral health diagnostics using machine learning. Patients with depression have been shown to exhibit reduced stress, monotony, and reduced volume in speech, consistent with clinical observations of depressed individuals. These patients also tend to speak at a slower rate, with longer pauses and lower volume than the general population. Furthermore, research has shown that vocal biomarkers recognized by machine learning models are positively correlated with and potentially useful for detecting mental illnesses such as depression. Researchers have devoted significant efforts to studying the acoustic and semantic aspects and correlates of depression and its speech. AI technologies in this field can be divided into three main categories: 1. Semantic-based: Automatic speech recognition (ASR) is applied to convert speech data into text strings, which are then subjected to natural language processing (NLP) to build a natural language-based classification model. 2. Acoustic-based: Extract acoustic features directly from the audio and build a classification model based on them. These features can be either manually designed features, such as rhythmic features / spectral-based correlation features, or latent feature embeddings through pre-trained models. 3. Multimodal-based: There have been some attempts to combine these two modalities to create multimodal AI models and improve the accuracy of the evaluation. These technologies and related research often utilize user voice recordings related to specific topics as model input. For example, responses to specific, fixed questions about rest schedules or descriptions of health conditions. However, it is noteworthy that major depressive disorder (MDD) can lead to a numbness to normal emotions, particularly sadness, fear, anger, and shame. Many previous studies have suggested that depressed patients, unlike non-depressed individuals, particularly need to utilize emotional preferences and emotion regulation strategies. Unfortunately, existing voice-based AI technologies have not paid sufficient attention to users' emotional preferences and emotion regulation abilities during the entire process, from collecting training data to training models and applying models to practical applications. Previous research has shown that the human voice conveys a lot of emotional information. For example, some prior art literature has shown that it is possible to detect not only basic emotional tendencies in a voice (e.g., positive vs. negative emotions, excitement vs. calmness, etc.), but also subtle emotional nuances. This could potentially improve the accuracy level of such AI models by proactively covering more emotional effects or better utilizing emotional information from voice (especially emotional valence). This patent describes a system for building such technology and methods for using it.
[0003] The aim of this invention is to assess mental health status based on vocal-based biomarkers combined with emotion determination / recognition. Another objective of the present invention is to utilize induced emotions and emotion recognition to eliminate bias in mental health assessment based on vocal-based biomarkers. It is yet another object of the present invention to provide enhanced filtering techniques in mental health assessments based on vocalization-based biomarkers. It is yet another object of the present invention to ensure a relatively high level of input quality in the speech signal while conducting mental health assessments based on speech-based biomarkers. Yet another object of the present invention is to enhance understanding of vocal behavior during mental health assessment based on vocalization-based biomarkers.
[0004] The present invention provides a multimodal system for speech-based mental health assessment with emotional stimuli and methods for its use. According to the present invention, there is provided a multimodal system for audio-based mental health assessment with emotional stimuli, said system comprising: - Task Construction Module (configured to construct tasks to capture acoustic, linguistic, and emotional features of the user's voice); - a stimulus output module configured to receive data from the task construction module, the stimulus output module including one or more stimuli to be presented to the user based on the constructed task, for inducing one or more types of triggers for the user's actions, the triggers being in the form of input responses; - an answer receiving module (configured to present one or more stimuli to a user based on a task constructed from the aforementioned stimulus output module and to receive answers corresponding to one or more formats); - Functional modules include: o a feature construction module (configured to define, for each of the aforementioned constructed tasks, features defined in terms of a learnable heuristic evaluation task); a feature extraction module configured to extract one or more defined features related to said constructed task from said received corresponding answers, using a learnable heuristic evaluation model taking into account at least one selected from the ranked constructed tasks; a feature fusion module configured to fuse two or more defined features to obtain a fused feature; - The autoencoder leverages the fused features described above to define the relationship between: o The speech modality relating to said feature fusion module works in conjunction with said answer receiving module to extract high level features obtained from said answer and output extracted high level text features. The text modality related to the feature fusion module works in conjunction with the answer receiving module to extract high-level features from the answer and output the extracted high-level speech features. The autoencoder is configured to receive extracted high-level text features and extracted high-level audio features in parallel from the audio modality and the text modality, and to output a shared expression feature dataset for emotion classification correlated with the mental health assessment. In at least one embodiment of the system, the constructed task is a pronunciation task or a writing task. In at least one embodiment of the system, the task construction module includes a first ranking module for ranking the constructed tasks in order of difficulty, thereby assigning a first value to each constructed task. In at least one embodiment of the system, the task construction module includes a first ranking module configured to rank the constructed tasks in order of difficulty, thereby assigning a first value to each constructed task, where the constructed tasks are stimuli corresponding to analyzed answers and classified into one of an assessed valence level (positive valence, negative valence, neutral valence). In at least one embodiment of the system, the task construction module includes a second ranking module configured to rank the complexity of the constructed tasks in order of complexity, thereby assigning a second value to each constructed task. In at least one embodiment of the system, the task construction module includes a second ranking module configured to rank the emotional expectations in terms of complexity and answer vectors of the constructed tasks, and thereby create a data collection pool based on the ranked constructed tasks. In at least one embodiment of the system, the constructed task is selected from several task groups (a cognitive task of counting numbers within a predetermined time, a task of pronouncing vowels within a predetermined time, a task of pronouncing words containing voiced and unvoiced sounds within a predetermined time, a task of reading words within a predetermined time, a task of reading a paragraph within a predetermined time, a task of reading a paragraph with phonemic and emotional complexity, a task related to open-ended questions with emotional variations, and a task related to open-ended tasks to be performed within a predetermined time). In at least one embodiment of the system, the constructed task includes one or more questions as stimuli, each of which is assigned a question embedding represented by a vector from 0 to N. Based on these question embeddings, a question-specific feature extractor is trained to extract word, phoneme, and syllable-level embeddings from the questions, and these extracted embeddings are forced to align and undergo mid-level feature fusion. In at least one embodiment of the system, the one or more stimuli are selected from a group of stimuli consisting of audio stimuli, video stimuli, combined audio-video stimuli, text stimuli, multimedia stimuli, physiological stimuli, and combinations thereof. In at least one embodiment of the system, the one or more stimuli include a stimulus vector tailored to elicit a text response vector, an audio response vector, a video response vector, a multimedia response vector, a physiological response vector, or a combination thereof, that responds to the stimulus vector. In at least one embodiment of the system, the one or more stimuli are configured to be analyzed through a first vector engine to identify a configuration vector and determine a ranked base state that correlates with the stimulus vector. In at least one embodiment of the system, one or more of the answers are selected, and the answers are selected from a group of answers consisting of audio answers, video answers, combined audio and video answers, text answers, multimedia answers, physiological answers, and combinations thereof. In at least one embodiment of the system, the one or more stimuli include response vectors that correlate to audio response vectors and / or video response vectors elicited in response to stimulus vectors from the stimulus output module. In at least one embodiment of the system, the one or more stimuli are configured to be analyzed through a first vector engine to identify a configuration vector and determine a ranked base state that correlates with the stimulus vector. In at least one embodiment of the system, the answer receiving module includes a sentence reading module configured to perform a sentence reading task within a time period preconfigured by a user. In at least one embodiment of the system, the feature construction module is a Geneva Minimal Acoustic Parameter Set (GeMAPS) based feature construction module. In at least one embodiment of the system, the feature construction module is a Geneva Minimal Acoustic Parameter Set (GeMAPS) based feature construction module, the feature construction module comprising: - Analyzes the audio using a set of 62 parameters. - Provides a 3-frame long symmetric moving average filter to smooth over time (this smoothing is performed within the voiced regions of the previous answer for pitch, jitter, and shimmer). - Arithmetic mean and coefficient of variation are applied as functions to 18 low-level descriptors (LLDs) to generate 36 parameters. - Apply 8 functions to the volume. - Apply 8 functions to the pitch. - Apply 8 functions to pitch. - Determine the Hammerberg index. - Determine the spectral features (for all unvoiced segments, this is done by looking at the spectral tilt between 0-500Hz and 500-1500Hz). - From the answers given above, determine the temporal features in continuous voiced and unvoiced regions. - Determines Viterbi-based smoothing of the F0 contour (prevents errors from causing a single voiced frame to be dropped). In at least one embodiment of the system, the feature construction module is a Geneva Minimal Acoustic Parameter Set (GeMAPS) based feature construction module comprising a set of low-level descriptors (LLDs) for analyzing spectral, pitch, and temporal features in the speech response, the features being selected from a group of features consisting of: Mel-Frequency Cepstral Coefficients (MFCCs) and their first and second derivatives. · Pitch and pitch variability. Energy and energy entropy. · Spectral centroid, broadness, and flatness. Spectral tilt. Spectral roll-off. Spectral variability. Zero crossing rate. Shimmer, jitter, and harmonic-to-noise ratio. Voicing probability (based on pitch) Temporal features such as loudness peak rate, mean duration and standard deviation in continuous voiced and unvoiced regions In at least one embodiment of the system, the feature construction module is a Geneva Minimal Acoustic Parameter Set (GeMAPS) based feature construction module, the feature construction module being configured with a set of frequency-related parameters selected from the group consisting of: Pitch (measured as the logarithmic fundamental frequency (F0) on a chromatic frequency scale, starting at 27.5 Hz (semitone 0)) Jitter (deviation in the length of individual consecutive F0 periods). Formant 1, 2, 3 Frequencies (center frequencies for the first, second, and third formants) Formant 1 (bandwidth related to the first formant) Energy-related parameters. Amplitude-related parameters. Shimmer (difference in peak amplitude between successive F0 intervals) Loudness (an estimate of perceived signal strength from the auditory spectrum) Harmonic-to-noise ratio (the ratio of the energy containing harmonic components to the energy containing noise-like components) Spectral (balance) parameters Alpha ratio (ratio of the total energy between 50-1000 Hz and 1-5 kHz) Hammerberg index (ratio of the strongest energy peak in the 0-2 kHz range to the strongest peak in the 2-5 kHz range) Spectral tilt (the slope of the linear regression of the logarithmic power spectrum within two specified bands: 0-500 Hz and 500-1500 Hz) Relative energy of formants 1, 2, and 3 (the ratio of the energy associated with the spectral harmonic peaks at the center frequencies of the first, second, and third formants to the energy associated with the spectral peak at the fundamental frequency (F0)) Harmonic difference H1-H2 (the ratio of the energy associated with the first harmonic (H1) of the fundamental frequency (F0) to the energy associated with the second harmonic (H2)) Harmonic Difference H1‐A3 (the ratio of the energy associated with the first harmonic (H1) of the fundamental frequency (F0) to the energy associated with the highest harmonic in the range of the third formant (A3)) In at least one embodiment of the system, the feature construction module is a Geneva Minimal Acoustic Parameter Set (GeMAPS) based feature construction module that utilizes higher order spectral analysis (HOSA) functions that utilize two or more component frequencies to achieve bispectral frequencies. In at least one embodiment of the system, the feature construction module is a Geneva Minimal Acoustic Parameter Set (GeMAPS) based feature construction module that utilizes higher order spectral analysis (HOSA) functions that utilize two or more component frequencies to achieve a bispectrum frequency, which utilizes a third accumulator to analyze the relationship between frequency components in a signal and examine nonlinear signals related to the answer. In at least one embodiment of the system, the feature extraction module fuses high-level feature embeddings by leveraging mid-level fusion. In at least one embodiment of the system, the feature extraction module includes the ability to extract specialized linguistic features, the extent to which each stimulus can be learned, making it highly suitable for a variety of linguistic tasks. In at least one embodiment of the system, the feature extraction module includes the ability to extract specialized emotional features and learn degrees of each stimulus, making it highly suitable for a variety of emotional tasks. In at least one embodiment of the system, the feature fusion module includes a speech module configured to perform high-level feature fusion and to classify sentiment from one or more responses by utilizing at least autoencoder-based feature fusion. In at least one embodiment of the system, the feature fusion module includes a speech module configured to perform high-level feature fusion and utilize at least autoencoder-based feature fusion to classify emotions from one or more answers, and the speech modality utilizes question-specific feature extraction to extract high-level features from time-frequency domain relationships in the answers and output the extracted high-level speech features. In at least one embodiment of the system, the feature fusion module includes a text module configured to provide high-level feature fusion capabilities and utilize at least autoencoder-based feature fusion to classify sentiment from one or more responses. In at least one embodiment of the system, the feature fusion module includes a text module configured to extract high-level features and utilize at least autoencoder-based feature fusion to classify sentiment from one or more responses, and the text modality outputs extracted high-level audio features utilizing a bidirectional long short-term memory network and an attention mechanism. In at least one embodiment of the system, the feature fusion module includes a speech module: - A module that uses extracted features from pre-trained models in acoustic feature embedding. - A module that compares the extracted features in the spectral domain. - A module that determines the characteristics of vocal tract coordination. - A module for determining features in repeated quantification analysis. - A module for determining features related to the number of bigrams and the bigram duration corresponding to the vocal landmarks. - A module that fuses the aforementioned features in an autoencoder. In at least one embodiment of the system, the feature extraction module includes an vocalization landmark extractor configured to determine event markers associated with the answer, the determination of the event markers correlating with positions on a timeline of acoustic events from the answer, including determining timestamp boundaries indicating abrupt changes in the acoustic answer, in a frame-independent manner. In at least one embodiment of the system, the feature extraction module includes a vocalization landmark extractor configured to identify event markers associated with the response, each having a start and end value and selected from a group consisting of glottal-based landmarks (g), periodicity-based landmarks (p), sonorant-based landmarks (s), fricative-based landmarks (f), voiced fricative-based landmarks (v), and burst-based landmarks (v). Each of these landmarks is used to identify a point in time at which a distinct abrupt articulatory event occurs, which correlates with abrupt changes in power across multiple frequency ranges and multiple time scales. In at least one embodiment of the system, the autoencoder comprises a multimodal, multiquestion input fusion architecture that includes: - One or more encoders that map one or more of the specific features mentioned above, combined with a task type, to a low-dimensional representation (each task is multiplied by a learnable degree encoding matrix based on the task type, and these degrees are correlated with mental health assessments). - One or more decoders that map said one or more specific features to said low-dimensional representation (configured to output a mental health assessment). - The aforementioned autoencoder is trained to minimize the reconstruction error between the input task and the decoder's output by utilizing a loss function. The present invention provides a multimodal use method for audio-based mental health assessment with emotional stimuli, said use method including: - Building tasks to capture the acoustic, linguistic, and emotional characteristics of a user's voice. - Receive data containing one or more stimuli based on the constructed task, and generate triggers that induce one or more types of user actions by presenting the task to the user (the triggers being in the form of input responses). - Presenting one or more stimuli to the user based on the constructed task and receiving corresponding responses in one or more formats. - For each of the tasks constructed above, defining features defined in terms of a learnable heuristic-type evaluation task. - extracting one or more defined features related to said constructed tasks from said received corresponding answers by utilizing a learnable heuristic evaluation model that considers at least one selected from the ranked constructed tasks. - Fusing two or more defined features to obtain a fused feature. - Leveraging the aforementioned fused characteristics, define the relationship between: o The speech modality relating to said feature fusion module works in conjunction with said answer receiving module to extract high level features obtained from said answer and output extracted high level text features. The text modality related to the feature fusion module works in conjunction with the answer receiving module to extract high-level features from the answer and output the extracted high-level speech features. The step of defining the relationship is configured to receive extracted high-level text features and extracted high-level audio features in parallel from the audio modality and the text modality, and output a shared expression feature dataset for emotion classification correlated with the mental health assessment. In at least one embodiment of the method of use, the constructed task includes one or more questions as stimuli, each of which is assigned a question embedding represented by a vector from 0 to N. Based on these question embeddings, a question-specific feature extraction module is trained to extract word embeddings, phoneme embeddings, and syllable-level embeddings from the questions, and these extracted embeddings are forced to align, as well as perform mid-level feature fusion. In at least one embodiment of the method of use, the step of defining the relationships comprises a multimodal, multiquestion input fusion architecture including: - Through an encoder, one or more specific features mentioned above are combined with the task type and mapped to the low-dimensional representation mentioned above. (Each task is multiplied by a learnable degree encoding matrix based on the task type, and these degrees are correlated with mental health assessments.) - Through a decoder, the one or more specific features are mapped to the low-dimensional representation, and a mental health assessment is output. - By utilizing a loss function, it is trained to minimize the reconstruction error between the input task and the decoder's output. In at least one embodiment of the method of use, the step of extracting high-level speech features includes extracting high-level features from time-frequency domain relationships in the response, thereby outputting the extracted high-level speech features. In at least one embodiment of the method of use, extracting said high-level text features comprises simulating intra-modal dynamics using a bidirectional long short-term memory network and an attention mechanism to output the extracted high-level text features. [Brief explanation of the drawings]
[0005] The present invention is now disclosed with reference to the accompanying drawings, in which: Figure 1 shows a high-level block diagram of the computing environment. Figure 2 shows the system for collecting training data with induced emotions. Figure 3 shows a sample emotion-elicitation question set that presents several tasks and questions to elicit emotion-based answers. Figure 4 shows the known interactions between phonetic components that form useful phonetic criteria. Figure 5 shows one example of a fixed reading sentence. Figure 6 shows a phoneme map without tones based on the user reading the sentence. Figure 7 shows another example of a fixed reading passage. Figure 8 shows a graph of phonemes versus occurrences based on a user reading a sentence. Figure 9 shows the content regarding the representation of HOSA (Higher Order Spectral Analysis) functions for different mental states of users. Figure 10 shows the autocorrelation and cross-correlation between the first and second delta MFCCs (Mel-frequency cepstral coefficients) extracted from a 10-second audio file, illustrating our framework for extracting delayed correlations from acoustic files that reflect psychomotor retardation. Figure 11 shows a schematic block diagram of an autoencoder. 12A-12H show various graphs for at least one type of corresponding task (question), where the original label is selected from "health" or "depression," in a non-limiting exemplary embodiment. The graphs show phonemes and segments correlated with the model predictions. Areas highlight verbal utterances characterized by high activation for different users at different start times. Figure 13 shows the flowchart. Figure 14 shows a high-level flowchart of a voice-based mental health assessment with emotional stimuli.
[0006] The present invention provides a multimodal system for speech-based mental health assessment with emotional stimuli and methods for its use. Mental health issues are closely linked to patients' emotions. When determining whether a user's input is useful to the system, the user's emotions are at least as important as the length of the response, and in some circumstances may be even more important. The present disclosure may be embodied as a system, method of use, computer program product, or program / product related to a mobile device. The computer program product may include a computer-readable storage medium (or media) containing computer-readable program instructions for a processor to perform aspects of the present disclosure. Aspects of the disclosed embodiments may include a tangible computer-readable medium storing software instructions that, when executed by one or more processors, are configured and capable of performing one or more methods, operations, etc., in accordance with the disclosed embodiments. Aspects of the disclosed embodiments may also include logic and instructions executed by one or more processors configured as special purpose processors based on programmed software instructions that, when executed, perform one or more operations in accordance with the disclosed embodiments. When describing this invention, the following definitions apply throughout (including those above): "Computer" may mean one or more devices or one or more systems that can accept structured input, process the input in accordance with prescribed rules, and produce the results of the processing as output. Examples of computers include computers, fixed or portable computers, computers with single processors, multiple processors, or multi-core processors (which may operate in parallel or non-parallel), general-purpose computers, supercomputers, mainframes, superminicomputers, minicomputers, workstations, microcomputers, servers, client servers, interactive televisions, web appliances, communications devices with Internet access, hybrid computer-interactive televisions, tablets (PCs), personal digital assistants (PDAs), mobile phones, application-specific hardware for emulating a computer or software (e.g., digital signal processors (DSPs), field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific instruction set processors (ASIPs), chips, chipsets), systems-on-chips (SoCs), multiprocessor systems-on-chips (MPSoCs), optical computers, quantum computers, biological computers, and devices (typically including input, output, memory, arithmetic, logic, and control units) that accept data, process the data according to one or more stored software programs, and produce a result. "Software" can refer to the rules established for operating a computer or part of a computer. Examples of software may include code segments, instructions, applets, pre-compiled code, compiled code, interpreted code, computer programs, and programmed logic. "Computer-readable medium" may refer to any storage device used to store data that can be accessed by a computer. Examples of computer-readable medium may include magnetic hard disks, floppy disks, optical disks such as CD-ROMs and DVDs, magnetic tape, memory chips, or any other type of medium capable of storing machine-readable instructions. "Computer system" may refer to a system containing one or more computers, each of which may contain computer-readable media containing software that operates that computer. Examples of computer systems include a distributed computer system that processes information across a network of connected computer systems, a system in which two or more computer systems connected over a network send and receive information between the computer systems, and one or more devices or systems (typically including input, output, memory, arithmetic, logic, and control units) that may accept data and process the data in accordance with one or more stored software programs to produce a result. A "network" may refer to multiple computers and related devices that may be connected by communications facilities. A network may include permanent connections such as cables, as well as temporary connections performed through telephone or other communications links. A network may also include wired connections (e.g., coaxial cable, twisted pair, fiber optics, waveguide, etc.) or wireless connections (e.g., radio frequency waveforms, free-space optical waveforms, acoustic waveforms, satellite communications, etc.). Examples of networks include the Internet, intranets, local area networks (LANs), wide area networks (WANs), and combinations of networks (e.g., the Internet and an intranet). Typical examples of networks include Internet Protocol (IP), Asynchronous Transfer Mode (ATM), Synchronous Optical Network (SONET), User Datagram Protocol (UDP), and IEEE 802.x, and may operate using a variety of protocols. The terms "data" and "data item" as used in this entry refer to a sequence of bits. Thus, a data item may be the contents of a file, part of a file, a page in memory, an object in an object-oriented program, a digital message, a digitally scanned image, part of a video or audio signal, or any other entity that can be represented by a sequence of bits. The term "data processing" as used in this entry refers to the processing of a data item and may depend on the type of data item being processed. For example, a data processor for a digital image may be different from a data processor for an audio signal. The terms "first," "second," etc., as used herein do not denote any order, priority, quantity, or importance, but are used merely to distinguish one element from another. Furthermore, the terms "1" and "one" as used hereinafter do not denote a limitation of quantity, but rather denote the presence of one or more of the referenced item. Figure 1 shows a conceptual block diagram of a computing environment including one or more network client devices (112, 114, 116, 118) connected to a network server (100) through a network, and one or more databases (122, 124, 126, 128) connected to the network server (100). Although speech production may seem simple, it is actually a highly complex process involving the coordination of cognitive and physiological behaviors within the brain, requiring a complex series of actions to be effective. Speech production typically originates from physiological processes in the body, beginning with the generation of potential energy in the lungs and changes in air pressure within the vocal tract. When speech sounds are produced, the lungs expel air, and the velocity of the air also influences the regularity of the vocal cords. As the harmonic, rich sound energy emanating from the glottis passes through the vocal tract and larynx, movements of the pharynx, oral cavity, nasal cavity, and speech organs (e.g., tongue, teeth, lips, jaw, and soft palate) alter the amplitude of the harmonic sounds of speech, acting as sound filters. Depression and psychomotor retardation are associated with vocal dysfunction, reduced harmonic formant amplitude, and other physical abnormalities in some depressed individuals. This can result in poor laryngeal control due to psychomotor retardation, resulting in a "breathiness" in the speech of patients. This contrasts with the speech of healthy individuals. Vocal intensity has been shown to be a strong indicator of the severity of depression. Depressed individuals often speak with reduced vocal intensity and may appear to speak in a monotone voice. Many researchers have studied how people with depression use language in two ways: (a) by listening to recordings of patients speaking, and (b) by analyzing the texts they have written. People with depression often have problems with language skills and may use inappropriate or unclear words, leave things unfinished, or repeat the same words or phrases over and over again. To identify depression, it is important to identify potentially abnormal phonological elements, which also interact with linguistic information (e.g., the meanings of words and phrases) to express an individual's emotional state. Previous subjective assessments of depressed individuals have often focused on the patient's speech patterns and associated behaviors, based on preconceived notions of how depressed people behave emotionally. Figure 2 shows the system for collecting training data with induced emotions. In at least one embodiment, the client devices (112, 114, 116, 118) are communicatively coupled to a stimulus output module (202) configured to provide one or more stimuli to a user corresponding to the client device. In at least one embodiment, the stimulus output module (202) is configured to provide output stimuli to the user to elicit a user response in the form of an input response. The output stimuli may include audio stimuli, video stimuli, combined audio-visual stimuli, text stimuli, multimedia stimuli, physiological stimuli, or combinations thereof, and the like. The stimuli may include stimulus vectors tailored to elicit a text response vector, an audio response vector, a video response vector, a multimedia response vector, or a physiological response vector, or any combination thereof, that responds to the stimulus vector. One or more stimuli are analyzed through a first vector engine (232) to identify constituent vectors and determine ranked baseline states that correlate with the stimulus vectors. In some embodiments, the stimulus output module (202) is used to send audio-based tasks to subjects (users) with specific mental disorders based on specific criteria (e.g., DSM-V diagnoses) and healthy subjects (users) without mental disorders. This is performed through the network server (100) or through client devices (112, 114, 116, 118) corresponding to the users or subjects. In preferred embodiments, these stimuli are vector-structured to elicit specific emotional responses in the subjects, including, but not limited to: - Emotional response due to happiness - Emotional reactions due to sadness - Neutral emotional response In some embodiments, the stimulus output module (202) includes stimuli that elicit specific user or subject actions, such as imperative questions that force the user to repeat a certain number of words or vowel utterances. Embodiments of the present disclosure may include a task construction module configured to construct a task to capture acoustic, linguistic, and emotional characteristics of a user's verbal utterances. Data from the task construction module is provided to a stimulus output module (202). Because the present invention deals with vocal-based biomarkers, it may be advantageously utilized for tasks involving such vocalization tasks. In at least one embodiment, the task construction module includes a first ranking module for ranking the constructed tasks in order of difficulty, thereby assigning a value to each constructed task. In a preferred embodiment, the recording of the patient's responses requires positive stimulation of the user (speaker) according to different emotional valences (ranked levels: positive, negative, neutral) to better capture the user's emotional preferences and utilization of emotion regulation strategies related to their mental state. In at least one embodiment, the task construction module includes a second ranking module configured to assess the complexity of the questions and the emotional expectation of the ranked questions to create a data collection pool. In multi-class classification, depression levels have an ordinal relationship, and therefore, in the vocalization tasks for different depression levels from the task construction module, a learnable "degree" is set so that the loss for each question and each depression level can be optimized. Figure 3 shows a sample emotion-elicitation task set (question set) that presents several speech tasks for eliciting emotion-based responses. Below, we present the emotion-eliciting patterns for each speech task, in at least one implementation. Well-designed tasks not only adequately cover different emotions, but also provide good phonemic coverage in the target language. [Table 2] In at least one embodiment, the client devices (112, 114, 116, 118) are communicatively coupled to at least one response receiving module (204) configured to present one or more stimuli from the stimulus output module (202) to a user of the corresponding client device and to receive response input in one or more forms. At least one embodiment of the response receiving module (204) provides a receiving module (204a) configured to capture a user's input response to an output stimulus. The input response may include audio input, video input, combined audio and video input, multimedia input, or other similar input modalities. The input response may include a response vector associated with an audio response vector or a video response vector elicited in response to a stimulus vector associated with the stimulus output module (202). Embodiments of the present disclosure may also include a data pre-processing module. In some embodiments of the task construction module, question construction may include simple cognitive tasks consisting of open-ended questions such as number counting, vowel pronunciation, word pronunciation with voiced and unvoiced components, paragraph reading with phonemic and emotional complexity, emotional transitions, and cognitive sentence generation, the answers of which may be collected through a response receiving module (204). Embodiments of the receiving module (204a) may include constructing and evaluating answers, recording answers, and processing answer vectors until boundary conditions are met to allow them to be used as eligible data points for training. The answer vectors may also include flashcards for specific data collection tasks, and may also collect metadata about the user's recordings, including details about task completion and qualifications. Embodiments of the response receiving module (204) may also include a first measurement module configured to analyze, measure, and output as a first measurement data set including at least: - Content comprehension measured in terms of answer vectors. - The accuracy of the answer content measured in terms of the answer vector. - The quality of the acoustic signal measured with respect to the response vector. Embodiments of the response receiving module (204) may also include a second measurement module configured to analyze, measure, and output as a second measurement data set including at least: - Silence measured from the response vector. - Signal-to-noise ratio measured from the response vector. - Articulation clarity measured from response vectors. - Activity index measured from the response vector. It is based on a threshold value that is preset for the second measurement data mentioned above. In a preferred embodiment, segments of the response vector that measure a signal-to-noise ratio of 15 or greater are used as qualifying speech samples by the system and method of the present invention. In at least one embodiment of the receiving module (204a), at least one acoustic receiving module (204a.1) is provided, configured to capture a user's acoustic input responses to one or more output stimuli in the form of response acoustic signals. The response acoustic signals include response acoustic vectors that are correlated with the stimulus vectors of the stimulus output module (202) by a second vector engine (234). The second vector engine determines constituent response acoustic vectors and determines a first state of the user in relation to the stimulus vectors of the stimulus output module (202). In at least one embodiment of the receiving module (204a), at least one text receiving module (204a.2) is provided, communicatively coupled to the acoustic receiving module (204a.1) and configured to obtain a user's acoustic input response from the acoustic receiving module (204a.1) and transcribe it into text via a transcription engine to provide a response text signal. The response text signal includes a response text vector, which is correlated with the stimulus vector of the stimulus output module (202) by a third vector engine (236). This second vector engine identifies a constituent response text vector and identifies a second state of the user in relation to the stimulus vector of the stimulus output module (202). At least one embodiment of the response receiving module (204) includes a physiological response receiving module (204b) configured to sense one or more physiological response signals of the user to the output stimuli of the stimulus output module (202) through one or more physiological sensors. The physiological response signals may include physiological vectors that correlate to the physiological signals in response to the output stimuli of the stimulus output module (202). These vectors are analyzed through a fourth vector engine (238) configured to determine a third state of the user in relation to the stimulation vectors of the stimulus output module (202). In at least one embodiment of the response receiving module (204), the neurological response receiving module is provided, configured to sense one or more neurological response signals of a user to the output stimuli of the stimulus output module (202) through one or more neurological sensors. The neurological response signals may include neurological vectors that correlate to the neurological signals in response to the output stimuli of the stimulus output module (202). These vectors are analyzed through a fifth vector engine configured to determine a fourth state of the user in relation to the stimulation vectors of the stimulus output module (202). Embodiments relating to a vector engine form an engine for identifying vocal biomarkers. Typically, a vocal biomarker identification engine uses a three-stage approach. - Use knowledge of the vocal domain to develop speech data collection protocols and increase the presence and robustness of vocal biomarkers associated with depression. - Identifying the optimal set of features that capture subtle acoustic differences between depressed and non-depressed individuals. (These features should be relevant for good detection and robust to noise, making them useful for detecting depression in natural environments.) - Require multidimensional analysis, as the status of vocal biomarkers in voice systems is not limited to binary "yes" or "no" values (which will require pre-determining the appropriate tasks and the time to perform them so that vocal biomarkers remain visible or present when using automated detection systems for people with depression). The time required for detection, or the Task-Based Detection Threshold (TBDT), may vary depending on gender, severity of depression, and age. Therefore, to ensure that the system and its use have a stable detection range for symptoms, a procedure must be implemented to identify the detection threshold and detection reliability measure score (DCMS) by arranging the speech tasks in an efficient order. Embodiments of the present disclosure may also include a protocol construction engine. Automated speech processing using machine learning is becoming increasingly popular in digital healthcare and has great potential as a non-invasive and remote medical screening tool. However, there is a need to better understand the protocols used in speech processing and perform measurements that will help create new protocols with specific criteria. Figure 4 shows the known interactions between phonetic components that form useful phonetic criteria. Healthcare professionals use speech-language assessments to screen, diagnose, and monitor patients for a variety of physical disorders. During these assessments, clinicians observe the patient's speech production, including articulation, breathing, phonation, and voice quality, as well as the patient's own language skills, including grammar, pragmatics, memory, and expressiveness. Abnormal speech and language symptoms are often early indicators of various physical disorders and illnesses. At least one embodiment of the response receiving module (204) comprises a text-to-speech module that allows a user to read text based on predefined tasks or prompts. Benefits of utilizing text-to-speech protocols include ease of use, reproducibility, the ability to provide clear reference points, and the limited range and controlled variation of sounds used. Furthermore, they are relatively easy to integrate into digital smart device applications. When selecting a text-to-speech protocol for analyzing a patient's health status, it is important to carefully consider factors such as the speaker's background, the specific disease to focus on, the time required for the task, and the number of samples required. In at least one embodiment of the sentence reading module, emotion-based, anchored sentence reading samples are provided. Figure 5 shows one example of a fixed reading sentence. Figure 6 shows a phoneme map without tones based on the user reading the sentence. Figure 7 shows another example of a fixed reading passage. Figure 8 shows a graph of phonemes versus occurrences based on a user reading a sentence. An embodiment of the present disclosure includes a feature module with a feature construction function for each constructed task, where the features are defined in terms of a learnable heuristic-type evaluation task. At least one implementation of the feature construction function utilizes the Geneva Minimal Acoustic Parameter Set (GeMAPS), a set of 62 parameters used for speech analysis. A three-frame symmetric moving average filter is used for time smoothing, and this smoothing is performed only within the voiced regions of pitch, jitter, and shimmer. Arithmetic mean and coefficient of variation are applied as functions across all 18 low-level descriptions (LLDs) to generate 36 parameters. Eight additional functions are applied to loudness and pitch, including the arithmetic mean of alpha ratio, Hammerberg exponent, and spectral slope (0-500 Hz and 500-1500 Hz) across all unvoiced segments. Temporal features such as the frequency of loudness peaks, the mean length and standard deviation of consecutive voiced and unvoiced regions, and the number of consecutive voiced regions per second are also included. There is no minimum length imposed on voiced or unvoiced regions, and it performs Viterbi-based smoothing of the F0 (fundamental frequency) contours to prevent errors from causing a single voiced frame to be dropped. eGeMAPS (Geneva Minimal Acoustic Parameter Set) is a feature set used in the analysis of speech, audio, and music. It is used in openSMILE (Open Source Multimodal Interface for Language and Emotion Recognition, Technical University of Munich). eGeMAPS is a subset of the larger Geneva Minimal Acoustic Parameter Set (GeMAPS) and is specifically designed for the task of emotion recognition. It contains a set of low-level descriptors (LLDs) for analyzing the spectral, pitch, and temporal characteristics of speech signals. Features include: Mel-Frequency Cepstral Coefficients (MFCCs) and their first and second derivatives Pitch and pitch variability Energy and energy entropy Spectral centroid, broadness, and flatness Spectral tilt Spectral Roll-Off Spectral variability Zero crossing rate Shimmer, jitter, and harmonic-to-noise ratio Voicing probability (based on pitch) Temporal features such as loudness peak rate, mean duration and standard deviation in continuous voiced and unvoiced regions In total, 87 features are included. eGeMAPS is designed as a minimal feature set that can provide robust analytical performance for a wide range of situations. It demonstrates robust performance regardless of recording conditions and speaker characteristics, and has also been shown to be effective in multiple emotion recognition tasks. Amplitude-related parameters Pitch (measured as the logarithmic fundamental frequency (F0) on a chromatic frequency scale, starting at 27.5 Hz (semitone 0)) Jitter (deviation in the length of individual consecutive F0 periods) Formant 1, 2, 3 frequencies (center frequencies for the first, second, and third formants) Formant 1 (bandwidth related to the first formant) Energy / amplitude related parameters Shimmer (difference in peak amplitude between successive F0 intervals) Loudness (an estimate of perceived signal strength from the auditory spectrum) Harmonic-to-Noise (HNR) ratio (the ratio of the energy containing harmonic components to the energy containing noise-like components) Spectral (balance) parameters Alpha ratio (ratio of total energy between 50-1000 Hz and 1-5 kHz) Hammerberg index (ratio of the strongest energy peak in the 0-2 kHz range to the strongest peak in the 2-5 kHz range) Spectral tilt (the slope of the linear regression of the logarithmic power spectrum within two specified bands: 0-500 Hz and 500-1500 Hz) Relative energy of formants 1, 2, and 3 (the ratio of the energy associated with the spectral harmonic peaks at the center frequencies of the first, second, and third formants to the energy associated with the spectral peak at the fundamental frequency (F0)) Harmonic difference H1-H2 (the ratio of the energy associated with the first harmonic (H1) of the fundamental frequency (F0) to the energy associated with the second harmonic (H2)) Harmonic Difference H1‐A3 (the ratio of the energy associated with the first harmonic (H1) of the fundamental frequency (F0) to the energy associated with the highest harmonic in the range of the third formant (A3)) At least one implementation of the feature construction function uses higher-order spectral analysis (HOSA) functions. Higher-order spectral analysis (HOSA) functions are functions of two or more component frequencies, as opposed to a power spectrum, which is a function of a single frequency. These spectral analysis functions can be used to identify phase coupling between Fourier components, making them particularly useful for detecting and characterizing nonlinearities in a system. To achieve this, the amplitudes in the higher-order spectral analysis functions are normalized by the power at the component frequencies. The normalized spectral analysis function, also known as the nth-order coherence index, is a function that combines the nth-order cumulant spectrum and the power spectrum. The bispectrum is a method that uses third-order cumulants to analyze the relationships between frequency components in a signal and is particularly useful in investigating nonlinear signals. The bispectrum provides information about the phase relationships between frequency components, which is significant because it contains information not available in the spectral domain and is therefore more informative than the power spectrum. Higher-order statistics are highly effective in studying nonlinear signals because they can capture the relationships between phase components. The bispectrum reveals information not available in the spectral domain, making it one of the best methods for this purpose. Figure 9 shows the content regarding the representation of HOSA (Higher Order Spectral Analysis) functions for different mental states of users. Embodiments of the present disclosure may include a feature module including a feature extractor with a variable / heuristic evaluation model that considers ranked tasks, at least one selected from a set of ranked questions. In preferred embodiments, the degree of learnability varies from question to question or task to task. In preferred embodiments, the feature extractor uses fusion at intermediate levels and fusion of feature embeddings at higher levels. In this paragraph, we describe an embodiment of the feature extractor. When training a task (e.g., acoustic questions), each stimulus (e.g., question) may be assigned a dedicated feature extractor with a learnable degree. In a preferred embodiment, linguistic tasks / stimuli (questions) may be assigned dedicated linguistic feature extractors with a degree of learnability for each stimulus (e.g., question). In a preferred embodiment, emotion tasks / stimuli (questions) may be assigned a dedicated emotion feature extractor with a degree of learnability for each stimulus (e.g., question). In a preferred embodiment, the selector is configured to randomly select a sample of answer vectors based on pre-set thresholds such as signal-to-noise ratio, clarity of the answer (speech), etc. Embodiments of the present disclosure may include a data loader configured to randomly upsample or downsample answer vectors until the training data batches are balanced. The data loader is constructed to randomly upsample or downsample as appropriate until the training data batches are balanced. Embodiments may also include systems and methods configured to use soft labels instead of hard labels in classification problems due to data ordinality. Despite attempts to utilize unitary emotion feature representations, existing techniques have proven insufficient to achieve recognition due to poor feature differentiation and a lack of ability to effectively capture the dynamic interplay between emotions in speech recognition tasks. Accordingly, embodiments of the present disclosure may include a feature module comprising a feature fusion module configured to fuse the various output stimuli and user states described above. Typically, the feature fusion module works in conjunction with an answer receiving module (204). In a preferred embodiment, the feature fusion module is configured to classify mental health conditions from multiple types of answer vectors by utilizing an audio modality with high-level feature extraction and at least autoencoder-based feature fusion. In a preferred embodiment, the audio modality uses question-specific feature extraction to extract high-level features from time-frequency domain relationships in the answer vectors, with the output being the extracted high-level audio features. In a preferred embodiment, the feature fusion module is configured to classify sentiment from multiple types of response vectors by leveraging a text modality with high-level feature extraction and at least autoencoder-based feature fusion. In a preferred embodiment, the text modality uses a bidirectional long short-term memory network and an attention mechanism to simulate intra-modal dynamics, with the output being extracted high-level text features. In a preferred embodiment of the speech modality feature fusion module, the system and its method extract high-level embeddings, i.e., high-dimensional vector embeddings, from pre-trained models, such as huBert, Wav2vec, and Whisper, using raw speech input. Similarly, the system and its method extracts acoustic feature embeddings (from GeMAPS and emobase egemap) and compares them in the spectral domain. The system and its method also extracts log-mel spectrograms, MFCCs (Mel Frequency Cepstral Coefficients), as well as higher-order spectral features (HOSA), psychomotor delays, and neuromuscular activity features by utilizing vocal tract tuning and iterative quantitative analysis. Bigram counts and bigram durations are calculated using vocal landmarks that indicate pronunciation effectiveness. All features are fused using an autoencoder, and feature nonlinearities are preserved during feature fusion in the latent feature space. Previous research has shown that leveraging lexical information improves the performance of emotional valence estimation. Lexical information is obtained from a pre-trained acoustic model, and the learned representations can improve the performance of emotional valence estimation from vocalizations. Our system and its usage explores leveraging the representations from the pre-trained model and improving the performance of depression biomarker estimation from vocal signals while evaluating psychomotor retardation through task-specific feature extraction functions such as neuromuscular coordination features. Our system and its usage also explores the fusion of representations to improve the performance of depression biomarker estimation. Human vocal communication is roughly composed of two layers: the linguistic layer, which conveys messages in the form of words and their meanings, and the paralinguistic layer, which conveys how words are spoken, such as the expressiveness and emotional tone of the voice. Given the self-supervised learning architecture of pre-trained models and the existence of large-scale vocalization datasets made publicly available, we can speculate that the representations produced by these pre-trained models may contain lexical information that can help make emotional valence estimation run more smoothly. In some embodiments of the present invention, we explore a multimodal granularity framework that allows systems and their usage to extract speech embeddings at various subword levels. The figures show that embeddings extracted from related models are typically frame-level embeddings and are effective in capturing rich frame-level information. However, they lack the ability to capture segment-level information useful for identifying depression biomarkers. Therefore, the systems and usage of the present invention introduce segment-level embeddings, which include not only frame-level embeddings but also word, phoneme, and syllable-level embeddings that are closely related to phonology. Phonology contains information about the rhythm (cadence) of a language's speech and can convey speech characteristics (e.g., depression status). As a result, segment-level embeddings may be highly useful in multimodal depression biomarker recognition. By using the forced alignment method, we can obtain the temporal boundaries of phonemes, which can then be grouped to obtain syllable boundaries. After the forced alignment information is provided, we can extract the relevant speech segments corresponding to those linguistic units. Embodiments of the present disclosure may also include vocal landmark extraction functionality. Articulatory landmarks are event markers related to the production of linguistic speech. They provide information about articulatory events, such as vocal fold vibration, solely by relying on the location of acoustic events in time, such as consonant closures and releases, nasal closures and releases, velocities, and vowel maxima. Unlike frame-based processing frameworks, landmark methods can detect timestamp boundaries that indicate abrupt changes in speech intelligibility without relying on frames. This approach offers an alternative to frame-based processing and avoids its drawbacks by focusing on acoustically measurable changes in linguistic speech. VTH (Vocal Tract Adjustment) employs six landmarks, each with a start and end state. These landmarks, "g (larynx)," "p (periodicity)," "s (rumble)," "f (fricative)," "v (voiced fricative)," and "b (plosive)," are used to identify the time points of different sudden phonation events. These are detected by observing significant evidence of abrupt changes in energy intensity (i.e., rise or fall) associated with voicing across multiple frequency ranges and multiple time scales. In landmarks, "s" and "v" are associated with voiced sounds, while "f" and "b" are associated with unvoiced sounds. Embodiments of the present disclosure may also include a vocal tract adjustment engine. VTC features have been shown to have the ability to capture psychomotor activity associated with depression, and have been the most successful in two A VEC challenges for predicting depression severity, based on the observation that vocal tract parameters in depressed patients are less "tunable" (correlated) compared to healthy speakers. [Table 3] Figure 10 shows the autocorrelation and cross-correlation between the first and second delta MFCCs (Mel-frequency cepstral coefficients) extracted from a 10-second audio file, illustrating our framework for extracting delayed correlations from acoustic files that reflect psychomotor retardation. Figure 11 shows a schematic block diagram of an autoencoder. An embodiment of the present disclosure may include an autoencoder that defines, in coordination with a stimulus output module (202), a relationship between the audio modality of a feature fusion module operating in conjunction with an answer receiving module (204) and the text modality of the feature fusion module operating in conjunction with the answer receiving module (204). After the aforementioned processing, i.e., processing of the audio modality and the text modality, is completed, the autoencoder is configured to output a shared expressive feature dataset for emotion classification by feeding the extracted high-level text features and the extracted high-level audio features to the autoencoder in parallel. What is unique about this system and its usage is that it can measure the accuracy of the autoencoder in reconstructing the shared expressive feature dataset and minimize reconstruction errors, as well as evaluate the performance of depression detection in the emotion recognition module (300). This is a significant departure from previous systems and usages that focus solely on learning high-level features from input data using autoencoders. Autoencoders excel at capturing low-level shared representations in a nonlinear manner. This significant strength is leveraged in the current inventive system and usage, where fixed, closed tasks (e.g., paragraph reading, counting, or phoneme pronunciation tasks) are evaluated using feature extractors with optimal specific questions, while open-ended tasks (questions) are evaluated in both audio and text modalities while optimizing losses. Autoencoders are a type of neural network that can be used for feature fusion of audio data. Autoencoders can be used for multimodal feature fusion by combining data from multiple modalities (e.g., audio and text representations) into a single joint representation. This can be done by training an autoencoder to encode the input data into a low-dimensional space and then decoding it back to the original space. The encoded representation is a bottleneck, or compact feature representation, that captures the most important information from the input data. A detailed overview of autoencoder architectures used for multimodal feature fusion of speech data 1. Collect and preprocess audio and other modality data. 2. Build an autoencoder architecture with an encoder and decoder. 3. The encoder takes input data from speech or other modalities and compresses it into a low-dimensional representation (the bottleneck or latent representation). 4. The decoder takes the bottleneck representation and reconstructs the original input data. 5. Train the autoencoder by providing input data from both modalities as input to the encoder and using the original input data as the target output for the decoder. 6. After training, the bottleneck representation is used as the fused feature representation for both modalities. 7. This fused feature representation can then be used for more advanced analysis such as classification and clustering. Intermediate fusion strategies leverage prior knowledge to learn marginal representations for each modality, discover intra-modality correlations, and either learn joint representations based on their content or perform predictions directly. While intermediate fusion strategies can mitigate dimensional imbalance between modalities by forcing marginal representations to be similar in scale, excessive reduction of the dimensionality of larger modalities can result in significant loss of important information when the imbalance is very large. However, if the input features of the lower-dimensional modalities are selected based on prior knowledge, the imbalance does not necessarily lead to poor performance. Each speech task in the current invention's dataset contributes very specific emotional derivatives. For example, the words in paragraph 1 consist of mainly positive emotion words combined with neutral emotion, which leads healthy individuals to have a positive intention when speaking. In other words, the words in the paragraph are spoken in combination with different emotions (positive, negative, neutral), which allows us to grasp the emotional information associated with the acoustic information, which is an important indicator for detecting differences in phonemes and emotional information between healthy individuals and patients. To use different questions in the same model, the inventors assign each question a question embedding, which is a vector from 0 to N. This specific question feature extractor knows in advance which word, phoneme, and syllable level embeddings need to be extracted from the data, and is forced to align for mid-level feature fusion. Prior knowledge can come from linguists and phonetic experts about specific tasks and task sequences validated from a model evaluation process, which is performed by developing an integrated gradient utilizing SHAP values on the raw speech signal. FIG. 12A shows a graph for at least one type of corresponding task (question), in a non-limiting exemplary embodiment, where the original label is "healthy." The graph shows phonemes and segments that are positively correlated with the model prediction. Areas highlight linguistic utterances characterized by high activation in the audio file. FIG. 12B shows a graph for at least one type of task (question) with an original label of "depression" in a non-limiting exemplary embodiment. The graph shows phonemes and segments that are positively correlated with the model prediction. The area highlights verbal utterances with high activation characteristics for the second user (start time 0 seconds). FIG. 12C shows a graph for at least one type of task (question), where the original label is "healthy," in a non-limiting exemplary embodiment. The graph shows phonemes and segments that are positively correlated with the model prediction. The area highlights verbal utterances with high activation characteristics for a third user (start time 0 seconds). FIG. 12D shows a graph for at least one type of task (question) with an original label of "depression" in a non-limiting exemplary embodiment. The graph shows phonemes and segments that are positively correlated with the model prediction. The area highlights verbal utterances with high activation characteristics for the fourth user (start time 1 second). In a non-limiting exemplary embodiment, Figure 12E shows a graph for at least one type of task (question), where the original label is "healthy." The graph shows phonemes and segments that are positively correlated with the model prediction. The area highlights verbal utterances with high activation characteristics for the fifth user (start time 0 seconds). FIG. 12F shows a graph for at least one type of task (question) with an original label of "depression" in a non-limiting exemplary embodiment. The graph shows phonemes and segments that are positively correlated with the model prediction. The area highlights verbal utterances with high activation characteristics for user 6 (start time 0 seconds). FIG. 12G shows a graph for at least one type of task (question) with an original label of "healthy" in a non-limiting exemplary embodiment, where the graph shows phonemes and segments that are positively correlated with the model prediction. The area highlights verbal utterances with high activation characteristics for user 7 (start time 10 seconds). Figure 12H shows a graph for at least one type of task (question) with an original label of "depression" in a non-limiting exemplary embodiment. The graph shows phonemes and segments that are positively correlated with the model prediction. The area highlights verbal utterances with high activation characteristics for user 8 (start time 0 seconds). The mathematical formula for the autoencoder-based multimodal multi-question-input fusion architecture can be expressed as follows: Let X1, X2, X3, ..., Xn be the n modalities (speech, text, etc.) and Q1, Q2, Q3, ..., Qm be the m questions. Each question is multiplied by the learnability of the encoded matrix based on the question type. The encoder function E maps the n modalities and m questions to a low-dimensional representation Z. Z = E(X1, X2, X3, ..., Xn, Q1, Q2, Q3, ..., Qm) The decoder function D is Map the low-dimensional representation Z back to the original input space. X'1, X'2, X'3, ..., X'n = D(Z) Autoencoders are typically trained to minimize the reconstruction error between the original input and the decoder output by utilizing a loss function such as mean squared error (MSE) or cross-entropy. L = 1 / nm Σ (X'i - Xi)^2 or L = - 1 / nm Σ Xi * log(X'i) The low-dimensional representation Z is a fused feature representation of n modalities and m questions, which can be used for further analysis such as classification and clustering. The encoder degree is also used as a feature extractor for the downstream task of depression classification. Autoencoder-based feature fusion is obtained in the following way. First, hidden representations for the audio input and each text input are obtained using HuBERT and BERT models, respectively, and psychomotor delay features. Since the dataset contains similar answers to many questions, a degree representing the question ID is also input into the overall model. The question ID degree is element-wise multiplied by the hidden text features. The resulting text and audio embeddings are concatenated by forced alignment, and this concatenated feature is used to train an autoencoder model to obtain a fused representation from the bottleneck layer. The trained autoencoder degrees (up to the bottleneck layer) are loaded into one model, and a classification function head is connected to the bottleneck layer to train the classification model. In some embodiments, the subject (user) accesses data from the stimulus output module (202) through a client device (112, 114, 116, 118) in response to a response receiving module (204) and records their response. In at least one embodiment, the various vector engines (232, 234, 236, 238) are configured as part of a network server (100) and are further communicatively connected to an emotion recognition module (300) and an automatic speech recognition module (400). In a preferred embodiment, responses from the client devices (112, 114, 116, 118) are sampled at a sampling rate of 16 kHz and transmitted back to the network server (100), from where they are sent to the emotion recognition module (300). The emotion recognition module (300) generates multiple emotion classes through one or more emotion vectors, alternately pairs them with one or more stimuli, and outputs them along with their respective confidence scores. If the emotion does not match the target elicited emotion at a predetermined confidence score, the user is prompted to respond again. In some embodiments, a first predefined stimulus and a first defined emotion vector are presented to the user through the stimulus output module (202), and the response and associated response vector are recorded through the response receiving module (204). In some embodiments, a second predefined stimulus and a first predefined emotion vector are presented to the user through the stimulus output module (202), and the response and associated response vector are recorded through the response receiving module (204). Thereby, a specific predefined stimulus is composed of one or more predefined emotion vectors, various combinations of stimuli and emotion vectors are presented to the user through the stimulus output module (202), and responses and associated response vectors are recorded through the response receiving module (204). In a preferred embodiment, responses at the client devices (112, 114, 116, 118) are sampled at a sampling rate of 16 kHz and transmitted back to the network server (100), which then transmits the response data to the automatic speech recognition module (400). The response data is also transcribed using automatic speech recognition technology. In at least one embodiment, the response data, associated transcripts, and recognized emotions are stored on a server in a training storage or training database (122). In at least one embodiment, the iteration and prioritization module is configured to prioritize the stimuli based on predefined rules applied to the received answers, such as distinguishing between different emotional stimuli and neutral stimuli, presenting happy before sad, etc. In such cases, the stored data should reflect the order of questions regarding the stimuli. In at least one embodiment, the networked server (100) is communicatively coupled to a multimodal mental health classifier (500) (the aforementioned autoencoder), which is an artificial intelligence (AI) module configured to be trained to classify mental health issues based on speech recordings and text responses. Artificial intelligence ("AI") modules may be trained on structured and unstructured datasets of stimuli and answers. Training may be performed using supervised learning, unsupervised learning, or a combination of both. Machine learning (ML) and AI algorithms may be utilized to learn from the various modules. The AI module may learn rules (defined by confidence scores) about a user's mental health status by querying a vector engine to identify responses to stimuli, stimulus order, response latency, etc. Deep learning models, neural networks, deep belief networks, decision trees, genetic algorithms, and other ML and AI models may be used alone or in combination to learn from the solution grid. Artificial Intelligence (“AI”) modules may consist of a sequential combination of CNNs (Convolutional Neural Networks) and LSTMs (Long Short-Term Memory). By capturing high-level representations of the raw waveform and both short-term and long-term temporal variations, this method is used to accurately build relationships between vocal biomarkers and language models specific to users (subjects) suffering from a specific mental health condition (depression). The model encodes temporal cues associated with depression, manifested by motor cortex involvement, while taking into account deep versus shallow features. As mentioned in the data collection section, data was collected based on a different order of questions and stimuli. The training process described above was repeated for this specialized dataset to ensure that performance gains from induced emotions were not influenced by the leading questions. Given the current size of the data, no clear results or trends were observed. Our system and its use may one day yield new findings as more data is collected on different sequences and combinations. When this happens, comparing the results from these different "experiments" will allow us to efficiently select the optimal sequence of stimuli. Figure 13 shows the flowchart. Step 1: Present stimuli to the user to elicit an acoustic response. Step 2: Store emotion-related vectors across different stimuli. Step 3: Record the user's acoustic response. Step 4: Transcribe and record the text version of the user's recorded acoustic responses. Step 4a: Record the user's physiological responses, if desired. Step 5: Extract vectors from the user's acoustic responses and further analyze them to classify the emotion recognition information (obtained through an emotion detection model that produces multi-class emotions). Step 6: Emotion Recognition To perform emotion recognition from the perspective of emotion recognition information, we extract vectors from the user's text responses and proceed with the analysis. Step 6a: If necessary, extract vectors from the user's physiological responses for further analysis, and perform emotion recognition (obtained through an emotion detection model that produces multi-class emotions) in terms of emotion recognition information. Step 7: Emotion recognition information from the text responses and emotion signal vectors from the training dataset are compared by correlation with the emotion vectors associated with different stimuli to obtain the first spatial vector distance and the first degree-ranked vector difference. Step 7a.1: Optionally, first proceed with the analysis of the training dataset for a given stimulus, and then extract the most relevant emotion from the patient's responses to that stimulus using pre-defined percentiles, along with a confidence score threshold for that emotion. Step 7a.2: For a given stimulus, record the emotion derived from the user's answer and a threshold confidence score for that emotion in a database. STEP 7a.3: At runtime, for each input, obtain a sentiment class and / or confidence score and compare it with the data recorded in STEP 7a.2. Step 8: Emotion recognition information from the speech responses and emotion signal vectors from the training dataset are compared by correlation with the emotion vectors associated with different stimuli to obtain a second spatial vector distance and a second graded vector difference. Step 8a: Optionally, compare the emotion recognition information from the physiological responses and the emotion signal vectors from the training dataset by correlation with the vectors related to emotions associated with different stimuli to obtain a third spatial vector distance and a third degree-ranked vector difference. Step 8b: Optionally, as an alternative to Steps 7 and 8, use emotion recognition information from different modalities (acoustic, text, physical responses) to obtain emotion output, which is then compared with the emotion signal vector. Step 9: Intelligently integrate the first spatial vector distance, the first degree vector difference, the second spatial vector distance, and the second degree vector difference to obtain a first confidence score. Step 10: If necessary, intelligently integrate the first spatial vector distance, the first degree vector difference, the second spatial vector distance, the second degree vector difference, the third spatial vector distance, and the third degree vector difference to obtain a second confidence score. Step 11: Repeatedly present the stimuli (Step 1) and obtain a transcribed text response from the further acoustic response with respect to a first confidence score and a second confidence score based on the perceived emotion. Figure 14 shows a high-level flowchart of a voice-based mental health assessment with emotional stimuli. In a non-limiting exemplary embodiment, to capture the impact of evoked emotions (e.g., stimuli presented through multimedia output stimuli) on the performance of the invented system, a model is trained for each evoked emotion using a given dataset. A comparison (which can be performed using either n-fold cross-validation or a dedicated validation set) is then conducted to compare the performance of different models, and the top three questions (stimuli) that showed the best performance are selected. An alternative approach is to train a single model on all data combined and then run the model on a validation set consisting of responses to only one stimulus. A single metric (e.g., AUROC) is then used to compare the performance of the models for each stimulus, and the top three stimuli are selected. Using their client devices, users respond to stimuli one by one in their preferred environment. The system and its usage of the present invention performs some preprocessing and validation (e.g., vocal activation, checking for slackness, etc.) on each received response, and then passes it through the emotion recognition model (200) for emotion recognition. If the recognized emotion and its reliability score match the evoked emotion, the system and its usage of the present invention proceeds to the next stage of questioning. If the emotion does not match, the system and its usage of the present invention re-presents the same stimulus to prompt the user for an appropriate response. This process continues. If the user attempts to respond to the same stimulus but exceeds a certain threshold and emotion recognition still fails, the system and its usage of the present invention utilizes and presents similar alternative stimuli under the same evoked emotion. Furthermore, if the number of response attempts exceeds a certain threshold, the system and its usage of the present invention can proceed with evaluation while warning the user that the collected data is unreliable. The collected responses and corresponding stimulus records are sent to a multimodal mental health classifier (500) which assesses the user's current mental health status and associated risk quotient. In a preferred embodiment, classification results with a higher degree of latitude for sadness-inducing emotions are prioritized and given a higher weight in the calculation of the final result, and an assessment report is generated and transmitted back to the user and provider in real-time voice or text (through the client device) or non-real-time (e.g., email, message, etc.). The results for all evoked emotions (responses to stimuli) are recorded in the evaluation results storage or evaluation results database (124). Table 1 below shows the model performance metrics (averaged over 10 runs) for the top evoked emotion questions. By leveraging evoked emotions and emotion recognition, our system and its usage outperform other multimodal techniques, improving performance by over 10% when compared to models with hand-designed features. [Table 1] A technical advantage of the present invention is that by utilizing a combination of emotion elicitation and emotion recognition when processing vocal input from a user, it is possible to achieve significantly greater sensitivity and specificity than previously possible, thereby providing a high-quality voice-based mental health assessment and monitoring system. The technological advancement of this invention is its ability to provide an optimal multimodal architecture for speech-based mental health assessment systems and their uses, which not only includes emotionally charged elements but also combines acoustic characteristics with verbal speech and linguistic characteristics. While certain specific examples have been disclosed in this detailed illustrative description, it will be apparent that various modifications may be made by those skilled in the art without departing from the spirit of the invention and the scope of the following claims. It should further be clearly understood that the foregoing description is to be interpreted purely as illustrative of the present invention and not as limiting.
Claims
1. 1. A multimodal system for audio-based mental health assessment with emotional stimuli, comprising: a task construction module configured to construct a set of tasks for capturing acoustic, linguistic, and emotional features of a user's voice; a stimulus output module that outputs data including one or more stimuli that are presented to the user to elicit a behavioral action based on the constructed set of tasks; a response receiving module configured to present one of a plurality of stimuli to a user and receive a response corresponding to one or more formats in response to the presentation of the stimuli; a feature construction module that defines, for each task, features defined in terms of a learnable heuristic evaluation; a feature extraction module that extracts defined features from the received corresponding answers in association with the constructed tasks using a learnable heuristic evaluation model that takes into account at least one selected from the ranked tasks; and a feature fusion module for fusing the extracted two or more defined features to obtain a fused feature; a functional module comprising: using the fused features to define a relationship between the speech modality and the text modality of the feature fusion module; The speech modality cooperates with the answer receiving module to extract high-level features from the answer and output extracted high-level text features; The text modality works in cooperation with the answer receiving module to extract high-level features from the answer and output the extracted high-level speech features. an autoencoder; the autoencoder is configured to receive and fuse extracted high-level text features and extracted high-level audio features in parallel from the audio modality and the text modality to output a shared expression feature dataset used for emotion classification correlated with mental health assessment. Multimodal systems.
2. 10. The multimodal system of claim 1, The task construction module classifies the analyzed response stimuli into one of positive valence, negative valence, or neutral valence, which is the level of evaluated valence, and includes a first ranking module that ranks the constructed tasks in order of difficulty.
3. 10. The multimodal system of claim 1, The task construction module includes a second ranking module configured to rank the constructed tasks in order of complexity.
4. 10. The multimodal system of claim 1, Each constructed task is selected from a group of tasks including cognitive tasks of counting numbers within a predetermined time, tasks of pronouncing vowels within a predetermined time, tasks of pronouncing words containing voiced and unvoiced sounds within a predetermined time, tasks of reading words within a predetermined time, tasks of reading paragraphs within a predetermined time, tasks of reading paragraphs with phonemic and emotional complexity, tasks related to open-ended questions with emotional variations, and tasks related to open tasks to be performed within a predetermined time, in a multimodal system.
5. 10. The multimodal system of claim 1, The stimulus output module selects at least one stimulus from a group of stimuli consisting of audio stimuli, video stimuli, audio-visual stimuli, text stimuli, multimedia stimuli, physiological stimuli, and combinations thereof.
6. 10. The multimodal system of claim 1, A multimodal system, wherein the one or more stimuli include information indicating a stimulus vector adjusted to elicit, as a response to the stimulus vector, a text response vector indicating a response to a text stimulus, an audio response vector indicating a response to an audio stimulus, a video response vector indicating a response to a video stimulus, a multimedia response vector indicating a response to a multimedia stimulus, and / or a physiological response vector indicating a response to a physiological stimulus.
7. 10. The multimodal system of claim 1, A multimodal system, wherein the one or more stimuli are analyzed through a first vector engine configured to identify constituent vectors, and based on the constituent vectors of the stimuli, a basic state correlated with the stimulus vector of the stimuli is determined.
8. 10. The multimodal system of claim 1, A multimodal system, wherein the one or more stimuli include response vector information correlated with an audio response vector indicating an audio response and / or a visual response vector indicating a visual response elicited in response to a stimulus vector from the stimulus output module.
9. 10. The multimodal system of claim 1, The multimodal system, wherein the answer receiving module (204) includes a text reading module configured to perform a text reading task within a time period preconfigured by a user.
10. 10. The multimodal system of claim 1, the feature construction module has a Geneva Minimal Acoustic Parameter Set (GeMAPS) based feature construction function for analyzing speech responses; It analyzes the voice using a set of 62 parameters, providing a 3 frame long symmetric moving average filter to smooth over time, said smoothing being performed within voiced regions of the answer for pitch, jitter, and shimmer; The arithmetic mean and coefficient of variation were applied as a function to 18 low-level descriptors (LLDs), generating 36 parameters. Apply eight functions to the volume, Applying eight functions to pitch, determining the arithmetic mean of the alpha ratios; Determine the Hammerberg Index, determining spectral features for all unvoiced segments with reference to spectral slopes between 0 and 500 Hz and between 500 and 1500 Hz; determining temporal features in successive voiced and unvoiced regions; A multimodal system that determines Viterbi-based smoothing of the fundamental frequency contour to prevent errors from causing a single voiced frame to be dropped.
11. 10. The multimodal system of claim 1, the feature construction module has a Geneva Minimal Acoustic Parameter Set (GeMAPS) based feature construction function for analyzing speech responses; a set of low-level descriptors (LLDs) for analyzing spectral, pitch, and temporal features in said speech response; The features are: Mel-frequency cepstral coefficients (MFCCs) and their first and second derivatives, pitch and pitch variability, Energy and energy entropy, spectral centroid, broadness, and flatness; Spectral tilt, Spectral roll-off, Spectral variability, Zero crossing rate, Shimmer, jitter, harmonic to noise ratio, Pitch-based voicing probability, Temporal features in terms of loudness peak rate, mean duration and standard deviation in continuous voiced and unvoiced regions; A multimodal system, selected from a set of features including:
12. 10. The multimodal system of claim 1, the feature construction module has a Geneva Minimal Acoustic Parameter Set (GeMAPS) based feature construction function for analyzing speech responses; The features are: Pitch measured on a chromatic frequency scale, starting at 27.5 Hz (semitone 0) and scaling logarithmically with the fundamental frequency, jitter, which indicates the deviation in the length of each successive fundamental frequency period; Formant 1, 2, and 3 frequencies indicating the center frequencies for the first, second, and third formants; Formant 1, which indicates the bandwidth for the first formant; Energy-related parameters, amplitude-related parameters, Shimmer, which indicates the difference in peak amplitude between successive fundamental frequency periods; loudness, which gives an estimate of perceived signal strength from the auditory spectrum; Harmonic to noise ratio, which indicates the ratio of the energy containing harmonic components to the energy containing noise components. Spectral balance parameters, Alpha ratio, which indicates the ratio of the total energy between 50 and 1000 Hz and between 1 and 5 kHz; the Hammerberg index, which indicates the ratio between the most intense energy peak in the 0-2 kHz region and the most intense peak in the 2-5 kHz region; Spectral tilt, which indicates the slope of the linear regression of the logarithmic power spectrum within two specified bands: 0 to 500 Hz and 500 to 1500 Hz; the relative energies of formants 1, 2, and 3, which indicate the ratio of the energy associated with the spectral harmonic peaks at the center frequencies of the first, second, and third formants to the energy associated with the spectral peak at the fundamental frequency; a harmonic difference (H1-H2) that indicates the ratio of the energy associated with the first harmonic (H1) of the fundamental frequency to the energy associated with the second harmonic (H2); a harmonic difference (H1-A3) that indicates the ratio of the energy associated with the first harmonic (H1) of the fundamental frequency to the energy associated with the highest harmonic (A3) in the range of the third formant; A multimodal system, wherein the frequency-related parameters are selected from the following:
13. 10. The multimodal system of claim 1, The feature construction module includes a Geneva Minimal Acoustic Parameter Set (GeMAPS)-based feature construction function for analyzing speech responses, the feature construction function utilizing higher-order spectral analysis (HOSA) functions that utilize two or more component frequencies to achieve bispectral frequencies, which use third-order products to analyze the relationship between frequency components in a signal and examine nonlinear signals related to responses, a multimodal system.
14. 10. The multimodal system of claim 1, A multimodal system, wherein the feature fusion module integrates high-level feature embeddings obtained from the audio and text modalities using mid-level fusion.
15. 10. The multimodal system of claim 1, The feature fusion module includes a speech module configured to classify emotions from one or more answers, and the speech modality uses a question-specific feature extraction module to extract high-level features from time-frequency domain relationships in the answers and output the extracted high-level speech features.
16. 10. The multimodal system of claim 1, The feature fusion module includes a text module configured to classify sentiment from one or more responses using at least autoencoder-based feature fusion, and the text modalities use a bidirectional long short-term memory network and an attention mechanism to simulate intra-modality dynamics and output extracted high-level text features.
17. 10. The multimodal system of claim 1, The feature fusion module: Use the extracted features in acoustic feature embedding from the pre-trained model; The extracted features are compared in the spectral domain. Determine the characteristics of vocal tract coordination; determining features in replicate quantification analyses; determining features relating to the number of bigrams and the bigram duration associated with the vocalization landmarks; Fuse the features with an autoencoder. A multimodal system, including a speech module.
18. 10. The multimodal system of claim 1, The feature extraction module: a vocal landmark extraction function configured to determine an event marker associated with the answer; A multimodal system in which the determination of the event markers is correlated with a position on a timeline relative to an acoustic event from the response, the determination including determining timestamp boundaries that indicate abrupt changes in the acoustic response, and is performed frame-independently.
19. 10. The multimodal system of claim 1, the feature extraction module includes a vocal landmark extraction function; configured to identify an event marker associated with the answer; A multimodal system in which each event marker has a start value and an end value and is selected from a group of landmarks consisting of glottal-based landmarks, periodic-based landmarks, sonorant-based landmarks, fricative-based landmarks, voiced fricative-based landmarks, and burst-based landmarks, and each landmark is used to identify points in time at which different abrupt articulatory events occur and correlate with abrupt changes in power across multiple frequency ranges and multiple time scales.
20. 10. The multimodal system of claim 1, The autoencoder is constructed from a multi-modal and multi-question input fusion architecture; one or more encoders that map one or more specific features to a low-dimensional representation in combination with a task type, where each task is multiplied by an encoding matrix of learnable degrees based on the task type, and these degrees correlate with mental health assessments; one or more decoders that map the one or more particular features to the reduced dimensional representation, the decoders being configured to output a mental health assessment; The autoencoder is trained to minimize the reconstruction error between the input task and the decoder's output using a loss function.
21. 1. A multimodal method for audio-based mental health assessment involving emotional stimuli, comprising: constructing a set of tasks to capture acoustic, linguistic, and emotional characteristics of a user's voice; receiving data including one or more stimuli to be presented to the user to elicit a behavioral action based on the constructed set of tasks; presenting one or more stimuli to the user based on each configured task; receiving corresponding answers in one or more formats in response to presentation of the stimuli; defining, for each task, features defined in terms of a learnable heuristic evaluation; extracting defined features from the received corresponding answers in association with the constructed task using a learnable heuristic evaluation model that takes into account at least one selected from the ranked tasks; fusing the extracted two or more defined features to obtain a fused feature; - defining a relationship between an output of the audio modality representing the extracted high-level features and an output of the text modality representing the extracted high-level features based on the fused features; receiving and fusing extracted high-level text features and extracted high-level audio features in parallel from the audio modality and the text modality according to the relationship definition step, and outputting a shared expression feature dataset used for emotion classification correlated with mental health assessment; A multimodal method comprising:
22. 22. A multimodal method according to claim 21, comprising: The constructed task set includes one or more questions as stimuli, and in order to use different questions in the same model, each question is assigned a question embedding represented by a vector from 0 to N; A multimodal method that specializes the feature extraction module to the question by training it based on the question embedding, extracts word embeddings, phoneme embeddings, and syllable-level embeddings from the question, and performs mid-level feature fusion by forcibly aligning the extracted embeddings.
23. 22. A multimodal method according to claim 21, comprising: the step of defining the relationships includes building a multi-modal and multi-question input fusion architecture; Through an encoder, one or more specific features are combined with task types and mapped to a low-dimensional representation, where each task is multiplied by an encoding matrix of learnable degrees based on the task type, and these degrees correlate with mental health assessments; and mapping the one or more particular features to the low-dimensional representation via a decoder to output a mental health assessment. A multimodal method that uses a loss function to train to minimize the reconstruction error between the input task and the decoder's output.
Citation Information
Patent Citations
How to control time width in speech synthesis
JP2005539261A
Behavior analysis method and apparatus
JP2012000449A
A System for Speech-Based Assessment of Patient Mental Status
JP2017532082A
Early detection system of depression, anxiety, premature dementia or suicide by Artificial intelligence-based speech analysis
KR1020190081626A
Systems for speech-based assessment of a patient's state-of-mind
US20180214061A1