Depression risk assessment method, device, equipment and program product
By acquiring voice and electrocardiogram data, and using correlation selection algorithms and neural networks to process depressive symptoms, the subjectivity and prediction accuracy problems of depression diagnosis in existing technologies have been solved, and a more accurate depression risk assessment has been achieved.
Patent Information
- Application Number
- CN202510739787.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-10-28
Smart Images

Figure CN120837073A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the fields of artificial intelligence and healthcare, and specifically to a method, device, equipment, and program product for assessing depression risk. Background Technology
[0002] Currently, the diagnosis of depression mainly relies on subjective assessments by clinicians and self-reported scales by patients. This approach has several limitations, including high subjectivity, dependence on patient cooperation, diagnostic delays, and potential cultural biases. Furthermore, traditional machine learning methods struggle to effectively model the complex nonlinear relationship between physiological characteristics and depressive symptoms, often resulting in limited predictive accuracy. Summary of the Invention
[0003] In view of the above problems, this disclosure provides a method, device, equipment and procedure for assessing depression risk.
[0004] According to a first aspect of this disclosure, a method for assessing depression risk is provided, comprising: acquiring speech data and electrocardiogram (ECG) data for language tasks, wherein the speech data includes multiple speech sentences, the multiple language tasks represent multiple language emotion types, and the ECG data represents the biometrics of the subject during the expression of the speech sentences; using a relevance selection algorithm, determining depressive performance features representing depressive mood tendencies from multiple heart rate features extracted based on the ECG data and multiple acoustic features extracted based on the speech sentences; processing the depressive performance features using a feature neural network to obtain a depression score for the speech sentences, the depression score being used to identify key speech sentences from the multiple speech sentences; processing multiple depressive performance features related to the key speech sentences in the same language task using a feature sentence network to obtain initial contribution values; processing the initial contribution values corresponding to each of the multiple language tasks using a task scoring network to obtain target contribution values for the multiple language tasks, the target contribution values for the multiple language tasks representing the degree of contribution of the multiple language tasks to the depressive mood tendencies of the subject; and determining a depression risk assessment result for the subject based on the depression score and the target contribution values.
[0005] The second aspect of this disclosure provides a depression risk assessment device, comprising: an acquisition module for acquiring speech data and electrocardiogram (ECG) data for language tasks, wherein the speech data includes multiple speech sentences, the multiple language tasks represent multiple language emotion types, and the ECG data represents the biometrics of the subject during the expression of the speech sentences; a depression performance feature determination module for determining depression performance features representing a depressive mood tendency from multiple heart rate features extracted from the ECG data and multiple acoustic features extracted from the speech sentences using a correlation selection algorithm; a depression scoring module for processing the depression performance features using a feature neural network to obtain a depression score for the speech sentences, the depression score being used to determine key speech sentences from the multiple speech sentences; an initial contribution value acquisition module for processing depression performance features related to key speech sentences in the same language task using a feature sentence network to obtain an initial contribution value; a target contribution value acquisition module for processing the initial contribution values corresponding to each of the multiple language tasks using a task scoring network to obtain a target contribution value, the target contribution value representing the degree of contribution of multiple language tasks to the depressive mood tendency of the subject; and an assessment result module for determining a depression risk assessment result for the subject based on the depression score and the target contribution value.
[0006] A third aspect of this disclosure provides an electronic device comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.
[0007] A fourth aspect of this disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.
[0008] The fifth aspect of the present disclosure further provides a computer program product, comprising a computer program or instructions, which implement the steps of the above method when executed by a processor.
[0009] According to embodiments of this disclosure, by utilizing a correlation selection algorithm, representative depressive features can be selected from multiple heart rate and acoustic features, effectively capturing abnormal manifestations of patients in multiple system functions. Using a feature neural network to process the depressive features yields a depression score for speech sentences. By using the depression score, speech sentences that contribute to the depressive mood of the test subject can be preferentially selected, making it easier to capture subtle feature differences in the semantic expression stage. Using a feature sentence network to process multiple depressive features related to key speech sentences yields initial contribution values. Using a task scoring network to process the initial contribution values corresponding to multiple language tasks yields target contribution values for multiple language tasks. The target contribution values of multiple language tasks represent the degree to which multiple language tasks collectively contribute to the depressive mood of the test subject. This stage, by constructing a processing flow from depressive features to key speech sentences to language tasks, can accurately capture the complete analysis chain from microscopic speech features to macroscopic emotional expression. Finally, based on the depression score and target contribution values, a comprehensive assessment of whether the test subject suffers from depression is made, improving prediction accuracy. Attached Figure Description
[0010] The foregoing contents, other objects, features, and advantages of this disclosure will become clearer from the following description of embodiments of this disclosure with reference to the accompanying drawings.
[0011] Figure 1 The diagram illustrates an application scenario of the depression risk assessment method and apparatus according to embodiments of this disclosure.
[0012] Figure 2 A flowchart of a depression risk assessment method according to an embodiment of this disclosure is shown.
[0013] Figure 3 A structural diagram of a feature neural network according to an embodiment of the present disclosure is shown.
[0014] Figure 4 A structural diagram of a feature sentence network according to an embodiment of the present disclosure is shown.
[0015] Figure 5 A structural diagram of a depression assessment model according to an embodiment of this disclosure is shown.
[0016] Figure 6 A graph comparing normal average values with depression is shown according to an embodiment of this disclosure.
[0017] Figure 7 Shape function graphs of four key depressive symptoms according to embodiments of the present disclosure are shown.
[0018] Figure 8 A structural block diagram of a depression risk assessment device according to an embodiment of the present disclosure is shown. Detailed Implementation
[0019] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.
[0020] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0021] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0022] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0023] Depression, a common mental disorder, severely impacts patients' quality of life and social functioning, and has become a significant challenge in global public health. Currently, the diagnosis of depression mainly relies on subjective assessments by clinicians and patient self-report scales (such as the PHQ-9 and HAMD). This approach has several limitations, including strong subjectivity, dependence on patient cooperation, diagnostic delays, and potential cultural biases.
[0024] With the development of computer science and biomedical engineering, methods that analyze various physiological and behavioral characteristics of individuals to assist in the diagnosis of depression are gradually gaining attention. The pathogenesis of depression is complex, involving dysfunction across multiple systems and processes. Existing research has demonstrated that speech features can characterize the patterns of speech changes during emotional arousal; studies have found that depressed patients exhibit specific patterns in speech rate, pitch changes, and energy distribution. In addition to acoustic indicators, physiological markers, especially heart rate variability, have also become important biological indicators of depression through different measurement dimensions. Depressed patients typically exhibit an imbalance between increased sympathetic nervous system activity and decreased parasympathetic nervous system activity. Both have potential diagnostic value for depression. However, previous studies often only used one type of indicator for analysis, resulting in incomplete analysis.
[0025] Furthermore, traditional machine learning methods (such as support vector machines, random forests, and logistic regression) struggle to effectively model the complex nonlinear relationships between speech and heart rate variability features and depressive symptoms, often resulting in limited prediction accuracy and reduced clinical application value. Existing research largely focuses on statistical analysis of overall speech features, neglecting sentence-level semantic structure and differences in emotional expression, failing to fully capture the subtle changing patterns exhibited by depressed patients when faced with different emotional types of tasks. In addition, while deep learning models possess strong expressive capabilities in feature extraction and complex pattern recognition, their inherent "black box" nature makes it difficult for clinicians to understand and trust the model's decision-making basis, severely limiting the widespread application and ethical acceptance of these technologies in clinical practice. Interpretability is crucial for clinical practice.
[0026] In view of this, the present disclosure provides a method, device, and apparatus for assessing depression risk. The method includes: acquiring speech data and electrocardiogram (ECG) data for language tasks, wherein the speech data includes multiple speech sentences, the multiple language tasks represent multiple language emotion types, and the ECG data represents the biometrics of the subject during the expression of the speech sentences; using a relevance selection algorithm, determining depressive performance features representing depressive mood tendencies from multiple heart rate features extracted from the ECG data and multiple acoustic features extracted from the speech sentences; processing the depressive performance features using a feature neural network to obtain a depression score for the speech sentences, the depression score being used to identify key speech sentences from the multiple speech sentences; processing multiple depressive performance features related to the key speech sentences in the same language task using a feature sentence network to obtain initial contribution values; processing the initial contribution values corresponding to each of the multiple language tasks using a task scoring network to obtain target contribution values for the multiple language tasks, the target contribution values representing the degree of contribution of the multiple language tasks to the depressive mood tendencies of the subject; and determining the depression risk assessment result for the subject based on the depression score and the target contribution values.
[0027] It should be noted that the depression risk assessment method and device provided in this disclosure can be used in the field of artificial intelligence, or in any field other than artificial intelligence, such as the medical and health field. Therefore, the application field of the depression risk assessment method and device provided in this disclosure is not limited.
[0028] In the technical solutions disclosed herein, the user information (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0029] In scenarios involving automated decision-making using personal information, the methods, devices, and systems provided in this disclosure all offer users corresponding entry points for choosing to agree to or reject the automated decision-making results. If the user chooses to reject, the process proceeds to the expert decision-making stage. Here, "automated decision-making" refers to the activity of automatically analyzing and evaluating an individual's behavioral habits, interests, or economic, health, and credit status through computer programs, and then making a decision. Here, "expert decision-making" refers to the activity of making decisions by personnel who specialize in a particular field, possess specialized experience, knowledge, and skills, and have reached a certain level of professional expertise.
[0030] Figure 1 An application scenario diagram of the depression risk assessment method according to an embodiment of this disclosure is shown.
[0031] like Figure 1 As shown, the application scenario according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0032] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).
[0033] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0034] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0035] It should be noted that the depression risk assessment method provided in this embodiment can generally be executed by server 105. Correspondingly, the depression risk assessment device provided in this embodiment can generally be located in server 105. The depression risk assessment method provided in this embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the depression risk assessment device provided in this embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.
[0036] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0037] Figure 2 A flowchart of a depression risk assessment method according to an embodiment of this disclosure is shown.
[0038] like Figure 2 As shown, the depression risk assessment method of this embodiment includes operations S210 to S260, and the depression risk assessment method can be performed by an electronic device.
[0039] In operation S210, voice data and electrocardiogram data for language tasks are acquired. The voice data includes multiple voice sentences.
[0040] In operation S220, a correlation selection algorithm is used to determine depressive features that characterize depressive mood tendencies from multiple heart rate features extracted from electrocardiogram data and multiple acoustic features extracted from speech sentences.
[0041] In operation S230, the feature neural network is used to process the depressive performance features to obtain a depression score for the speech sentence.
[0042] In operation S240, the feature sentence network is used to process multiple depressive manifestation features related to key speech sentences in the same language task to obtain initial contribution values.
[0043] In operation S250, the task scoring network is used to process the initial contribution values of each of the multiple language tasks to obtain the target contribution values of the multiple language tasks.
[0044] In operation S260, the depression risk assessment results for the test subject are determined based on the depression score and target contribution value.
[0045] According to embodiments of this disclosure, multiple language tasks represent multiple language emotion types, such as positive, negative, and neutral emotions. This allows test subjects to complete language tasks with different emotion types. Language tasks may also include image description and text reading tasks, but are not limited to these. Embodiments of this disclosure do not limit the specific language tasks. Electrocardiogram (ECG) data can represent the biometrics of the test subject during speech expression. High-quality microphones and portable dynamic ECG recorders can be used for recording. During data acquisition, a quiet, interference-free environment is ensured, and the microphone is kept at an appropriate distance from the test subject to capture clear speech signals. The ECG acquisition equipment uses standard lead positions to ensure the quality and stability of the ECG signal. All data acquisition processes must be approved by an ethics committee and the informed consent of the test subject must be obtained.
[0046] According to embodiments of this disclosure, the acquired speech data can be segmented into several speech sentences for each language task using specialized software combined with manual annotation, based on semantic and pause features. First, an automatic segmentation function is used to identify possible sentence boundaries based on energy changes and pauses; then, professionals manually correct the automatic segmentation results to ensure the semantic integrity and boundary accuracy of the sentence segmentation.
[0047] According to embodiments of this disclosure, a correlation selection algorithm can be used to determine depressive features characterizing depressive mood tendencies from multiple heart rate features extracted from electrocardiogram data and multiple acoustic features extracted from speech sentences.
[0048] According to embodiments of this disclosure, in order to uniformly represent the input data format in multiple stages, the speech or text records of each test subject under all language tasks are organized into a three-dimensional tensor. , where T represents the total number of tasks, S represents the number of sentences in each task, and F represents the feature dimension of a single sentence. For each subject i, the formal definition is shown in formula (1).
[0049] (1);
[0050] in, Indicates the subject In the The first task, the first The fth depressive feature extracted from the sentence.
[0051] According to embodiments of this disclosure, a depression score for speech sentences can be obtained by processing depressive features using a feature neural network. To maintain the stability and consistency of the model, the depression score can be used to preferentially select multiple speech sentences with high depression scores in each language task for subsequent analysis. Key speech sentences are then identified from the multiple speech sentences using the depression score. Specifically:
[0052] The feature neural network adopts a two-layer structure to balance model complexity and expressive power, as shown in Equation (2).
[0053] (2);
[0054] in, and For network weights, and For bias parameters, This is the exponential unit activation function, used in the first layer. The linear unit activation function was modified for use in the second layer. To prevent overfitting, dropout layers with appropriate proportions were added after each layer of the network.
[0055] According to embodiments of this disclosure, by utilizing a feature neural network to process depressive manifestation features, subtle differences between speech sentences and electrocardiogram data can be captured from a more granular dimension, and its output prediction is shown in Equation (3).
[0056] (3);
[0057] in The depression score of subject i. It is a feature neural network used to process the f-th depressive symptom feature. This is a global bias term. This is the output mapping function.
[0058] According to embodiments of this disclosure, an initial contribution value can be obtained by using a feature sentence network to process multiple depressive features related to key speech sentences in the same language task. The N sentences with the highest depression scores from each language task can be selected as key speech sentences, where N is greater than or equal to 0. The initial contribution value can characterize the degree to which each key speech sentence contributes to the depressive affective tendency of the test subject.
[0059] According to embodiments of this disclosure, a task scoring network is used to process the initial contribution values corresponding to multiple language tasks to obtain target contribution values for multiple language tasks. The target contribution values of multiple language tasks characterize the degree to which multiple language tasks collectively contribute to the depressive affective tendency of the test subject. By using a task scoring network, it is possible to model the overall differences between different language tasks. This task scoring network can not only capture the differences between depressive manifestation features within key speech sentences and the expression differences between key speech sentences within language tasks, but also further integrate key information from each task stage.
[0060] According to embodiments of this disclosure, a depression risk assessment result for a test subject can be determined based on a depression score and a target contribution value. The test subject's voice and electrocardiogram data can be processed using a depression risk assessment method to obtain a depression score and a target contribution value. Based on the predicted depression score and target contribution value, the test subject is classified into depression risk categories. If both the depression score and the target contribution value are less than a first preset threshold, the test subject is considered to have no depression risk; if both the depression score and the target contribution value are greater than or equal to the first preset threshold, the test subject is considered to have depression risk.
[0061] According to embodiments of this disclosure, the depression risk assessment results are only used as intermediate information for determining the depression status, and all the above steps are performed by a computer or other device.
[0062] Figure 3 A structural diagram of a feature neural network according to an embodiment of the present disclosure is shown.
[0063] like Figure 3As shown, the feature neural network comprises multiple feature subnetworks. Each feature subnetwork can include a two-layer structure consisting of a 32-node execution unit activation function and a 16-node modified linear unit activation function, used to balance model complexity and expressive power. Each speech sentence can include multiple depressive manifestation features (i.e., input features). Each input feature in the speech sentence corresponds to a feature subnetwork. For example, if the input feature is X1310, the corresponding feature subnetwork is 311; if the input feature is X2320, the corresponding feature subnetwork is 321; and if the input feature is X3330, the corresponding feature subnetwork is 331. The data output from each feature subnetwork can be additively integrated to obtain the depression score for that speech sentence.
[0064] According to embodiments of this disclosure, when training a feature neural network, mean squared error can be used as the main loss function, while L2 parameter regularization and feature output constraints are added. The AdamW optimizer is used for training, with a learning rate of 0.001 and a batch size of 256. To further improve the accuracy of depression assessment and diagnosis, the analysis granularity can be refined from the language task level to the speech sentence level, and various machine learning models can be built on this basis, including traditional machine learning methods and sentence-based feature neural network models. Experimental results are shown in Table 1:
[0065] Table 1. Correlation of depression scores and accuracy in predicting depression among different models.
[0066]
[0067] As shown in Table 1, all models based on speech sentence data exhibited significant performance improvements, with the feature neural network model demonstrating the most outstanding performance. Regarding the correlation between depression scores and data, the feature neural network model achieved a score of 0.851, a 0.007 improvement compared to the best-performing traditional model at the sentence level, Ridge Regression. In terms of depression prediction accuracy, the feature neural network model achieved 95.0%, a 0.7% improvement over Ridge Regression's 94.3%.
[0068] According to embodiments of this disclosure, by comparing the performance of language task-level and speech sentence-level models, this embodiment finds that refining the analysis granularity to the speech sentence level can significantly improve the model's prediction performance. As shown in Table 2:
[0069] Table 2 Performance Comparison of Language Task-Level and Sentence-Level Models
[0070]
[0071] The results show that refining the analysis granularity to the sentence level can significantly improve the model's predictive performance. Various models generally outperform task-level data on sentence-level data, verifying the effectiveness of sentence-level granularity in modeling the micro-features of depression.
[0072] According to embodiments of this disclosure, by utilizing a correlation selection algorithm, representative depressive features can be selected from multiple heart rate and acoustic features, effectively capturing abnormal manifestations of patients in multiple system functions. Using a feature neural network to process the depressive features yields a depression score for speech sentences. By using the depression score, speech sentences that contribute to the depressive mood of the test subject can be preferentially selected, making it easier to capture subtle feature differences in the semantic expression stage. Using a feature sentence network to process multiple depressive features related to key speech sentences yields initial contribution values. Using a task scoring network to process the initial contribution values corresponding to multiple language tasks yields target contribution values for multiple language tasks. The target contribution values of multiple language tasks represent the degree to which multiple language tasks collectively contribute to the depressive mood of the test subject. This stage, by constructing a processing flow from depressive features to key speech sentences to language tasks, can accurately capture the complete analysis chain from microscopic speech features to macroscopic emotional expression. Finally, based on the depression score and target contribution values, a comprehensive assessment of whether the test subject suffers from depression is made, improving prediction accuracy.
[0073] According to embodiments of this disclosure, the initial contribution value is obtained by using a feature sentence network to process multiple depressive manifestation features related to key speech sentences in the same language task. This includes: extracting features from multiple depressive manifestation features related to key speech sentences in the same language task to obtain sentence contribution values; and fusing the sentence contribution values corresponding to each of the multiple key speech sentences in the same language task to obtain the initial contribution value.
[0074] According to embodiments of this disclosure, feature extraction can be performed on multiple depressive manifestation features related to key speech sentences in the same language task to obtain sentence contribution values.
[0075] According to embodiments of this disclosure, the sentence contribution values corresponding to multiple key speech sentences in the same language task can be fused to obtain an initial contribution value. Specifically, a sentence-level subnetwork can be designed to model the contribution degree of each key speech sentence to the depressive emotional tendency of the test subject, as shown in formulas (4) and (5).
[0076] (4);
[0077] (5);
[0078] in, This represents the sentence contribution value of the s-th key speech sentence of the test subject i. Representing sentence-level networks, The sentence representation of the s-th key speech sentence of the test object i. and These are the weight parameters for the sentence-level subnetwork. and For the corresponding bias term, This indicates a modified linear unit activation function. This represents the initial contribution value. This is a global bias term. This is the output mapping function.
[0079] Figure 4 A structural diagram of a feature sentence network according to an embodiment of the present disclosure is shown.
[0080] like Figure 4 As shown, depressive features can be processed using a feature neural network to obtain a depression score for a speech sentence. This depression score can be used for pre-screening from multiple speech sentences. In this embodiment, taking the selection of the four key speech sentences 410 with the highest predicted scores as an example, in the first stage, a feature-level sub-network 420 can be used to process multiple depressive features in each key speech sentence separately to obtain feature contribution values. In the second stage, a sentence-level sub-network 430 can be used to aggregate the feature contribution values corresponding to the multiple depressive features associated with the key speech sentences to obtain a sentence contribution value. Finally, the sentence contribution values corresponding to multiple key speech sentences in the same language task can be fused to obtain an initial contribution value 440.
[0081] In training the feature sentence network model, an optimizer (Adaptive Moment Estimation Weight Decay, AdamW) was used with a learning rate of 0.0008 and a batch size of 256. Experimental results show that the feature sentence network model achieved a correlation index of 0.856 for depression scores, which is 0.005 higher than the feature neural network model (0.851). Regarding the accuracy of depression classification, the feature sentence network model reached 95.2%, also showing a slight improvement compared to the feature neural network model (95.0%). This result indicates that introducing a sentence-level feature aggregation mechanism helps enhance the model's ability to model local semantic structures and improve overall prediction performance.
[0082] According to embodiments of this disclosure, feature extraction is performed on multiple depressive manifestation features related to a key speech sentence in the same language task to obtain a sentence contribution value. This includes: extracting features from multiple depressive manifestation features related to the key speech sentence to obtain feature contribution values; and aggregating the feature contribution values corresponding to each of the multiple depressive manifestation features related to the key speech sentence to obtain a sentence contribution value.
[0083] According to embodiments of this disclosure, multiple depressive manifestation features related to key speech sentences can be extracted to obtain feature contribution values, as shown in formula (6).
[0084] (6);
[0085] in, This represents the feature contribution value of the f-th depressive symptom in the s-th key speech sentence. and For network weights, and For bias parameters, For exponential unit activation functions, To modify the activation function of the linear unit
[0086] The feature contribution values corresponding to multiple depressive manifestations associated with the key speech sentence are aggregated to obtain the sentence contribution value, as shown in formula (7).
[0087] (7);
[0088] in, The sentence contribution value of the s-th key speech sentence for the i-th subject reflects the degree of contribution of each depressive manifestation feature to the depressive affective tendency of the key speech sentence, and captures the fine-grained differences within the key speech sentence. This is a sentence bias term.
[0089] According to embodiments of this disclosure, a task scoring network is used to process the initial contribution values corresponding to multiple language tasks to obtain target contribution values. This includes: for each language task, feature extraction is performed on the initial contribution value of the language task to obtain a task contribution value, whereby the task contribution value represents the degree of contribution of the language task to the depressive affective tendency of the test subject; and the task contribution values corresponding to multiple language tasks are weighted and aggregated to obtain target contribution values for multiple language tasks.
[0090] According to embodiments of this disclosure, for each language task, the initial contribution value of the language task is feature-extracted to obtain the task contribution value. Specifically, a task scoring network can be introduced to integrate the representation differences between different language tasks, as shown in formula (8).
[0091] (8);
[0092] in, This represents the task contribution value of the t-th language task for the test object i. and For the weight parameters of the task scoring network, and For the corresponding bias parameters.
[0093] Finally, the target contribution value is obtained by weighted aggregation of the task contribution values corresponding to each of the multiple language tasks, as shown in formula (9).
[0094] (9);
[0095] in, This represents the target contribution value for multiple language tasks. This is a global bias term.
[0096] Figure 5 A structural diagram of a depression assessment model according to an embodiment of this disclosure is shown.
[0097] like Figure 5 As shown, the experimenter data 510 can be acquired first, and then key speech sentences 520 can be selected from the experimenter data using a feature neural network. The four key speech sentences with the highest prediction scores can be selected for each language task. In the first stage, a feature-level sub-network 530 can be used to process multiple depressive manifestation features in each key speech sentence of each language task to obtain feature contribution values. In the second stage, a sentence-level sub-network 540 can be used to aggregate the feature contribution values corresponding to multiple depressive manifestation features related to the key speech sentences to obtain sentence contribution values. Each language task can obtain four sentence contribution values. Finally, the sentence contribution values corresponding to multiple key speech sentences in the same language task can be fused to obtain an initial contribution value 440. The task scoring network 550 is then used to process the initial contribution values corresponding to multiple language tasks to obtain the target contribution value 560 for multiple language tasks.
[0098] In training the depression assessment model, the AdamW optimizer can be used with a learning rate of 0.0005 and a batch size of 256. Experimental results show that the performance gradually improves at different stages, as shown in Table 3.
[0099] Table 3 Performance Comparison of Different Models
[0100] Model Depression score correlation Accuracy of predicting depression Feature Neural Network Model 0.851 0.950 Feature Sentence Network Model 0.856 0.952 Depression assessment model 0.862 0.958
[0101] As shown in Table 3, the performance of the three models—feature neural network model, feature sentence network model, and depression assessment model—shows a gradual improvement trend: Depression Assessment Model > Feature Sentence Network Model > Feature Neural Network Model. In depression prediction, the correlation coefficient of the depression assessment model is 0.862, which is 0.006 and 0.011 higher than the feature sentence network model (0.856) and the feature neural network model (0.851), respectively. In the depression classification task, the accuracy of the depression assessment model reaches 95.8%, making it the best performing model among all models.
[0102] It should be noted that although the improvement in prediction accuracy of the depression assessment model is relatively limited, this structure enhances the model's ability to express multi-stage features while ensuring performance stability by introducing sentence-level and task-level modeling mechanisms, thus laying a structural foundation for subsequent phased quantitative analysis.
[0103] According to embodiments of this disclosure, the depression assessment model includes a feature sentence network and a task scoring network. The method further includes: processing target depression performance features using the depression assessment model to determine the target contribution value of the target depression performance features, wherein the depression performance features include the target depression performance features; processing multiple depression performance features related to a target key speech sentence using the depression assessment model to determine the target contribution value of the target key speech sentence, wherein the key speech sentence includes the target key speech sentence; processing multiple depression performance features related to a target language task using the depression assessment model to determine the target contribution value of the target language task, wherein the language task includes the target language task; and determining the depression analysis result based on the target contribution values corresponding to the target depression performance features, the target key speech sentence, and the target language task, respectively.
[0104] According to embodiments of this disclosure, by processing target depressive features using a depression assessment model, the target contribution value of the target depressive features can be determined. The depressive features may include the target depressive features. Specifically, to calculate the independent contribution of a specific type of feature f (i.e., the target depressive feature) to the depressive mood tendency of the test subject, only the value of the target depressive feature can be retained across all language tasks and all key speech sentences, while all other depressive features are set to zero. This operation is used to evaluate the independent influence of the target depressive feature on the model output, as shown in Equation (10).
[0105] (10);
[0106] in, The target contribution value represents the characteristic of depression. For indicator functions: when The value is 1 if the condition is met and 0 otherwise, and is used to retain only the input corresponding to the target depressive performance feature f in all speech tasks t and key speech sentences s. All other target depressive symptoms were set to zero.
[0107] In the feature generation stage, input features can be mapped to target contribution values. By iterating through the depressive symptoms from smallest to largest, the shape of this function, i.e., the shape function, can be plotted. By sorting and batch averaging the values of the depressive symptoms in all training samples, a smooth curve can be drawn to show the trend of the output target contribution value as the feature values increase from low to high.
[0108] To establish a reference standard and facilitate comparative analysis with patients with depression, the average contribution value of the normal population in the test set to each depressive symptom was calculated, as shown in formula (11).
[0109] (11);
[0110] Where N represents the number of testers, This represents the target contribution value corresponding to the target depressive manifestation characteristics.
[0111] According to embodiments of this disclosure, by using a depression assessment model to process multiple depressive features related to a target key speech phrase, the target contribution value of the target key speech phrase can be determined. The key speech phrase may include the target key speech phrase. In order to quantify the independent contribution of the target key speech phrase to the depressive mood tendency of the test subject, only all depressive representation features of the s-th key speech phrase in the t-th language task can be retained, while the feature values of all depressive representation features of all key speech phrases in all other language tasks of the test subject and all other key speech phrases in the same language task can be set to zero, as shown in formula (12).
[0112] (12);
[0113] in, This represents the target contribution value of the s-th key speech phrase in the t-th language task. It is an indicator function: it takes the value 1 if and only if t'=t and s'=s, otherwise it is 0. It is used to implement the operation of "only retaining all the depressive features of the key speech sentence s in task t, and setting the feature values of other depressive features in all other key speech sentences to zero".
[0114] According to embodiments of this disclosure, by using a depression assessment model to process multiple depressive features related to a target language task, the target contribution value of the target language task can be determined. The language task includes the target language task. Specifically, only all depressive features of all key phonological sentences in the t-th language task can be retained, while all depressive features in other language tasks can be set to zero, in order to assess the degree of independent contribution of the target language task to the depressive affective tendency of the test subject, as shown in formula (13).
[0115] (13);
[0116] in, This represents the target contribution value for the target language task. Used to select the t-th language task, all depressive features in other language tasks are set to zero.
[0117] Figure 6 A graph comparing normal average values with depression is shown according to an embodiment of this disclosure.
[0118] like Figure 6 As shown, to systematically demonstrate the interpretability advantages of the model in the diagnosis of depression, this embodiment selects the best-performing depression assessment model for interpretive analysis, aiming to identify potential abnormal quantitative representations. To enhance the specificity and intuitiveness of the analysis, a typical subject (a 16-year-old female with a depression score of 16 and a clinical diagnosis of moderate to severe depression) is selected for case analysis.
[0119] Characterization Phase: This analysis identifies potential key abnormal indicators by comparing the characteristic contribution values of depressed subjects and healthy groups across various characteristic types. For example... Figure 6 As shown in the figure, the bars represent the feature contribution values of the healthy group, and the line segments represent the feature contribution values of the individual subject. The figure reveals that several high-contribution features exist in the speech modality, playing a crucial role in the model's prediction of an increase in an individual's depression score. The model particularly relies on features such as the average fundamental frequency, the five-cycle average periodic jitter, and localized participle jitter, with contribution values reaching +1.82, +1.67, and +1.23 respectively, demonstrating strong model sensitivity.
[0120] The average fundamental frequency is typically controlled by vocal cord tension and the neural modulation system, and tends to increase during states of anxiety, tension, or high arousal. Therefore, the model's dependence on the average fundamental frequency reflects its successful capture of the potential association between changes in speech pitch and the risk of depression. Five-cycle average periodic jitter and localized decibel jitter reflect the stability of speech in terms of temporal periodicity and amplitude fluctuations, respectively. The former is closely related to glottal control and vocal muscle coordination, and abnormalities are often seen in large changes in speech rate and rhythmic irregularities; the latter is affected by respiratory regulation and the degree of glottal opening, and abnormalities often result in fluctuating volume and a lack of intonation. The model assigns high weights to these two features, suggesting their important role in identifying abnormal emotional expression in the subject and providing an interpretable physiological basis for understanding the speech behavior manifestations of depression.
[0121] In the ECG modality, the low-frequency power of heart rate variability also showed a certain positive contribution (+0.35). This indicator reflects the mixed regulation of the sympathetic and parasympathetic nervous systems, and is particularly closely related to blood pressure regulation, emotional arousal, and cognitive resource mobilization. The model assigns a certain weight to it, indicating that even if the overall physiological variability capacity declines, some sympathetic-related features still retain explanatory power in the model and may be used to identify an individual's level of emotional stress response. The total power of heart rate variability is the ECG feature with the largest negative contribution in the model (+1.59), reflecting the overall activity of the autonomic nervous system. In healthy individuals, the total power of heart rate variability is used to "reduce" the prediction of depression, but in this subject, its contribution is significantly weakened, suggesting that the model perceives a decline in its physiological regulatory capacity.
[0122] Sentence and Task Stages: To further analyze the explanatory power of the depression assessment model in the task and sentence stages, this embodiment selects the task with the highest contribution in this subject—"Describe a picture with neutral emotion"—for in-depth analysis. The total contribution value of this task is +2.82, which is much higher than the average reference value of the healthy group (1.82), indicating that the model identifies strong abnormal signals in this task.
[0123] The first sentence, "Looking around, looking around, everyone seems to have no idea what they're doing," received the highest contribution from the model (+1.26). This sentence not only exhibits repetitive and hesitant language, but also conveys confusion and uncertainty semantically. This is highly correlated with the increased average fundamental frequency (+1.82) and increased five-cycle average periodic jitter (+1.67) in the speech features, reflecting increased vocal cord tension and unstable vocal cycles. These features are typically prominent in states of emotional tension or anxiety. The third sentence, "I feel the whole scene is very chaotic," is a typical negative subjective evaluation, inconsistent with the actual content of the image, indicating that the subject tends to interpret neutral stimuli as negative situations. Combined with ECG feature analysis, this sentence expresses a positive low-frequency power contribution (+0.35) associated with heart rate variability, while the total power contribution of heart rate variability is weakened (+1.59), reflecting abnormal regulation of the autonomic nervous system in emotional arousal and cognitive resource mobilization. The model assigned a medium-to-high contribution (+0.67) to this sentence, demonstrating its ability to identify the association between emotional cognitive biases and abnormal physiological regulation. The second sentence, "The person standing by the stove had a very calm expression. I didn't know what to say. Oh, and a mouse ran away," exhibits semantic jumps and confused clues. At this point, the localized vibrato feature value increases (+1.23), reflecting instability in speech amplitude. This feature is associated with respiratory regulation and glottal control disorder, suggesting that the subject's speech control ability declines when cognitive processing load increases. The model assigns this sentence a moderate contribution (+0.55). In contrast, the fourth sentence, "Overall, it felt relatively simple and ordinary," exhibits semantic neutrality, with a negative contribution from the model (-0.22), indicating that the ability to maintain objective evaluation may be an important indicator of adolescent psychological resilience.
[0124] Figure 7 Shape function graphs of four key depressive symptoms according to embodiments of the present disclosure are shown.
[0125] like Figure 7As shown, the study selected four features with the most significant contributions for in-depth analysis: mean fundamental frequency, localized decibel jitter, five-cycle average jitter, and low-frequency power of heart rate variability. The results showed that the mean fundamental frequency had a slight impact on depression scores below 225 Hz, but its contribution increased rapidly and then stabilized above this threshold, indicating that a higher speech fundamental frequency may be associated with an increased risk of depression, consistent with the acoustic changes in mood fluctuations in depressed patients. Localized decibel jitter had no significant contribution in the low-value range, but showed a sharp upward trend above the critical value of 1.2, revealing a potential association between increased speech amplitude perturbation and the severity of depressive symptoms. Five-cycle average jitter showed a similar pattern, contributing weakly below 0.0125, but its positive contribution to the score significantly increased above this threshold, suggesting a close relationship between speech cycle instability and depression risk. The low-frequency power of heart rate variability had a small impact on depression scores below 0.045, but showed a significant negative contribution above this threshold, indicating that higher heart rate variability may predict a lower risk of depression.
[0126] To evaluate the consistency of this learning curve under different initial random seeds, an ensemble approach was adopted: for the depression assessment model, a semi-transparent blue line was used, with the darkest shade representing the mean. This method effectively presents the consistency and bias of the model learning the same shape function, providing a basis for evaluating the confidence of the shape function. Simultaneously, the study used pink bars in the same chart to display the normalized data density distribution; the shades of the color correspond to the amount of data, facilitating the assessment of the sufficiency of training data in each region.
[0127] According to embodiments of this disclosure, depression analysis results can be determined based on the target contribution values corresponding to the target depressive characteristics, target key phonological phrases, and target language tasks. This provides clinicians with a detailed report interpreting the predicted results, including analysis of the target depressive characteristics of the subject, the target contribution values corresponding to the target key phonological phrases and target language tasks, and clinical reference suggestions, assisting in the development of personalized intervention plans.
[0128] According to the embodiments of this disclosure, in order to maintain consistency among model structures, the three-stage models of this disclosure (feature neural network model, feature sentence network model, and depression assessment model) all adopt the same loss function design. Mean squared error is used as the main loss function during model training. At the same time, to improve the model's generalization performance and interpretability, two additional regularization terms are added: L2 parameter regularization and feature output constraint. The complete loss function is shown in formula (14).
[0129] (14);
[0130] in, This is the prediction error term; This is an L2 parameter regularization term to suppress excessively large network weights; This is a feature output constraint term to prevent the feature network output values from becoming too large, thereby improving the robustness of the model's interpretation. Among them... and This is a hyperparameter used to weigh the importance of each regularization loss in the overall loss function.
[0131] The model can be trained and optimized using the following methods:
[0132] The first step is cross-validation. A 5-fold cross-validation strategy is used to evaluate model performance. All data is divided into a training set (70%) and a fixed test set (30%). Within the training set, it is randomly divided into an 80% sub-training set and a 20% validation set, and the training is repeated 20 times under different random seeds to enhance the stability of the results. In each run, the validation set is used for early stopping to prevent overfitting.
[0133] Step 2: Ensemble learning strategy. For each fold, the models trained 20 times are ensembled for the final prediction to improve model stability and prediction accuracy.
[0134] Step 3: Optimizer and Learning Rate Settings. The AdamW optimizer is used for training. Appropriate initial learning rates are set for different models, the batch size is set to 256, the maximum number of training epochs is 100, and an early stopping strategy is adopted when the validation loss fails to improve for 30 consecutive epochs. The learning rate is decayed by multiplying by 0.995 after each training epoch.
[0135] Step 4: Hyperparameter Optimization. Systematic tuning of key hyperparameters of the model is performed, including the learning rate (range 0.0001-0.01), output penalty coefficient (range 0.01-0.2), weight decay coefficient (range 0.001-0.01), and Dropout rate at different stages (range 0.1-0.4). Hyperparameter optimization employs a grid search method, selecting parameters based on validation set performance to balance the model's predictive performance and generalization ability.
[0136] According to embodiments of this disclosure, using a correlation selection algorithm to determine depressive manifestation features characterizing depressive mood tendencies from multiple heart rate features extracted from electrocardiogram data and multiple acoustic features extracted from speech data includes: obtaining a self-rating scale for assessing the degree of depression; using the correlation selection algorithm to calculate the correlation coefficient between each feature in the feature set and the self-rating information of the mood in the self-rating scale, wherein the features in the feature set are heart rate features or acoustic features; and determining depressive manifestation features related to speech sentences from the feature set based on multiple correlation coefficients.
[0137] According to embodiments of this disclosure, a self-rating scale for assessing the degree of depression is obtained. Multidimensional features can be extracted from the collected speech and electrocardiogram (ECG) data, specifically including: acoustic features of three main dimensions extracted from the speech data: speech quality indicators, including average fundamental frequency, harmonic noise ratio, jitter variables (e.g., local jitter, relative average perturbation jitter, and five-point periodic perturbation quotient), and dimness variables (e.g., local dimness and decibel-scale dimness); spectral characteristics, including formant frequencies and their respective medians; and acoustic characteristics, including entropy, zero-crossing rate, and short-time energy. Simultaneously, heart rate features are extracted from the ECG data, mainly including three dimensions: time-domain indicators such as root mean square difference, standard deviation of the interval difference between adjacent ventricular depolarizations, and standard deviation of heartbeat intervals; frequency-domain indicators including low frequency, high frequency, and total power, characterizing the regulatory function of the autonomic nervous system model; and nonlinear domain indicators such as scatter plot parameters, approximate entropy, and sample entropy, used to reveal the complexity and self-similarity properties of heart rate specificity. All features are standardized to eliminate dimensional influences.
[0138] According to embodiments of this disclosure, a correlation-based selection algorithm can be used to identify the features most relevant to depression assessment. By calculating the correlation coefficient between each feature in the feature set and the self-rating emotion information in the self-rating scale, those features most relevant to depression can be identified. The features in the feature set are heart rate features or acoustic features. Based on multiple correlation coefficients, depressive manifestation features related to speech sentences are determined from the feature set.
[0139] According to embodiments of this disclosure, depressive features that characterize depressive mood tendencies are selected from multiple heart rate and acoustic features, while other interfering features are excluded. This can effectively capture abnormal performance of the subject in multiple system functions (heart rate, language).
[0140] According to embodiments of this disclosure, determining a key speech sentence from multiple speech sentences includes: sorting the depression scores corresponding to each of the multiple speech sentences to obtain a sorting result; and determining the key speech sentence from the multiple speech sentences based on the sorting result.
[0141] According to embodiments of this disclosure, each speech sentence has a corresponding depression score. The depression scores corresponding to all speech sentences are sorted to obtain a sorting result. Based on the sorting result, the speech sentences that contribute to the depressive emotional tendency of the test subject can be selected from multiple speech sentences and identified as key speech sentences. Using key speech sentences can more easily capture the subtle feature differences of the test subject in the semantic expression stage.
[0142] Figure 8 A structural block diagram of a depression risk assessment device according to an embodiment of the present disclosure is shown.
[0143] like Figure 8As shown, the depression risk assessment device of this embodiment includes an acquisition module 810, a depression performance characteristic determination module 820, a depression scoring module 830, an initial contribution value acquisition module 840, a target contribution value acquisition module 850, and an assessment result module 860.
[0144] The acquisition module 810 is used to acquire speech data and electrocardiogram data for language tasks. The speech data includes multiple speech sentences, and the multiple language tasks represent multiple language emotion types. The electrocardiogram data represents the biological characteristics of the test subject in the process of expressing speech sentences.
[0145] The depressive manifestation feature determination module 820 is used to determine depressive manifestation features that characterize depressive mood tendencies from multiple heart rate features extracted based on electrocardiogram data and multiple acoustic features extracted based on speech sentences using a correlation selection algorithm.
[0146] The depression scoring module 830 is used to process the features of depression using a feature neural network to obtain a depression score for a speech sentence. The depression score is used to identify key speech sentences from multiple speech sentences.
[0147] The initial contribution value is obtained from module 840, which uses the feature sentence network to process depressive performance features related to key speech sentences in the same language task to obtain the initial contribution value.
[0148] The target contribution value acquisition module 850 is used to process the initial contribution values of multiple language tasks using the task scoring network to obtain the target contribution value. The target contribution value represents the degree of contribution of multiple language tasks to the depressive affective tendency of the test subject.
[0149] The assessment results module 860 is used to determine the depression risk assessment results for the test subject based on the depression score and target contribution value.
[0150] According to embodiments of this disclosure, by utilizing a correlation selection algorithm, representative depressive features can be selected from multiple heart rate and acoustic features, effectively capturing abnormal manifestations of patients in multiple system functions. Using a feature neural network to process the depressive features yields a depression score for speech sentences. By using the depression score, speech sentences that contribute to the depressive mood of the test subject can be preferentially selected, making it easier to capture subtle feature differences in the semantic expression stage. Using a feature sentence network to process multiple depressive features related to key speech sentences yields initial contribution values. Using a task scoring network to process the initial contribution values corresponding to multiple language tasks yields target contribution values for multiple language tasks. The target contribution values of multiple language tasks represent the degree to which multiple language tasks collectively contribute to the depressive mood of the test subject. This stage, by constructing a processing flow from depressive features to key speech sentences to language tasks, can accurately capture the complete analysis chain from microscopic speech features to macroscopic emotional expression. Finally, based on the depression score and target contribution values, a comprehensive assessment of whether the test subject suffers from depression is made, improving prediction accuracy.
[0151] According to embodiments of this disclosure, the initial contribution value obtaining module 840 includes: a sentence contribution value unit and an initial contribution value unit.
[0152] The sentence contribution value unit is used to extract features from multiple depressive manifestation features related to key speech sentences in the same language task, and obtain the sentence contribution value.
[0153] The initial contribution value unit is used to merge the sentence contribution values corresponding to multiple key speech sentences in the same language task to obtain the initial contribution value.
[0154] According to embodiments of this disclosure, the sentence contribution value unit includes: a feature contribution value subunit and a sentence contribution value subunit.
[0155] The feature contribution value subunit is used to extract features from multiple depressive manifestation features related to key speech sentences to obtain feature contribution values.
[0156] The sentence contribution value subunit is used to aggregate the feature contribution values corresponding to multiple depressive manifestation features related to the key speech sentence to obtain the sentence contribution value.
[0157] According to embodiments of this disclosure, the target contribution value obtaining module 850 includes: a task contribution value unit and a target contribution value unit.
[0158] The task contribution value unit is used to extract features from the initial contribution value of each language task to obtain the task contribution value. The task contribution value represents the degree of contribution of the language task to the depressive affective tendency of the test subject.
[0159] The target contribution value unit is used to weight and aggregate the task contribution values corresponding to multiple language tasks to obtain the target contribution value.
[0160] According to embodiments of this disclosure, the depression assessment model includes a feature sentence network and a task scoring network.
[0161] According to an embodiment of this disclosure, the depression risk assessment device includes: a first module, a second module, a third module, and a depression analysis result module.
[0162] The first module is used to process target depressive features using a depression assessment model and determine the target contribution value of the target depressive features, which include the target depressive features.
[0163] The second module is used to process multiple depressive manifestation features related to the target key speech sentence using a depression assessment model, and to determine the target contribution value of the target key speech sentence, which includes the target key speech sentence.
[0164] The third module is used to process multiple depressive manifestations related to the target language task using a depression assessment model, and to determine the target contribution value of the target language task, which includes the target language task.
[0165] The depression analysis results module is used to determine the depression analysis results based on the target contribution values corresponding to the target depression manifestation characteristics, target key speech sentences, and target language tasks.
[0166] According to embodiments of this disclosure, the depressive manifestation feature determination module 820 includes: an acquisition unit, a correlation coefficient unit, and a depressive manifestation feature unit.
[0167] The acquisition unit is used to obtain a self-rating scale for assessing the degree of depression.
[0168] The correlation coefficient unit is used to calculate the correlation coefficient between each feature in the feature set and the emotional self-rating information in the self-rating scale using the correlation selection algorithm. The features in the feature set are heart rate features or acoustic features.
[0169] The depressive manifestation feature unit is used to determine depressive manifestation features related to speech sentences from a feature set based on multiple correlation coefficients.
[0170] According to embodiments of this disclosure, the depression scoring module 830 includes: a ranking result unit and a key speech sentence unit.
[0171] The sorting result unit is used to sort the depression scores corresponding to multiple speech sentences to obtain the sorting results.
[0172] Key speech sentence unit, used to identify key speech sentences from multiple speech sentences based on the ranking results.
[0173] According to embodiments of this disclosure, any multiple modules among the acquisition module 810, the depressive manifestation feature determination module 820, the depression scoring module 830, the initial contribution value acquisition module 840, the target contribution value acquisition module 850, and the evaluation result module 860 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in one module. According to embodiments of this disclosure, at least one of the acquisition module 810, the depressive manifestation feature determination module 820, the depression scoring module 830, the initial contribution value acquisition module 840, the target contribution value acquisition module 850, and the evaluation result module 860 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the following modules can be implemented, at least partially, as a computer program module: the acquisition module 810, the depression symptom determination module 820, the depression scoring module 830, the initial contribution value acquisition module 840, the target contribution value acquisition module 850, and the evaluation result module 860. When the computer program module is run, it can perform the corresponding function.
[0174] Those skilled in the art will appreciate that the features described in the various embodiments of the present disclosure may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in the present disclosure. In particular, the features described in the various embodiments of the present disclosure may be combined and / or coupled in various ways without departing from the spirit and teachings of the present disclosure. All such combinations and / or couplings fall within the scope of the present disclosure.
[0175] The above describes the embodiments of the present disclosure. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present disclosure, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present disclosure.
Claims
1. A method for assessing depression risk, characterized in that, include: Acquire speech data and electrocardiogram (ECG) data for language tasks. The speech data includes multiple speech sentences, and the multiple language tasks represent multiple language emotion types. The ECG data represents the biometrics of the test subject in the process of expressing the speech sentences. Using a correlation selection algorithm, depressive features characterizing depressive mood tendencies are determined from multiple heart rate features extracted based on the electrocardiogram data and multiple acoustic features extracted based on the speech sentences. The depressive features are processed using a feature neural network to obtain a depression score for the speech sentences, and the depression score is used to identify key speech sentences from multiple speech sentences; The feature sentence network is used to process multiple depressive manifestation features related to the key speech sentence in the same language task to obtain an initial contribution value; The initial contribution values of each of the multiple language tasks are processed using a task scoring network to obtain the target contribution values of the multiple language tasks. The target contribution values of the multiple language tasks represent the degree of contribution of the multiple language tasks to the depressive affective tendency of the test subject. Based on the depression score and the target contribution value, the depression risk assessment result for the subject to be tested is determined.
2. The method according to claim 1, characterized in that, The process of using a feature sentence network to process multiple depressive manifestation features related to the key speech sentence in the same language task, and obtaining initial contribution values, includes: Feature extraction is performed on multiple depressive manifestation features related to the key speech sentence in the same language task to obtain sentence contribution values; The sentence contribution values corresponding to multiple key speech sentences in the same language task are merged to obtain an initial contribution value.
3. The method according to claim 2, characterized in that, The step of extracting features from multiple depressive manifestation features related to the key speech sentence in the same language task to obtain sentence contribution values includes: Feature extraction is performed on multiple depressive manifestation features related to the key speech sentence to obtain feature contribution values; The feature contribution values corresponding to each of the multiple depressive manifestation features associated with the key speech sentence are aggregated to obtain the sentence contribution value.
4. The method according to claim 1, characterized in that, The step of using a task scoring network to process the initial contribution values corresponding to each of the multiple language tasks to obtain the target contribution value includes: For each language task, feature extraction is performed on the initial contribution value of the language task to obtain the task contribution value, which represents the degree of contribution of the language task to the depressive affective tendency of the test subject. The target contribution value is obtained by weighted aggregation of the contribution values of each of the multiple language tasks.
5. The method according to claim 1, characterized in that, The depression assessment model includes the feature sentence network and the task scoring network, and the method further includes: The depression assessment model is used to process target depressive features to determine the target contribution value of the target depressive features, wherein the depressive features include the target depressive features. The depression assessment model is used to process multiple depression manifestation features related to the target key speech sentence to determine the target contribution value of the target key speech sentence, wherein the key speech sentence includes the target key speech sentence; The depression assessment model is used to process multiple depression manifestation features related to the target language task to determine the target contribution value of the target language task, wherein the language task includes the target language task. The depression analysis results are determined based on the target contribution values corresponding to the target depressive manifestation features, the target key speech phrases, and the target language task.
6. The method according to claim 1, characterized in that, The method of using a correlation selection algorithm to determine depressive features characterizing depressive mood tendencies from multiple heart rate features extracted from the electrocardiogram data and multiple acoustic features extracted from the speech data includes: Obtain a self-rating scale to assess the degree of depression; Using the aforementioned correlation selection algorithm, the correlation coefficient between each feature in the feature set and the emotional self-rating information in the self-rating scale is calculated, wherein the feature in the feature set is the heart rate feature or the acoustic feature; Based on multiple correlation coefficients, depressive features associated with the spoken sentences are determined from the feature set.
7. The method according to claim 1, characterized in that, The step of determining the key speech sentence from the plurality of speech sentences includes: The depression scores corresponding to each of the multiple spoken sentences are sorted to obtain the sorting results; Based on the sorting results, key speech sentences are determined from the multiple speech sentences.
8. A depression risk assessment device, characterized in that, include: The acquisition module is used to acquire speech data and electrocardiogram (ECG) data for language tasks. The speech data includes multiple speech sentences, and the multiple language tasks represent multiple language emotion types. The ECG data represents the biometric characteristics of the test subject in the process of expressing the speech sentences. The module for determining depressive manifestation features is used to determine depressive manifestation features that characterize depressive mood tendencies from multiple heart rate features extracted based on the electrocardiogram data and multiple acoustic features extracted based on the speech sentences using a correlation selection algorithm. A depression scoring module is used to process the depression performance features using a feature neural network to obtain a depression score for the speech sentence, and the depression score is used to identify key speech sentences from multiple speech sentences; The initial contribution value acquisition module is used to process the depressive performance features related to the key speech sentence in the same language task using the feature sentence network to obtain the initial contribution value; The target contribution value acquisition module is used to process the initial contribution values corresponding to each of the multiple language tasks using a task scoring network to obtain a target contribution value, wherein the target contribution value represents the degree of contribution of the multiple language tasks to the depressive affective tendency of the test subject. The assessment results module is used to determine the depression risk assessment results for the test subject based on the depression score and the target contribution value.
9. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.