Multi-modal depression detection model training method and system, computer equipment and computer program product

By constructing a multimodal depression detection model and using audio and video data for individual baseline modeling, the problems of high cost and low accuracy of traditional depression detection are solved, and accurate depression risk detection for the elderly are achieved.

CN120388749APending Publication Date: 2025-07-29SHENZHEN KESI CHUANGDONG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510355180.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

Traditional depression testing methods have high cost, low frequency and unstable accuracy for the elderly, making it difficult to adapt to the daily life scenarios of elderly people living alone.

Method used

By obtaining audio and video data of the elderly, multimodal data is constructed, individual baseline modeling is performed, and depression prediction network is trained to generate target depression detection models.

Benefits of technology

Accurate and efficient depression risk detection for the elderly is achieved, which can accurately capture user status changes and predict depression risks, improving the accuracy and stability of the detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388749A_ABST
    Figure CN120388749A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of emotion recognition, and provides a multi-modal depression detection model training method and system, computer equipment and a computer program product, and the method comprises the steps: obtaining the multi-modal data of a target user based on the audio data and video data of the target user; performing individual baseline modeling for the target user based on the multi-modal data; and training a depression prediction network based on the multi-modal data and the individual baseline to obtain a target depression detection model. According to the method, the multi-modal data of the target user is obtained through the audio data and the video data, the obtained multi-modal data can more comprehensively reflect the user state, then the individual baseline is established for the target user based on the multi-modal data, the normal states of different individuals are accurately mastered, and the user experience is improved. And finally, training the depression prediction network based on the multi-modal data and the individual baseline to obtain a target depression detection model capable of accurately capturing the state change of the user and accurately predicting the depression risk of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of emotion recognition, and particularly relates to a method, system, computer device and computer program product for training a multimodal depression detection model. Background Art

[0002] With the acceleration of the global aging process, the mental health problems of elderly people living alone are becoming increasingly prominent. Especially the high potential incidence of depression in the elderly population has had a serious negative impact on the quality of life and sense of happiness of the elderly in their daily lives. Traditional depression detection methods, such as questionnaires and interviews, although effective to a certain extent, have many limitations, especially in adapting to the daily life scenarios of the elderly living alone.

[0003] On the one hand, traditional depression detection methods require professional personnel to communicate face-to-face with the elderly, increasing the implementation cost and also restricting the frequency and scope of detection; on the other hand, limited by the communication and expression ability of the elderly, the detection accuracy is not stable enough.

[0004] Therefore, there is an urgent need for a depression detection model to accurately and efficiently detect the potential depression risk of the target group including the elderly. Summary of the Invention

[0005] In view of this, embodiments of this application provide a method, system, computer device and computer program product for training a multimodal depression detection model, so as to provide a depression detection model that can accurately and efficiently detect the depression risk of the target group including the elderly.

[0006] The first aspect of the embodiments of this application provides a method for training a multimodal depression detection model, including:

[0007] Obtaining the multimodal data of the target user based on the audio data and video data of the target user, and the time stamps of the multimodal data correspond to the audio data and the video data;

[0008] Performing individual baseline modeling for the target user based on the multimodal data;

[0009] Training a depression prediction network based on the multimodal data and the individual baseline to obtain a target depression detection model.

[0010] In one implementation of the first aspect, the multimodal data includes speech data, semantic data, body data, face data and rPPG data;

[0011] The obtaining the multimodal data of the target user based on the audio data and video data of the target user includes:

[0012] Obtain the speech data and semantic data of the target user based on the audio data, where the speech data includes high-order speech feature vectors and the semantic data includes high-order semantic feature vectors;

[0013] Perform target segmentation on the target user based on the video data to obtain a body video and a face video;

[0014] Obtain the body data of the target user based on the body video, where the body data includes high-order body feature vectors;

[0015] Obtain the face data of the target user based on the face video, where the face data includes high-order face feature vectors;

[0016] Obtain the rPPG data of the target user based on the face video and the rPPG algorithm, where the rPPG data includes high-order rPPG feature vectors.

[0017] In one implementation manner of the first aspect, training a depression prediction network based on the multi-modal data and the individual baseline to obtain a target depression detection model includes:

[0018] Perform high-level feature fusion on the high-order speech feature vectors, the high-order semantic feature vectors, the high-order body feature vectors, the high-order face feature vectors, and the high-order rPPG feature vectors with the same time stamp to obtain a first fusion feature vector corresponding to the time stamp;

[0019] Concatenate multiple first fusion feature vectors to obtain a segment-level fusion vector;

[0020] Fuse each first fusion feature vector in the segment-level fusion vector with the position feature vector, the pattern feature vector, and the time feature vector corresponding to the first fusion feature vector to obtain a second fusion feature vector, and concatenate the second fusion feature vectors to obtain a target segment-level vector. The position feature vector includes the position information of the first fusion feature vector in the segment-level fusion vector, the pattern feature vector includes the acquisition method of the first fusion feature vector, and the time feature vector includes the time stamp information of the first fusion feature vector;

[0021] Train a depression prediction network based on the target segment-level vector and the individual baseline to obtain a target depression detection model.

[0022] In one implementation manner of the first aspect, the segment-level fusion vector includes the first fusion feature vectors whose time stamps belong to multiple time periods of multiple days;

[0023] Training a depression prediction network based on the target segment-level vector and the individual baseline to obtain a target depression detection model includes:

[0024] Input the target segment-level vector into the depression prediction network to obtain a hidden state sequence, where the hidden state sequence includes a plurality of hidden state vectors, and the hidden state vectors correspond one-to-one with the second fusion feature vectors in the target segment-level vector;

[0025] Based on the hidden state sequence and the individual baseline, perform depression risk prediction to obtain a predicted risk value;

[0026] Calculate a loss function based on the predicted risk value and the actual risk value, and perform backpropagation to update the parameters of the depression prediction network;

[0027] When the preset training termination condition is reached, output the depression prediction network as the target depression detection model.

[0028] In one implementation manner of the first aspect, the training of the depression prediction network based on the target segment-level vector and the individual baseline to obtain the target depression detection model includes:

[0029] Input multiple second fusion feature vectors corresponding to the same true value in the target segment-level vector into the depression prediction network to obtain a daily hidden state sequence, where the daily hidden state sequence includes a plurality of daily hidden state vectors;

[0030] Perform pooling on the daily hidden state sequence to obtain a daily vector;

[0031] Based on the daily vector and the individual baseline, perform depression risk prediction to obtain a predicted risk value;

[0032] Calculate a loss function based on the predicted risk value and the true value, and perform backpropagation to update the parameters of the depression prediction network;

[0033] When the preset training termination condition is reached, output the depression prediction network as the target depression detection model.

[0034] In one implementation manner of the first aspect, the segment-level fusion vector includes the first fusion feature vectors of multiple time periods on the same day;

[0035] The training of the depression prediction network based on the target segment-level vector and the individual baseline to obtain the target depression detection model includes:

[0036] Input the first fusion feature vector with the earliest timestamp into the depression prediction network to obtain a first hidden state, and output a first daily prediction in combination with the individual baseline;

[0037] Input the first fused feature vector corresponding to the next timestamp into the depression prediction network, update the previous hidden state to a new hidden state, and combine the individual baseline to output a new daily prediction;

[0038] Repeat the steps of inputting the first fused feature vector corresponding to the next timestamp into the depression prediction network, updating the previous hidden state to a new hidden state, and combining the individual baseline to output a new daily prediction until all the fused feature vectors of the same day are traversed, and output the final daily prediction.

[0039] In an implementation manner of the first aspect, training the depression prediction network based on the multi-modal data and the individual baseline to obtain a target depression detection model includes:

[0040] Segment the multi-modal data of M days of the target user into M - N + 1 training samples through a sliding window of length N, where each training sample includes N days of multi-modal data;

[0041] Predict the prediction risk value of the last day in each training sample based on each training sample and the individual baseline respectively;

[0042] Calculate the loss function based on the prediction risk value of the last day in each training sample and the corresponding true value, and perform backpropagation to update the parameters of the depression prediction network;

[0043] When the preset training termination condition is reached, output the depression prediction network as the target depression detection model.

[0044] A multi-modal depression detection model training system provided in the second aspect of the embodiments of the present application includes:

[0045] A multi-modal data extraction module, configured to obtain the multi-modal data of the target user based on the audio data and video data of the target user, and the timestamps of the multi-modal data correspond to the audio data and the video data;

[0046] An individual baseline module, configured to perform individual baseline modeling for the target user based on the multi-modal data;

[0047] A training module, configured to train a depression prediction network based on the multi-modal data and the individual baseline to obtain a target depression detection model.

[0048] A computer device provided in the third aspect of the embodiments of the present application includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the steps of the method described in the first aspect are implemented.

[0049] A fourth aspect of the embodiments of the present application provides a computer program product, including a computer program, which, when run, causes the method described in the first aspect to be executed.

[0050] The beneficial effects of the first aspect of the embodiments of the present application are as follows: Multimodal data of the target user is obtained through audio data and video data, and the obtained multimodal data can more comprehensively reflect the user's state. Then, an individual baseline is established for the target user based on the multimodal data to achieve an accurate grasp of the normal state of different individuals. Finally, a depression prediction network is trained based on the multimodal data and the individual baseline to obtain a target depression detection model that can accurately capture changes in the user's state and accurately predict the user's depression risk.

[0051] It can be understood that the beneficial effects of the second to fourth aspects described above can refer to the relevant descriptions in the first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0053] Figure 1 It is a schematic flowchart of the implementation of the multimodal depression detection model training method provided by the embodiments of the present application;

[0054] Figure 2 It is a schematic diagram of the multimodal depression detection model training system provided by the embodiments of the present application;

[0055] Figure 3 It is a schematic diagram of the computer device provided by the embodiments of the present application;

[0056] Figure 4 It is a schematic diagram of the computer program product provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] In the following description, specific details such as specific system structures and technologies are proposed for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0058] It should be understood that when used in the specification of this application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0059] It should also be understood that the term "and / or" used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0060] As used in the specification of this application and the appended claims, the term "if" may be interpreted as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" may be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" depending on the context.

[0061] In addition, in the description of the specification of this application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0062] The reference to "one embodiment" or "some embodiments" etc. described in the specification of this application means that a specific feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0063] The multi-modal depression detection model training method provided by the embodiments of this application can be applied to computer devices such as mobile phones, tablet computers, laptop computers, Ultra-Mobile Personal Computers (UMPCs), netbooks, and Personal Digital Assistants (PDAs). The embodiments of this application do not impose any restrictions on the specific types of computer devices.

[0064] An embodiment of the present application provides a method for training a multi-modal depression detection model, which is used to train a depression detection model that can accurately detect the depression risk of a target group including the elderly. The method provided by the present application obtains the multi-modal data of the target user through audio data and video data. The obtained multi-modal data can more comprehensively reflect the user's state. Then, an individual baseline is established for the target user based on the multi-modal data to achieve an accurate grasp of the normal state of different individuals. Finally, a depression prediction network is trained based on the multi-modal data and the individual baseline to obtain a target depression detection model that can accurately capture the change of the user's state and accurately predict the depression risk of the target user.

[0065] As Figure 1 shown, an embodiment of the present application provides a method for training a multi-modal depression detection model, including:

[0066] Step S10, obtaining the multi-modal data of the target user based on the audio data and video data of the target user, and the time stamp of the multi-modal data corresponds to the audio data and the video data.

[0067] In application, the time stamp of the multi-modal data is determined by the acquisition time of the audio data and video data. The first audio data and the first video data of the target user can be actively acquired in an interactive manner. The acquisition types of the first audio data and the first video data are interactive. For example, the system regularly or when the user has a relatively high depression risk, initiates voice or text inquiries (such as "How are you feeling recently?", "Are you feeling tired today?"), and performs short-term video and audio recordings through the camera and microphone. The interactive acquisition frequency can be 1 to 2 times a day or once every other day. The interactive acquisition can be specifically triggered by the user actively or under preset conditions.

[0068] It can be understood that by actively acquiring the first audio data and the first video data in an interactive manner, the user is required to cooperate to a certain extent. Therefore, the quality of the acquired first audio data and the second video data is higher. For example, the first audio data has less noise, clearer human voices, and contains more "depression-related information", etc. The person in the first video data is more stable, the user is more likely to face the camera directly, and the facial data is clearer, etc.

[0069] In the application, the second audio data and the second video data of the target user can also be collected in a passive manner. The collection types of the second audio data and the second video data are passive. For example, through the camera and microphone running continuously with the user's permission, the front video data and audio data of the target user are automatically collected. Passive collection can be triggered by the camera when detecting a frontal face image of a person, or when detecting clear and continuous human voices, or recording 30s of audio-visual data every two hours. In the application, data processing such as data cleaning, low-confidence data elimination, and data augmentation also needs to be performed on the second audio data and the second video data collected in a passive manner to improve the data quality.

[0070] It can be understood that although the quality of the second audio data and the second video data collected in a passive manner is lower than that of the first audio data and the first video data collected in an active manner. For example, the second audio data includes more noise, contains less effective "depression-related information", has a lower signal-to-noise ratio, and in the second video data, the person may move, the clarity of the person is affected by distance and light, the face is skewed, etc. However, by performing data processing operations such as data cleaning on the second audio data and the second video data, the data quality can be improved while retaining the advantages of user imperceptibility and silence in passive collection, and more convenient data collection can be achieved. Step S20, based on the multimodal data, perform individual baseline modeling for the target user.

[0071] In the application, after extracting a large amount of multimodal data from a large amount of audio data and video data over multiple days, an individual baseline belonging to the target user for the user state during the "normal" or "healthy" period of the target user is established. It can be understood that compared with all users adopting the same health standard, by establishing an individual baseline, the individual differences between users and the differences in different periods of the same user can be used to accurately grasp and evaluate the state of the target user.

[0072] Step S30, based on the multimodal data and the individual baseline, train a depression prediction network to obtain a target depression detection model.

[0073] In the application, the target depression detection model is used to input the audio data and video data of the target user and output the depression risk prediction value of the target user. Specifically, the target depression detection model can be deployed on computer devices such as smart home devices, mobile phones, smart wearable devices, and computers.

[0074] In an application, after the deployment of the target depression detection model is completed, it is possible to accurately capture changes in the user's state and achieve efficient and accurate depression detection for target users including the elderly. When the depression risk of the target user is detected to be higher than the interference threshold, an interference instruction is sent to the interference device to implement depression interference for the target user and reduce the depression risk of the target user. Specifically, the interference device includes smart home devices such as music players, aroma diffusers, massage chairs, companion robots, and other devices.

[0075] In one embodiment, the multimodal data includes voice data, semantic data, body data, face data, and rPPG data;

[0076] Step S10 of obtaining the multimodal data of the target user based on the audio data and video data of the target user includes:

[0077] Step S101, obtaining the voice data and semantic data of the target user based on the audio data, where the voice data includes high-order voice feature vectors, and the semantic data includes high-order semantic feature vectors.

[0078] In an application, high-order voice feature vectors and high-order semantic feature vectors are extracted from the audio data through a voice feature encoder. The high-order voice feature vectors correspond to information such as speech rate, pitch, fundamental frequency range, voice clarity, and sound intensity, and the semantic features correspond to information such as lexical negativity and semantic emotion.

[0079] Step S102, performing target segmentation on the target user based on the video data to obtain a body video and a face video.

[0080] Step S103, obtaining the body data of the target user based on the body video, where the body data includes high-order body feature vectors.

[0081] In an application, high-order body feature vectors are extracted from the body video through a body feature encoder. The high-order body feature vectors correspond to action information such as the user's movement slowness, exercise amount, action stereotypy, action mutability, and action regularity.

[0082] Step S104, obtaining the face data of the target user based on the face video, where the face data includes high-order face feature vectors.

[0083] In an application, high-order face feature vectors in the face video are obtained through a face feature encoder. The high-order face feature vectors correspond to information such as eye movement data, microexpressions, and macroexpressions.

[0084] Step S105, obtaining the rPPG data of the target user based on the face video and the rPPG algorithm, where the rPPG data includes high-order rPPG feature vectors.

[0085] In the application, face detection is performed based on a face video, and face key points (such as the corners of the eyes, the tip of the nose, the corners of the mouth, etc.) are tracked to perform region of interest (ROI) positioning and motion correction. The average pixel values of the R, G, and B channels are calculated within the ROI to obtain a three-dimensional signal that changes over time. This signal reflects the photoplethysmogram signal of the facial skin area; in order to eliminate the influence of head movement and light changes on the signal, further correction and normalization processing of the signal are required, which can be specifically achieved through some image processing techniques (such as affine transformation, histogram equalization, etc.); then, PCA (principal component analysis), ICA (independent component analysis), or more advanced methods for removing motion artifacts are used to filter out the main pulse components to obtain a preliminary rPPG (remote Photoplethysmography) waveform. High-order rPPG feature vectors are extracted from the rPPG waveform through an rPPG encoder. The high-order rPPG feature vectors include features such as heart rate, heart rate variability, pulse wave conduction characteristics, atrial fibrillation information, premature beat information, blood pressure, and blood sugar.

[0086] In one embodiment, step S30 of training a depression prediction network based on the multimodal data and the individual baseline to obtain a target depression detection model includes:

[0087] Step S31: Perform high-level feature fusion on the high-order speech feature vectors, the high-order semantic feature vectors, the high-order body feature vectors, the high-order face feature vectors, and the high-order rPPG feature vectors with the same time stamp to obtain a first fusion feature vector corresponding to the time stamp.

[0088] In the application, high-level feature fusion includes any one of vector splicing, additive fusion, and mid-term fusion.

[0089] Step S32: Splice a plurality of the first fusion feature vectors to obtain a segment-level fusion vector.

[0090] In the application, the first fusion feature vectors in the segment-level fusion vector are arranged in the order of the corresponding time stamps. In the case where the segment-level fusion vector has a fixed length, null feature vectors are used for padding, and the null feature vectors are skipped in subsequent calculations.

[0091] Step S33: Fuse each first fusion feature vector in the segment-level fusion vector with the corresponding position feature vector, mode feature vector, and time feature vector of the first fusion feature vector to obtain a second fusion feature vector, and splice the second fusion feature vectors to obtain a target segment-level vector. The position feature vector includes the position information of the first fusion feature vector in the segment-level fusion vector, the mode feature vector includes the acquisition method of the first fusion feature vector, and the time feature vector includes the timestamp information of the first fusion feature vector.

[0092] In applications, when the dimensions of the feature vectors to be fused are the same, each first fusion feature vector in the segment-level fusion vector can be fused with the corresponding position feature vector, mode feature vector, and time feature vector of the first fusion feature vector through an addition operation.

[0093] In applications, each first fusion feature vector in the segment-level fusion vector can also be fused with the corresponding position feature vector, mode feature vector, and time feature vector of the first fusion feature vector by first splicing and then performing a fully connected mapping.

[0094] It can be understood that through the end-to-end multi-modal fusion technology of this application, the most discriminative deep features can be automatically extracted from the original multi-modal data (speech data, semantic data, body data, face data, and rPPG data), overcoming the problems of complex feature engineering and modality fragmentation in traditional methods, and improving the generalization ability and performance of the model. At the same time, based on the end-to-end Transformer architecture, high-order speech, semantic, body, face, and rPPG feature vectors are fused. First, for multi-modal features with the same timestamp, accurate multi-modal data fusion and sequence modeling are achieved through position feature vectors (PositionEmbedding), mode feature vectors (Mode Embedding, such as interactive or passive), and time feature vectors (TimeEmbedding). The interaction relationships between various features are automatically learned through the multi-head attention mechanism, thereby improving the accuracy of the model in capturing depressive features.

[0095] Step S34: Train a depression prediction network based on the target segment-level vector and the individual baseline to obtain a target depression detection model.

[0096] In one embodiment, the segment-level fusion vector includes the first fusion feature vectors whose timestamps belong to multiple time periods of multiple days;

[0097] The step S34 of training a depression prediction network based on the target segment-level vector and the individual baseline to obtain a target depression detection model includes:

[0098] Step S3411: Input the target segment-level vector into the depression prediction network to obtain a hidden state sequence, where the hidden state sequence includes multiple hidden state vectors, and the hidden state vectors correspond one-to-one to the second fusion feature vectors in the target segment-level vector.

[0099] Step S3412: Based on the hidden state sequence and the individual baseline, perform depression risk prediction to obtain a predicted risk value.

[0100] Step S3413: Calculate a loss function based on the predicted risk value and the actual risk value, and perform backpropagation to update the parameters of the depression prediction network.

[0101] Step S3414: When the preset training termination condition is reached, output the depression prediction network as the target depression detection model.

[0102] It can be understood that, compared with the traditional model training based on short-term unimodal data, the obtained model has the disadvantages of being unable to reflect the long-term trend changes of individual emotions and being vulnerable to interference by accidental factors. Training the depression detection model based on multi-modal data and individual baselines in multiple time periods of multiple days significantly improves the ability of the trained depression detection model to capture fluctuations in the depression risk state, and can effectively distinguish temporary mood fluctuations from persistent depression tendencies, thereby improving the accuracy and stability of the detection model.

[0103] In one embodiment, step S34, based on the target segment-level vector and the individual baseline, trains a depression prediction network to obtain a target depression detection model, including:

[0104] Step S3421: Input multiple second fusion feature vectors corresponding to the same true value in the target segment-level vector into the depression prediction network to obtain a daily hidden state sequence, where the daily hidden state sequence includes multiple daily hidden state vectors.

[0105] In applications, the true value is the actual depression risk value of the user. The daily true value can be obtained by having the user fill out a standard quantification form such as the HAMD (Hamilton Depression Rating Scale) or PHQ-9 (Patient Health Questionnaire - 9) once a day. All data collected on the same day is associated with the daily true value of that day; the weekly true value is obtained by having the user fill out a standard quantification form once a week and is associated with all target segment-level vectors of that week; for interactively obtained data, a segment-level true value is generated based on the mood score therein and is associated with the corresponding target segment-level vector. In applications, for interactively collected target segment-level vectors, segment-level true values are associated, and for passively collected target segment-level vectors, daily true values or weekly true values are associated, so that each target segment-level vector corresponds to one true value.

[0106] Step S3422, perform pooling on the daily hidden state sequence to obtain a daily-level vector.

[0107] Step S3423, based on the daily-level vector and the individual baseline, perform depression risk prediction to obtain a predicted risk value.

[0108] Step S3424, calculate a loss function based on the predicted risk value and the true value, and perform backpropagation to update the parameters of the depression prediction network.

[0109] Step S3425, when reaching the preset training termination condition, output the depression prediction network as the target depression detection model.

[0110] In one embodiment, the segment-level fusion vector includes the first fusion feature vectors of multiple time periods on the same day of the time stamps;

[0111] The step S34, based on the target segment-level vector and the individual baseline, train a depression prediction network to obtain a target depression detection model, including:

[0112] Step S3431, input the first fusion feature vector with the first time stamp into the depression prediction network to obtain a first hidden state, and output a first daily-level prediction in combination with the individual baseline;

[0113] Step S3432, input the first fusion feature vector corresponding to the next time stamp into the depression prediction network, update the previous hidden state to a new hidden state, and output a new daily-level prediction in combination with the individual baseline;

[0114] Step S3433, repeat the previous step S3432 until all the fusion feature vectors on the same day are traversed, and output the final daily-level prediction.

[0115] In application, the target depression detection model obtained by the above method can perform depression detection in real time, without having to wait until the end of the day or after collecting all the measurement data of the day to perform the depression detection at that time, improving the user's depression detection experience.

[0116] In one embodiment, the step S30, based on the multimodal data and the individual baseline, train a depression prediction network to obtain a target depression detection model, including:

[0117] Step S301, divide the M-day multimodal data of the target user into M - N + 1 training samples through a sliding window of length N, where each training sample includes N-day multimodal data.

[0118] In an application, M and N are positive integers, and M is not less than N. Each of the training samples includes multi-modal data for N consecutive days.

[0119] Step S302: Based on each of the training samples and the individual baseline, predict the predicted risk value for the last day in the training sample.

[0120] Step S303: Calculate a loss function based on the predicted risk value and the corresponding true value for the last day in each training sample, and perform backpropagation to update the parameters of the depression prediction network.

[0121] Step S304: When a preset training termination condition is reached, output the depression prediction network as the target depression detection model.

[0122] In an application, the target depression detection model trained by the above method can output the depression detection result for the last day in N days through N days of data. By using multi-day and multi-segment data, it is possible to capture more subtle changes in depression state information, improving the accuracy of depression detection.

[0123] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the order of execution. The order of execution of each process should be determined based on its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0124] The embodiment of the present application also provides a multi-modal depression detection model training system for executing the steps in the embodiment of the above multi-modal depression detection model training method. The multi-modal depression detection model training system can be a virtual appliance in a computer device, run by the processor of the computer device, or the computer device itself.

[0125] As Figure 2 shown, the multi-modal depression detection model training system 200 provided by the embodiment of the present application includes:

[0126] A multi-modal data extraction module 201 for obtaining the multi-modal data of the target user based on the audio data and video data of the target user, where the time stamp of the multi-modal data corresponds to the audio data and the video data;

[0127] An individual baseline module 202 for performing individual baseline modeling for the target user based on the multi-modal data;

[0128] A training module 203 for training a depression prediction network based on the multi-modal data and the individual baseline to obtain a target depression detection model.

[0129] In one embodiment, the multimodal data includes speech data, semantic data, body data, face data, and rPPG data;

[0130] The multimodal data extraction module 201 is configured to:

[0131] Obtain the speech data and semantic data of the target user based on the audio data, where the speech data includes high-order speech feature vectors and the semantic data includes high-order semantic feature vectors;

[0132] Perform target segmentation on the target user based on the video data to obtain a body video and a face video;

[0133] Obtain the body data of the target user based on the body video, where the body data includes high-order body feature vectors;

[0134] Obtain the face data of the target user based on the face video, where the face data includes high-order face feature vectors;

[0135] Obtain the rPPG data of the target user based on the face video and the rPPG algorithm, where the rPPG data includes high-order rPPG feature vectors.

[0136] In one embodiment, the training module 203 is configured to:

[0137] Perform high-level feature fusion on the high-order speech feature vectors, the high-order semantic feature vectors, the high-order body feature vectors, the high-order face feature vectors, and the high-order rPPG feature vectors with the same timestamp to obtain a first fusion feature vector corresponding to the timestamp;

[0138] Concatenate multiple first fusion feature vectors to obtain a segment-level fusion vector;

[0139] Fuse each first fusion feature vector in the segment-level fusion vector with the position feature vector, the pattern feature vector, and the time feature vector corresponding to the first fusion feature vector to obtain a second fusion feature vector, and concatenate the second fusion feature vectors to obtain a target segment-level vector. The position feature vector includes the position information of the first fusion feature vector in the segment-level fusion vector, the pattern feature vector includes the acquisition method of the first fusion feature vector, and the time feature vector includes the timestamp information of the first fusion feature vector;

[0140] Train a depression prediction network based on the target segment-level vector and the individual baseline to obtain a target depression detection model.

[0141] In one embodiment, the segment-level fusion vector includes the first fusion feature vectors whose timestamps belong to multiple time periods of multiple days, and the training module 203 is configured to:

[0142] Input the target segment-level vector into the depression prediction network to obtain a hidden state sequence, where the hidden state sequence includes multiple hidden state vectors, and the hidden state vectors correspond one-to-one with the second fusion feature vectors in the target segment-level vector;

[0143] Perform depression risk prediction based on the hidden state sequence and the individual baseline to obtain a predicted risk value;

[0144] Calculate a loss function based on the predicted risk value and the actual risk value, and perform backpropagation to update the parameters of the depression prediction network;

[0145] When the preset training termination condition is reached, output the depression prediction network as the target depression detection model.

[0146] In one embodiment, the training module 203 is configured to:

[0147] Input multiple second fusion feature vectors corresponding to the same true value in the target segment-level vector into the depression prediction network to obtain a daily hidden state sequence, where the daily hidden state sequence includes multiple daily hidden state vectors;

[0148] Perform pooling on the daily hidden state sequence to obtain a daily vector;

[0149] Perform depression risk prediction based on the daily vector and the individual baseline to obtain a predicted risk value;

[0150] Calculate a loss function based on the predicted risk value and the true value, and perform backpropagation to update the parameters of the depression prediction network;

[0151] When the preset training termination condition is reached, output the depression prediction network as the target depression detection model.

[0152] In application, when multiple second fusion feature vectors in the target segment-level vector correspond to one true value, through the above method, multiple second fusion feature vectors on the same day only generate one prediction, realizing direct supervision for daily-level labels, and only performing one loss backpropagation during training, enabling the target depression detection model to automatically learn how to find the information most relevant to daily-level depression detection among multiple target segment-level vectors in the segment-level fusion vector.

[0153] In one embodiment, the segment-level fusion vector includes the first fusion feature vectors whose timestamps belong to multiple time periods of the same day, and the training module 203 is configured to:

[0154] Training a depression prediction network based on the target segment-level vector and the individual baseline to obtain a target depression detection model, including:

[0155] Input the first fusion feature vector with the earliest timestamp into the depression prediction network to obtain a first hidden state, and combine it with the individual baseline to output a first daily-level prediction;

[0156] Input the first fusion feature vector corresponding to the next timestamp into the depression prediction network, update the previous hidden state to a new hidden state, and combine it with the individual baseline to output a new daily-level prediction;

[0157] Repeat the above step until all the fusion feature vectors of the same day are traversed, and output the final daily-level prediction.

[0158] In one embodiment, the training module 203 is used for:

[0159] Divide the multi-modal data of M days of the target user into M - N + 1 training samples through a sliding window of length N, where each training sample includes N days of multi-modal data;

[0160] Predict the predicted risk value of the last day in each training sample based on each training sample and the individual baseline respectively;

[0161] Calculate a loss function based on the predicted risk value of the last day in each training sample and the corresponding true value, and perform backpropagation to update the parameters of the depression prediction network;

[0162] When the preset training termination condition is reached, output the depression prediction network as the target depression detection model.

[0163] In an application, each module in the multi-modal depression detection model training system can be a software program module, can also be implemented by different logic circuits integrated in a processor, or can be implemented by multiple distributed processors.

[0164] Figure 3 This is a schematic structural diagram of a computer device provided in an embodiment of the present application. As Figure 3 shown, the computer device 3 in this embodiment includes: at least one processor 30( Figure 3 only one is shown in the figure), a memory 31, and a computer program 32 stored in the memory 31 and executable on the at least one processor 30. When the processor 30 executes the computer program 32, it implements the steps in any of the above-mentioned embodiments of the multi-modal depression detection model training method.

[0165] The computer device may include, but is not limited to, a processor 30 and a memory 31. Those skilled in the art can understand that Figure 3 merely an example of the computer device 3, which does not constitute a limitation on the computer device 3, and may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0166] The processor 30 may be a central processing unit (CPU), and the processor 30 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0167] In some embodiments, the memory 31 may be an internal storage unit of the computer device 3, such as the hard disk or memory of the computer device 3. In other embodiments, the memory 31 may also be an external storage device of the computer device 3, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 3. Further, the memory 31 may also include both the internal storage unit and the external storage device of the computer device 3. The memory 31 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program. The memory 31 may also be used to temporarily store data that has been output or will be output.

[0168] It should be noted that for the content such as information interaction and execution process between the above-mentioned device / units, since it is based on the same concept as the method embodiment of the present application, for its specific functions and the technical effects brought, reference may be specifically made to the method embodiment part, and details are not described herein again.

[0169] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0170] An embodiment of this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the foregoing method embodiments can be implemented.

[0171] An embodiment of this application provides a computer program product, including a computer program. When the computer program is run, the steps in the foregoing method embodiments for training various multi-modal depression detection models are executed.

[0172] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of this application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps in the foregoing method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the device / computer equipment, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0173] In the above embodiments, the descriptions of the various embodiments each have their own emphasis. For parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0174] Those of ordinary skill in the art will realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0175] In the embodiments provided in this application, it should be understood that the disclosed computer devices and methods can be implemented in other ways. For example, the computer device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, indirect couplings or communication connections of devices or units, and can be in electrical, mechanical or other forms.

[0176] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0177] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included in the protection scope of this application.

Claims

1. A training method for a multi-modal depression detection model, characterized in that Including: Obtaining multimodal data of the target user based on the audio data and video data of the target user, where the timestamps of the multimodal data correspond to the audio data and the video data; Performing individual baseline modeling for the target user based on the multimodal data; Training a depression prediction network based on the multimodal data and the individual baseline to obtain a target depression detection model.

2. The multi-modal depression detection model training method according to claim 1, wherein The multimodal data includes speech data, semantic data, body data, face data, and rPPG data; The obtaining of the multimodal data of the target user based on the audio data and video data of the target user includes: Obtaining the speech data and semantic data of the target user based on the audio data, where the speech data includes high-order speech feature vectors and the semantic data includes high-order semantic feature vectors; Performing target segmentation on the target user based on the video data to obtain a body video and a face video; Obtaining the body data of the target user based on the body video, where the body data includes high-order body feature vectors; Obtaining the face data of the target user based on the face video, where the face data includes high-order face feature vectors; Obtaining the rPPG data of the target user based on the face video and the rPPG algorithm, where the rPPG data includes high-order rPPG feature vectors.

3. The multi-modal depression detection model training method according to claim 2, wherein The training of the depression prediction network based on the multimodal data and the individual baseline to obtain a target depression detection model includes: Performing high-level feature fusion on the high-order speech feature vectors, high-order semantic feature vectors, high-order body feature vectors, high-order face feature vectors, and high-order rPPG feature vectors with the same timestamp to obtain a first fusion feature vector corresponding to the timestamp; Concatenating multiple of the first fusion feature vectors to obtain a segment-level fusion vector; Fusing each first fusion feature vector in the segment-level fusion vector with the corresponding position feature vector, pattern feature vector, and time feature vector of the first fusion feature vector to obtain a second fusion feature vector, and concatenating the second fusion feature vectors to obtain a target segment-level vector, where the position feature vector includes the position information of the first fusion feature vector in the segment-level fusion vector, the pattern feature vector includes the acquisition method of the first fusion feature vector, and the time feature vector includes the timestamp information of the first fusion feature vector; Training a depression prediction network based on the target segment-level vector and the individual baseline to obtain a target depression detection model.

4. The multimodal depression detection model training method according to claim 3, wherein, The segment-level fusion vector includes the first fusion feature vectors whose timestamps belong to multiple time periods of multiple days; The training of the depression prediction network based on the target segment-level vector and the individual baseline to obtain a target depression detection model includes: Inputting the target segment-level vector into the depression prediction network to obtain a hidden state sequence, where the hidden state sequence includes multiple hidden state vectors, and the hidden state vectors correspond one-to-one to the second fusion feature vectors in the target segment-level vector; Performing depression risk prediction based on the hidden state sequence and the individual baseline to obtain a predicted risk value; Calculate a loss function based on the predicted risk value and the actual risk value, and perform backpropagation to update the parameters of the depression prediction network; When a preset training termination condition is reached, output the depression prediction network as the target depression detection model.

5. The multi-modal depression detection model training method according to claim 3, wherein, The training of the depression prediction network based on the target segment-level vector and the individual baseline to obtain the target depression detection model includes: Input multiple second fusion feature vectors corresponding to the same true value in the target segment-level vector into the depression prediction network to obtain a daily hidden state sequence, where the daily hidden state sequence includes multiple daily hidden state vectors; Perform pooling on the daily hidden state sequence to obtain a daily vector; Perform depression risk prediction based on the daily vector and the individual baseline to obtain a predicted risk value; Calculate a loss function based on the predicted risk value and the true value, and perform backpropagation to update the parameters of the depression prediction network; When a preset training termination condition is reached, output the depression prediction network as the target depression detection model.

6. The method for training a multi-modal depression detection model according to claim 3, wherein, The segment-level fusion vector includes the first fusion feature vectors of multiple time periods whose timestamps belong to the same day; The training of the depression prediction network based on the target segment-level vector and the individual baseline to obtain the target depression detection model includes: Input the first fusion feature vector with the first timestamp into the depression prediction network to obtain a first hidden state, and combine it with the individual baseline to output a first daily prediction; Input the first fusion feature vector corresponding to the next timestamp into the depression prediction network, update the previous hidden state to a new hidden state, and combine it with the individual baseline to output a new daily prediction; Repeat the step of inputting the first fusion feature vector corresponding to the next timestamp into the depression prediction network, updating the previous hidden state to a new hidden state, and combining it with the individual baseline to output a new daily prediction until all the fusion feature vectors of the same day are traversed, and output the final daily prediction.

7. The training method of the multimodal depression detection model according to any one of claims 1 to 6, characterized in that, The training of the depression prediction network based on the multimodal data and the individual baseline to obtain the target depression detection model includes: Divide the M-day multimodal data of the target user into M - N + 1 training samples through a sliding window of length N, where each training sample includes N-day multimodal data; Predict the predicted risk value of the last day in each training sample based on each training sample and the individual baseline; Calculate a loss function based on the predicted risk value of the last day in each training sample and the corresponding true value, and perform backpropagation to update the parameters of the depression prediction network; When a preset training termination condition is reached, output the depression prediction network as the target depression detection model.

8. A multi-modal depression detection model training system, characterized in that, It includes: A multimodal data extraction module for obtaining the multimodal data of the target user based on the audio data and video data of the target user, where the timestamp of the multimodal data corresponds to the audio data and the video data; An individual baseline module for performing individual baseline modeling for the target user based on the multimodal data; A training module, configured to train a depression prediction network based on the multimodal data and the individual baseline to obtain a target depression detection model.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer program product, characterized in that, It includes a computer program, which when run, causes the method according to any one of claims 1 to 7 to be executed.