A multimodal mental health prediction method and system

By combining deep learning networks and multimodal fusion technology, using multimodal data to predict mental health, the problem of low accuracy of mental health prediction in the existing technology is solved, and more efficient and accurate mental health recognition is achieved.

CN118136256BActive Publication Date: 2025-05-16ZHAOQING MEDICAL COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410359437.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-27
Publication Date
2025-05-16
Estimated Expiration
2044-03-27

AI Technical Summary

Technical Problem

The accuracy of the prior art is not high enough when using multimodal data fusion technology to predict mental health.

Method used

A multimodal mental health prediction method is proposed, combining deep learning networks and multimodal fusion technology, and features are extracted and fusion by collecting speech data, text data, expression image data and physiological data, building a mental health prediction model, and using composite emotion analysis units to obtain time factors to improve recognition accuracy.

Benefits of technology

It improves the efficiency and accuracy of mental health testing, overcomes the problem that existing methods have too coarse classification of mental health types, dynamically adjusts feature weights to accurately assist detection, ensures data reliability and safety, and thus improves the accuracy of mental health prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118136256B_ABST
    Figure CN118136256B_ABST
Patent Text Reader

Abstract

The present invention proposes a multimodal mental health prediction method, which includes: collecting multimodal data of the person to be tested, and preprocessing the multimodal data to obtain a training set; wherein the multimodal data includes voice data, text data, expression image data and physiological data; extracting features from the multimodal data, fusing the features of the multimodal data to form fusion features; constructing a mental health prediction model based on the fusion features, and using the fusion features to train the mental health model; using the trained mental health prediction model to predict the user's mental health, and obtaining corresponding mental health analysis results. The present invention combines the use of deep learning networks and multimodal fusion technology to improve the efficiency and accuracy of mental health testing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence, and in particular relates to a multimodal mental health prediction method and system. Background Art

[0002] The impact of mental health on individuals and society is profound and extensive. In the case of mental health, it will directly affect the quality of life of individuals. When individuals are in a state of mental health, they are more likely to feel satisfied, happy and meaningful. They are better able to cope with stress, challenges and changes, thereby achieving a better balance and satisfaction in life. On the contrary, psychological problems such as anxiety and depression may cause individuals to feel pain, helplessness and despair, seriously affecting their quality of life and happiness. People with mental health are more likely to establish good and stable relationships with others, communicate and interact better, and thus gain better support and a sense of belonging in social environments. On the contrary, psychological problems may cause individuals to feel difficult, lonely and excluded in social situations, further exacerbating their psychological problems. When individuals are generally in a state of mental health, society is more likely to remain harmonious, stable and prosperous. On the contrary, psychological problems may lead to an increase in social conflicts, crimes and violence, and have a negative impact on social stability and development. Maintaining and promoting mental health is essential to the health, happiness and development of individuals and society. We should attach importance to the importance of mental health and take positive measures to maintain and promote personal mental health. At the same time, we can also promote social stability and development by improving the level of social mental health.

[0003] At present, mental health detection is mostly done through psychological tests or related instruments to test changes in physiological signs. Furthermore, the prediction of users' mental health using multimodal data fusion technology is not accurate enough. Summary of the invention

[0004] The purpose of this invention is to propose a multimodal mental health prediction method and system, which uses deep learning networks and multimodal fusion technology, adds a composite emotion analysis unit to further analyze emotions and obtain time factors, and further improves the accuracy of mental health recognition.

[0005] In order to achieve the above object, a first aspect of the present invention provides a multimodal mental health prediction method, the method comprising:

[0006] S1. Collect multimodal data of the person to be tested, and preprocess the multimodal data to obtain a training set; wherein the multimodal data includes voice data, text data, expression image data and physiological data;

[0007] S2, extracting features from the multimodal data, and fusing the features of the multimodal data to form fusion features;

[0008] S3, construct a mental health prediction model based on the fusion features, and use the fusion features to train the mental health model;

[0009] S4, using the trained mental health prediction model to predict the user's mental health and obtain corresponding mental health analysis results;

[0010] The specific steps of S2 are:

[0011] S21, inputting text data and expression image data into an encoder; wherein, using a BERT encoder to extract features from the text data to obtain text unimodal features It is expressed as follows:

[0012]

[0013] Among them, Con i Represents the input text, i represents the i-th obtained text, that is, Con i =Con1, Con2, Con3...Con i , W1 represents the weight parameter in the text data extraction process;

[0014] The expression image data image i Input DNN convolution encoder to obtain image unimodal features It is expressed as follows:

[0015]

[0016] That is, i∈[1,2,3,4....];

[0017] S22, inputting the voice data and physiological data into the long short-term memory neural network to obtain composite emotion change characteristics;

[0018] S23. Based on the three modal features of text unimodal features, image unimodal features and compound emotion change features, the three modal features are fused in pairs to obtain corresponding bimodal fusion features.

[0019] Furthermore, the method further includes constructing a screening model based on a deep learning network in the preprocessing step, setting the mental health prediction condition as a hard constraint, and the method includes:

[0020] Define a feasibility function f(s,a) to indicate whether taking action a under state s satisfies the hard constraint. Map the mental health prediction condition of the current user into the feasibility function and input the mental health prediction condition as the same action. When the hard constraint is met, f(s,a)=1 and output the mental health prediction condition of the current user; otherwise, f(s,a)=0 and return the data.

[0021] According to the above scheme, further, the mental health prediction condition is a condition in which voice data, text data, expression image data and physiological data coexist, and the current condition is satisfied within a specified data range.

[0022] According to the above scheme, further, the physiological data includes brain wave signals, psychological test scores, number of blinks, heart rate and skin electrical signals;

[0023] Among them, before preprocessing, the physiological data is first subjected to line filtering, noise reduction, and baseline drift removal;

[0024] According to the above scheme, further, the S22 specifically includes the following steps:

[0025] Each voice data is converted into text, and then each voice data is processed by frame processing to obtain the voice features of each frame data, and at the same time, the physiological features of the physiological data in each time scale are extracted by using a time sliding window method; wherein, the embedding layer encodes each voice data into an emotion expression vector and inputs it into the feature extraction layer, and the emotion expression vector uses the emotion expression in the voice data preprocessed according to the predicted emotion expression dictionary statistics, and encodes it into an emotion expression vector, and then the emotion expression vector is input into the feature extraction layer, and the emotion cause feature unit is used to extract the event cause through a long short-term memory neural network;

[0026] The composite emotion knowledge unit is combined with the long short-term memory neural network to extract the composite emotion change features of voice features and physiological features, and the composite emotion change feature mood is obtained. i .

[0027] According to the above scheme, further, the S23 specifically includes the following steps:

[0028] The three modal features of text unimodal features, image unimodal features and composite emotion change features are fused in pairs, as shown below:

[0029]

[0030]

[0031]

[0032] in, Represents the fusion features of text data and expression image data, Represents the fusion features of text data and compound emotion change features, It represents the fusion feature of expression image data and compound emotion change features; t represents the time step, which represents a certain time period; Indicates the characteristics of complex emotional changes;

[0033] The three fusion features are encoded through the long short-term memory neural network to obtain the temporal characteristics of the dual modality, which is expressed as follows:

[0034]

[0035] Where M2 represents the sequence of dual-modal fusion, M2 = {(Con i , image i ), (Con i , motion i ),(image i , motion i )}, k represents the fusion of the two modal vectors at the current time step t, represents the output of the long-short time neural network, Represents the memory unit of the long short-term memory network.

[0036] According to the above solution, further, S3 specifically includes:

[0037] The bimodal fusion features and the three modal features of text unimodal features, image unimodal features and compound emotion change features are input into the corresponding three-layer deep neural network. The bimodal features are used as the training set, and the three modal features of text unimodal features, image unimodal features and compound emotion change features are used as the test set for training to obtain the mental health prediction model.

[0038] According to the above solution, further, S5 is specifically:

[0039] The dual-modal fusion features are fused to form composite fusion features;

[0040] The composite fusion features are input into the fully connected layer and then linearly transformed to obtain the output value of the mental health category;

[0041] According to the output value of the mental health category, the probability value of the corresponding mental health category is calculated by the Softmax function, and the category with the largest probability value is taken as the result of mental health identification, which is expressed as follows:

[0042]

[0043] Among them, P represents the probability value of the corresponding mental health category, V t Represents the element of the mental health category at the current time step t output by the connection layer, n represents the total mental health category, j represents the index variable used to traverse the output values ​​of all mental health categories, V j An element representing the mental health category corresponding to the index variable.

[0044] According to the above scheme, further, the mental health prediction model uses cross entropy as the loss function L, and the Adam optimizer updates the model parameters, specifically including:

[0045]

[0046] in, Denotes the predicted label, y t represents the training label, γ represents the hyperparameter for adjusting the composite emotion change feature loss function, Represents the loss function of the composite emotion change feature at the current time step t;

[0047] The back-propagation algorithm is used to derive the loss function and obtain the gradient of the loss function. The gradient of the loss function is repeatedly used and the parameters of the deep learning network are updated along the gradient direction to optimize the parameters of the deep learning network.

[0048] In a second aspect of the present invention, a multimodal mental health prediction system is provided, the system comprising:

[0049] A data collection module, used to collect multimodal data of the person to be tested, and pre-process the multimodal data to obtain a training set; wherein the multimodal data includes voice data, text data, expression image data and physiological data;

[0050] The feature extraction module is used to extract features from multimodal data and fuse the features of multimodal data to form fused features;

[0051] A model building module, used to build a mental health prediction model based on the fusion features and train the mental health model using the fusion features;

[0052] The mental health classification module is used to predict the user's mental health using the trained mental health prediction model and obtain the corresponding mental health analysis results;

[0053] The specific steps of the feature extraction module 502 are as follows:

[0054] S21, inputting text data and expression image data into an encoder; wherein, using a BERT encoder to extract features from the text data to obtain text unimodal features It is expressed as follows:

[0055]

[0056] Among them, Con i Represents the input text, i represents the i-th obtained text, that is, Con i =Con1, Con2, Con3...Con i ; The expression image data image i Input VGG16 convolution encoder to obtain image unimodal features It is expressed as follows:

[0057]

[0058] That is, i∈[1,2,3,4…];

[0059] S22, inputting the voice data and physiological data into the long short-term memory neural network to obtain composite emotion change characteristics;

[0060] S23. Based on the three modal features of text unimodal features, image unimodal features and compound emotion change features, the three modal features are fused in pairs to obtain corresponding bimodal fusion features.

[0061] The beneficial technical effects of the present invention are at least as follows:

[0062] (1) The present invention improves the efficiency and accuracy of mental health testing by combining deep learning networks and multimodal fusion technology. Deep learning continuously optimizes the mental health model here to further improve the accuracy of recognition. This combination utilizes the powerful data processing and pattern recognition capabilities of deep learning, as well as the high efficiency of reinforcement learning in decision-making and strategy optimization.

[0063] (2) At the same time, the technology of composite emotion classification unit is used in the middle part to overcome the problem of coarse granularity of mental health type classification in existing mental health detection methods.

[0064] (3) Weights are also added to each encoder, which can dynamically adjust the weights of various features to accurately assist the type of mental health detection.

[0065] (4) At the same time, a feasibility function is designed in the model to further screen the data preprocessing, ensuring the reliability and security of the data, thereby improving the accuracy of mental health prediction.

[0066] (5) Through multi-task feature extraction, the problem of neural network gradient messages is avoided and the effect of priority recognition of subtle psychological changes is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] The present invention is further described using the accompanying drawings, but the embodiments in the accompanying drawings do not constitute any limitation to the present invention. A person skilled in the art can obtain other drawings based on the following drawings without creative work.

[0068] Figure 1 This is a flow chart of a multimodal mental health prediction method of the present invention.

[0069] Figure 2 This is a framework diagram of a multimodal mental health prediction system of the present invention. DETAILED DESCRIPTION

[0070] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.

[0071] like Figure 1 As shown, an embodiment of the present invention provides a multimodal mental health prediction method, the method comprising:

[0072] S1. Collect multimodal data of the person to be tested, and preprocess the multimodal data to obtain a training set; wherein the multimodal data includes voice data, text data, expression image data and physiological data;

[0073] Among them, the text data uses GloVe vector as the feature input word vector, and the dimension is 300; the speech data is captured at a frequency of 30 frames per second, and the dimension is 74.

[0074] S2, extracting features from the multimodal data, and fusing the features of the multimodal data to form fusion features;

[0075] S3, construct a mental health prediction model based on the fusion features, and use the fusion features to train the mental health model;

[0076] S4. Use the trained mental health prediction model to predict the user's mental health and obtain corresponding mental health analysis results.

[0077] Among them, the results of mental health analysis include depression, anger, happiness, happiness and other moods.

[0078] The specific steps of S2 are:

[0079] S21, inputting text data and expression image data into an encoder; wherein, using a BERT encoder to extract features from the text data to obtain text unimodal features It is expressed as follows:

[0080]

[0081] Among them, Con i Represents the input text, i represents the i-th obtained text, that is, Con i =Con1, Con2, Con3...Con i , W1 represents the weight parameter in the text data extraction process;

[0082] The expression image data image i Input DNN convolution encoder to obtain image unimodal features It is expressed as follows:

[0083]

[0084] That is, i∈[1,2,3,4…];

[0085] S22, inputting the voice data and physiological data into the long short-term memory neural network to obtain composite emotion change characteristics;

[0086] S23. Based on the three modal features of text unimodal features, image unimodal features and compound emotion change features, the three modal features are fused in pairs to obtain corresponding bimodal fusion features.

[0087] Furthermore, the method further includes constructing a screening model based on a deep learning network in the preprocessing step, setting the mental health prediction condition as a hard constraint, and the method includes:

[0088] Define a feasibility function f(s,a) to indicate whether taking action a under state s satisfies the hard constraint. Map the mental health prediction condition of the current user into the feasibility function and input the mental health prediction condition as the same action. When the hard constraint is met, f(s,a)=1 and output the mental health prediction condition of the current user; otherwise, f(s,a)=0 and return the data.

[0089] Furthermore, the mental health prediction condition is a condition in which voice data, text data, expression image data and physiological data coexist, and the current condition is satisfied within a specified data range.

[0090] Furthermore, the physiological data includes brain wave signals, psychological test scores, blink times, heart rate and skin electrical signals;

[0091] Among them, before preprocessing, the physiological data is first subjected to line filtering, noise reduction, and baseline drift removal;

[0092] Furthermore, the S22 specifically includes the following steps:

[0093] Each voice data is converted into text, and then each voice data is processed by frame processing to obtain the voice features of each frame data, and at the same time, the physiological features of the physiological data in each time scale are extracted by using a time sliding window method; wherein, the embedding layer encodes each voice data into an emotion expression vector and inputs it into the feature extraction layer, and the emotion expression vector uses the emotion expression in the voice data preprocessed according to the predicted emotion expression dictionary statistics, and encodes it into an emotion expression vector, and then the emotion expression vector is input into the feature extraction layer, and the emotion cause feature unit is used to extract the event cause through a long short-term memory neural network;

[0094] The composite emotion knowledge unit is combined with the long short-term memory neural network to extract the composite emotion change features of voice features and physiological features, and the composite emotion change feature motion is obtained. i .

[0095] Furthermore, the S23 specifically includes the following steps:

[0096] The three modal features of text unimodal features, image unimodal features and composite emotion change features are fused in pairs, as shown below:

[0097]

[0098]

[0099]

[0100] in, Represents the fusion features of text data and expression image data, Represents the fusion features of text data and compound emotion change features, It represents the fusion feature of expression image data and compound emotion change features; t represents the time step, which represents a certain time period; Indicates the characteristics of complex emotional changes;

[0101] The three fusion features are encoded through the long short-term memory neural network to obtain the temporal characteristics of the dual modality, which is expressed as follows:

[0102]

[0103] Where M2 represents the sequence of dual-modal fusion, M2 = {(Con i , image i ), (Con i , motion i ),(image i , motion i)}, k represents the fusion of the two modal vectors at the current time step t, represents the output of the long-short time neural network, Represents the memory unit of the long short-term memory network.

[0104] Furthermore, the S3 specifically includes:

[0105] The bimodal fusion features and the three modal features of text unimodal features, image unimodal features and compound emotion change features are input into the corresponding three-layer deep neural network. The bimodal features are used as the training set, and the three modal features of text unimodal features, image unimodal features and compound emotion change features are used as the test set for training to obtain the mental health prediction model.

[0106] Furthermore, the S5 is specifically:

[0107] The dual-modal fusion features are fused to form composite fusion features;

[0108] The composite fusion features are input into the fully connected layer and then linearly transformed to obtain the output value of the mental health category;

[0109] According to the output value of the mental health category, the probability value of the corresponding mental health category is calculated by the Softmax function, and the category with the largest probability value is taken as the result of mental identification, which is expressed as follows:

[0110]

[0111] Among them, P represents the probability value of the corresponding mental health category, V t Represents the element of the mental health category at the current time step t output by the connection layer, n represents the total mental health category, j represents the index variable used to traverse the output values ​​of all mental health categories, V j An element representing the mental health category corresponding to the index variable.

[0112] Furthermore, the mental health prediction model uses cross entropy as the loss function L, and the Adam optimizer updates the model parameters, specifically including:

[0113]

[0114] in, Denotes the predicted label, y t represents the training label, γ represents the hyperparameter for adjusting the composite emotion change feature loss function, Represents the loss function of the composite emotion change feature at the current time step t;

[0115] The back-propagation algorithm is used to derive the loss function and obtain the gradient of the loss function. The gradient of the loss function is repeatedly used and the parameters of the deep learning network are updated along the gradient direction to optimize the parameters of the deep learning network.

[0116] like Figure 2 As shown, an embodiment of the present invention further provides a multimodal mental health prediction system, the system comprising:

[0117] The data collection module 501 is used to collect multimodal data of the person to be tested, and pre-process the multimodal data to obtain a training set; wherein the multimodal data includes voice data, text data, expression image data and physiological data;

[0118] The feature extraction module 502 is used to extract features from the multimodal data and fuse the features of the multimodal data to form fused features;

[0119] A model building module 503 is used to build a mental health prediction model based on the fusion features and train the mental health model using the fusion features;

[0120] The mental health classification module 504 is used to predict the user's mental health using the trained mental health prediction model to obtain the corresponding mental health analysis results;

[0121] The specific steps of the feature extraction module 502 are as follows:

[0122] S21, inputting text data and expression image data into an encoder; wherein, using a BERT encoder to extract features from the text data to obtain text unimodal features It is expressed as follows:

[0123]

[0124] Among them, Con i Represents the input text, i represents the i-th obtained text, that is, Con i =Con1, Con2, Con3...Con i ; The expression image data image i Input VGG16 convolution encoder to obtain image unimodal features It is expressed as follows:

[0125]

[0126] That is, i∈[1,2,3,4…];

[0127] S22, inputting the voice data and physiological data into the long short-term memory neural network to obtain composite emotion change characteristics;

[0128] S23. Based on the three modal features of text unimodal features, image unimodal features and compound emotion change features, the three modal features are fused in pairs to obtain corresponding bimodal fusion features.

[0129] In summary, the present invention uses multiple convolutional networks for extraction, and each feature is extracted efficiently and accurately. Multimodal data (such as text, voice, image, etc.) provides rich information about the mental health status of an individual. Each modality has its unique characteristics and advantages. By fusing this information through a multi-convolutional network, the complementarity of various modalities can be fully utilized, so as to more comprehensively understand the mental health status of an individual. At the same time, the information loss or noise interference that may be caused by single-modal data is reduced. By combining information from multiple modalities, the accuracy and reliability of mental health identification can be improved. Different mental health problems may show different characteristics in data of different modalities. By using a multi-convolutional network for multimodal identification, suitable modalities can be flexibly selected for combination and analysis according to specific problems and needs, so as to better adapt to different application scenarios. Convolutional neural network is an important tool in the field of deep learning, with powerful feature extraction and classification capabilities. By constructing a multi-convolutional network structure, useful features can be more effectively extracted from multimodal data, and high-precision mental health identification can be achieved. Using a multi-convolutional network for multimodal mental health identification not only helps to improve the recognition accuracy, but also helps to promote the development of the field of mental health research. The present invention provides a new means for subsequent research to explore and understand mental health issues.

[0130] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0131] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a division of logical functions. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be electrical, mechanical or other forms.

[0132] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0133] Although embodiments of the present invention have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.

Claims

1. A multimodal mental health prediction method, characterized in that: The method comprises: S1. Collect multimodal data of the person to be tested, and preprocess the multimodal data to obtain a training set; wherein the multimodal data includes voice data, text data, expression image data and physiological data; S2, extracting features from the multimodal data, and fusing the features of the multimodal data to form fusion features; S3, construct a mental health prediction model based on the fusion features, and use the fusion features to train the mental health model; S4, using the trained mental health prediction model to predict the user's mental health and obtain corresponding mental health analysis results; The specific steps of S2 are: S21, inputting text data and expression image data into an encoder; wherein, using a BERT encoder to extract features from the text data to obtain text unimodal features It is expressed as follows: Among them, Con i Represents the input text, i represents the i-th obtained text, that is, Con i =Con1,Con2,Con3...Con i , W1 represents the weight parameter in the text data extraction process; The expression image data image i Input DNN convolution encoder to obtain image unimodal features It is expressed as follows: That is, i∈[1, 2, 3, 4....]; S22, inputting the voice data and physiological data into the long short-term memory neural network to obtain composite emotion change characteristics; S23, based on the three modal features of text unimodal features, image unimodal features and composite emotion change features, two-by-two modal fusions are performed to obtain corresponding bimodal fusion features; Wherein, the specific steps of S22 include: Each voice data is converted into text, and then each voice data is processed by frame processing to obtain the voice features of each frame data, and at the same time, the physiological features of the physiological data in each time scale are extracted by using a time sliding window method; wherein, the embedding layer encodes each voice data into an emotion expression vector and inputs it into the feature extraction layer, and the emotion expression vector uses the emotion expression in the voice data preprocessed according to the predicted emotion expression dictionary statistics, and encodes it into an emotion expression vector, and then the emotion expression vector is input into the feature extraction layer, and the emotion cause feature unit is used to extract the event cause through a long short-term memory neural network; The composite emotion knowledge unit is combined with the long short-term memory neural network to extract the composite emotion change features of voice features and physiological features, and the composite emotion change feature motion is obtained. i ; The specific steps of S23 include: The three modal features of text unimodal features, image unimodal features and composite emotion change features are fused in pairs, as shown below: in, Represents the fusion features of text data and expression image data, Represents the fusion features of text data and compound emotion change features, It represents the fusion feature of expression image data and compound emotion change features; t represents the time step, which represents a certain time period; Indicates the characteristics of complex emotional changes; The three fusion features are encoded through the long short-term memory neural network to obtain the temporal characteristics of the dual modality, which is expressed as follows: Where M2 represents the sequence of dual-modal fusion, M2 = {(Con i , image i ), (Con i , motion i ),(image i , motion i )}, k represents the fusion of the two modal vectors at the current time step t, represents the output of the long-short time neural network, Represents the memory unit of the long short-term memory network; The S3 specifically includes: The bimodal fusion features and the three modal features of text unimodal features, image unimodal features and compound emotion change features are input into the corresponding three-layer deep neural network. The bimodal features are used as the training set, and the three modal features of text unimodal features, image unimodal features and compound emotion change features are used as the test set for training to obtain a mental health prediction model. The S4 is specifically: Fusing the dual-modal fusion features to form composite fusion features; The composite fusion features are input into the fully connected layer and then linearly transformed to obtain the output value of the mental health category; According to the output value of the mental health category, the probability value of the corresponding mental health category is calculated by the Softmax function, and the category with the largest probability value is taken as the result of mental health identification, which is expressed as follows: Among them, P represents the probability value of the corresponding mental health category, V t Represents the element of the mental health category at the current time step t output by the connection layer, n represents the total mental health category, j represents the index variable used to traverse the output values ​​of all mental health categories, V j An element representing the mental health category corresponding to the index variable.

2. A multimodal mental health prediction method according to claim 1, characterized in that: The method further includes constructing a screening model based on a deep learning network in the preprocessing step, setting the mental health prediction condition as a hard constraint, and the method includes: Define a feasibility function f(s,a) to indicate whether taking action a under state s satisfies the hard constraint. Map the mental health prediction condition of the current user into the feasibility function and input the mental health prediction condition as the same action. When the hard constraint is met, f(s,a)=1 and output the mental health prediction condition of the current user; otherwise, f(s,a)=0 and return the data.

3. A multimodal mental health prediction method according to claim 2, characterized in that: The mental health prediction condition is a condition in which voice data, text data, expression image data and physiological data coexist, and the current condition is satisfied within a specified data range.

4. A multimodal mental health prediction method according to claim 1, characterized in that: The physiological data include brain wave signals, psychological test scores, blink times, heart rate and skin electrical signals; Among them, before preprocessing, the physiological data is first subjected to line filtering, noise reduction and baseline drift removal.

5. A multimodal mental health prediction method according to claim 1, characterized in that: The mental health prediction model uses cross entropy as the loss function L, and the Adam optimizer updates the model parameters, specifically including: in, Denotes the predicted label, y t represents the training label, γ represents the hyperparameter for adjusting the composite emotion change feature loss function, Represents the loss function of the composite emotion change feature at the current time step t; The back-propagation algorithm is used to derive the loss function and obtain the gradient of the loss function. The gradient of the loss function is repeatedly used and the parameters of the deep learning network are updated along the gradient direction to optimize the parameters of the deep learning network.

6. A multimodal mental health prediction system, characterized in that: The system comprises: A data collection module, used to collect multimodal data of the person to be tested, and pre-process the multimodal data to obtain a training set; wherein the multimodal data includes voice data, text data, expression image data and physiological data; The feature extraction module is used to extract features from multimodal data and fuse the features of multimodal data to form fused features; A model building module, used to build a mental health prediction model based on the fusion features and train the mental health model using the fusion features; The mental health classification module is used to predict the user's mental health using the trained mental health prediction model and obtain the corresponding mental health analysis results; Among them, the specific steps of the feature extraction module are: S21, inputting text data and expression image data into an encoder; wherein, using a BERT encoder to extract features from the text data to obtain text unimodal features It is expressed as follows: Among them, Con i Represents the input text, i represents the i-th obtained text, that is, Con i =Con1,Con2,Con3...Con i ; The expression image data image i Input VGG16 convolution encoder to obtain image unimodal features It is expressed as follows: That is, i∈[1,2,3,4…]; S22, inputting the voice data and physiological data into the long short-term memory neural network to obtain composite emotion change characteristics; S23, based on the three modal features of text unimodal features, image unimodal features and composite emotion change features, two-by-two modal fusions are performed to obtain corresponding bimodal fusion features; Wherein, the specific steps of S22 include: Each voice data is converted into text, and then each voice data is processed by frame processing to obtain the voice features of each frame data, and at the same time, the physiological features of the physiological data in each time scale are extracted by using a time sliding window method; wherein, the embedding layer encodes each voice data into an emotion expression vector and inputs it into the feature extraction layer, and the emotion expression vector uses the emotion expression in the voice data preprocessed according to the predicted emotion expression dictionary statistics, and encodes it into an emotion expression vector, and then the emotion expression vector is input into the feature extraction layer, and the emotion cause feature unit is used to extract the event cause through a long short-term memory neural network; The composite emotion knowledge unit is combined with the long short-term memory neural network to extract the composite emotion change features of voice features and physiological features, and the composite emotion change feature motion is obtained. i ; The specific steps of S23 include: The three modal features of text unimodal features, image unimodal features and composite emotion change features are fused in pairs, as shown below: in, Represents the fusion features of text data and expression image data, Represents the fusion features of text data and compound emotion change features, It represents the fusion feature of expression image data and compound emotion change features; t represents the time step, which represents a certain time period; Indicates the characteristics of complex emotional changes; The three fusion features are encoded through the long short-term memory neural network to obtain the temporal characteristics of the dual modality, which is expressed as follows: Where M2 represents the sequence of dual-modal fusion, M2 = {(Con i , image i ), (Con i , motion i ),(image i , motion i )}, k represents the fusion of the two modal vectors at the current time step t, represents the output of the long-short time neural network, Represents the memory unit of the long short-term memory network; The specific steps of the model building module include: The bimodal fusion features and the three modal features of text unimodal features, image unimodal features and compound emotion change features are input into the corresponding three-layer deep neural network. The bimodal features are used as the training set, and the three modal features of text unimodal features, image unimodal features and compound emotion change features are used as the test set for training to obtain a mental health prediction model. The specific steps of the mental health classification module are: Fusing the dual-modal fusion features to form composite fusion features; The composite fusion features are input into the fully connected layer and then linearly transformed to obtain the output value of the mental health category; According to the output value of the mental health category, the probability value of the corresponding mental health category is calculated by the Softmax function, and the category with the largest probability value is taken as the result of mental health identification, which is expressed as follows: Among them, P represents the probability value of the corresponding mental health category, V t Represents the element of the mental health category at the current time step t output by the connection layer, n represents the total mental health category, j represents the index variable used to traverse the output values ​​of all mental health categories, V j An element representing the mental health category corresponding to the index variable.

Citation Information

Patent Citations

  • Multi-mode based emotion recognition method

    CN108805089A

  • Hierarchical multi-modal sentiment analysis method based on multi-task learning

    CN114973045A