A depression emotion recognition method and device based on multi-modal body surface information fusion
Patent Information
- Application Number
- CN202311511936.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-13
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2043-11-13
AI Technical Summary
[0004]有鉴于此,本申请提供了一种基于多模态体表信息融合的抑郁情绪识别方法、系统、终端及计算机可读存储介质,以解决现有技术中通过人脸表情、文本、或者语音进行抑郁识别带来的模态单一,识别结果不够准确,降低抑郁症识别正确率的问题
[0043] The beneficial effects of this application are as follows: Unlike existing technologies, this application is an objective and efficient depressive mood recognition technology based on multimodal body surface information fusion. By acquiring perceptual data of multiple modalities of depressive mood, such as facial, eye, gait, voice, tongue image, pulse wave, and exhaled breath, it makes the multimodal body surface information of depressive mood more comprehensive, allowing for the full extraction of key information conducive to identifying depressive mood, thereby improving the accuracy of depressive mood recognition. Secondly, this application inputs the perceptual data into a multimodal fusion model to perform multimodal information fusion calculations on the perceptual data, obtaining fused multimodal data. This enables the fusion of different modalities of body surface data, and the fusion model can also complete missing modalities. Thirdly, this application inputs the multimodal data into a recognition network model to classify the multimodal data and obtain the depressive mood recognition result, improving the accuracy of the depressive mood recognition result.
Smart Images

Figure CN117481652B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method, system, terminal, and computer-readable storage medium for recognizing depressive mood based on multimodal body surface information fusion. Background Technology
[0002] Depression is an affective disorder characterized primarily by depressed mood, and it is one of the most common mental disorders in modern society. Therefore, developing efficient and objective early screening technologies for depression can effectively prevent the worsening of the condition and also effectively prevent self-harming behaviors caused by depression.
[0003] In existing technologies, depression screening techniques are mostly based on questionnaires and interviews, which are overly complex and lack objectivity. Furthermore, depressed patients sometimes exhibit negative treatment responses (especially those unwilling to cooperate with treatment or children), for example, concealing their true condition when completing questionnaires. To address this issue, some studies have attempted to detect depression through facial expression and speech detection, identifying depressive moods through video, voice, and text. However, existing depression recognition methods based on facial expressions, text, or voice often employ a single modality, resulting in incomplete information. Some patents employ depression recognition based on multiple information sources, but these only utilize 2-3 modalities, still losing some crucial information related to depressive moods. Additionally, existing multimodal data fusion models have shortcomings, simply splicing information or using simplistic voting mechanisms in the decision-making process, leading to inaccurate results and reduced accuracy in depression identification. Summary of the Invention
[0004] In view of this, this application provides a method, system, terminal and computer-readable storage medium for recognizing depressive mood based on multimodal body surface information fusion, in order to solve the problem that existing technologies that use facial expressions, text or voice for depression recognition have a single modality, inaccurate recognition results and reduce the accuracy of depression recognition.
[0005] This application proposes a method for recognizing depressive mood based on multimodal body surface information fusion, wherein the method for recognizing depressive mood based on multimodal body surface information fusion includes:
[0006] Acquire perceptual data on multimodal body surface information related to depressive mood;
[0007] The sensed data is input into a multimodal fusion model, and multimodal information fusion calculation is performed on the sensed data to obtain fused multimodal data.
[0008] The multimodal data is input into a recognition network model to classify the multimodal data and obtain the recognition result of depressive mood.
[0009] Optionally, the multimodal body surface perception data includes one or more of the following: static images and dynamic video data of the tongue, video data of gait and eyes, time-series signal data of pulse waves, time-series signal data of speech information, and waveform data of odor concentration;
[0010] The static images and dynamic video data of the tongue are obtained by a tongue imaging device that senses the tongue.
[0011] The gait and eye video data are obtained by recording the dynamic changes in gait and eye movement using a camera;
[0012] The time-series signal data of the pulse wave is obtained by recording the pulse wave at the wrist using a multi-channel pulse diagnostic instrument;
[0013] The time-series signal data of the voice information is obtained by recording voice information with a microphone;
[0014] The waveform data of the odor concentration is obtained by recording the composition of exhaled gas using an electronic nose.
[0015] Optionally, the multimodal fusion model includes a multimodal data encoding unit, a cross-modal learning unit, a multi-scale fusion unit, and a modality conversion unit;
[0016] The step of inputting the perceived data into a multimodal fusion model and performing multimodal information fusion calculations on the perceived data to obtain fused multimodal data specifically includes:
[0017] The body surface sensing data of each single modality in the sensing data is input into the multimodal data encoding unit to obtain the body surface sensing data set of the encoded sensing data.
[0018] The body surface perception data set of each modality is input into the cross-modal learning unit to obtain the cross-modal learning feature set of the multimodal body surface perception data of all modalities;
[0019] Each factor of the body surface perception dataset is input into the multi-scale fusion unit to obtain the first fusion feature;
[0020] The first fused feature and the cross-modal learning feature set are input into the modality conversion unit to obtain the fused multimodal data.
[0021] Optionally, the step of inputting the body surface perception data set of each modality into the cross-modal learning unit to obtain the cross-modal learning feature set of the multimodal body surface perception data of all modalities specifically includes:
[0022] The body surface perception data sets of the first and second modalities are processed by convolution algorithm to obtain the first feature map set and the second feature map set respectively;
[0023] Numerical vector normalization is performed on the first feature map set and the second feature map set to obtain the weights of the first mode and the second mode, respectively.
[0024] Based on the weights, the first feature map set, and the second feature map set, a cross-modal learning feature set of the multimodal body surface perception data of the first modality is obtained;
[0025] The body surface perception data set for each modality is processed sequentially to obtain the cross-modal learning feature set of the multimodal body surface perception data for all modalities.
[0026] Optionally, the step of inputting each factor of the body surface perception data set into the multi-scale fusion unit to obtain the first fusion feature specifically includes:
[0027] Each factor of the body surface perception data set is concatenated to obtain the concatenated features of each modality;
[0028] The spliced features are processed by one-dimensional convolution and three-dimensional convolution algorithms to obtain a first feature map and a second feature map.
[0029] The first feature map and the second feature map are subjected to nonlinear and fusion processing to obtain the first fused feature.
[0030] Optionally, the step of inputting the first fused feature and the cross-modal learning feature set into the modality conversion unit to obtain the fused multimodal data specifically includes:
[0031] The first fused feature is split into feature vectors after multi-scale transformation of each modality to obtain a feature vector group;
[0032] Based on the feature vector group and the cross-modal learning feature set, a loss function is used to constrain the data, resulting in the fused multimodal data.
[0033] Optionally, the step of inputting the multimodal data into a recognition network model to classify the multimodal data and obtain the recognition result of depressive mood specifically includes:
[0034] The multimodal data is segmented into body surface information multimodal vectors, and the body surface information multimodal vectors are input into a one-dimensional convolution algorithm for embedding to obtain body surface information embedded multimodal vectors.
[0035] The surface information is embedded into a multimodal vector and processed by intramodal multi-head attention and a fully connected module to obtain a fully connected multimodal vector.
[0036] The fully connected multimodal vector is output to the depression regression sub-network algorithm for classification, and the identification result of depressive mood is obtained.
[0037] This application also proposes a depressive mood recognition system based on multimodal body surface information fusion, wherein the depressive mood recognition system based on multimodal body surface information fusion includes:
[0038] The data perception module is used to acquire perception data of multimodal body surface information related to depressive mood;
[0039] The multimodal fusion module is used to input the perceived data into the multimodal fusion model, perform multimodal information fusion calculation on the perceived data, and obtain fused multimodal data;
[0040] The emotion recognition module is used to input the multimodal data into the recognition network model, classify the multimodal data, and obtain the recognition result of depressive emotion.
[0041] This application also proposes a terminal, the terminal comprising: a memory, a processor, and a depressive mood recognition program based on multimodal body surface information fusion stored in the memory and executable on the processor, wherein when the depressive mood recognition program based on multimodal body surface information fusion is executed by the processor, it implements the steps of the depressive mood recognition method based on multimodal body surface information fusion as described above.
[0042] This application also proposes a computer-readable storage medium storing a depressive mood recognition program based on multimodal body surface information fusion, wherein when the depressive mood recognition program based on multimodal body surface information fusion is executed by a processor, it implements the steps of the depressive mood recognition method based on multimodal body surface information fusion as described above.
[0043] The beneficial effects of this application are as follows: Unlike existing technologies, this application is an objective and efficient depressive mood recognition technology based on multimodal body surface information fusion. By acquiring perceptual data of multiple modalities of depressive mood, such as facial, eye, gait, voice, tongue image, pulse wave, and exhaled breath, it makes the multimodal body surface information of depressive mood more comprehensive, allowing for the full extraction of key information conducive to identifying depressive mood, thereby improving the accuracy of depressive mood recognition. Secondly, this application inputs the perceptual data into a multimodal fusion model to perform multimodal information fusion calculations on the perceptual data, obtaining fused multimodal data. This enables the fusion of different modalities of body surface data, and the fusion model can also complete missing modalities. Thirdly, this application inputs the multimodal data into a recognition network model to classify the multimodal data and obtain the depressive mood recognition result, improving the accuracy of the depressive mood recognition result.
[0044] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Attached Figure Description
[0045] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0046] Figure 1 This is a flowchart of a preferred embodiment of the depressive mood recognition method based on multimodal body surface information fusion in this application;
[0047] Figure 2 This is a flowchart of the data perception process for the depressive mood recognition method based on multimodal body surface information fusion proposed in this application;
[0048] Figure 3 This is a flowchart of the multimodal fusion process of the depressive mood recognition method based on multimodal body surface information fusion proposed in this application;
[0049] Figure 4 This is a flowchart of the emotion recognition process of the depressive emotion recognition method based on multimodal body surface information fusion in this application;
[0050] Figure 5 This is a flowchart of the multi-scale fusion unit of the depressive mood recognition method based on multimodal body surface information fusion in this application;
[0051] Figure 6 This is a schematic diagram of a preferred embodiment of the depressive mood recognition system based on multimodal body surface information fusion in this application;
[0052] Figure 7 This is a schematic diagram of the operating environment of a preferred embodiment of the terminal of this application. Detailed Implementation
[0053] To enable those skilled in the art to better understand the technical solutions of this application, the following detailed description, in conjunction with the accompanying drawings and specific embodiments, provides a method, system, terminal, and computer-readable storage medium for recognizing depressive mood based on multimodal body surface information fusion. It is understood that the described embodiments are merely some embodiments of this application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0054] The terms "first," "second," etc., used in this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0055] This application provides a method, system, terminal, and computer-readable storage medium for recognizing depressive mood based on multimodal body surface information fusion, in order to solve the problem that existing technologies that rely on facial expressions, text, or voice for depression recognition suffer from modality limitations, inaccurate recognition results, and reduced accuracy in depression recognition.
[0056] Please see Figures 1 to 5 , Figure 1 This is a flowchart of a preferred embodiment of the depressive mood recognition method based on multimodal body surface information fusion in this application; Figure 2 This is a flowchart of the data perception process for the depressive mood recognition method based on multimodal body surface information fusion proposed in this application; Figure 3 This is a flowchart of the multimodal fusion process of the depressive mood recognition method based on multimodal body surface information fusion proposed in this application; Figure 4 This is a flowchart of the emotion recognition process of the depressive emotion recognition method based on multimodal body surface information fusion in this application; Figure 5 This is a flowchart of the multi-scale fusion unit of the depressive mood recognition method based on multimodal body surface information fusion proposed in this application.
[0057] This application proposes a method for recognizing depressive mood based on multimodal body surface information fusion, wherein, for example... Figure 1 As shown, the method for recognizing depressive mood based on multimodal body surface information fusion includes the following steps:
[0058] Step S100: Obtain perceptual data of multimodal body surface information related to depressive mood.
[0059] Specifically, research on depression in both Western and Traditional Chinese Medicine has shown that depression manifests in a patient's demeanor, behavior, or other physical characteristics, such as facial features, eye movements, gait, voice, tongue appearance, pulse wave, and exhaled breath, which differ from those of healthy individuals. Therefore, identifying depressive mood based on physical characteristics is feasible, and this information is readily available. Accordingly, multimodal physical perception data is obtained to capture various modalities of depressive mood. This multimodal data includes data on the face, eyes, gait, voice, tongue appearance, pulse wave, and exhaled breath. This provides a more comprehensive understanding of the various modalities of depressive mood, allowing for the extraction of key information conducive to identifying depressive mood, thereby improving the accuracy of depressive mood identification.
[0060] Step S100: Acquire perceptual data of multimodal body surface information related to depressive mood, specifically including: acquiring perceptual data of multiple modal body surface information related to depressive mood collected by the sensing device.
[0061] Specifically, using sensing devices, multimodal body surface perception data is acquired to obtain various modal body surface information related to depressive mood. This means comprehensive perception of multimodal body surface information related to depression. Images and videos of the face, tongue, and eyes can be perceived through devices such as cameras and tongue diagnostic instruments, while pulse waves, exhaled air, and voice information can be acquired through pulse diagnostic instruments, electronic noses, and microphones. Multimodal data perception is achieved through independent sensing sub-modules, and different data formats are output.
[0062] Among them, such as Figure 2 and Figure 6 As shown, the multimodal body surface perception data includes one or more of the following: static images and dynamic video data of the tongue, video data of gait and eyes, time-series signal data of pulse waves, time-series signal data of speech information, and waveform data of odor concentration.
[0063] The static images and dynamic video data of the tongue are obtained by a tongue imaging device that senses the tongue.
[0064] The gait and eye video data are obtained by recording the dynamic changes in gait and eye movement using a camera;
[0065] The time-series signal data of the pulse wave is obtained by recording the pulse wave at the wrist using a multi-channel pulse diagnostic instrument;
[0066] The time-series signal data of the voice information is obtained by recording voice information with a microphone;
[0067] The waveform data of the odor concentration is obtained by recording the composition of exhaled gas using an electronic nose.
[0068] Specifically, a tongue imager is used to perceive the tongue, outputting static images and dynamic video data of the tongue. A camera records dynamic changes in gait and eye movements, outputting video data of gait and eye movements. A multi-channel pulse oximeter records the pulse wave at the wrist, outputting time-series pulse wave signal data. A microphone records speech information, outputting time-series speech signal data. An electronic nose records the composition of exhaled air, outputting waveform data of odor concentration. During the perception process, the above comprehensive data are acquired to obtain multimodal body surface perception data {X1, X2, ..., X...} of various modalities of depressive mood. n}, where n is the number of modalities, ensuring the quality of multimodal body surface data from the source.
[0069] Step S200: Input the perceived data into the multimodal fusion model, perform multimodal information fusion calculation on the perceived data, and obtain the fused multimodal data.
[0070] Specifically, such as Figure 3 As shown, based on the multimodal fusion model, multimodal information fusion calculation is performed on the perceived data to obtain fused multimodal data. That is, the effective identification of depressive mood can be achieved through comprehensive perception and fusion calculation of multimodal body surface information, and the missing modalities can also be filled.
[0071] The multimodal fusion model includes a multimodal data encoding unit, a cross-modal learning unit, a multi-scale fusion unit, and a modality conversion unit.
[0072] Step S200: Inputting the perceived data into a multimodal fusion model, performing multimodal information fusion calculation on the perceived data, and obtaining fused multimodal data, specifically including:
[0073] The body surface sensing data of each single modality in the sensing data is input into the multimodal data encoding unit to obtain the body surface sensing data set of the encoded sensing data.
[0074] The body surface perception data set of each modality is input into the cross-modal learning unit to obtain the cross-modal learning feature set of the multimodal body surface perception data of all modalities;
[0075] Each factor of the body surface perception dataset is input into the multi-scale fusion unit to obtain the first fusion feature;
[0076] The first fused feature and the cross-modal learning feature set are input into the modality conversion unit to obtain the fused multimodal data.
[0077] Specifically, the multimodal body surface sensing data {X1,X2,…,X...} will be used to...n The surface perception data of each single modality is input into the multimodal data encoding unit to obtain the surface perception data set of the encoded multimodal surface perception data. The surface perception data set of each modality is input into the cross-modal learning unit to obtain the cross-modal learning feature set of the multimodal surface perception data of all modalities. Each factor of the surface perception data set is input into the multi-scale fusion unit to obtain the first fusion feature. The first fusion feature and the cross-modal learning feature set are input into the modality conversion unit to obtain the second fusion feature, and then the fused multimodal data is obtained.
[0078] Specifically, the step of inputting the surface perception data of each single modality in the multimodal surface perception data into the multimodal data encoding unit to obtain the encoded surface perception data set of the multimodal surface perception data includes:
[0079] The surface sensing data of each single modality in the multimodal surface sensing data is input into the corresponding encoder for encoding, thereby obtaining the surface sensing data set of the encoded multimodal surface sensing data.
[0080] Specifically, the surface sensing data for each single modality in the multimodal surface sensing data is encoded separately using the corresponding encoder, i.e., the multimodal data encoding E = {E1, E2, ..., E...} n E1 is an encoder for the body surface data X1, which is a k-layer neural network, k∈R. After being encoded by encoder E1, the diagnostic data X1 is transformed into a vector f1; in turn, the multimodal body surface data {X1,X2,…,X...} are processed. n After processing by E, it becomes {f1,f2,…,f} n}, thus obtaining the encoded multimodal body surface information body surface perception data set, the encoded data F={f1,f2,…,f n It can output cross-modal learning units and multi-scale fusion units respectively.
[0081] Specifically, inputting the body surface perception data set of each modality into the cross-modal learning unit to obtain the cross-modal learning feature set of the multimodal body surface perception data of all modalities includes:
[0082] Taking the first mode and the second mode as examples, the body surface perception data sets of the first mode and the second mode are processed by convolution algorithm to obtain the first feature map set and the second feature map set respectively;
[0083] Numerical vector normalization is performed on the first feature map set and the second feature map set to obtain the weights of the first mode and the second mode, respectively.
[0084] Based on the weights, the first feature map set, and the second feature map set, a cross-modal learning feature set of the multimodal body surface perception data of the first modality is obtained;
[0085] The body surface perception data set for each modality is processed sequentially to obtain the cross-modal learning feature set of the multimodal body surface perception data for all modalities.
[0086] Specifically, the encoded body surface perception data sets of the first and second modalities are processed by three convolutional block algorithms to obtain the first feature map set and the second feature map set, which is the information f encoded by the first modality i. i The first feature map {q} is obtained after three convolutional blocks. i ,k i ,v i The information f encoded by the second mode j j After passing through three more convolutional blocks, the second feature map set {q} is obtained. j ,k j ,v j The first and second feature maps are subjected to numerical vector normalization to obtain the weights between the first and second modes, respectively. This allows the use of att methods. ij =softmax(q) i ek j Calculate the weights between the first modality i and the second modality j to obtain the cross-modal attention weight att. ij Then, cross-modal feature f′ i It can be determined by the weights and v of the first feature map set. j The product is obtained by multiplying and then adding them together, i.e.:
[0087]
[0088] The same processing is performed on the features of each modality in turn, and finally the cross-modal learning feature set {f1',f2',f3',…,f7',f8',f...} of the multimodal body surface perception data is obtained. n '}.
[0089] Specifically, inputting each factor of the body surface perception data set into the multi-scale fusion unit to obtain the first fusion feature includes:
[0090] Each factor of the body surface perception data set is concatenated to obtain the concatenated features of each modality;
[0091] The spliced features are processed by one-dimensional convolution and three-dimensional convolution algorithms to obtain a first feature map and a second feature map.
[0092] The first feature map and the second feature map are subjected to nonlinear and fusion processing to obtain the first fused feature.
[0093] Specifically, such as Figure 5 As shown, each factor of the encoded body surface perception data set is concatenated sequentially to form f, resulting in the concatenated features of each modality, f = {f1, f2, ..., f...}. n The concatenated feature f is processed by 1×1×1 one-dimensional convolution and 3×3×3 three-dimensional convolution algorithms to obtain the first feature map s. a =W sa *f and the second feature map s b =W sb *f, which performs nonlinear and fusion processing on the first and second feature maps to obtain f. sa =σ(s a f,f sb =σ(s b Furthermore, we arrive at the first fusion feature, which is the fusion feature F. s =f sa +f sb .
[0094] Specifically, inputting the first fused feature and the cross-modal learning feature set into the modality conversion unit to obtain the fused multimodal data includes:
[0095] The first fused feature is split into feature vectors after multi-scale transformation of each modality to obtain a feature vector group;
[0096] Based on the feature vector group and the cross-modal learning feature set, a loss function is used to constrain the data, resulting in the fused multimodal data.
[0097] Specifically, in the modality transformation unit, the relationships between modalities are learned, and the first fused feature is decomposed into feature vectors after multi-scale transformation of each modality, resulting in a feature vector set, namely fs1, fs2, ..., fs n Let P be the relational parameter of the i-th mode. i ={a i ,b i ,...,h i ,...,p i}, Then the i-th mode can be represented as F si =P i e Q i T And the feature vector values F of each mode si The true value f of the eigenvector after multi-scale transformation snThe data are compared and constrained using a loss function to obtain the fused multimodal data.
[0098] The loss function is calculated based on the probability distribution of using these feature values, as follows:
[0099]
[0100] Where Lcor is a constraint condition, the optimal fused multimodal data f is obtained. s1 ,f s2 ,...,f sn Both P() and Q() are probability functions.
[0101] Based on the cross-modal learning feature set f s ={f'1,f'2,f'3,...,f'7,f'8,f' n} and feature vector group {f s1 ,f s2 ,...,f sn}, to obtain the fused multimodal data, and then f, the fused multimodal data s The output is sent to the emotion recognition module to learn a non-linear representation model. In the case of a missing modality, the missing modality is filled in by the non-linear representation of other modalities.
[0102] Step S300: Input the multimodal data into the recognition network model, classify the multimodal data, and obtain the recognition result of depressive mood.
[0103] Specifically, such as Figure 4 As shown, by inputting multimodal data into the recognition network model and classifying the multimodal data, the recognition results of depressive mood can be obtained. This can fully extract key information that is conducive to the recognition of depressive mood, thereby improving the recognition performance of depressive mood.
[0104] Specifically, step S300 involves inputting the multimodal data into a recognition network model to classify the multimodal data and obtain the recognition result of depressive mood.
[0105] The multimodal data is segmented into body surface information multimodal vectors, and the body surface information multimodal vectors are input into a one-dimensional convolution algorithm for embedding to obtain body surface information embedded multimodal vectors.
[0106] The surface information is embedded into a multimodal vector and processed by intramodal multi-head attention and a fully connected module to obtain a fully connected multimodal vector.
[0107] The fully connected multimodal vector is output to the depression regression sub-network algorithm for classification, and the identification result of depressive mood is obtained.
[0108] Specifically, based on the multimodal data, the multimodal data f s Segmented into multimodal vectors of body surface information Multimodal vector of body surface information As input, the data is embedded into a one-dimensional convolutional algorithm to obtain a multimodal vector of body surface information. This multimodal vector is then processed by intramodal multi-head attention and a fully connected module to obtain a fully connected multimodal vector. This fully connected multimodal vector is then output to a depression regression sub-network algorithm for classification. The depression regression sub-network algorithm concatenates the outputs of the multimodal encoder sub-network and then uses a fully connected module to regress the severity of depression. Binary cross-entropy is used as the loss function for this depression regression to obtain the identification result of depressive mood.
[0109] Please see Figures 6 to 7 , Figure 6 This is a schematic diagram of a preferred embodiment of the depressive mood recognition system based on multimodal body surface information fusion in this application; Figure 7 This is a schematic diagram of the operating environment of a preferred embodiment of the terminal of this application.
[0110] In some embodiments, such as Figure 6 As shown, based on the above-mentioned depressive mood recognition method based on multimodal body surface information fusion, this application also proposes a depressive mood recognition system based on multimodal body surface information fusion, wherein the depressive mood recognition system based on multimodal body surface information fusion includes:
[0111] Data perception module 51 is used to acquire perception data of multimodal body surface information related to depressive mood;
[0112] The multimodal fusion module 52 is used to input the perceived data into the multimodal fusion model, perform multimodal information fusion calculation on the perceived data, and obtain fused multimodal data;
[0113] The emotion recognition module 53 is used to input the multimodal data into the recognition network model, classify the multimodal data, and obtain the recognition result of depressive emotion.
[0114] In some embodiments, such as Figure 7 As shown, based on the above-mentioned method and system for recognizing depressive mood based on multimodal body surface information fusion, this application also proposes a terminal, which includes: a memory 10, a processor 20, and a display 30. Figure 7 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0115] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard drive or memory. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, the memory 20 may include both internal and external storage units. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code installed on the terminal. The memory 20 can also be used to temporarily store data that has been output or will be output.
[0116] In one embodiment, the memory 20 stores a depressive mood recognition program 40 based on multimodal body surface information fusion. The depressive mood recognition program 40 based on multimodal body surface information fusion can be executed by the processor 10 to realize the depressive mood recognition method based on multimodal body surface information fusion in this application.
[0117] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in the memory 20 or process data, such as executing the depressive mood recognition method based on multimodal body surface information fusion.
[0118] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface. The components 10-30 of the terminal communicate with each other via a system bus.
[0119] This application also proposes a computer-readable storage medium storing a depressive mood recognition program based on multimodal body surface information fusion, wherein when the depressive mood recognition program based on multimodal body surface information fusion is executed by a processor, it implements the steps of the depressive mood recognition method based on multimodal body surface information fusion as described above.
[0120] In summary, this application acquires multimodal body surface perception data of various modalities related to depressive mood, such as facial, eye, gait, voice, tongue appearance, pulse wave, and exhaled breath data. This makes the multimodal body surface information of depressive mood more comprehensive, allowing for the extraction of key information conducive to identifying depressive mood, thereby improving the accuracy of depressive mood recognition. Secondly, this application performs multimodal information fusion calculations on the perception data based on a multimodal fusion model to obtain fused multimodal data, enabling the fusion of body surface data from different modalities. This fusion model can also complete missing modalities. Thirdly, this application classifies the multimodal data based on a recognition network model to obtain the recognition result of depressive mood, improving the accuracy of the depressive mood recognition result. This method perceives multimodal body surface information related to emotions, then quantifies the body surface data and performs fusion calculations to diagnose depression. It is an objective method for diagnosing depression, avoiding subjective factors in the judgment of depression. Unlike facial expressions and voice recognition for depression, this application comprehensively acquires multimodal body surface information related to depression, which can fully extract information related to depression, thereby improving the recognition rate of depressive mood. It is also a convenient, non-invasive, and user-friendly early warning technology for depression, which can be used for large-scale screening of high-incidence groups of depression such as students and adolescents. In addition, the multimodal body surface information fusion model can also realize the fusion of multi-source heterogeneous multimodal body surface data, improve the fusion performance of multimodal data, and thus improve the accuracy of depression recognition.
[0121] It should be noted that the various optional implementation methods described in the embodiments of this application can be combined with each other or implemented individually, and the embodiments of this application do not limit this.
[0122] In the description of this application, it should be understood that the terms "upper," "lower," "left," and "right," etc., indicating orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or a specific orientational structure and operation. Therefore, they should not be construed as limitations on this application. Furthermore, "first" and "second" are only for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "multiple" means two or more.
[0123] In the description of this application, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0124] The above embodiments are described with reference to the accompanying drawings. Other different forms and embodiments are also feasible without departing from the principles of this application, and therefore this application should not be construed as limiting the embodiments set forth herein. Rather, these embodiments are provided to make this application complete and perfect, and to convey the scope of this application to those skilled in the art. In the drawings, component dimensions and relative dimensions may be exaggerated for clarity. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. The terms “comprising” and / or “including”, when used in this specification, indicate the presence of said features, integers, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, elements, components, and / or groups thereof. Unless otherwise shown, numerical ranges, when stated, include the upper and lower limits of the range and any subranges therebetween.
[0125] The above description is only a part of the embodiments of this application and does not limit the scope of protection of this application. Any equivalent device or equivalent process transformation made based on the content of this application specification and drawings, or direct or indirect application in other related technical fields, are similarly included in the patent protection scope of this application.
Claims
1. A method for recognizing depressive mood based on multimodal body surface information fusion, characterized in that, include: Acquire perceptual data on multimodal body surface information related to depressive mood; The sensed data is input into a multimodal fusion model, and multimodal information fusion calculation is performed on the sensed data to obtain fused multimodal data. The multimodal data is input into the recognition network model, and the multimodal data is classified to obtain the recognition result of depressive mood; The multimodal fusion model includes a multimodal data encoding unit, a cross-modal learning unit, a multi-scale fusion unit, and a modality conversion unit; The step of inputting the perceived data into a multimodal fusion model and performing multimodal information fusion calculations on the perceived data to obtain fused multimodal data specifically includes: The body surface sensing data of each single modality in the sensing data is input into the multimodal data encoding unit to obtain the body surface sensing data set of the encoded sensing data. The body surface perception data set of each modality is input into the cross-modal learning unit to obtain the cross-modal learning feature set of the body surface perception data set of all modalities; Each factor of the body surface perception dataset is input into the multi-scale fusion unit to obtain the first fusion feature; The first fused feature and the cross-modal learning feature set are input into the modality conversion unit to obtain the fused multimodal data; The step of inputting each factor of the body surface perception data set into the multi-scale fusion unit to obtain the first fusion feature specifically includes: Each factor of the body surface perception data set is concatenated to obtain the concatenated features of each modality; The spliced features are processed by one-dimensional convolution and three-dimensional convolution algorithms to obtain a first feature map and a second feature map. The first feature map and the second feature map are subjected to nonlinear and fusion processing to obtain the first fused feature; The step of inputting the first fused feature and the cross-modal learning feature set into the modality conversion unit to obtain the fused multimodal data specifically includes: The first fused feature is split into feature vectors after multi-scale transformation of each modality to obtain a feature vector group; Based on the feature vector group and the cross-modal learning feature set, a loss function is used to constrain the data, resulting in the fused multimodal data.
2. The method for recognizing depressive mood based on multimodal body surface information fusion according to claim 1, characterized in that, The multimodal body surface information perception data includes one or more of the following: static images and dynamic video data of the tongue, video data of gait and eyes, time-series signal data of pulse waves, time-series signal data of speech information, and waveform data of odor concentration. The static images and dynamic video data of the tongue are obtained by a tongue imaging device that senses the tongue. The gait and eye video data are obtained by recording the dynamic changes in gait and eye movement using a camera; The time-series signal data of the pulse wave is obtained by recording the pulse wave at the wrist using a multi-channel pulse diagnostic instrument; The time-series signal data of the voice information is obtained by recording voice information with a microphone; The waveform data of the odor concentration is obtained by recording the composition of exhaled gas using an electronic nose.
3. The method for recognizing depressive mood based on multimodal body surface information fusion according to claim 1, characterized in that, The step of inputting the body surface perception data set of each modality into the cross-modal learning unit to obtain the cross-modal learning feature set of the body surface perception data set of all modalities specifically includes: The body surface perception data sets of the first and second modalities are processed by convolution algorithm to obtain the first feature map set and the second feature map set respectively; Numerical vector normalization is performed on the first feature map set and the second feature map set to obtain the weights of the first mode and the second mode, respectively. Based on the weights, the first feature map set, and the second feature map set, a cross-modal learning feature set of the body surface perception data set of the first modality is obtained; The body surface perception data set for each modality is processed sequentially to obtain a cross-modal learning feature set for the body surface perception data set of all modalities.
4. The method for recognizing depressive mood based on multimodal body surface information fusion according to claim 1, characterized in that, The step of inputting the multimodal data into a recognition network model, classifying the multimodal data, and obtaining the recognition result of depressive mood specifically includes: The multimodal data is segmented into body surface information multimodal vectors, and the body surface information multimodal vectors are input into a one-dimensional convolution algorithm for embedding to obtain body surface information embedded multimodal vectors. The surface information is embedded into a multimodal vector and processed by intramodal multi-head attention and a fully connected module to obtain a fully connected multimodal vector. The fully connected multimodal vector is output to the depression regression sub-network algorithm for classification, and the identification result of depressive mood is obtained.
5. A depressive mood recognition system based on multimodal body surface information fusion, characterized in that, The steps for implementing the depressive mood recognition method based on multimodal body surface information fusion as described in any one of claims 1-4, wherein the depressive mood recognition system based on multimodal body surface information fusion comprises: The data perception module is used to acquire perception data of multimodal body surface information related to depressive mood; The multimodal fusion module is used to input the perceived data into the multimodal fusion model, perform multimodal information fusion calculation on the perceived data, and obtain fused multimodal data; The emotion recognition module is used to input the multimodal data into the recognition network model, classify the multimodal data, and obtain the recognition result of depressive emotion.
6. A terminal, characterized in that, The terminal includes: a memory, a processor, and a depressive mood recognition program based on multimodal body surface information fusion stored in the memory and executable on the processor. When the depressive mood recognition program based on multimodal body surface information fusion is executed by the processor, it implements the steps of the depressive mood recognition method based on multimodal body surface information fusion as described in any one of claims 1-4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a depressive mood recognition program based on multimodal body surface information fusion, which, when executed by a processor, implements the steps of the depressive mood recognition method based on multimodal body surface information fusion as described in any one of claims 1-4.
Citation Information
Patent Citations
Multi-modal information fusion-based depression evaluation system and equipment
CN115064246A