Depression quantification method and device based on multi-modal feature adaptation and electronic equipment

By performing dimensionality reduction and feature extraction on multimodal data, the problem of incomplete data application in the quantitative analysis of depressive mood in existing technologies has been solved, and efficient and accurate quantitative analysis of depressive mood has been achieved.

CN115758114BActive Publication Date: 2026-03-20ANHUI IFLYHEALTH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-11
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

When using existing technologies to quantitatively analyze depressive mood based on emotion recognition, the large amount of data generated from face-to-face consultations cannot be directly applied, resulting in incomplete information and inaccurate analysis results.

Method used

By acquiring data from at least two modalities to be identified, dimensionality reduction is performed based on the correlation between the data to be identified and low-dimensional features. Emotional features are then extracted using a general modality feature extraction framework or a modality-specific feature extraction framework to conduct quantitative analysis of depressive mood.

Benefits of technology

It enables quantitative analysis of depressive mood based on complete, unsegmented data to be identified, ensuring the reliability and accuracy of the analysis, avoiding memory pressure, and providing comprehensive emotional information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115758114B_ABST
    Figure CN115758114B_ABST
Patent Text Reader

Abstract

This invention provides a method, apparatus, and electronic device for quantitative depression analysis based on multimodal feature adaptation. The method includes: acquiring data to be identified from at least two modalities; reducing the dimensionality of the data to be identified based on the correlation between the data to be identified and low-dimensional features, and extracting features from the dimensionality-reduced data to obtain the emotional features of the data to be identified; wherein the dimensionality of the low-dimensional features is lower than the feature dimension of the data to be identified; and performing quantitative analysis of depressive mood based on the emotional features of the data to be identified from at least two modalities. The method, apparatus, electronic device, and storage medium provided by this invention avoid the memory pressure caused by directly extracting features from high-dimensional data, and can directly perform quantitative analysis of depressive mood based on complete, unsegmented data to be identified, thereby ensuring the reliability and accuracy of quantitative analysis of depressive mood.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of quantitative detection of depression, and particularly relates to a depression quantitative method and device based on multi-modal feature adaptation and electronic equipment. BACKGROUND

[0002] The main purpose of emotion analysis is to analyze the public's views on certain products, events, people or ideas.

[0003] In the prior art, although some methods perform multi-modal fusion when performing emotion analysis, these methods are based on modal-specific extractors to extract features, and then perform explicit fusion. Such operation is limited by the structural characteristics of the feature extractor of some modal and the high-dimensional audio and video data of the modal itself. The input of the model can only be input in the form of interview segments, and the overall information is lost.

[0004] However, due to the limitation of display memory, existing methods need to extract artificial features for use on high-dimensional data. Especially in the quantitative analysis of depression based on emotion recognition, the large amount of data generated by face-to-face consultation cannot be directly applied, resulting in incomplete information for quantitative analysis and inaccurate analysis results. SUMMARY

[0005] The present application provides a depression quantitative method and device based on multi-modal feature adaptation and electronic equipment to solve the problem that in the prior art, the large amount of data generated by face-to-face consultation cannot be directly applied in the quantitative analysis of depression based on emotion recognition, resulting in incomplete information for quantitative analysis and inaccurate analysis results.

[0006] The present application provides a depression quantitative method based on multi-modal feature adaptation, comprising:

[0007] Obtaining at least two modal to-be-recognized data;

[0008] Based on the correlation between the to-be-recognized data and low-dimensional features, the to-be-recognized data is reduced in dimension, and the data obtained by dimension reduction is feature-extracted to obtain the emotion features of the to-be-recognized data; the dimension of the low-dimensional features is lower than the feature dimension of the to-be-recognized data;

[0009] Based on the emotion features of the to-be-recognized data of the at least two modal, depression emotion quantitative analysis is performed.

[0010] According to the depression quantitative method based on multi-modal feature adaptation provided by the present application, the to-be-recognized data is reduced in dimension based on the correlation between the to-be-recognized data and low-dimensional features, and the data obtained by dimension reduction is feature-extracted to obtain the emotion features of the to-be-recognized data, comprising:

[0011] extract emotion features of the to-be-identified data based on the universal modal feature extraction framework;

[0012] The universal modal feature extraction framework is obtained by jointly training an emotion recognition framework based on sample data of at least two modalities and emotion labels corresponding to the sample data, and is used to reduce dimensionality of the to-be-identified data based on correlation between the to-be-identified data and low-dimensional features obtained by training, and extract features from the data obtained by reducing dimensionality.

[0013] According to the depression quantitative method based on multi-modal feature self-adaption provided by the application, the universal modal feature extraction framework extracts emotion features of the to-be-identified data, which comprises:

[0014] The universal modal feature extraction framework extracts emotion features of the to-be-identified data of the at least two modalities based on the same universal modal feature extraction framework.

[0015] Or,

[0016] The universal modal feature extraction framework extracts emotion features of the to-be-identified data of the at least two modalities based on the universal modal feature extraction framework corresponding to the at least two modalities respectively.

[0017] According to the depression quantitative method based on multi-modal feature self-adaption provided by the application, the universal modal feature extraction framework comprises at least two extraction layers.

[0018] The universal modal feature extraction framework extracts emotion features of the to-be-identified data, which comprises:

[0019] The universal modal feature extraction framework extracts emotion features of the to-be-identified data, which comprises:

[0020] The universal modal feature extraction framework extracts emotion features of the to-be-identified data, which comprises:

[0021] According to the depression quantitative method based on multi-modal feature self-adaption provided by the application, in the case that the universal modal feature extraction framework extracts emotion features of the to-be-identified data of the at least two modalities based on the same universal modal feature extraction framework, the determination step of the universal modal feature extraction framework comprises:

[0022] determine a sample data set, wherein the sample data set comprises sample data of at least two modalities, and the sample data corresponds to an emotion label;

[0023] perform modal data discard on the sample data set to obtain a modal incomplete data set;

[0024] train a joint emotion recognition framework based on the modal incomplete data set to obtain the universal modal feature extraction framework.

[0025] According to the depression quantitative method based on multi-modal feature self-adaption provided in the application, the emotion features of the to-be-recognized data of the at least two modalities are used for depression emotion quantitative analysis, which comprises:

[0026] perform feature fusion on the emotion features of the to-be-recognized data of the at least two modalities to obtain fused features;

[0027] perform emotion classification on the to-be-recognized data of the at least two modalities based on the fused features, and / or determine the emotion intensity of the to-be-recognized data of the at least two modalities corresponding to a preset emotion.

[0028] According to the depression quantitative method based on multi-modal feature self-adaption provided in the application, the at least two modalities comprise at least two of video, voice, text, behavior, expression, body index and physiological data.

[0029] The application further provides a depression quantitative device based on multi-modal feature self-adaption, which comprises:

[0030] an acquisition unit configured to acquire to-be-recognized data of at least two modalities;

[0031] a determination emotion feature unit configured to perform dimension reduction on the to-be-recognized data based on the correlation between the to-be-recognized data and low-dimensional features, and perform feature extraction on the data obtained by the dimension reduction to obtain emotion features of the to-be-recognized data, wherein the dimension of the low-dimensional features is lower than the feature dimension of the to-be-recognized data;

[0032] a depression emotion quantitative analysis unit configured to perform depression emotion quantitative analysis based on the emotion features of the to-be-recognized data of the at least two modalities.

[0033] The application further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the depression quantitative method based on multi-modal feature self-adaption as described above when executing the program.

[0034] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the depression quantification method based on multimodal feature adaptation as described above.

[0035] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the depression quantification method based on multimodal features as described above.

[0036] The present invention provides a method, device, and electronic device for quantitative depression based on multimodal feature adaptation. It reduces the dimensionality of the data to be identified based on the correlation between the data to be identified and low-dimensional features, and then extracts features and identifies emotions from the reduced data. This avoids the memory pressure caused by directly extracting features from high-dimensional data. It allows for quantitative analysis of depressive emotions directly based on complete, unsegmented data to be identified, ensuring that the quantitative analysis of depressive emotions can reference complete and comprehensive emotional information provided by the data to be identified, thereby ensuring the reliability and accuracy of the quantitative analysis of depressive emotions. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0038] Figure 1 This is a flowchart illustrating the depression quantification method based on multimodal feature adaptation provided by the present invention.

[0039] Figure 2 This is one of the flowcharts for extracting emotional features from data to be identified provided by the present invention;

[0040] Figure 3 This is the second flowchart illustrating the process of extracting emotional features from data to be identified, provided by the present invention.

[0041] Figure 4 This is a schematic diagram of the general modal feature extraction framework provided by the present invention;

[0042] Figure 5 This is the third flowchart illustrating the process of extracting emotional features from data to be identified provided by the present invention;

[0043] Figure 6 This is a flowchart illustrating the determination steps of the general modal feature extraction framework provided by the present invention;

[0044] Figure 7is a flowchart of step 130 in the depression quantitative method based on multi-modal feature self-adaption provided by the present application.

[0045] Figure 8 is a structural schematic diagram of the depression quantitative device based on multi-modal feature self-adaption provided by the present application.

[0046] Figure 9 is a structural schematic diagram of the electronic device provided by the present application. DETAILED DESCRIPTION

[0047] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described below in connection with the drawings in the present application. Obviously, the described embodiments are some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0048] In the related art, depression quantitative detection mainly relies on doctors to conduct face-to-face consultation according to different scales, which is extremely low in efficiency and difficult to benefit the public.

[0049] Although some automatic methods can assist doctors in depression quantitative analysis and detection, some of these methods are based on single-modal information for emotion recognition, such as text, audio, and single-modal information cannot fully represent the overall state of a person, which inevitably causes information loss and limits the checking performance of the method.

[0050] Some other methods use multi-modal information for emotion recognition, but they are based on interview segment input and require all modal information when used.

[0051] To this end, the present application provides a depression quantitative method based on multi-modal feature self-adaption, Figure 1 is a flowchart of the depression quantitative method based on multi-modal feature self-adaption provided by the present application, as Figure 1 shown, the method comprises:

[0052] Step 110, obtaining at least two modalities of to-be-recognized data.

[0053] Specifically, the to-be-recognized data is data that needs to be recognized, and the to-be-recognized data can be data of the entire interview process. The at least two modalities here refer to two or more modalities, which can include the voice modality and the text modality, or the text modality and the video modality, or the voice modality, the text modality, the video modality, the behavior modality, the expression modality, the body index and the physiological data, etc., which are not limited by the embodiments of the present application.

[0054] Correspondingly, the to-be-recognized data herein can include voice data and text data, can also include text data and video data, and can further include voice data, text data, video data, behavior data, expression data, body index data, physiological data, and the like, and embodiments of the present application do not make specific limitation thereon.

[0055] The voice data herein can be obtained through a sound pickup device. The sound pickup device herein can be a smart phone, a tablet computer, and can also be a smart electric appliance such as a sound box, a television, an air conditioner, and the like. After the sound pickup device obtains voice data through a microphone array, the voice data can be amplified and de-noised, and embodiments of the present application do not make specific limitation thereon.

[0056] The text data herein can be directly input by a user, can also be obtained after voice transcription is performed on audio data collected, and can further be obtained by performing OCR (Optical Character Recognition) on an image collected by an image collection device such as a scanner, a mobile phone, a camera, and the like.

[0057] The video data herein can be a video pre-recorded and stored, and can also be a video stream collected in real time, and embodiments of the present application do not make specific limitation thereon.

[0058] The behavior data herein can be behavior data pre-collected and stored, and can also be behavior data collected in real time, and can include interpersonal interaction situation data and the like.

[0059] The expression data herein can be an expression video pre-recorded and stored, can also be an expression video stream collected in real time, and can further be an expression image collected by an image collection device such as a tablet computer, a mobile phone, a camera, and the like, and embodiments of the present application do not make specific limitation thereon.

[0060] The body index data herein can be body index related data directly input by a user, and can include sleep situation data, anxiety situation data, and attention concentration situation data and the like.

[0061] The physiological data herein can be physiological related data directly input by a user, and can include hand numbness frequency data, tremor frequency data, and general weakness frequency data and the like, and embodiments of the present application do not make specific limitation thereon.

[0062] In step 120, the to-be-recognized data is reduced in dimension based on a correlation between the to-be-recognized data and a low-dimensional feature, and feature extraction is performed on the data obtained after the reduction in dimension, to obtain an emotion feature of the to-be-recognized data. The low-dimensional feature has a dimension lower than a feature dimension of the to-be-recognized data.

[0063] Specifically, due to the limitation of the display memory, and the feature dimension of the to-be-recognized data is usually high, the to-be-recognized data cannot be completely recognized. To solve this problem, the embodiment of the present application realizes data dimension reduction based on correlation, thereby providing conditions for emotion feature extraction and emotion recognition based on complete to-be-recognized data.

[0064] Further, after obtaining the to-be-recognized data of at least two modalities, the to-be-recognized data can be dimension reduced based on the correlation between the to-be-recognized data and the low-dimensional features.

[0065] Here, the feature dimension of the to-be-recognized data is related to the modality to which the to-be-recognized data belongs. For example, the to-be-recognized data of the video modality is an image sequence composed of multiple frames of images, each frame of image corresponds to a spatial dimension, and each frame of image in the image sequence has a front-back relationship and also corresponds to a time dimension, so the feature dimension of the to-be-recognized data of the video modality is usually high. Here, the feature dimension of the to-be-recognized data of various modalities can be pre-set.

[0066] The low-dimensional feature refers to a feature with a feature dimension lower than that of the to-be-recognized data, which can specifically be an implicit space matrix with a dimension lower than that of the to-be-recognized data. For example, the low-dimensional feature can be an implicit space matrix with a size of N*D, and the to-be-recognized data can be a matrix with a size of M*C, N*D

[0067] Here, the data dimension reduction based on the correlation between the to-be-recognized data and the low-dimensional features can be implemented by applying an attention mechanism.

[0068] For example, the to-be-recognized data and the low-dimensional features can be input into a cross-attention (Cross-Attention) module together, and the low-dimensional features can be used to extract useful feature information from the to-be-recognized data, so as to realize data dimension reduction of the to-be-recognized data.

[0069] After dimension reduction of the to-be-recognized data, the dimension-reduced data can be subjected to feature extraction to obtain the emotion features of the to-be-recognized data. Here, the dimension-reduced data can be subjected to feature extraction using a feature extraction framework. The feature extraction framework can be a Transformer or an MLP (Multi-Layer Perceptron), and the embodiment of the present application does not make a specific limitation.

[0070] Understandably, the approach of first reducing dimensionality and then extracting features in step 120 avoids the memory pressure caused by directly extracting features from high-dimensional data, making it possible to extract emotional features and perform quantitative analysis of depressive mood based on complete, unsegmented data to be identified.

[0071] Step 130: Quantitative analysis of depressive mood is performed based on the emotional characteristics of the data to be identified in the at least two modalities.

[0072] Specifically, after obtaining the emotional features of the data to be identified based on at least two modalities, a quantitative analysis of depressive mood can be performed based on the emotional features of the data to be identified based on at least two modalities. This quantitative analysis of depressive mood may include emotion classification, or it may include determining the emotional intensity of a preset emotion corresponding to the data to be identified based on at least two modalities. It may also include both emotion classification and determining the emotional intensity of a preset emotion corresponding to the data to be identified based on at least two modalities; the embodiments of the present invention do not specifically limit this.

[0073] Taking depressive mood as an example, mood classification can be to analyze whether the test subject has depressive mood, and to determine the intensity of the preset mood corresponding to the data to be identified in at least two modalities. Mood intensity can be to determine the severity of the depressive mood of the test subject, namely, four levels of depressive mood: none, mild, moderate and severe.

[0074] The method provided in this invention reduces the dimensionality of the data to be identified based on the correlation between the data to be identified and low-dimensional features, and then performs feature extraction and emotion recognition on the data obtained from the dimensionality reduction. This avoids the memory pressure caused by directly extracting features from high-dimensional data, and allows for quantitative analysis of depressive mood directly based on complete, unsegmented data to be identified. This ensures that the quantitative analysis of depressive mood can refer to the complete and comprehensive emotional information provided by the data to be identified, thereby ensuring the reliability and accuracy of the quantitative analysis of depressive mood.

[0075] Based on any of the above embodiments, the multimodal feature-adaptive quantitative depression method provided by the embodiments of the present invention can be applied to the identification and quantitative analysis of depressive mood. Furthermore, the severity of depressive mood obtained from the quantitative analysis of depressive mood in the embodiments of the present invention can be applied to clinical diagnosis as a reference factor for doctors in diagnosing depression. In addition, the severity of depressive mood obtained from the quantitative analysis of depressive mood in the embodiments of the present invention can also be applied to medical record quality inspection, comparing the severity of depressive mood obtained from automated analysis with the severity of depression diagnosed by doctors in the medical records, thereby verifying the quality of the medical records. The embodiments of the present invention do not specifically limit this application.

[0076] Based on the above embodiments, step 120 includes:

[0077] extracting emotion features of the to-be-identified data based on a general modal feature extraction framework;

[0078] The general modal feature extraction framework is obtained by jointly training an emotion recognition framework based on sample data of at least two modalities and emotion labels corresponding to the sample data.

[0079] Specifically, in order to extract emotion features of the to-be-identified data, the general modal feature extraction framework needs to be obtained by the following steps before step 120 is performed:

[0080] The sample data of at least two modalities can be collected in advance, and the sample data of at least two modalities can be provided with corresponding emotion labels. In addition, an emotion recognition framework can be constructed, and the emotion recognition framework is used to recognize the emotion features of the to-be-identified data extracted by the general modal feature extraction framework.

[0081] In addition, an initial general modal feature extraction framework can be constructed. The initial general modal feature extraction framework can functionally include two parts of dimension reduction and feature extraction. The dimension reduction part is used to reduce the dimension of the to-be-identified data based on the correlation between the to-be-identified data and the low-dimensional features obtained by training, and the feature extraction part is used to extract features from the data obtained by dimension reduction. Subsequently, the initial general modal feature extraction framework can be trained based on the sample data of at least two modalities and the emotion labels corresponding to the sample data, and the emotion recognition framework is used to recognize the emotion features of the to-be-identified data extracted by the general modal feature extraction framework.

[0082] In this process, the initial general modal feature extraction framework and the initial emotion recognition framework can be used as an initial recognition model, and the initial recognition model is an initial model for training the general modal feature extraction framework. After obtaining the initial recognition model including the initial general modal feature extraction framework and the initial emotion recognition framework, the initial recognition model can be trained using the sample data of at least two modalities collected in advance and the emotion labels corresponding to the sample data:

[0083] First, the sample data of at least two modalities is input into the initial general modal feature extraction framework, and the initial general modal feature extraction framework extracts features from the sample data of at least two modalities to obtain and output initial emotion features of the sample data of at least two modalities. It can be understood that the initial general modal feature extraction framework is an initial model before the initial recognition model is trained, and in order to distinguish from the emotion features output by the initial recognition model, the emotion features output by the initial general modal feature extraction framework are referred to as initial emotion features.

[0084] Secondly, the initial emotion features of the at least two modalities are input into an initial emotion recognition framework of the initial recognition model, and the initial emotion recognition framework performs emotion recognition on the initial emotion features of the at least two modalities to obtain and output emotion recognition results of the at least two modalities.

[0085] After obtaining the emotion recognition results of the at least two modalities based on the initial recognition model, the emotion recognition results of the at least two modalities can be compared with emotion labels corresponding to sample data of the at least two modalities pre-labeled, a loss function value is calculated according to the difference between the two, and the initial recognition model is iterated in parameters based on the loss function value as a whole, and the initial recognition model after completing the parameter iteration is recorded as an emotion recognition model.

[0086] It can be understood that the greater the difference between the emotion recognition results of the at least two modalities and the emotion labels corresponding to the sample data of the at least two modalities pre-labeled, the greater the loss function value; the smaller the difference between the emotion recognition results of the at least two modalities and the emotion labels corresponding to the sample data of the at least two modalities pre-labeled, the smaller the loss function value.

[0087] The emotion recognition model after parameter iteration has the same structure as the initial recognition model, so the emotion recognition model can be divided into two parts, the initial general modality feature extraction framework after parameter iteration and the initial emotion recognition framework after parameter iteration. For the initial general modality feature extraction framework after parameter iteration, this part can be directly used as a general modality feature extraction framework.

[0088] Here, the general modality feature extraction framework is a model trained to extract emotion features of to-be-recognized data. Like the initial general modality feature extraction framework, the general modality feature extraction framework can also functionally include two parts of dimension reduction and feature extraction.

[0089] That is, the feature extraction part is part of the general modality feature extraction framework, and in the training process of the modality feature extraction framework, the feature extraction part learns to extract emotion features of to-be-recognized data to extract emotion features that can be used for subsequent quantitative analysis of depressive emotions.

[0090] Based on the above embodiment, the general modality feature extraction framework is used to extract emotion features of the to-be-recognized data, including:

[0091] Based on the same general modality feature extraction framework, emotion features of the to-be-recognized data of the at least two modalities are extracted;

[0092] Or,

[0093] extract emotion features of the to-be-identified data of the at least two modalities based on the general modal feature extraction framework corresponding to each of the at least two modalities.

[0094] Specifically, Figure 2 is one of the flowcharts provided by the present application for extracting emotion features of to-be-identified data, as shown in Figure 2 emotion features of the to-be-identified data of the at least two modalities can be extracted based on the same general modal feature extraction framework, that is, the to-be-identified data of the at least two modalities can be input into the same general modal feature extraction framework, and emotion features of the to-be-identified data of the at least two modalities can be extracted.

[0095] For example, the voice data can be input into the general modal feature extraction framework, and emotion features of the voice data can be extracted; the expression data can also be input into the same general modal feature extraction framework, and emotion features of the expression data can be extracted; and the behavior data can also be input into the same general modal feature extraction framework, and emotion features of the behavior data can be extracted. The present application embodiment does not make specific limitations in this regard.

[0096] In this process, the general modal feature extraction framework can randomly discard some modal information during training, for example, 10% of the to-be-identified data of the modalities, wherein the to-be-identified data of the modalities missing can use a zero matrix as the input of the same general modal feature extraction framework.

[0097] It can be understood that even if there is a case of missing to-be-identified data of some modalities, the general modal feature extraction framework can implicitly learn the emotion information between the to-be-identified data of the at least two modalities, and subsequent emotion recognition based on the emotion information learned by the same general modal feature extraction framework improves the accuracy and reliability of subsequent emotion recognition.

[0098] Figure 3 is another flowchart provided by the present application for extracting emotion features of to-be-identified data, as shown in Figure 3 emotion features of the to-be-identified data of the at least two modalities can also be extracted based on the general modal feature extraction framework corresponding to each of the at least two modalities.

[0099] That is, each modality can correspond to a dedicated general modal feature extraction framework, and the modalities and the general modal feature extraction frameworks correspond one-to-one.

[0100] For example, the voice data can be input into the general modal feature extraction framework corresponding to the voice data, to extract the emotion features of the voice data, the text data can be input into the general modal feature extraction framework corresponding to the text data, to extract the emotion features of the text data, and the video data can be input into the general modal feature extraction framework corresponding to the video data, to extract the emotion features of the video data. After the voice features, the text features and the video features are extracted in the general modal feature extraction framework corresponding to each modality, the voice features, the text features and the video features can be fused, and the voice features, the text features and the video features can be spliced, or the voice features, the text features and the video features can be weighted by using an attention mechanism and then spliced, which is not limited in the embodiments of the present application.

[0101] It can be understood that after the voice features, the text features and the video features are fused, the fused features obtained by the fusion can be used as the emotion features of the to-be-recognized data of at least two modalities.

[0102] Based on the above embodiments, Figure 4 is a structural schematic diagram of the general modal feature extraction framework provided by the present application, as Figure 4 indicated, the general modal feature extraction framework includes at least two extraction layers.

[0103] Specifically, the general modal feature extraction framework can include at least two extraction layers, that is, the general modal feature extraction framework can include two or more extraction layers, and the at least two extraction layers are all used to apply the correlation between the to-be-recognized data and the low-dimensional features of the current extraction layer to reduce the dimension of the to-be-recognized data, and to extract features from the data obtained by the dimension reduction.

[0104] Figure 5 is the third flowchart of extracting the emotion features of the to-be-recognized data provided by the present application, as Figure 5 indicated, based on the general modal feature extraction framework, the emotion features of the to-be-recognized data are extracted, including:

[0105] In step 210, based on the current extraction layer in the at least two extraction layers, the correlation between the to-be-recognized data and the low-dimensional features of the current extraction layer is applied to reduce the dimension of the to-be-recognized data, and features are extracted from the data obtained by the dimension reduction, to obtain current features;

[0106] In step 220, based on the current features, the low-dimensional features of the next extraction layer of the current extraction layer are determined, and the next extraction layer is taken as the current extraction layer, to return to the feature extraction based on the current extraction layer, until the current extraction layer is the last extraction layer in the at least two extraction layers, and the current features are taken as the emotion features.

[0107] Specifically, in the process of extracting the emotion feature of the to-be-identified data, first, the extraction layer ranked first among the at least two extraction layers can be taken as the current extraction layer, and the flow of extracting the emotion feature of the to-be-identified data is executed:

[0108] The low-dimensional feature learned by the model training in advance and the to-be-identified data can be input into the current extraction layer among the at least two extraction layers, then the correlation between the to-be-identified data and the low-dimensional feature of the current extraction layer can be applied based on the cross attention module (Cross Attention) to reduce the dimension of the to-be-identified data, and the feature extraction is performed on the data obtained by the dimension reduction using the feature encoding model to obtain the current feature, where the feature encoding model can be LatentTransformer.

[0109] After the feature extraction of the current extraction layer is completed, the next extraction layer of the current extraction layer, that is, the extraction layer ranked second, can be taken as the current extraction layer, and the flow of extracting the emotion feature of the to-be-identified data is returned to be executed:

[0110] That is, after the current feature output by the extraction layer ranked first is obtained, the current feature and the to-be-identified data can be input into the current extraction layer, the correlation between the to-be-identified data and the low-dimensional feature of the current extraction layer can be applied based on the cross attention module to reduce the dimension of the to-be-identified data, and the feature extraction is performed on the data obtained by the dimension reduction using the feature encoding model to obtain the current feature.

[0111] The flow of extracting the emotion feature of the to-be-identified data with the extraction layer ranked third as the current extraction layer is similar to the flow of extracting the emotion feature of the to-be-identified data with the extraction layer ranked second as the current extraction layer, which will not be described here.

[0112] By analogy, until the current extraction layer is the last extraction layer among the at least two extraction layers, the last extraction layer is the last extraction layer of the at least two extraction layers.

[0113] After the current feature is extracted in the last extraction layer, the current feature obtained in the last extraction layer can be taken as the emotion feature.

[0114] The method provided by the embodiment of the application extracts the emotion feature of the to-be-identified data through the general modal feature extraction framework, can retain the complete information of the emotion feature because the to-be-identified data can be the data of the entire interview process, and improves the accuracy of subsequent emotion recognition.

[0115] The existing emotion recognition method based on multiple modalities has strong data dependence and needs to input all modalities to make the final judgment. However, in actual clinical practice, the modalities are often missing, so that the emotion recognition cannot be performed.

[0116] To solve the above problems, the embodiment of the present application discards modal data of a sample data set to obtain a modal incomplete data set.

[0117] Based on the above embodiment, Figure 6 is a flowchart of the determination step of the general modal feature extraction framework provided by the present application, as Figure 6 shown, in the case of extracting emotion features of the to-be-identified data of at least two modalities based on the same general modal feature extraction framework, the determination step of the general modal feature extraction framework comprises:

[0118] Step 610, determining a sample data set, wherein the sample data set comprises sample data of at least two modalities, and emotion labels corresponding to the sample data;

[0119] Step 620, discarding modal data of the sample data set to obtain a modal incomplete data set;

[0120] Step 630, training a joint emotion recognition framework based on the modal incomplete data set to obtain the general modal feature extraction framework.

[0121] Specifically, in the case of extracting emotion features of the to-be-identified data of at least two modalities based on the same general modal feature extraction framework, in the case that some to-be-identified data of modalities may be missing, a sample data set can be determined, wherein the sample data set can comprise sample data of at least two modalities, and emotion labels corresponding to the sample data. The emotion labels here can be depression emotion intensity labels.

[0122] For example, the sample data set can comprise speech sample data and depression emotion intensity labels corresponding to the speech sample data, and text sample data and depression emotion intensity labels corresponding to the text sample data. It can also comprise text sample data and depression emotion intensity labels corresponding to the text sample data, and video sample data and depression emotion intensity labels corresponding to the video sample data. It can also comprise speech sample data and depression emotion intensity labels corresponding to the speech sample data, text sample data and depression emotion intensity labels corresponding to the text sample data, video sample data and depression emotion intensity labels corresponding to the video sample data, behavior sample data and depression emotion intensity labels corresponding to the behavior sample data, expression sample data and depression emotion intensity labels corresponding to the expression sample data, and physiological sample data and depression emotion intensity labels corresponding to the physiological sample data. The embodiment of the present application does not make specific limitations thereto.

[0123] After the sample data set is determined, modal data discard can be performed on the sample data set to obtain a modal incomplete data set. The modal data discard here refers to randomly discarding some modal sample data, for example, 10% of the sample data and the depression emotion intensity labels corresponding to the sample data can be randomly discarded. The modal incomplete data set here refers to the data set obtained after modal data discard is performed on the sample data set.

[0124] After the modal incomplete data set is obtained, the universal modal feature extraction framework can be obtained by training the emotion recognition framework based on the modal incomplete data set.

[0125] For example, an initial universal modal feature extraction framework can be constructed, and then the initial universal modal feature extraction framework can be trained based on the sample data in the modal incomplete data set and the depression emotion intensity labels corresponding to the sample data in the modal incomplete data set, and the initial universal modal feature extraction framework after the training is completed can be used as the universal modal feature extraction framework.

[0126] In addition, in the process of training the emotion recognition framework based on the modal incomplete data set, a cross entropy loss function (Cross Entropy Loss Function) can be used, a mean squared error loss function (Mean Squared Error, MSE) can also be used, and a stochastic gradient descent method can also be used to update the parameters of the initial universal modal feature extraction framework, and the embodiments of the present application do not make specific limitations.

[0127] The method provided by the embodiments of the present application trains the universal modal feature extraction framework based on the modal incomplete data set and the emotion recognition framework, so that the emotion features can also be extracted by the universal modal feature extraction framework for the modal incomplete data to be recognized, and the accuracy of subsequent depression emotion quantitative analysis is improved.

[0128] Based on the above embodiments, Figure 7 is a flowchart of step 130 in the depression quantitative method based on multi-modal feature self-adaption provided by the present application, as Figure 7 shown, step 130 includes:

[0129] Step 131, performing feature fusion on the emotion features of the at least two modal data to be recognized to obtain fused features;

[0130] Step 132, performing emotion classification on the at least two modal data to be recognized based on the fused features, and / or determining the emotion intensity of the at least two modal data to be recognized corresponding to the preset emotion.

[0131] Specifically, after obtaining the emotion features of the to-be-identified data of at least two modalities, the emotion features of the to-be-identified data of at least two modalities can be fused to obtain fused features. The feature fusion can splice the emotion features of the to-be-identified data of at least two modalities, and can also splice the emotion features of the to-be-identified data of at least two modalities after weighting by using an attention mechanism. The embodiments of the present application do not make specific limitations in this regard.

[0132] It can be understood that the fused features herein are emotion features of the to-be-identified data of at least two modalities of the voice modality, the text modality, the video modality, the behavior modality, the expression modality, the body index, and the physiological data.

[0133] After obtaining the fused features, the emotion of the to-be-identified data of at least two modalities can be classified based on the fused features using an MLP (Multi-Layer Perceptron), and

[0134] And / or, determining the emotion intensity of the to-be-identified data of at least two modalities corresponding to the preset emotion.

[0135] The MLP can include an input layer, a hidden layer, and an output layer, the number of layers of the hidden layer can be set to 1-2 layers, and the number of nodes of each layer can be set to 16-32. Wherein, the emotion of the to-be-identified data of at least two modalities can be classified using Sigmoid as an activation function to classify whether there is a depression emotion, or using a Softmax activation function to classify whether there is a depression emotion, and the embodiments of the present application do not make specific limitations in this regard.

[0136] Taking depression emotion as an example, using HAMD-17 (Hamilton Depression Scale), the depression emotion corresponds to 17 sub-scenes, wherein, a score can be regressed for each sub-scene, and then the regression scores of each sub-scene are added to obtain a final score, and the final score is used to determine the emotion intensity of the to-be-identified data of at least two modalities corresponding to the preset emotion.

[0137] For example, determining the emotion intensity of the to-be-identified data of at least two modalities corresponding to the preset emotion can use an emotion intensity of 0-7 as no depression emotion, an emotion intensity of 7-17 as mild depression emotion, an emotion intensity of 17-24 as moderate depression emotion, and an emotion intensity of 24-54 as severe depression emotion, and the embodiments of the present application do not make specific limitations in this regard.

[0138] Correspondingly, the emotion of the to-be-identified data of at least two modalities can be classified as no depression emotion with an emotion intensity of 0-7, and depression emotion with an emotion intensity of more than 7, and the embodiments of the present application do not make specific limitations in this regard.

[0139] The method provided by the embodiment of the present application can perform emotion classification on the to-be-identified data of at least two modalities based on fusion features, and / or determine the emotion intensity of the to-be-identified data of at least two modalities corresponding to a preset emotion, thereby facilitating the determination of the emotion classification and / or the emotion intensity of the to-be-identified data when performing quantitative analysis of depressive emotion, and improving the convenience of subsequent quantitative analysis of depressive emotion.

[0140] According to the above embodiment, the at least two modalities include at least two of video, voice, text, behavior, expression, body index, and physiological data.

[0141] Specifically, the at least two modalities can include a video modality and a voice modality, or a voice modality and a text modality, or a video modality, a voice modality, a text modality, a behavior modality, an expression modality, a body index, and physiological data, and the embodiment of the present application does not make specific limitation thereon.

[0142] The method provided by the embodiment of the present application can obtain richer modal information by including at least two of video, voice, text, behavior, expression, body index, and physiological data, thereby improving the accuracy of subsequent quantitative analysis of depressive emotion.

[0143] According to any of the above embodiments, a depressive quantitative method based on multi-modal feature self-adaption includes the following steps:

[0144] In the first step, to-be-identified data of at least two modalities is obtained, and the at least two modalities include at least two of video, voice, and text.

[0145] In the second step, emotion features of the to-be-identified data of the at least two modalities are extracted based on a same general modal feature extraction framework, or emotion features of the to-be-identified data of the at least two modalities are extracted based on general modal feature extraction frameworks corresponding to the at least two modalities respectively.

[0146] The general modal feature extraction framework is obtained based on sample data of the at least two modalities and emotion labels corresponding to the sample data, and is trained by combining an emotion recognition framework. The general modal feature extraction framework is used to reduce the dimension of the to-be-identified data based on the correlation between the to-be-identified data and low-dimensional features obtained by training, and to extract features from the data obtained by reducing the dimension. The dimension of the low-dimensional features is lower than the dimension of the features of the to-be-identified data.

[0147] The general modal feature extraction framework includes at least two extraction layers.

[0148] Then, the emotion features of the to-be-identified data are extracted based on the general modal feature extraction framework, including:

[0149] The dimensionality reduction can be performed on the to-be-identified data based on a correlation between the to-be-identified data and low-dimensional features of a current extraction layer of the at least two extraction layers, and feature extraction is performed on the data obtained after the dimensionality reduction to obtain current features;

[0150] Then, the low-dimensional features of a next extraction layer of the current extraction layer can be determined based on the current features, and the next extraction layer is taken as the current extraction layer, and the feature extraction based on the current extraction layer is returned until the current extraction layer is the last extraction layer of the at least two extraction layers, and the current features are taken as emotion features.

[0151] In the third step, emotion features of the to-be-identified data of the at least two modalities can be fused to obtain fused features.

[0152] In the fourth step, emotion classification can be performed on the to-be-identified data of the at least two modalities based on the fused features, and / or the emotion intensity of a preset emotion corresponding to the to-be-identified data of the at least two modalities can be determined.

[0153] The depression quantitative device based on the multi-modal feature self-adaption provided by the present application is described below, and the depression quantitative device based on the multi-modal feature self-adaption described below can be correspondingly referred to the depression quantitative method based on the multi-modal feature self-adaption described above.

[0154] Based on any of the above embodiments, the present application provides a depression quantitative device based on multi-modal feature self-adaption, Figure 8 is a structural schematic diagram of the depression quantitative device based on multi-modal feature self-adaption provided by the present application, as Figure 8 shown, the device comprises:

[0155] The acquisition unit 810 is configured to acquire to-be-identified data of at least two modalities.

[0156] The emotion feature determination unit 820 is configured to perform dimensionality reduction on the to-be-identified data based on a correlation between the to-be-identified data and low-dimensional features, and perform feature extraction on the data obtained after the dimensionality reduction to obtain emotion features of the to-be-identified data; the dimensionality of the low-dimensional features is lower than the feature dimensionality of the to-be-identified data.

[0157] The depression emotion quantitative analysis unit 830 is configured to perform depression emotion quantitative analysis based on the emotion features of the to-be-identified data of the at least two modalities.

[0158] The device provided in this invention reduces the dimensionality of the data to be identified based on the correlation between the data to be identified and low-dimensional features, and then performs feature extraction and emotion recognition on the data obtained from the dimensionality reduction. This avoids the memory pressure caused by directly extracting features from high-dimensional data. It can directly perform quantitative analysis of depressive mood based on complete, unsegmented data to be identified, thereby ensuring that the quantitative analysis of depressive mood can refer to the complete and comprehensive emotional information provided by the data to be identified, thus ensuring the reliability and accuracy of the quantitative analysis of depressive mood.

[0159] Based on any of the above embodiments, determining the emotion feature unit specifically includes:

[0160] An emotion feature subunit is defined for extracting emotion features from the data to be identified based on a general modality feature extraction framework.

[0161] The general modality feature extraction framework is trained by combining sample data from at least two modalities and the emotion labels corresponding to the sample data with an emotion recognition framework. The general modality feature extraction framework is used to reduce the dimensionality of the data to be identified based on the correlation between the data to be identified and the low-dimensional features obtained from the training, and to extract features from the data obtained from the dimensionality reduction.

[0162] Based on any of the above embodiments, the determination of the emotion feature subunit is specifically used for:

[0163] Based on the same general modality feature extraction framework, emotional features of the data to be identified in at least two modalities are extracted;

[0164] or,

[0165] Based on the general modality feature extraction framework corresponding to the at least two modalities respectively, the emotion features of the data to be identified in the at least two modalities are extracted.

[0166] Based on any of the above embodiments, the general modal feature extraction framework includes at least two extraction layers;

[0167] The specific use of the emotional characteristic subunit is as follows:

[0168] Based on the current extraction layer among the at least two extraction layers, the correlation between the data to be identified and the low-dimensional features of the current extraction layer is applied to reduce the dimensionality of the data to be identified, and feature extraction is performed on the data obtained after dimensionality reduction to obtain the current features;

[0169] Based on the current feature, determine the low-dimensional feature of the next extraction layer of the current extraction layer, and take the next extraction layer as the current extraction layer, return to feature extraction based on the current extraction layer until the current extraction layer is the last extraction layer in the at least two extraction layers, and take the current feature as the emotion feature.

[0170] Based on any of the above embodiments, in the case of extracting emotion features of the at least two modalities of to-be-recognized data based on the same general modal feature extraction framework, the determination step of the general modal feature extraction framework includes:

[0171] Determine a sample data set, which includes sample data of at least two modalities and emotion labels corresponding to the sample data;

[0172] Discard modal data of the sample data set to obtain a modal incomplete data set;

[0173] Based on the modal incomplete data set, train a joint emotion recognition framework to obtain the general modal feature extraction framework.

[0174] Based on any of the above embodiments, the depression emotion quantitative analysis unit specifically includes:

[0175] Feature fusion is performed on the emotion features of the at least two modalities of to-be-recognized data to obtain fused features;

[0176] Based on the fused features, emotion classification is performed on the at least two modalities of to-be-recognized data, and / or the emotion intensity of the at least two modalities of to-be-recognized data corresponding to a preset emotion is determined.

[0177] Based on any of the above embodiments, the modalities include at least two of video, voice, text, behavior, expression, body index, and physiological data.

[0178] Figure 9 An example of an entity structure diagram of an electronic device is shown in FIG. 1. Figure 9As shown, the electronic device can include a processor 910, a communications interface 920, a memory 930, and a communications bus 940, wherein the processor 910, the communications interface 920, and the memory 930 complete mutual communication through the communications bus 940. The processor 910 can invoke a logic instruction in the memory 930 to execute a depression quantification method based on multi-modal feature adaptation, which includes: obtaining to-be-recognized data of at least two modalities; performing dimension reduction on the to-be-recognized data based on a correlation between the to-be-recognized data and a low-dimensional feature, and performing feature extraction on the data obtained by the dimension reduction to obtain emotion features of the to-be-recognized data; the dimension of the low-dimensional feature is lower than the feature dimension of the to-be-recognized data; and performing depression emotion quantification analysis based on the emotion features of the to-be-recognized data of the at least two modalities.

[0179] In addition, the logic instruction in the memory 930 described above can be implemented in the form of a software function unit and sold or used as an independent product, and can be stored in a computer-readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0180] On the other hand, the present application also provides a computer program product, which includes a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute the depression quantification method based on multi-modal feature adaptation provided by the above-mentioned method, which includes: obtaining to-be-recognized data of at least two modalities; performing dimension reduction on the to-be-recognized data based on a correlation between the to-be-recognized data and a low-dimensional feature, and performing feature extraction on the data obtained by the dimension reduction to obtain emotion features of the to-be-recognized data; the dimension of the low-dimensional feature is lower than the feature dimension of the to-be-recognized data; and performing depression emotion quantification analysis based on the emotion features of the to-be-recognized data of the at least two modalities.

[0181] In yet another aspect, the present application also provides a non-transitory computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the multi-modal feature adaptive depression quantification method provided by the above method, the method comprising: obtaining to-be-identified data of at least two modalities; performing dimension reduction on the to-be-identified data based on the correlation between the to-be-identified data and low-dimensional features, and performing feature extraction on the data obtained by the dimension reduction to obtain emotional features of the to-be-identified data; the dimension of the low-dimensional features is lower than the feature dimension of the to-be-identified data; and performing depression emotion quantification analysis based on the emotional features of the to-be-identified data of the at least two modalities.

[0182] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0183] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and necessary universal hardware platforms, and of course can also be realized by hardware. Based on such understanding, the above technical solutions, essentially or in the form of software products, can be embodied in a computer software product, which can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.

[0184] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A quantitative method for depression based on multimodal feature adaptation, characterized in that, include: Acquire data to be identified in at least two modalities; Based on the correlation between the data to be identified and the low-dimensional features, the data to be identified is dimensionality reduced, and features are extracted from the dimensionality-reduced data to obtain the emotional features of the data to be identified; the dimension of the low-dimensional features is lower than the feature dimension of the data to be identified; different modalities of the data to be identified correspond to the same low-dimensional features, or correspond to different low-dimensional features; the low-dimensional features are latent space matrices with dimensions lower than the feature dimension of the data to be identified. Based on the emotional characteristics of the data to be identified in at least two modalities, a quantitative analysis of depressive mood is performed. The process involves reducing the dimensionality of the data to be identified based on the correlation between the data to be identified and low-dimensional features, and then extracting features from the dimensionality-reduced data to obtain the emotional features of the data to be identified, including: Based on a general modality feature extraction framework, the emotional features of the data to be identified are extracted; The general modality feature extraction framework is trained by combining sample data from at least two modalities and the emotion labels corresponding to the sample data with an emotion recognition framework. The general modality feature extraction framework is used to reduce the dimensionality of the data to be identified based on the correlation between the data to be identified and the low-dimensional features obtained by training, and to extract features from the data obtained by dimensionality reduction. The general modal feature extraction framework includes at least two extraction layers; The emotional features of the data to be identified are extracted using the general modality feature extraction framework, including: Based on the current extraction layer among the at least two extraction layers, the correlation between the data to be identified and the low-dimensional features of the current extraction layer is applied to reduce the dimensionality of the data to be identified, and feature extraction is performed on the data obtained after dimensionality reduction to obtain the current features; Based on the current feature, determine the low-dimensional feature of the next extraction layer of the current extraction layer, and take the next extraction layer as the current extraction layer. Return to the current extraction layer for feature extraction until the current extraction layer is the last extraction layer among the at least two extraction layers, and take the current feature as the emotion feature.

2. The method for quantitative depression based on multimodal feature adaptation according to claim 1, characterized in that, The emotional features of the data to be identified are extracted using the general modality feature extraction framework, including: Based on the same general modality feature extraction framework, emotional features of the data to be identified in at least two modalities are extracted; or, Based on the general modality feature extraction framework corresponding to the at least two modalities respectively, the emotion features of the data to be identified in the at least two modalities are extracted.

3. The method for quantitative depression based on multimodal feature adaptation according to claim 2, characterized in that, In the case of extracting emotional features of the at least two modalities of the data to be identified based on the same general modality feature extraction framework, the step of determining the general modality feature extraction framework includes: Determine a sample dataset, which includes sample data from at least two modalities and the sentiment labels corresponding to the sample data; Modal data is discarded from the sample dataset to obtain a modally incomplete dataset; Based on the incomplete modality dataset, the general modality feature extraction framework is obtained by training in conjunction with the emotion recognition framework.

4. The method for quantitative depression based on multimodal feature adaptation according to any one of claims 1 to 3, characterized in that, The quantitative analysis of depressive mood based on the emotional features of the data to be identified from at least two modalities includes: The emotional features of the data to be identified from at least two modalities are fused to obtain fused features; Based on the fusion features, emotion classification is performed on the data to be identified in the at least two modalities, and / or the emotion intensity of the preset emotion corresponding to the data to be identified in the at least two modalities is determined.

5. The method for quantitative depression based on multimodal feature adaptation according to any one of claims 1 to 3, characterized in that, The at least two modalities include at least two of video, voice, text, behavior, facial expressions, body indicators, and physiological data.

6. A depression quantification device based on multimodal feature adaptation, characterized in that, include: The acquisition unit is used to acquire data to be identified in at least two modalities; An emotion feature unit is defined to reduce the dimensionality of the data to be identified based on the correlation between the data to be identified and the low-dimensional features, and to extract features from the dimensionality-reduced data to obtain the emotion features of the data to be identified; the dimension of the low-dimensional features is lower than the feature dimension of the data to be identified; different modalities of the data to be identified correspond to the same low-dimensional features or different low-dimensional features; the low-dimensional features are latent space matrices with dimensions lower than the feature dimension of the data to be identified. A quantitative analysis unit for depressive mood is used to perform quantitative analysis of depressive mood based on the emotional characteristics of the data to be identified in the at least two modalities. The emotional feature determination unit is specifically used for: Based on a general modality feature extraction framework, the emotional features of the data to be identified are extracted; The general modality feature extraction framework is trained by combining sample data from at least two modalities and the emotion labels corresponding to the sample data with an emotion recognition framework. The general modality feature extraction framework is used to reduce the dimensionality of the data to be identified based on the correlation between the data to be identified and the low-dimensional features obtained by training, and to extract features from the data obtained by dimensionality reduction. The general modal feature extraction framework includes at least two extraction layers; The emotional features of the data to be identified are extracted using the general modality feature extraction framework, including: Based on the current extraction layer among the at least two extraction layers, the correlation between the data to be identified and the low-dimensional features of the current extraction layer is applied to reduce the dimensionality of the data to be identified, and feature extraction is performed on the data obtained after dimensionality reduction to obtain the current features; Based on the current feature, determine the low-dimensional feature of the next extraction layer of the current extraction layer, and take the next extraction layer as the current extraction layer. Return to the current extraction layer for feature extraction until the current extraction layer is the last extraction layer among the at least two extraction layers, and take the current feature as the emotion feature.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the depression quantification method based on multimodal feature adaptation as described in any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the multimodal feature-adaptive method for quantifying depression as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Emotion analysis method and device and electronic equipment

    CN115035438A