Multi-modal psychological data analysis method and system based on multi-dimensional consistency constraint
By introducing gradient direction consistency and spatial structure consistency strategies in multimodal psychological data analysis, the optimization deviation and noise problems between modes are solved, and more accurate mental health analysis is achieved.
Patent Information
- Application Number
- CN202510342204.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-18
AI Technical Summary
The existing multimodal psychological data analysis methods have deviations and interactive noise problems in the cross-modal optimization process, resulting in inconsistent information between modals and affecting the accuracy of mental health analysis.
Using a method based on multi-dimensional consistency constraints, the initial features of each modal are extracted through a feature extractor, and a common feature extractor with shared parameters is used to process the modal features, and combining gradient direction consistency and spatial structure consistency strategies to ensure the consistency of each modal gradient and spatial structure, and finally predict through the MLP layer.
The accuracy of multimodal psychological data analysis is improved, the inhibition of dominant modes is avoided, the information interaction and sample correlation between modes are enhanced, and the overall performance of the model is improved.
Smart Images

Figure CN120337123A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and relates to a method and system for psychological data analysis, in particular to a multi-modal psychological data analysis method and system based on multi-dimensional consistency constraints. Background Art
[0002] Mental health status affects cognitive function, behavioral response and social adaptation ability. Psychological data analysis plays an important role in the evolution of artificial intelligence systems. Multi-modal psychological data analysis has attracted great interest from researchers due to its wide application in different fields, such as clinical diagnosis, psychological intervention and health management. As Figure 1 (a) shows, multi-modal psychological data analysis predicts mental health status by integrating multiple modal data extracted from video clips of text, audio and visual information.
[0003] Text information data conveys mental health status through abstract words, sound information data conveys mental health status through fluctuating audio signals, and visual information data expresses mental health status through important facial expressions. These different information transmission methods lead to differences in the representation of mental health status in different modes. At the same time, multiple modes work together to serve the same mental health status. Existing multi-modal psychological data analysis methods construct modal separation encoders to maintain the uniqueness of each mode and construct modal invariant encoders to extract common information between modes. The above methods focus on obtaining mode-specific information for mutual learning and distillation between modes, while ignoring the differences between multiple modes under common goal learning.
[0004] The purpose of multi-modal learning is to obtain a unified representation of human mental health status from different modes. However, due to the heterogeneity of modes, the unified representation is biased. Modal heterogeneity leads to different update directions of modes, which tend to dominate the dominant mode (the mode that has the greatest impact on mental health analysis). Some studies have shown that the dominant mode interferes with other learning directions. In other words, during training, the dominant mode affects the gradient direction of other modes, resulting in sub-optimal or poor performance of other modes. As Figure 1 (b) shows, the optimization of vision and audio may stop before reaching the optimum. Therefore, these modes are not optimized in their respective optimal directions, and the learned unified information cannot fully reflect the consistency goals of all modes, which has a counterproductive effect on enhancing the representation of modes. In addition, as Figure 1As shown in (c), some methods attempt to facilitate interaction between modalities by aligning or extracting information from different modalities. However, these methods ignore that direct alignment or distillation under modality heterogeneity introduces interaction noise. There are inherent differences between different modalities, resulting in inconsistent data distributions and expressions. Therefore, this leads to a certain amount of noise in the interaction between modalities.
[0005] In summary, it is crucial to design a new method for multi-modal psychological data analysis to solve the above problems. Summary of the Invention
[0006] To solve the problems of cross-modal optimization bias and interaction noise encountered in multi-modal psychological data analysis and improve the accuracy of multi-modal mental health analysis, the present invention provides a multi-modal psychological data analysis method and system based on multi-dimensional consistency constraints.
[0007] The technical solution adopted by the method of the present invention is: a multi-modal psychological data analysis method based on multi-dimensional consistency constraints, comprising the following steps:
[0008] Step 1: For the psychological interview video of the person to be tested, extract the modal data x m from the interview video, including text data L, video data V, and audio data A; where m ∈ {L, V, A};
[0009] Step 2: Apply different feature extractors with non-shared parameters to process the three modal data respectively to obtain the initial features of the three modalities;
[0010] Step 3: Input the obtained initial features into a convolutional neural network to extract the unimodal specific information X m of each modality; retain the time information to make the output consistent in dimension;
[0011] Step 4: Input the unimodal specific feature X m of each modality into a common feature extractor φ θ with shared parameters θ to obtain the features of each modality C and the common psychological data information X
[0012] Step 5: Use the unimodal specific information X m and the common psychological data information X C as the fusion information X′ = [X L ; X V ; X A ; X C , input it into the MLP layer for prediction to obtain the multi-modal psychological data analysis result.
[0013] Preferably, in step 2, a text feature extractor is used to extract text feature data from the consultation video;
[0014] The text feature extractor consists of a word embedding layer and a Transformer encoder layer connected in sequence; the word embedding layer consists of word vector embedding, position encoding, and segment embedding, and is used to encode the input text sequence into a word vector representation while retaining position and sentence segment information; the Transformer encoder layer consists of a multi-head self-attention mechanism and a feed-forward neural network, where the multi-head self-attention mechanism is used to capture global dependencies in the text, and the feed-forward neural network is used to further extract high-level features.
[0015] Preferably, in step 2, a video feature extractor is used to extract video feature data from the consultation video;
[0016] The video feature extractor consists of a Patch segmentation layer, a Patch embedding layer, a Swin Transformer Block layer, and a hierarchical Swin Transformer encoder layer connected in sequence; the Patch segmentation layer divides the input image into small blocks of a preset size and converts the two-dimensional image into a sequence form for input; the Patch embedding layer maps each small block to a feature vector of a fixed dimension through a linear mapping layer to form an initial image feature representation; the hierarchical Swin Transformer encoder layer consists of several stages, and each stage consists of a Patch Merging layer and a Swin Transformer Block layer connected in sequence; the Patch Merging layer is used for downsampling to reduce the resolution of the feature map and adjust the number of channels at the same time to save computational effort and maintain information integrity; the Swin Transformer Block layer consists of a first LayerNorm layer, a W-MSA layer, a second LayerNorm layer, a first MLP layer, a third LayerNorm layer, a SW-MSA layer, a fourth LayerNorm layer, and a second MLP layer connected in sequence. Among them, after the input of the first LayerNorm layer is connected with the output of the W-MSA layer in a residual connection, it is input into the second LayerNorm layer; after the input of the second LayerNorm layer is connected with the output of the first MLP layer in a residual connection, it is input into the third LayerNorm layer; after the input of the third LayerNorm layer is connected with the output of the SW-MSA layer in a residual connection, it is input into the fourth LayerNorm layer; after the input of the fourth LayerNorm layer is connected with the output of the second MLP layer in a residual connection, it outputs.
[0017] Preferably, in step 2, an audio feature extractor is used to extract audio feature data from the consultation video;
[0018] The audio feature extractor consists of a feature extractor, a Transformer encoder, and a quantization module. The feature extractor consists of several 1D convolutional layers and is used to extract low-level time-frequency features from the original audio waveform. The Transformer encoder consists of multiple self-attention mechanisms and feed-forward networks and is used to model long-term dependencies and generate high-level speech representations. The quantization module uses the Gumbel Softmax method to achieve discretization, thus ensuring that the entire process can still be optimized end-to-end.
[0019] Preferably, in step 3, a convolutional neural network for text data feature extraction is used to extract the unimodal specific information of the initial text features; a convolutional neural network for video data feature extraction is used to extract the unimodal specific information of the initial video features; a convolutional neural network for audio data feature extraction is used to extract the unimodal specific information of the initial audio features.
[0020] One-dimensional convolutional networks are used for each modality data to keep the dimensions consistent, and then two convolutional layers are used as common feature encodings. The convolutional neural networks are all trained networks, and spatial consistency loss and gradient consistency loss are added during training.
[0021] Preferably, in step 4, the common feature extractor φ θ consists of two sequentially connected convolutional layers, and a Relu layer is set in the middle as a shared space.
[0022] Preferably, in step 4, the common psychological data information
[0023] Preferably, in step 4, the common feature extractor φ θ is a trained network;
[0024] The loss function used during training is:
[0025]
[0026] where,
[0027]
[0028] where, y i represents the emotion label of the i-th sample in the training sample set, where negative numbers indicate negative, 0 indicates neutral, and positive numbers indicate positive;
[0029]
[0030] Calculate the gradients of each sample on each modality where, represents the gradient with respect to the common feature extractor φθ The weight θ m The first derivative, where e represents the loss function; N is the number of training samples; a single gradient is the weight θ θ of the common feature extractor φ for the i-th sample from modality m ∈ {L, V, A} m The first derivative, so the gradient for each modality is Then calculate the gradient center of all modalities where |m| represents the number of modalities;
[0031]
[0032] First, calculate the structural correlation of samples within each modality; during the training of a batch, the correlation between each sample and other samples where is the real number field, represents an N×N matrix, <,> calculates the cosine similarity between samples; deduce the spatial structure matrix from the obtained structural correlation
[0033]
[0034] where τ is the temperature to expand the difference between correlations, () ≠1 represents removing the value 1 from the matrix, changing the size of the matrix from N×N to N×(N - 1); in Align the spatial structures of samples in multiple modalities to ensure a consistent representation for all modalities.
[0035] Preferably, before calculating the balanced gradient modality, normalize each modality gradient with L2 normalization ∥·∥2; at the same time, align the gradient direction of each modality with the balanced modality gradient direction using cosine similarity; use the Euclidean distance to align the spatial structures of samples in different modalities and calculate the average value of all samples as the optimization objective.
[0036] The technical solution adopted by the system of the present invention is: A multi-modal psychological data analysis system based on multi-dimensional consistency constraints includes:
[0037] One or more processors;
[0038] A storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the multi-modal psychological data analysis method based on multi-dimensional consistency constraints as described above.
[0039] Compared with the prior art, the beneficial effects of the present invention include:
[0040] (1) The present invention proposes a Gradient Direction Consistency strategy (GDC) to achieve unbiased optimization between modes. GDC constructs a balanced learning direction between different modalities to encourage maximizing the optimization of each modality, avoiding the suppression of the dominant modality.
[0041] (2) The present invention proposes a Spatial Structure Consistency strategy (SSC) to achieve effective interaction between modalities. By aligning the spatial structures of samples in different modes, SSC can not only learn the interaction information between modes but also learn the correlation between samples.
[0042] (3) The present invention conducts a large number of experiments, which strongly prove the effectiveness of this method in multi-modal mental health analysis. Compared with existing methods, this method achieves superior performance. Brief Description of the Drawings
[0043] The following uses embodiments and specific implementation manners to further illustrate the technical solutions of the present invention. Additionally, during the process of illustrating the technical solutions, some drawings are also used. For those skilled in the art, without creative efforts, other drawings and the intent of the present invention can also be obtained based on these drawings.
[0044] Figure 1 It is the principle of multi-modal mental data analysis in the background technology of the present invention, where yellow, green, and blue respectively represent text information, vision, and audio information. (a) is a process of multi-modal mental health analysis. (b) shows the gradient directions of different modes during the training process. (c) demonstrates the spatial structures of samples in different modes.
[0045] Figure 2 It is a schematic diagram of the network structure of the video feature extractor in the embodiment of the present invention;
[0046] Figure 3 It is a schematic diagram of the structure of the audio feature extractor in the embodiment of the present invention;
[0047] Figure 4 It is a schematic diagram of the training principle of the common feature extractor in the embodiment of the present invention, where Gradient and Structure Consistency (GSCon) learns more effective common information at the overall and individual levels. At the overall level, GDC makes the gradient update directions between modes more consistent. At the individual level, SSC is consistent with the spatial structures of samples in different modes;
[0048] Figure 5This is a schematic diagram of the similarity of the optimization direction between each modality and the central gradient during the training of the common feature extractor in an embodiment of the present invention; wherein the yellow, green, and blue rectangles respectively represent the directional similarity of the language, vision, and audio with the central gradient during the training process. The orange and purple lines represent the average similarity of the modalities in the basic network and the proposed GSCon. The horizontal axis represents the training time, and the vertical axis represents the similarity. The effect is best under color display;
[0049] Figure 6 A visualization diagram of the spatial structure of samples in an embodiment of the present invention; the yellow, green, and blue curves represent text, visual, and audio information, respectively. The spatial structure of samples in different modes in GSCon is more consistent than in the baseline;
[0050] Figure 7 A schematic diagram of learning common information in an embodiment of the present invention, wherein the horizontal axis represents the results of fusion information and common information on the baseline and GSCon respectively;
[0051] Figure 8 The temperature τ schematic diagram is discussed in the embodiment of the present invention, and the present invention ultimately uses the average comprehensive performance of ACC7, ACC2, and f1 as the main selection criterion. DETAILED DESCRIPTION
[0052] In order to facilitate those skilled in the art to understand and implement the present invention, the present invention is further described in detail below in conjunction with embodiments. It should be understood that the implementation examples described herein are only used to illustrate and explain the present invention and are not used to limit the present invention.
[0053] This embodiment provides a multimodal psychological data analysis method based on multi-dimensional consistency constraints, including the following steps:
[0054] Step 1: Extract the modal data x from the psychological consultation video of the subject m , including text data L, video data V and audio data A; where m∈{L,V,A};
[0055] Step 2: Apply different feature extractors with non-shared parameters to process the three modal data to obtain the initial features of the three modalities;
[0056] In one embodiment, a text feature extractor is used to extract text feature data in the medical consultation video;
[0057] The text feature extractor consists of a sequentially connected Embedding Layer and Transformer Encoder Layers; the Embedding Layer consists of TokenEmbeddings, Position Embeddings, and Segment Embeddings, and is used to encode the input text sequence into a word vector representation while retaining position and sentence segment information; the Transformer Encoder Layers consist of a Multi-Head Self-Attention and a Feed-Forward Neural Network, where the Multi-Head Self-Attention is used to capture global dependencies in the text, and the Feed-Forward Neural Network is used to further extract high-level features.
[0058] Please refer to Figure 2 , in one implementation, a video feature extractor is used to extract video feature data from the consultation video;
[0059] The video feature extractor consists of a sequentially connected Patch Partition Layer, Patch Embedding Layer, Swin Transformer Block layer, and Hierarchical Swin Transformer Encoder Layers. The Patch Partition Layer divides the input image into small patches (Patches) with a preset size of 4×4 and converts the two-dimensional image into a sequence form for input. The Patch Embedding Layer maps each patch to a feature vector of a fixed dimension through a Linear Embedding Layer to form an initial image feature representation. The Hierarchical Swin Transformer Encoder Layers consist of several stages, and each stage is composed of a sequentially connected Patch Merging layer and Swin Transformer Block layer. The Patch Merging layer is used for downsampling to reduce the resolution of the feature map and adjust the number of channels simultaneously to save computational effort and maintain information integrity. The Swin Transformer Block layer is composed of a sequentially connected first LayerNorm layer, W-MSA layer, second LayerNorm layer, first MLP layer, third LayerNorm layer, SW-MSA layer, fourth LayerNorm layer, and second MLP layer. Among them, after the input of the first LayerNorm layer is connected with the output of the W-MSA layer in a residual connection, it is input into the second LayerNorm layer. After the input of the second LayerNorm layer is connected with the output of the first MLP layer in a residual connection, it is input into the third LayerNorm layer. After the input of the third LayerNorm layer is connected with the output of the SW-MSA layer in a residual connection, it is input into the fourth LayerNorm layer. After the input of the fourth LayerNorm layer is connected with the output of the second MLP layer in a residual connection, it outputs.
[0060] Among them, W-MSA (Window Multi-Head Self-Attention) performs attention calculation within a local window, reducing the computational complexity; SW-MSA (Shifted Window Multi-Head Self-Attention) introduces cross-window information interaction through window sliding, enhancing the global modeling ability; MLP (Multi-Layer Perceptron) extracts non-linear features; LayerNorm performs layer normalization to maintain training stability; the residual connection helps the gradient flow and prevents degradation.
[0061] In one implementation, the video feature extractor is composed of a pre-trained MTCNN face detection algorithm and the OpenFace2.0 toolkit.
[0062] The MTCNN (Multi-task Cascaded Convolutional Networks) face detection algorithm consists of three cascaded networks (P-Net, R-Net, O-Net), which gradually screen the face regions in the image. P-Net (Proposal Network): Quickly generates a large number of face candidate boxes and outputs the face confidence and bounding box regression parameters. R-Net (Refine Network): Further screens the candidate boxes output by P-Net, removes the misdetected regions, and finely adjusts the bounding box positions. O-Net (Output Network): Finally confirms the face region and extracts the key point positions (such as eyes, nose, corners of the mouth, etc.). The OpenFace2.0 toolkit consists of facial landmark detection and tracking, head pose estimation, eye gaze estimation, and facial expression recognition, and is used to extract a set of 68 facial landmarks, 17 facial action units, head pose, head direction, and eye gaze.
[0063] The video feature extractor is a trained video feature extractor; the training set of the first video feature extractor uses the MIntRec dataset, and the training set of the second video feature extractor uses the CMU-MOSI and CH-SIMS datasets.
[0064] Please refer to Figure 3 , in one implementation, an audio feature extractor is used to extract the audio feature data in the interrogation video.
[0065] The audio feature extractor consists of a Feature Encoder, a Transformer Encoder, and a Quantization Module. The Feature Encoder is composed of several 1D Convolutional Layers and is used to extract low-level time-frequency features from the original audio waveform. The Transformer Encoder consists of multiple layers of Multi-Head Self-Attention and FeedForward Network and is used to model long-term dependencies and generate high-level speech representations. The Quantization Module uses the Gumbel Softmax method to achieve discretization, thereby ensuring that the entire process can still be optimized end-to-end. The specific implementation method includes introducing G codebooks, each codebook containing V discrete units. By selecting the corresponding vectors from multiple codebooks and concatenating these vectors together, a complete quantization representation q is finally formed.
[0066] In one implementation, the audio feature extractor consists of pre-information extraction and subsequent signal analysis. The pre-information extraction consists of pitch tracking, polarity detection, and glottal closure instant (GCI) determination and is used to obtain the core information required for pitch synchronous analysis. The subsequent signal analysis consists of spectral envelope estimation and formant tracking, sine modeling, glottal analysis, and phase processing to ensure high-performance output of the final analysis result.
[0067] The audio feature extractor is a trained audio feature extractor. The first type of audio feature extractor uses the MIntRec dataset for training, and the second type of audio feature extractor uses the CMU-MOSI and CH-SIMS datasets for training.
[0068] Step 3: Input the obtained initial features into a convolutional neural network to extract the unimodal-specific information X of each modality m ; Preserve the time information to make the output consistent in dimension;
[0069] In one implementation, a convolutional neural network for text data feature extraction is used to extract the unimodal-specific information of text initial features; a convolutional neural network for video data feature extraction is used to extract the unimodal-specific information of video initial features; a convolutional neural network for audio data feature extraction is used to extract the unimodal-specific information of audio initial features.
[0070] For each modality data, a one-dimensional convolutional network is used to keep the dimension consistent, and then two convolutional layers are used as a common feature encoder. The convolutional neural networks are all trained networks. During training, spatial consistency loss and gradient consistency loss are added to the loss function.
[0071] Step 4: Input the unimodal-specific feature X of each modality m into the common feature extractor φ with shared parameter θ θ to obtain the feature of each modality and the common mental health data information X of all modalities C ;
[0072] In one implementation, the common feature extractor φ θ consists of two convolutional layers connected in sequence, with a Relu layer set in the middle as the shared space.
[0073] In one implementation, the common mental health data information
[0074] In one implementation, see Figure 4 , the common feature extractor φ θ , is a trained network; based on the traditional optimization function, this embodiment adds constraints on the inconsistent multi-modal gradient directions and the inconsistent sample space structures. A sample consists of text, visual, and audio information, and the specific information of each modality is obtained through the feature extractor and the convolutional neural network. Subsequently, multiple specific features are sent to the common feature extractor to obtain the common features. During training, this embodiment adds constraints on the inconsistent multi-modal gradient directions and the inconsistent sample space structures, and adds the penalties brought by the inconsistent multi-modal gradient directions and the inconsistent sample space structures to the common feature extractor to obtain more effective common information.
[0075] The gradient direction consistency strategy of this embodiment is characterized in that, first, calculate the gradient of each cross-modal sample; then calculate the gradient center of all modalities; before calculating the balanced gradient modality, normalize each modality gradient with L2 normalization ∥·∥2; and avoid aligning the gradient directions in a pairwise manner; use the cosine similarity to align the gradient direction of each modality with the balanced modality gradient direction.
[0076] The space structure consistency strategy of this embodiment is characterized in that, first, put different modalities into the common extractor to obtain the common information; then, based on the common information, calculate the structural correlation of the samples within each modality, and eliminate the self-correlation of the samples from the obtained correlation matrix to obtain the space structure matrix of the samples under different modalities; finally, use the Euclidean distance to align the space structures of the samples under different modalities, and calculate the average value of all samples as the optimization target to ensure the consistent representation of all modalities.
[0077] The total loss function adopted during the training process is:
[0078]
[0079] Among them, α is the balance parameter;
[0080]
[0081] Among them, y i represents the sentiment label of the i-th sample in the training sample set, where negative numbers indicate negative, 0 indicates neutral, and positive numbers indicate positive;
[0082]
[0083] Calculate the gradient of each sample on each modality Among them, represents the weight θ θ for the common feature extractor φ m of the first derivative, l represents the loss function; N is the number of training samples; a single gradient is the first derivative of the weight θ θ for the common feature extractor φ m from the i-th sample of modality m ∈ {L, V, A}, so the gradient of each modality is Then calculate the gradient center of all modalities Among them, |m| represents the number of modalities;
[0084]
[0085] First, calculate the structural correlation of the samples within each modality; during the training of a batch, the correlation between each sample and other samples Among them, is the real number field, represents an N×N matrix, represents the spatial structure matrix of the samples in modality m ∈ {L, V, A}, <,> calculates the cosine similarity between samples; the self-correlation of the samples is eliminated from the obtained correlation matrix, and the structural correlation is deduced from the spatial structure matrix
[0086]
[0087] Among them, τ is the temperature, which is used to expand the difference between correlations; () ≠1 represents removing the value 1 from the matrix, changing the size of the matrix from N×N to N×(N - 1); in align the spatial structures of the samples in multiple modalities to ensure a consistent representation of all modalities.
[0088] Step 5: The unimodal specific information X m and the common mental health data information XC As the fused information X′ = [X L ; X V ; X A ; X C , it is input into the MLP layer for prediction to obtain the multi-modal psychological data analysis result.
[0089] This embodiment also provides a multi-modal psychological data analysis system based on multi-dimensional consistency constraints, including:
[0090] One or more processors;
[0091] A storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the multi-modal psychological data analysis method based on multi-dimensional consistency constraints.
[0092] The following further elaborates on the present invention through specific experiments.
[0093] The deep learning framework used in this experiment is Pytorch. The hardware environment of the experiment is an RTX3090 GPU with 24GB of memory.
[0094] Considering fair comparison and privacy issues, features are extracted from all modal information used in the experiment. In the MIntRec and CMU-MOSI data sets, text information is extracted using a pre-trained text feature extractor (hereinafter denoted as: BERT-base-uncased model). For audio information, an audio feature extractor (hereinafter denoted as: Wav2Vec2.0) is used, and the MIntRec data set is adopted. The video information in MIntRec is extracted using a video feature extractor (hereinafter denoted as: Swin-Transformer model), and each video frame in the visual data of the CMU-MOSI and CH-SIMS data sets is encoded by Facet. During the network training process, Adam is used as the optimizer in this experiment, and the learning rate is 1e-4. The evaluation of the test set is performed using the model that shows the best performance on the validation set during the training phase. The temperature τ and the balance parameter in the experiment are set to 0.07 and 0.5 respectively. In particular, the temperature τ and the balance parameter α are determined using the CMU-MOSI data set without additional adjustment to the CH-SIMS data set. In multi-modal intent understanding, the parameters τ and α are set to 0.03 and 0.1 respectively.
[0095] CMU-MOSI Dataset: This dataset is a multi-modal mental health analysis dataset, containing 2,199 monologue video clips, among which there are 1,284 training samples, 229 validation samples, and 686 test samples. Acoustic and visual information are obtained at sampling rates of 12.5hz and 15hz respectively. The sample labels in the CMU-MOSI dataset are marked with mental health state scores from -3 to 3, where -3 represents a strong negative state, 0 represents a normal state, and 3 represents a strong positive state.
[0096] CH-SIMS Dataset: This dataset is a multi-modal sentiment analysis dataset, sourced from movies, TV series, and variety shows, including 2,281 samples of movie review video clips. These clips are divided into training, validation, and test sets, containing 1,368, 456, and 457 samples respectively. The samples in the CH-SIMS dataset are labeled with sentiment scores from -1 to 1, using 5 intervals to allow for more fine-grained sentiment expression.
[0097] MIntRec Dataset: This dataset is a multi-modal intent understanding dataset, from the TV series "Superstore", to understand human intentions in the real world. It includes 2,224 samples, which consist of text, video, and audio. MIntRec is divided into training, validation, and test sets, consisting of 1,334, 445, and 445 samples respectively. The intent categories in MIntRec are divided into two levels: coarse-grained and fine-grained. The coarse-grained categories include "expressing emotions or attitudes" and "achieving goals", while the fine-grained categories contain 20 different intent classifications.
[0098] In addition, the experiment uses binary precision (ACC2), 7-class precision (ACC7), and F1 score as evaluation metrics for CMU-MOSI. The 7-class accuracy divides the mental health state range [-3, 3] into 7 parts: strongly negative, negative, weakly negative, neutral, weakly positive, positive, strongly positive. Binary precision divides the mental health state into two categories: positive state and negative state. The experiment uses binary precision (ACC2), 5-class precision (ACC5), and F1 score as evaluation metrics for CH-SIMS.
[0099] To verify the effectiveness of the present invention, the experiment compared the retrieval results of the present invention with existing methods (the best results are marked here with Bold ), and the results are shown in Tables 1 and 2 below.
[0100] Table 1 Comparison of CMU-MOSI Dataset. The average value is the average of ACC7, ACC2, f1
[0101]
[0102]
[0103] Table 2: Comparison of CH-SIMS datasets. The average value is the average of ACC5, ACC2, and f1
[0104]
[0105]
[0106] Tables 1 and 2 respectively compare and analyze the method proposed in the present invention with existing methods on the CMU-MOSI and CH-SIMS datasets. Existing methods include multimodal fusion methods and cross-modal learning methods. The experimental results show that the performance of the present invention has been significantly improved compared with existing methods. Whether it is a feature fusion method or a cross-modal attention method, its goal is to learn more superior common information and more unique unimodal information. However, these methods cannot fully solve the optimization deviation of common information and the spatial structure of samples between multiple modalities. In contrast, the method proposed in the present invention adjusts the multimodal gradient descent direction and the spatial structure of learning samples respectively. This unique method allows for a more effective learning process and improves the overall performance of the model. The consistency of the gradient direction ensures that the multimodal gradients converge in a more consistent direction during the learning process of common information. This alignment coordinates the multimodal information responses to a unified goal. On the other hand, the spatial structure consistency strategy mines more consistent sample association information from the mutual learning of multiple modalities. Making the spatial structure of samples more consistent can effectively avoid the feature difference problem caused by the expression differences between modalities and promote the development of multiple modalities towards a common goal. Whether it is the CMU-MOSI or CH-SIMS dataset, the present invention has achieved the best results under the comprehensive performance (average) of all indicators, and has also achieved better performance on finer indicators (ACC7 and ACC5).
[0107] This experiment further conducts ablation studies on the CMU-MOSI and CH-SIMS datasets. The results are shown in Table 3 below. In Table 3, Base is the baseline, GDC is the gradient direction consistency, and SSC is the spatial structure consistency. The average value represents the average of ACC7 (or ACC5), ACC2, and F1. The parameters in CH-SIMS are used in the CMU-MOSI experiment without any additional adjustment. The results of this experiment for each modality on the CMU-MOSI and CH-SIMS datasets are shown in Table 4. In Table 4, the average value is the average of ACC7, ACC2, and f1.
[0108] Table 3
[0109]
[0110]
[0111] Table 4
[0112]
[0113] To more strictly evaluate the effectiveness of the proposed GSCon method, a series of ablation experiments were conducted on the gradient direction consistency and spatial structure consistency strategies in this experiment. In addition to the ablation experiments, the prediction results in different ways, the learning of common information, and the parameter τ in spatial structure consistency were also discussed and analyzed.
[0114] Discuss different modules. The present invention improves the performance of the model at the overall level and the individual level. The performance improvement results in these two aspects are shown in the table. The results in Table 3 show that whether making the gradient descent directions (GDC) of each modality more consistent as a whole or adjusting different modalities separately (SSC) is beneficial to improving the performance of multimodal mental health analysis. In addition, although the performance improvement of gradient direction consistency is not as high as that of spatial structure consistency compared with the basic network, the combination of these two elements will produce a synergistic effect. This synergy proves the complementarity of the overall and the individual, effectively improving the overall performance of the model. In particular, the spatial structure consistency performs well, indicating that more information can be obtained by aligning the spatial structures of samples than directly aligning the information of samples.
[0115] Due to the heterogeneity of modalities, there are differences in the gradient update directions of different modalities, although they have the same goal. This results in the dominant modality suppressing the optimization of other modalities. To address this problem, the present invention proposes a new method aiming to improve gradient coherence - the cross - modal descent direction, thereby promoting the acquisition of multimodal mental health analysis goals. As Figure 5 shown, this experiment examined the alignment of the optimization directions between different modalities and the central gradient during the training process. Some significant observations were obtained from the analysis of this experiment: 1) From the comparison of (a) and (b), the method GSCon proposed by the present invention shows greater volatility in the average similarity, consistently restricting the range of gradient changes of each modality during the training process. 2) In (c), compared with the method of the present invention, the training process of the basic network terminates earlier, indicating that the dominant modality suppresses the optimization of the subordinate modality, and the method of the present invention effectively alleviates this problem. 3) In addition, the average similarity index in (c) shows that GSCON is always superior to the basic network, indicating that during the entire training process, the alignment degree between the gradient direction of a specific modality and the central gradient is higher. This means that the method proposed by the present invention enhances the consistency of the gradient descent direction.
[0116] Due to the heterogeneity of modalities, alignment or distillation between modalities generates interacting noise. To address this issue, the present invention proposes a spatial structure consistency strategy to align the spatial structures of samples under different modalities to avoid interacting noise. Figure 6 Shows the correlation between a single sample and four other samples. The horizontal axis represents the other samples, and the vertical axis represents the correlation strength with the sample on the left. The higher the value, the stronger the correlation. It can be clearly seen from the figure that in the basic network, different modalities exhibit variations in sample correlation. However, in GSCon, the correlation distribution among samples under the three modalities remains consistent, indicating that the spatial structures of samples under different modalities are consistent.
[0117] The prediction of the model proposed by the present invention is a combination of specific information and common information of multiple modalities. To analyze the role of each modality in more detail, this experiment shows the output results of each modality in Table 4. Initially, in the basic network, text information dominated the performance of the model, while the performance of the other two modalities was insufficient. However, in the proposed method, visual, audio, and common information all exhibited commendable performance. In addition, it can be seen that the common information promoted the learning of each modality. Whether it is text, visual information, or audio information, each piece of information showed satisfactory results in the final prediction. Finally, the common information in the output results performed well in ACC7, which requires more fine-grained information. During the process of extracting common information, multiple different modality information effectively supplemented this information.
[0118] Common information plays a key role in multi-modal mental health analysis tasks, but its learning ability is often overlooked in most studies. Figure 7 Shows the role of common information in the overall performance. First, it can be inferred from the basic network that the overall prediction performance without common information slightly exceeds that with shared information (71.68% vs. 71.65%). This indicates that common information does not make a significant contribution to the final mental health analysis in the basic network. Together with Table 4, it is obvious that the performance of the basic network is mainly affected by text information. In contrast, the common information in the method of the present invention not only improves the final analysis performance but also promotes the expression of other modalities in mental health analysis. The overall performance of GSCon with common information is better than the prediction without common information. Combining Table 4, it is found that the learning of common information and the performance of other modalities have been improved.
[0119] This experiment further discusses the parameter τ in the spatial structure consistency of samples under different modalities. The experimental results of each parameter are as Figure 8As shown. The method finally takes the average comprehensive performance of ACC2, f1, and ACC7 as the main selection criterion. It can be seen from the figure that when the parameter τ is 0.09, although ACC7 performs better, the performance of ACC2 and F1 starts to decline. This indicates that fine-grained information affects the performance of binary classification and is not conducive to the expression of the overall model.
[0120] Multimodal intent understanding and analysis are inherently subjective processes that require comprehensive judgment and integration of information from various modalities. Please see Table 5, the comparison of the MIntRec dataset. The average values represent the averages of ACC, F1, Pre, and Rec. * indicates no data augmentation, which is consistent with our work. The best results are shown in bold.
[0121] Table 5
[0122]
[0123] The results in Table 5 demonstrate the effectiveness of the method of the present invention in multimodal intent understanding. The results show that the method of the present invention provides significant advantages in this field, highlighting its potential to improve performance in complex, multimodal scenarios. It should be noted that multimodal intent understanding poses significant challenges in data annotation, often leading to imbalanced datasets. The imbalance of data is not conducive to the performance of the ACC (accuracy) metric. In this work, the present invention focuses on achieving balance between different modalities rather than solving data imbalance. Therefore, the performance of ACC may not be optimal. In addition, the model shows good performance in terms of F1 score, Pre (precision), and Rec (recall) metrics, indicating its excellent overall effectiveness even without solving data imbalance.
[0124] The present invention solves the optimization bias caused by the dominant modality and enhances the spatial structure of samples between modalities. The method of the present invention considers both the overall level and the individual level. At the overall level, the present invention proposes a consistency in the gradient direction, aiming to align the gradient descent directions of multiple modalities with the balanced gradient direction. This alignment promotes the consistency of the gradient directions across multiple modalities and facilitates unbiased learning of more stable common information. At the individual level, the present invention introduces the consistency of the spatial structure, which can mutually learn the spatial structures of samples across multiple modalities. This structural mutual learning can simultaneously learn sample knowledge and sample association knowledge from different modalities. The present invention evaluates the method of the present invention on the multimodal datasets CMU-MOSI and CH-SIMS, and it proves that the gradient and structure consistency strategies achieve commendable and comparable performance.
[0125] The multi-modal mental health data analysis technology proposed by the present invention combines facial expressions, voice, and text data, and is widely used in fields such as clinical diagnosis, personalized treatment, remote monitoring, mental health education, and intelligent hardware and AI applications. It can help early identify mental problems, develop customized treatment plans, achieve remote psychological support, and enhance the user experience. In addition, this technology also plays an important role in preventing mental illnesses, emotion management, and social support platforms, promoting the innovation and development of mental health services.
[0126] It should be understood that the above-described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. In addition, the technical features in each embodiment or individual embodiment provided by the present invention can be arbitrarily combined with each other to form a feasible technical solution. This combination is not restricted by the order of steps and / or the mode of structural composition, but must be based on what can be achieved by those of ordinary skill in the art. When the combination of technical solutions results in contradictions or cannot be achieved, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0127] It should be understood that the above description of the preferred embodiment is relatively detailed, and thus it should not be considered as a limitation on the protection scope of the present invention's patent. Those of ordinary skill in the art, under the inspiration of the present invention and without departing from the scope protected by the claims of the present invention, can also make substitutions or modifications, which all fall within the protection scope of the present invention. The scope of protection requested by the present invention shall be subject to the appended claims.
Claims
1. A multi-modal psychological data analysis method based on multi-dimensional consistency constraints, characterized in that Including the following steps: Step 1: For the psychological interview video of the person to be tested, extract the modal data x in the interview video m , including text data L, video data V, and audio data A; where m ∈ {L, V, A}; Step 2: Process the three-modal data using different parameter-unshared feature extractors respectively to obtain the initial features of the three modalities; Step 3: Input the obtained initial features into a convolutional neural network to extract the unimodal specific information X of each modality m ; retain the time information to make the output consistent in dimension; Step 4: Input the unimodal-specific feature X of each modality m into the common feature extractor φ with shared parameter θ θ to obtain the feature of each modality and the common psychological data information X of all modalities C ; Step 5: Take the unimodal specific information X m and the common psychological data information X C as the fusion information X′ = [X L ; X V ; X A ; X C , input the MLP layer for prediction to obtain the multi-modal psychological data analysis results.
2. The multi-modal psychological data analysis method based on multi-dimensional consistency constraints according to claim 1, characterized in that: In Step 2, a text feature extractor is used to extract text feature data from the consultation video; The text feature extractor consists of a word embedding layer and a Transformer encoder layer connected in sequence; the word embedding layer consists of word vector embedding, position encoding, and segment embedding, and is used to encode the input text sequence into a word vector representation while retaining position and sentence segment information; the Transformer encoder layer consists of a multi-head self-attention mechanism and a feed-forward neural network, where the multi-head self-attention mechanism is used to capture global dependencies in the text, and the feed-forward neural network is used to further extract high-level features.
3. The multimodal psychological data analysis method based on multi-dimensional consistency constraints according to claim 1, wherein: In Step 2, a video feature extractor is used to extract video feature data from the consultation video; The video feature extractor consists of a Patch segmentation layer, a Patch embedding layer, a Swin Transformer Block layer, and a hierarchical Swin Transformer encoder layer connected in sequence; the Patch segmentation layer divides the input image into small blocks of a preset size and converts the two-dimensional image into a sequence form for input; the Patch embedding layer maps each small block to a feature vector of a fixed dimension through a linear mapping layer to form an initial image feature representation; the hierarchical Swin Transformer encoder layer consists of several stages, and each stage consists of a Patch Merging layer and a Swin Transformer Block layer connected in sequence; The Patch Merging layer is used for downsampling to reduce the resolution of the feature map and adjust the number of channels at the same time to save computational effort and maintain information integrity; the Swin Transformer Block layer consists of a first LayerNorm layer, a W-MSA layer, a second LayerNorm layer, a first MLP layer, a third LayerNorm layer, a SW-MSA layer, a fourth LayerNorm layer, and a second MLP layer connected in sequence. Among them, after the input of the first LayerNorm layer is connected to the output of the W-MSA layer with a residual connection, it is input into the second LayerNorm layer; after the input of the second LayerNorm layer is connected to the output of the first MLP layer with a residual connection, it is input into the third LayerNorm layer; after the input of the third LayerNorm layer is connected to the output of the SW-MSA layer with a residual connection, it is input into the fourth LayerNorm layer; after the input of the fourth LayerNorm layer is connected to the output of the second MLP layer with a residual connection, it outputs.
4. The multi-modal psychological data analysis method based on multi-dimensional consistency constraints according to claim 1, wherein: In Step 2, an audio feature extractor is used to extract audio feature data from the consultation video; The audio feature extractor consists of a feature extractor, a Transformer encoder, and a quantization module. The feature extractor consists of several 1D convolutional layers and is used to extract low-level time-frequency features from the original audio waveform. The Transformer encoder consists of multiple self-attention mechanisms and feed-forward networks and is used to model long-term dependencies and generate high-level speech representations. The quantization module uses the Gumbel Softmax method to achieve discretization, thus ensuring that the entire process can still be optimized end-to-end.
5. The multi-modal psychological data analysis method based on multi-dimensional consistency constraints according to claim 1, characterized in that: In step 3, a convolutional neural network for text data feature extraction is used to extract the unimodal specific information of the initial text features; a convolutional neural network for video data feature extraction is used to extract the unimodal specific information of the initial video features; a convolutional neural network for audio data feature extraction is used to extract the unimodal specific information of the initial audio features. Each type of modal data uses a one-dimensional convolutional network to keep the dimensions consistent, and then two convolutional layers are used as a common feature encoder. The convolutional neural networks are all trained networks. During training, the loss function includes spatial consistency loss and gradient consistency loss.
6. The multi-modal psychological data analysis method based on multi-dimensional consistency constraints according to claim 1, wherein: In step 4, the common feature extractor φ θ consists of two sequentially connected convolutional layers, with a Relu layer set in the middle as the shared space.
7. The multi-modal psychological data analysis method based on multi-dimensional consistency constraints according to claim 1, characterized in that: In step 4, the common psychological data information 8. The multimodal psychological data analysis method based on multi-dimensional consistency constraints according to any one of claims 1-7, characterized in that: In step 4, the common feature extractor φ θ , is a trained network; The loss function used during training is: where α is a balance parameter. where y i represents the sentiment label of the i-th sample in the training sample set, where a negative number indicates negative, 0 indicates neutral, and a positive number indicates positive; Calculate the gradients of each sample on each modality m ∈ {L, V, A}; where represents the first-order derivative of the weight θ θ of the common feature extractor φ m , represents the loss function; N is the number of training samples; the single gradient is the first-order derivative of the weight θ θ of the common feature extractor φ m from the i-th sample of modality m ∈ {L, V, A}, so the gradient of each modality is Then calculate the gradient centers of all modalities where |m| represents the number of modalities; First, calculate the structural correlation of samples within each modality; during the training of a batch, the correlation between each sample and other samples wherein is the real number field, represents an N×N matrix, and <,> calculates the cosine similarity between samples; deduce the spatial structure matrix from the obtained structural correlation where τ is the temperature to expand the difference in correlation, () ≠1 denotes removing the value 1 from the matrix, changing the size of the matrix from N×N to N×(N - 1); align the spatial structures of samples in multiple modalities of to ensure a consistent representation across all modalities.
9. The multi-modal psychological data analysis method based on multi-dimensional consistency constraints according to claim 8, wherein: Before calculating the balanced gradient modality, each modal gradient is normalized using L2 normalization ∥·∥2. At the same time, the cosine similarity is used to align the gradient direction of each modality with the balanced modality gradient direction. The Euclidean distance is used to align the spatial structures of samples in different modalities, and the average value of all samples is calculated as the optimization target.
10. A multi-modal psychological data analysis system based on multi-dimensional consistency constraints, characterized in that, Includes: One or more processors; A storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the multi-modal psychological data analysis method based on multi-dimensional consistency constraints as described in any one of claims 1 to 9.