Multi-modal emotion discrimination method based on human body postures
By integrating multimodal emotion discrimination method with visual and language modalities, using the comparative learning of video and text features, the problems of low accuracy and poor generalization of single mode recognition are solved, achieving more efficient emotion recognition and stronger adaptability.
Patent Information
- Application Number
- CN202510416952.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, in the emotion recognition based on human posture, the recognition accuracy of single modal information is difficult to improve, and the model is poor generalization, especially when facing different emotions or scenarios.
The multimodal emotion discrimination method is used to integrate vision and language modality, and the alignment and discrimination of multimodal features is achieved through comparative learning. The video encoding module and text encoding module are used to extract features, and the similarity calculation is performed through the comparative learning module to output the emotional discrimination results.
It improves the accuracy and generalization ability of the model, reduces the dependence on the data set, reduces the computational complexity, improves the accuracy of emotion recognition and the ability to adapt to different emotional scenarios.
Smart Images

Figure CN120340132A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of emotion discrimination in computer vision, and particularly relates to a multi-modal emotion discrimination method based on human posture. Background Art
[0002] Traditional machine learning algorithms usually select and extract posture emotion features manually. Since this process requires a large number of experiments, it is very time-consuming; moreover, the recognition effect depends on the selection and extraction of posture emotion features; therefore, the manual method has fewer applicable scenarios and poor generalization. The emotion recognition algorithm based on deep learning can obtain complex and deep posture emotion features through learning a large number of samples, but for the emotion recognition algorithm relying on single posture information, it is difficult to improve the accuracy.
[0003] To address the above problems, the prior art, on the basis of only using posture features for emotion recognition, introduces other visual modal information to assist in recognition, such as facial expressions or motion gaits. Although this improves the accuracy of emotion recognition to a certain extent, the model has a large amount of data and a large number of computational parameters; and when facing different emotions or scenarios, the model performance will decline due to a large variety of emotion types or data imbalance, and the generalization is poor. Summary of the Invention
[0004] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a multi-modal emotion discrimination method based on human posture, which effectively integrates visual and language modalities, and realizes the alignment and discrimination of multi-modal features through contrastive learning, can reduce data dependence, and improve the model accuracy and generalization ability.
[0005] To achieve the above purpose, the present invention is implemented by the following technical solutions:
[0006] The present invention provides a multi-modal emotion discrimination method based on human posture, and the method includes:
[0007] Obtain the human posture video sequence data to be discriminated;
[0008] Concatenate the pre-constructed text prompt template with various emotion labels to obtain text sequence data;
[0009] Input the video sequence data and the text sequence data into a pre-constructed and trained multi-modal emotion discrimination network model, where the multi-modal emotion discrimination network model includes a video encoding module, a video frame feature fusion module, a first linear module, a text encoding module, a text feature enhancement module, a second linear module, and a contrastive learning module;
[0010] Use the video encoding module to segment, encode, and extract features from the video sequence data to generate a series of video frame features;
[0011] Input a series of video frame features into the video frame feature fusion module for fusion operations and the first linear module for normalization processing to generate video fusion features;
[0012] Use the text encoding module to extract features from the text sequence data to generate text primary features;
[0013] Input the text primary features and a series of video frame features into the text feature enhancement module for fusion operations and the second linear module for normalization processing to generate text enhanced features;
[0014] Use the contrastive learning module to calculate the similarity between the video fusion features and the text enhanced features and output an emotion discrimination result.
[0015] Furthermore, the video encoding module includes a linear projection layer, a first token layer, a first image feature extraction module, a first root mean square normalization layer for images, a first temporal transformation layer, a second image feature extraction module, a second root mean square normalization layer for images, a first fuser, a second temporal transformation layer, a third image feature extraction module, a third root mean square normalization layer for images, a second fuser, a third temporal transformation layer, a fourth image feature extraction module, a fourth root mean square normalization layer for images, a third fuser, a fourth temporal transformation layer, a fifth image feature extraction module, a fifth root mean square normalization layer for images, a fourth fuser, a fifth temporal transformation layer, a sixth image feature extraction module, a sixth root mean square normalization layer for images, a fifth fuser, and a sixth temporal transformation layer;
[0016] The method of using the video encoding module to segment, encode, and extract features from the video sequence data to generate a series of video frame features includes:
[0017] After segmenting the frame images in the video sequence data into image patches, input them into the linear projection layer for transformation to obtain image patch encodings;
[0018] Use the first token layer to add position tokens and class tokens to the image patch encodings to obtain an encoded image patch sequence;
[0019] Input the encoded image patch sequence into the first image feature extraction module for feature extraction and the first root mean square normalization layer for images for normalization processing to obtain first image normalized features;
[0020] The first image normalized features are successively introduced with temporal information through the first time conversion layer, feature extracted by the second image feature extraction module, and normalized by the second image root mean square normalization layer to obtain the second image normalized features;
[0021] After the second image normalized features and the first image normalized features are weighted and fused by the first fuser, they are successively introduced with temporal information through the second time conversion layer, feature extracted by the third image feature extraction module, and normalized by the third image root mean square normalization layer to obtain the third image normalized features;
[0022] Repeat the above steps of weighted fusion, introducing temporal information, feature extraction, and normalization until the sixth image normalized features are obtained;
[0023] After the sixth image normalized features and the fifth image normalized features are weighted and fused by the fifth fuser, they are introduced with temporal information through the sixth time conversion layer, and a series of video frame features are finally obtained.
[0024] Further, the first image feature extraction module to the sixth image feature extraction module all include a linear layer, a first SiLU activation function, a forward one-dimensional convolutional layer, a forward state space module, a second SiLU activation function, a backward one-dimensional convolutional layer, a backward state space module, a linear layer, a third SiLU activation function, a first multiplier, a second multiplier, a sixth fuser, a fourth SiLU activation function, and a linear layer;
[0025] The first image feature extraction module to the sixth image feature extraction module perform the following processing steps on the input image features:
[0026] The input image features are respectively input into the linear layer and the linear layer for transformation to generate a first input feature and a second input feature;
[0027] The first input feature is successively processed by the first SiLU activation function and the forward one-dimensional convolutional layer, and the dependency relationship is captured through the forward state space module to obtain a forward output feature; and, the first input feature is successively processed by the second SiLU activation function and the backward one-dimensional convolutional layer, and the dependency relationship is captured through the backward state space module to obtain a backward output feature;
[0028] Meanwhile, after performing a non-linear transformation on the second input feature using the third SiLU activation function, the forward output feature and the backward output feature are respectively gated by the first multiplier and the second multiplier to obtain a forward extracted feature and a backward extracted feature;
[0029] The forward extracted feature and the backward extracted feature are fused by the sixth fuser, and then sequentially passed through the fourth SiLU activation function and the linear layer for transformation to finally obtain an image extraction feature.
[0030] Further, the video frame feature fusion module includes a second tagging layer and a Transformer encoding layer;
[0031] The method of sequentially inputting a series of video frame features into the video frame feature fusion module for fusion operation and the first linear module for normalization processing to generate a video fusion feature includes:
[0032] Using the second tagging layer to add a time tag before each video frame feature in a series of video frame features to obtain video frame features with time tags;
[0033] The video frame features with time tags are subjected to feature fusion in the time sequence dimension through the Transformer encoding layer to obtain fused video frame features;
[0034] The fused video frame features are linearly mapped through the first linear module to finally obtain a video fusion feature.
[0035] Further, the text encoding module includes a first text feature extraction module, a first text root mean square normalization layer, a second text feature extraction module, a second text root mean square normalization layer, a seventh fuser, a third text feature extraction module, a third text root mean square normalization layer, an eighth fuser, a fourth text feature extraction module, a fourth text root mean square normalization layer, a ninth fuser, a fifth text feature extraction module, a fifth text root mean square normalization layer, a tenth fuser, a sixth text feature extraction module, a sixth text root mean square normalization layer, and an eleventh fuser;
[0036] The method of using the text encoding module to extract features from the text sequence data to generate text primary features includes:
[0037] The text sequence data is sequentially input into the first text feature extraction module for feature extraction and the first text root mean square normalization layer for normalization processing to obtain a first text normalized feature;
[0038] The first text normalized feature is successively subjected to feature extraction by the second text feature extraction module and normalization processing by the second text root mean square normalization layer to obtain the second text normalized feature;
[0039] After the second text normalized feature and the first text normalized feature are weighted and fused by the seventh fuser, they are successively subjected to feature extraction by the third text feature extraction module and normalization processing by the third text root mean square normalization layer to obtain the third text normalized feature;
[0040] Repeat the above steps of weighted fusion, feature extraction, and normalization processing until the sixth text normalized feature is obtained;
[0041] The sixth text normalized feature and the fifth text normalized feature are weighted and fused by the eleventh fuser to finally obtain the text primary feature.
[0042] Further, the first text feature extraction module to the sixth text feature extraction module all include an input linear layer, a convolutional layer, a fifth SiLU activation function, a state space module, a sixth SiLU activation function, a twelfth fuser, a seventh SiLU activation function, and an output linear layer;
[0043] The first text feature extraction module to the sixth text feature extraction module perform the following processing steps on the input text feature:
[0044] Input the input text feature into the input linear layer for conversion to obtain an initial word vector;
[0045] After the initial word vector is successively subjected to multi-dimensional feature fusion by the convolutional layer, introduction of non-linear features by the fifth SiLU activation function, and capture of dependency relationships by the state space module, an intermediate word vector is obtained;
[0046] After the initial word vector is non-linearly transformed by the sixth SiLU activation function, it is residually connected with the intermediate word vector through the twelfth fuser to obtain a fused word vector;
[0047] The fused word vector is successively subjected to dimension conversion by the seventh SiLU activation function and the output linear layer to finally obtain the text extraction feature.
[0048] Further, the text feature enhancement module includes a multi-head attention mechanism and a feed-forward neural network;
[0049] The method of inputting the text primary feature and a series of video frame features into the text feature enhancement module for fusion operation and the second linear module for standardization processing to generate the text enhancement feature includes:
[0050] Using the primary text features as query signals, and a series of video frame features as keys and values, and processing them sequentially through the multi-head attention mechanism and the feed-forward neural network to obtain text-enhanced prompts;
[0051] Among them, the expression of the text-enhanced prompt is:
[0052] ,
[0053] ,
[0054] In the formula, is the primary text feature; is a series of video frame features; is the multi-head attention mechanism; is the feed-forward neural network; is the text-enhanced prompt;
[0055] After weighted fusion of the text-enhanced prompt and the primary text feature, and then linear mapping through the second linear module, the text-enhanced feature is finally obtained.
[0056] Furthermore, the method for calculating the similarity between the video fusion feature and the text-enhanced feature using the contrast learning module and outputting the emotion discrimination result includes:
[0057] Performing contrast learning on the video fusion feature and the text-enhanced feature in the mapped feature space;
[0058] Defining a similarity matrix and calculating the cosine similarity between the video fusion feature and the text-enhanced feature;
[0059] Among them, the expression of the similarity matrix is:
[0060] ,
[0061] In the formula, represents the cosine similarity between the th text-enhanced feature and the th video fusion feature pair; is the th text-enhanced feature vector; is the th video fusion feature vector;
[0062] Outputting the emotion label corresponding to the maximum similarity, which is the emotion discrimination result corresponding to the human body posture video to be discriminated.
[0063] Furthermore, the training method of the multi-modal emotion discrimination network model includes:
[0064] Construct a human body pose emotion dataset;
[0065] Divide the human body pose emotion dataset into a video training set and a video test set according to a ratio of 8:2;
[0066] Extract pose description words according to the emotion labels in the human body pose emotion dataset, and construct a text prompt template;
[0067] Concatenate the text prompt template with various emotion labels to obtain text training data;
[0068] Use the video training set and the text training data to train a pre-constructed multi-modal emotion discrimination network model, use the video test set and the text training data to test the trained multi-modal emotion discrimination network model, and use the multi-modal emotion discrimination network model with the best test result as the finally trained multi-modal emotion discrimination network model.
[0069] Furthermore, the training method of the multi-modal emotion discrimination network model further includes:
[0070] Construct a contrast loss function according to the losses of positive sample pairs and negative sample pairs;
[0071] Optimize the multi-modal emotion discrimination network model with the goal of minimizing the value of the contrast loss function;
[0072] Among them, the loss of the positive sample pair is:
[0073] ,
[0074] ,
[0075] In the formula, represents the loss between a text and a video pair belonging to the same emotion; is the Sigmoid activation function; represents the th text enhancement feature and the th cosine similarity between video fusion feature pairs; is the th text enhancement feature vector; is the th video fusion feature vector;
[0076] The loss of the negative sample pair is:
[0077] ,
[0078] In the formula, represents the loss between a text and a video pair belonging to different emotions; is a hyperparameter. When the cosine similarity is small, the loss of the negative sample pair is small. When the cosine similarity is large, the loss of the negative sample pair increases;
[0079] The expression of the contrast loss function is as follows:
[0080] ,
[0081] In the formula, is the contrast loss function; indicates that the th text and video pair belongs to the positive sample pair, indicates that the th text and video pair belongs to the negative sample pair.
[0082] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:
[0083] The technical solution provided by the present invention is to use a trained multi-modal emotion discrimination network model to discriminate the emotion of a human body posture video. By introducing the text modality auxiliary learning in the human body posture emotion discrimination task and using the text feature enhancement module to effectively fuse the visual modality and the language modality, it can efficiently capture rich emotion features, improve the accuracy of model discrimination, reduce the dependence of the model on the data set, and effectively reduce the computational complexity of the model. At the same time, through the contrast learning model, the alignment and discrimination of the visual modality and the language modality are realized, thereby improving the generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0084] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0085] Figure 1 is a flowchart of a multi-modal emotion discrimination method based on human body posture provided by an embodiment of the present invention;
[0086] Figure 2 is a schematic structural diagram of a multi-modal emotion discrimination network model provided by an embodiment of the present invention;
[0087] Figure 3 is a schematic structural diagram of a video coding module provided by an embodiment of the present invention;
[0088] Figure 4 is a schematic structural diagram of an image feature extraction module provided by an embodiment of the present invention;
[0089] Figure 5It is a schematic structural diagram of a video frame feature fusion module provided by an embodiment of the present invention;
[0090] Figure 6 It is a schematic structural diagram of a text encoding module provided by an embodiment of the present invention;
[0091] Figure 7 It is a schematic structural diagram of a text feature extraction module provided by an embodiment of the present invention;
[0092] Figure 8 It is a schematic structural diagram of a text feature enhancement module provided by an embodiment of the present invention. Specific implementation manner
[0093] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.
[0094] Embodiment 1:
[0095] This embodiment provides a multi-modal emotion discrimination method based on human postures. As Figure 1 shown, it is a flowchart of the method provided in this embodiment, which mainly includes the following steps:
[0096] Step S1: Obtain the human posture video sequence data to be discriminated;
[0097] Step S2: Concatenate the pre-constructed text prompt template with various emotion labels to obtain text sequence data;
[0098] Step S3: Input the video sequence data and the text sequence data into the pre-constructed and trained multi-modal emotion discrimination network model;
[0099] Step S4: Use the video encoding module to segment, encode, and extract features from the video sequence data to generate a series of video frame features ;
[0100] Step S5: Input a series of video frame features into the video frame feature fusion module for fusion operation and the first linear module for normalization processing to generate video fusion features ;
[0101] Step S6: Use the text encoding module to extract features from the text sequence data to generate text primary features ;
[0102] Step S7: The text primary features and a series of video frame features Input into the text feature enhancement module in sequence for fusion operation, and then into the second linear module for normalization to generate text enhanced features ;
[0103] Step S8: Use the contrastive learning module to calculate the similarity between the video fusion features and the text enhanced features and output the emotion discrimination result.
[0104] It should be noted that the human body posture video obtained in step S1 refers to a video that records the facial emotion changes and the upper body limb movements or gesture movements of the human body. After obtaining the video, the frame images to be discriminated can be extracted from the video, and multiple frame images can show the facial emotion changes in time sequence. In addition, since the emotion corresponding to the video cannot be determined during emotion discrimination, in step S2, various different emotion labels can be concatenated with the text prompt templates that have been constructed during the training process to generate different text data respectively. In practical applications, the emotion labels can include some or all of those used during the training process, or can include those not used during the training process; the text prompt templates can be the same or different.
[0105] In this embodiment, as Figure 2 shown, it is a schematic structural diagram of the multi-modal emotion discrimination network model provided in this embodiment, mainly including a video encoding module, a video frame feature fusion module, a first linear module, a text encoding module, a text feature enhancement module, a second linear module, and a contrastive learning module. Among them, the video encoding module is serially connected to the video frame feature fusion module and the first linear module in sequence. The output end of the text encoding module and the output end of the video encoding module are commonly connected to the input end of the text feature enhancement module. The output end of the text feature enhancement module is connected to the input end of the second linear module, and the output end of the second linear module and the output end of the first linear module are commonly connected to the input end of the contrastive learning module.
[0106] Specifically, the video sequence data is segmented, encoded, and feature-extracted through the video encoding module to obtain a series of video frame features ; then, through the video frame feature fusion module for fusion operation, the fused video frame features are obtained; then, through the first linear module for normalization, the video fusion features are obtained. The text sequence data is feature-extracted through the text encoding module to obtain the text primary features corresponding to various emotion labels ; then, through the text feature enhancement module for fusion operation with a series of video frame features , the text fusion features are obtained; then, through the second linear module for normalization, the text enhanced features Finally, the contrastive learning module is used to calculate the similarity between the video fusion features and the text-enhanced features and output the emotion label corresponding to the maximum similarity, which is the emotion discrimination result corresponding to the human body posture video to be discriminated.
[0107] Furthermore, as Figure 3 shown, it is a schematic structural diagram of the video encoding module provided in this embodiment, which mainly includes a linear projection layer, a first marking layer, a first image feature extraction module, a first image root mean square normalization layer, a first time conversion layer, a second image feature extraction module, a second image root mean square normalization layer, a first fusion device, a second time conversion layer, a third image feature extraction module, a third image root mean square normalization layer, a second fusion device, a third time conversion layer, a fourth image feature extraction module, a fourth image root mean square normalization layer, a third fusion device, a fourth time conversion layer, a fifth image feature extraction module, a fifth image root mean square normalization layer, a fourth fusion device, a fifth time conversion layer, a sixth image feature extraction module, a sixth image root mean square normalization layer, a fifth fusion device, and a sixth time conversion layer.
[0108] Among them, the linear projection layer is sequentially connected in series with the first marking layer, the first image feature extraction module, and the first image root mean square normalization layer; the first image root mean square normalization layer is sequentially connected in series with the first time conversion layer, the second image feature extraction module, and the second image root mean square normalization layer, and the output end of the second image root mean square normalization layer and the output end of the first image root mean square normalization layer are commonly connected to the input end of the first fusion device; the output end of the first fusion device is sequentially connected in series with the second time conversion layer, the third image feature extraction module, and the third image root mean square normalization layer, and the output end of the third image root mean square normalization layer and the output end of the second image root mean square normalization layer are commonly connected to the input end of the second fusion device; and so on, the output end of the fourth fusion device is sequentially connected in series with the fifth time conversion layer, the sixth image feature extraction module, and the sixth image root mean square normalization layer, and the output end of the sixth image root mean square normalization layer and the output end of the fifth image root mean square normalization layer are commonly connected to the input end of the fifth fusion device; the output end of the fifth fusion device is connected to the input end of the sixth time conversion layer.
[0109] Next, in combination with Figure 3 , a further detailed description will be given on how to use the video encoding module to segment, encode, and extract features from the video sequence data in step S4 to generate a series of video frame features :
[0110] After splitting the input image of each frame in the video sequence data into image blocks , each image block Input to the linear projection layer for linear transformation to obtain the image patch encoding ;
[0111] Use the first marking layer to encode each image patch Add a position marker and a class marker , to obtain the encoded image patch sequence , where the class marker is used to represent the emotion category represented by the image patch sequence;
[0112] Input each encoded image patch sequence sequentially into the first image feature extraction module for feature extraction and the first image root mean square normalization layer for normalization processing to obtain the first image normalized feature ;
[0113] The first image normalized feature sequentially passes through the first time conversion layer to introduce the temporal information between frames, the second image feature extraction module for feature extraction, and the second image root mean square normalization layer for normalization processing to obtain the second image normalized feature ;
[0114] The second image normalized feature and the first image normalized feature are weighted and fused through the first fuser to achieve residual connection, and then sequentially pass through the second time conversion layer to introduce the temporal information between frames, the third image feature extraction module for feature extraction, and the third image root mean square normalization layer for normalization processing to obtain the third image normalized feature ;
[0115] Repeat the above steps of weighted fusion, introducing the temporal information between frames, feature extraction, and normalization processing until the sixth image normalized feature ;
[0116] The sixth image normalized feature and the fifth image normalized feature are weighted and fused through the fifth fuser to achieve residual connection, and then pass through the sixth time conversion layer to introduce the temporal information between frames, and the frame image feature is output , and finally a series of video frame features are obtained .
[0117] It should be noted that the video encoding module includes image feature extraction modules, image root mean square normalization layers, fusers and A temporal transformation layer performs a total of times of feature extraction, times of normalization, times of residual connection, and times of temporal modeling on the encoded sequence of image blocks. Among them, the output of the th image feature extraction module is , the output of the th image root mean square normalization layer is , and the output of the th temporal transformation layer is . In addition, usually takes values from 6 to 12. In this embodiment, in order to reduce the amount of calculation and the number of parameters and reduce the complexity of the training model, . In practical applications, the amount of data calculation for emotion discrimination may vary, and the required model performance also differs. When using this method and its model, appropriate values can be selected according to specific scenarios to help the model adapt to different data distributions and improve the model's performance in specific scenarios.
[0118] As an embodiment, as shown in Figure 5 , it is a schematic structural diagram of the video frame feature fusion module provided in this embodiment, mainly including a second marking layer and a Transformer encoding layer connected in series. The following combines Figure 5 to further elaborate on how to input a series of video frame features into the video frame feature fusion module in sequence for fusion operations and the first linear module for normalization processing to generate the video fusion feature :
[0119] The second marking layer adds a time mark before each video frame feature in a series of video frame features to obtain the video frame features with time marks ;
[0120] The video frame features with time marks are subjected to feature fusion in the temporal dimension through the Transformer encoding layer to obtain the fused video frame features ;
[0121] The fused video frame features are linearly mapped through the first linear module to finally obtain the video fusion feature .
[0122] As an alternative embodiment, in order to introduce the text modality to assist learning in the task of human pose and emotion discrimination, this embodiment introduces a text encoding module and a text feature enhancement module. As Figure 6 shown, it is a schematic structural diagram of the text encoding module provided in this embodiment, mainly including: a first text feature extraction module, a first text root mean square normalization layer, a second text feature extraction module, a second text root mean square normalization layer, a seventh fuser, a third text feature extraction module, a third text root mean square normalization layer, an eighth fuser, a fourth text feature extraction module, a fourth text root mean square normalization layer, a ninth fuser, a fifth text feature extraction module, a fifth text root mean square normalization layer, a tenth fuser, a sixth text feature extraction module, a sixth text root mean square normalization layer, and an eleventh fuser.
[0123] Among them, the output end of the first text feature extraction module is connected to the input end of the first text root mean square normalization layer, the output end of the first text root mean square normalization layer is sequentially connected in series with the second text feature extraction module and the second text root mean square normalization layer, and the output end of the second text root mean square normalization layer and the output end of the first text root mean square normalization layer are commonly connected to the input end of the seventh fuser; the output end of the seventh fuser is sequentially connected in series with the third text feature extraction module and the third text root mean square normalization layer, and the output end of the third text root mean square normalization layer and the output end of the second text root mean square normalization layer are commonly connected to the input end of the eighth fuser; and so on, the output end of the tenth fuser is sequentially connected in series with the sixth text feature extraction module and the sixth text root mean square normalization layer, and the output end of the sixth text root mean square normalization layer and the output end of the fifth text root mean square normalization layer are commonly connected to the input end of the eleventh fuser.
[0124] Next, in combination with Figure 6 , a further detailed description will be given on how to use the text encoding module to extract features from the text sequence data and generate the text primary features in step S6:
[0125] Input the text sequence data into the first text feature extraction module for feature extraction and the first text root mean square normalization layer for normalization processing to obtain the first text normalized feature ;
[0126] The first text normalized feature is sequentially subjected to feature extraction by the second text feature extraction module and normalization processing by the second text root mean square normalization layer to obtain the second text normalized feature ;
[0127] Combine the second text normalized feature with the first text normalized feature After weighted fusion through the seventh fuser to achieve residual connection, it is then successively passed through the third text feature extraction module for feature extraction and the third text root mean square normalization layer for normalization processing to obtain the third text normalized feature ;
[0128] Repeat the above steps of weighted fusion, feature extraction, and normalization processing until the sixth text normalized feature is obtained ;
[0129] The sixth text normalized feature and the fifth text normalized feature are weighted fused through the eleventh fuser to achieve residual connection, and finally the text primary feature is obtained .
[0130] It should be noted that the text encoding module includes text feature extraction modules, text root mean square normalization layers, and fusers. The input text sequence data is subjected to a total of times of feature extraction, times of normalization, and times of residual connection. In addition, usually takes values from 6 to 12. In this embodiment, in order to reduce the amount of calculation and the number of parameters and reduce the complexity of the training model, . In practical applications, the amount of data calculation for emotion discrimination may vary, and the required model performance also differs. When using this method and its model, appropriate numbers can be selected according to specific scenarios to help the model adapt to different data distributions and improve the model performance in specific scenarios
[0131] In this embodiment, the image root mean square normalization layer and the text root mean square normalization layer are set to improve the stability of the features, so that the output features of each layer have a unified scale; and the fuser that performs weighted fusion through residual connection is set to enable the model to effectively transmit information between multiple layers and prevent information loss
[0132] Furthermore, as Figure 8 shown, it is the structural schematic diagram of the text feature enhancement module provided in this embodiment, which mainly includes a series-connected multi-head cross attention (MHCA) and a feed-forward neural network (FFN). The following combines Figure 8 , to describe how to combine the text primary feature and a series of video frame features Input them into the text feature enhancement module for fusion operations and the second linear module for normalization processing in sequence to generate text enhancement features The method is further described in detail as follows:
[0133] Using the primary text feature as the query signal, and a series of video frame features as keys and values, process them through the multi-head attention mechanism and the feed-forward neural network in sequence to obtain the text enhancement prompt ;
[0134] Among them, the expression of the text enhancement prompt is:
[0135] ,
[0136] ,
[0137] In the formula, is the primary text feature; is a series of video frame features; is the multi-head attention mechanism; is the feed-forward neural network; is the text enhancement prompt;
[0138] After weighted fusion of the text enhancement prompt and the primary text feature , the text fusion feature is obtained, and then through linear mapping by the second linear module, the text enhancement feature is finally obtained.
[0139] It should be noted that the second linear module and the first linear module need to share parameters and weights, aiming to ensure that the text enhancement feature and the video fusion feature are finally used for contrastive learning in a unified feature space.
[0140] In the multi-modal emotion discrimination network model provided in this embodiment, the text encoding module includes a multi-layer text feature extraction module, a text root mean square normalization layer, and a fuser; each layer can gradually extract richer emotion features through the text feature extraction module, ensure the consistency and stability of the features through the text root mean square normalization layer, and perform residual connection through the fuser, which can effectively alleviate the problem of gradient disappearance that may occur in the deep network and enhance the efficiency of feature transmission. Then, use the text feature enhancement module to effectively fuse the visual modality and the language modality; finally, through the linear mapping with shared weights, the text enhancement feature can be aligned with the video fusion feature, so as to achieve the multi-modal emotion discrimination task.
[0141] In this embodiment, in step S8, the contrast learning module performs the following processing steps on the video fusion feature and the text enhanced feature :
[0142] Perform contrast learning on the video fusion feature and the text enhanced feature within the mapped feature space;
[0143] Define a similarity matrix , and calculate the cosine similarity between the video fusion feature and the text enhanced feature ;
[0144] Among them, the expression of the similarity matrix is:
[0145] ,
[0146] In the formula, represents the cosine similarity between the th text enhanced feature and the th video fusion feature pair; is the th text enhanced feature vector; is the th video fusion feature vector;
[0147] Output the emotion label corresponding to the maximum similarity, which is the emotion discrimination result corresponding to the human body posture video to be discriminated.
[0148] The multi-modal emotion discrimination method based on human body posture provided in this embodiment discriminates the emotion of the human body posture video by using the trained multi-modal emotion discrimination network model, introduces text modality assisted learning in the human body posture emotion discrimination task, and effectively fuses the visual modality and the language modality by using the text feature enhancement module, which can efficiently capture rich emotion features, improve the accuracy of model discrimination, reduce the dependence of the model on the data set, and effectively reduce the computational complexity of the model; at the same time, the alignment and discrimination of the visual modality and the language modality are realized through the contrast learning model, thereby improving the generalization ability of the model.
[0149] This flowchart only shows the logical order of the method described in this embodiment. On the premise of non-conflict, in other possible embodiments of the present invention, it may be different from Figure 1Complete the steps shown or described in the order shown. The method for multi-modal emotion discrimination based on human body postures provided in this embodiment can be applied to a terminal and can be executed by a multi-modal emotion discrimination device based on human body postures. The device can be implemented in a software and / or hardware manner and can be integrated into the terminal. For example: any smart phone, tablet computer or computer device with communication functions.
[0150] Embodiment 2:
[0151] The present embodiment provides a method for multi-modal emotion discrimination based on human body postures, which is different from Embodiment 1 in that, as Figure 4 shown, it is a schematic structural diagram of the image feature extraction module provided in this embodiment, mainly including a linear layer, a first SiLU activation function, a forward one-dimensional convolutional layer, a forward state space module, a second SiLU activation function, a reverse one-dimensional convolutional layer, a reverse state space module, a linear layer, a third SiLU activation function, a first multiplier, a second multiplier, a sixth fuser, a fourth SiLU activation function, and a linear layer.
[0152] Among them, the first output end of the linear layer is sequentially connected in series with the first SiLU activation function, the forward one-dimensional convolutional layer, and the forward state space module, the second output end of the linear layer is sequentially connected in series with the second SiLU activation function, the reverse one-dimensional convolutional layer, and the reverse state space module, the output end of the linear layer is connected to the input end of the third SiLU activation function. The output end of the third SiLU activation function performs a gating operation on the output end of the forward state space module through the first multiplier and performs a gating operation on the output end of the reverse state space module through the second multiplier. The output end of the first multiplier and the output end of the second multiplier are commonly connected to the input end of the sixth fuser. The output end of the sixth fuser is sequentially connected in series with the fourth SiLU activation function and a linear layer.
[0153] Next, in combination with Figure 4 , the processing steps of the image feature extraction module will be further described in detail:
[0154] Input the input image features respectively into the linear layer and the linear layer for transformation to generate a first input feature and a second input feature ;
[0155] The first input feature Processed sequentially through the first SiLU activation function and the forward one-dimensional convolutional layer, and then the forward state space module is used to capture dependencies to obtain the forward output features ; and, the first input feature Processed sequentially through the second SiLU activation function and the backward one-dimensional convolutional layer, and then the backward state space module is used to capture dependencies to obtain the backward output features ;
[0156] Meanwhile, after the third SiLU activation function is used to perform a non-linear transformation on the second input feature , the forward output feature and the backward output feature are respectively gated through the first multiplier and the second multiplier, which helps to retain the information of the early layers and solve the problem of vanishing gradients, and the forward extracted feature and the backward extracted feature are respectively obtained;
[0157] The forward extracted feature and the backward extracted feature are fused through the sixth fuser, and then sequentially transformed through the fourth SiLU activation function and the linear layer, and finally the image extraction feature is obtained.
[0158] The image feature extraction module provided in this embodiment captures the context relationship and spatial information in visual data through bidirectional sequence modeling, can efficiently capture rich emotion features, and improve the accuracy of model discrimination.
[0159] As an embodiment, as shown in Figure 7 , it is a schematic structural diagram of the text feature extraction module provided in this embodiment, mainly including: an input linear layer, a convolutional layer, a fifth SiLU activation function, a state space module, a sixth SiLU activation function, a twelfth fuser, a seventh SiLU activation function, and an output linear layer. Among them, the first output end of the input linear layer is sequentially connected in series with the convolutional layer, the fifth SiLU activation function, and the state space module, the second output end of the input linear layer is connected to the input end of the sixth SiLU activation function, the output end of the sixth SiLU activation function and the output end of the state space module are commonly connected to the input end of the twelfth fuser, and the output end of the twelfth fuser is sequentially connected in series with the seventh SiLU activation function and the output linear layer.
[0160] Next, in combination with Figure 7 , the processing steps of the text feature extraction module will be further described in detail:
[0161] The input text feature It is input into the input linear layer for conversion to obtain the initial word vector ;
[0162] The initial word vector is successively passed through the convolutional layer for multi-dimensional feature fusion to avoid the limitations of independent calculation of word vectors, the fifth SiLU activation function is introduced to introduce non-linear features to further enrich the expression ability of the model, and the state space module captures the long-term and short-term dependencies in the sequence, and then the intermediate word vector is obtained to retain the context information in the time series;
[0163] After the initial word vector is subjected to non-linear transformation by the sixth SiLU activation function, it is connected with the intermediate word vector through the twelfth fuser for residual connection, which helps to maintain the gradient flow and avoid the problem of gradient disappearance, and the fused word vector is obtained;
[0164] The fused word vector is successively passed through the seventh SiLU activation function and the output linear layer for dimension conversion to maintain consistency with the dimension of the input text, and finally the text extraction feature is obtained.
[0165] It should be noted that the calculation principles of the state space module, the forward state space module, and the reverse state space module are the same. The original state space model (State Space Model, SSM) is mainly used in control system theory to describe the dynamic behavior of the system and estimate the system state based on observed data. Its mechanism can be roughly summarized into the following two equations:
[0166] State equation:
[0167] ,
[0168] In the formula, is the state at time, which is the prediction of the state at time based on the state at the previous time; is the state transition matrix and is a learnable parameter; is the state at the previous time; is the input matrix and is a learnable parameter; is the input at
[0169] Output equation:
[0170] ,
[0171] In the formula, is The output at a moment, used to predict the current moment of the predicted value; is the prediction matrix and is a learnable parameter.
[0172] Furthermore, since the input in the above state equation is a continuous signal, but in this embodiment, whether it is video sequence data or text sequence data, they are both discrete signals. Therefore, the zero-order hold technology is needed to convert the discrete signal into a continuous signal. Among them, the principle of the zero-order hold technology is: each time a discrete signal is received, the value is held until a new discrete signal is received; the time to save this value is represented by a learnable parameter step size denoted. With the continuous input signal, a continuous output can be generated, and the value is sampled only according to the input time step The sampled value is the finally discretized output. Therefore, the two equations of the action mechanism of the original state space model SSM are updated to:
[0173] State equation:
[0174] ,
[0175] ,
[0176] ,
[0177] Output equation:
[0178] ,
[0179] In the formula, is the discretized matrix; is the discretized matrix; is the identity matrix; is the discrete time step.
[0180] The video coding module and the text coding module provided in this embodiment are both extended based on the Mamba model architecture, and the core of the Mamba model is the state space model SSM, which can be used for processing sequence information together with Transformer and RNN. However, due to the problems of gradient explosion or gradient disappearance, RNN is difficult to effectively capture long-distance dependence relationships; although Transformer can process long sequence information, the computational cost is relatively large. The state space model SSM can efficiently process sequences with a scale ranging from tens of thousands to millions through an implicit state update mechanism and reduce the computational complexity. Therefore, it has more significant advantages in long sequence modeling.
[0181] Embodiment 3:
[0182] This embodiment provides a multi-modal emotion discrimination method based on human postures, and further details the method of training the multi-modal emotion discrimination network model in step S3 of Embodiment 1:
[0183] Step S301: Construct a human posture emotion dataset;
[0184] Step S302: Divide the human posture emotion dataset into a video training set and a video test set according to the ratio of 8:2, which are used for training and performance testing of the network model respectively;
[0185] Step S303: Extract posture description words according to the emotion labels in the human posture emotion dataset, and construct a text prompt template;
[0186] Step S304: Concatenate the text prompt template with various emotion labels to obtain text training data;
[0187] Step S305: Train the pre-constructed multi-modal emotion discrimination network model using the video training set and the text training data, test the trained multi-modal emotion discrimination network model using the video test set and the text training data, and use the multi-modal emotion discrimination network model with the optimal test result as the finally trained multi-modal emotion discrimination network model.
[0188] Specifically, the human posture emotion dataset constructed in step S301 includes a public dataset and a self-built dataset; among them, the public dataset refers to the GEMEP emotion dataset and the FABO emotion dataset after data screening, and the self-built dataset refers to the corresponding relationship between various common emotions and body movements summarized by the laboratory according to relevant psychological literature and body movements, and the self-built emotion dataset recorded according to the corresponding relationship.
[0189] It should be noted that the GEMEP emotion dataset is recorded by professional actors. When the actors perform each emotion, the changes in the upper body limb movements are recorded using camera equipment. The FABO emotion dataset is a video dataset of the expressions and gestures of 23 testers from different countries, different ages, and different genders collected under the guidance of professionals. When recording the gestures, the tested persons make corresponding gesture movements according to different emotion requirements.
[0190] Furthermore, the constructed human posture emotion dataset can overcome the situation of model overfitting caused by the small number of videos in current emotion classification. It is divided into 10 emotions in total: happy, surprised, bored, confused, nervous, anxious, angry, sad, fearful, and disgusted. The corresponding relationships between these emotions and body movements are as follows:
[0191] Happy: Clap hands, raise both hands, open arms, nod when laughing; Surprised: Raise both hands, cover mouth; Bored: Support face with one hand, let the arm hang naturally, open both hands; Confused: Lean forward, raise hand, scratch head; Nervous: Hand trembles, pick at fingers; Anxious: Clench fists tightly, fold arms, scratch head, put hands on hips, bite fingers, touch ears; Angry: Body trembles, clench fists tightly, shake fists, pat the table with both hands; Sad: Body curls up, cover face with both hands, arms wrap around body or shoulders, hands close or move slowly, hands touch head; Fear: Cross and move hands and arms, hold hands or arms tightly, elbows inwards, cover face; Disgust: Cover mouth with one hand, raise one hand, hands close to body, cover face with both hands.
[0192] In this embodiment, the text prompt template constructed in step S303 refers to extracting gesture description words according to the performance characteristics of each emotion in gestures to design the mapping relationship between emotion and gesture, and converting the gesture description words into text modal features through feature extraction for assisting subsequent emotion discrimination. Among them, the constructed text prompt template includes: The human in this video feels {label}; Look, the human feels {label}; A human expressing {label} through body language; This human seems to express a feeling of {label}; {label}; Emotion classification of {label}; The man feels {label}; The woman feels {label}.
[0193] The information of the text modality provided in this embodiment is generated based on the emotion labels of the constructed human body gesture emotion dataset, without the participation of an additional dataset in training. Compared with single-modal or traditional multi-modal emotion discrimination methods, this application improves the emotion discrimination accuracy while reducing the dependence on an additional dataset, effectively reducing the computational complexity of the model.
[0194] As an optional embodiment, the method for training the multi-modal emotion discrimination network model in step S305 further includes:
[0195] To ensure the alignment of text features and video features, a contrast loss function is constructed according to the losses of positive sample pairs and negative sample pairs , where positive sample pairs and negative sample pairs are respectively used to guide the network to learn;
[0196] To minimize the differences between different modalities, with the contrast loss function The minimum value is the target, which maximizes the cosine similarity of positive sample pairs and minimizes the cosine similarity of negative sample pairs, thereby optimizing the multi-modal emotion discrimination network model;
[0197] Among them, the loss of positive sample pairs is:
[0198] ,
[0199] ,
[0200] In the formula, represents the loss between text and video pairs belonging to the same emotion; is the Sigmoid activation function; represents the th text enhancement feature and the th cosine similarity between video fusion feature pairs; is the th text enhancement feature vector, ; is the th video fusion feature vector, ;
[0201] The loss of negative sample pairs is:
[0202] ,
[0203] In the formula, represents the loss between text and video pairs belonging to different emotions; is a hyperparameter. When the cosine similarity is small, the loss of negative sample pairs is small. When the cosine similarity is large, the loss of negative sample pairs increases;
[0204] The expression of the contrast loss function is:
[0205] ,
[0206] In the formula, is the contrast loss function; represents that the th text and video pair belongs to a positive sample pair, represents that the th text and video pair belongs to a negative sample pair.
[0207] In this embodiment, by minimizing the contrast loss function, the parameters of the text encoding module and the video encoding module are adjusted to ensure that in the common feature space, the text features and video features of the same emotion category are closer to each other, while the distance between the text features and video features of different emotion categories is farther.
[0208] The multi-modal emotion discrimination method based on human posture provided by the embodiments of the present invention effectively integrates visual modal features and language modal features by combining human posture information and text information. Among them, the changes in human postures complement the non-verbal information in the text, and the text content also makes up for the ambiguity and incompleteness in the posture information. While capturing richer emotional features, it reduces the problem of information loss in a single modality, thereby achieving more accurate emotion discrimination. In addition, through contrastive learning, the alignment of the visual modality and the language modality is achieved, improving the generalization ability of the model in different emotions and different scenarios, and avoiding the problem of performance degradation caused by data imbalance or a large variety of emotion types in traditional methods.
[0209] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. In addition, terms such as "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features.
[0210] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0211] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0212] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, thereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 steps for implementing the functions specified in one block or a plurality of blocks.
[0213] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims. All of these fall within the protection scope of the present invention.
Claims
1. A multi-modal emotion discrimination method based on human body postures, characterized in that, The method includes: Obtaining the human body posture video sequence data to be discriminated; Splicing the pre-constructed text prompt template with various emotion tags to obtain text sequence data; Inputting the video sequence data and the text sequence data into a pre-constructed and trained multi-modal emotion discrimination network model, where the multi-modal emotion discrimination network model includes a video encoding module, a video frame feature fusion module, a first linear module, a text encoding module, a text feature enhancement module, a second linear module, and a contrastive learning module; Using the video encoding module to segment, encode, and extract features from the video sequence data to generate a series of video frame features; Sequentially inputting a series of video frame features into the video frame feature fusion module for fusion operation and the first linear module for normalization processing to generate video fusion features; Using the text encoding module to extract features from the text sequence data to generate text primary features; Sequentially inputting the text primary features and a series of video frame features into the text feature enhancement module for fusion operation and the second linear module for normalization processing to generate text enhanced features; Using the contrastive learning module to calculate the similarity between the video fusion features and the text enhanced features and output the emotion discrimination result.
2. The multi-modal emotion discrimination method based on human body postures according to claim 1, wherein, The video encoding module includes a linear projection layer, a first token layer, a first image feature extraction module, a first image root mean square normalization layer, a first temporal transformation layer, a second image feature extraction module, a second image root mean square normalization layer, a first fuser, a second temporal transformation layer, a third image feature extraction module, a third image root mean square normalization layer, a second fuser, a third temporal transformation layer, a fourth image feature extraction module, a fourth image root mean square normalization layer, a third fuser, a fourth temporal transformation layer, a fifth image feature extraction module, a fifth image root mean square normalization layer, a fourth fuser, a fifth temporal transformation layer, a sixth image feature extraction module, a sixth image root mean square normalization layer, a fifth fuser, and a sixth temporal transformation layer; The method of using the video encoding module to segment, encode, and extract features from the video sequence data to generate a series of video frame features includes: After segmenting the frame images in the video sequence data into image patches, inputting them into the linear projection layer for transformation to obtain image patch encodings; Using the first token layer to add position tokens and class tokens to the image patch encodings to obtain an encoded image patch sequence; Sequentially inputting the encoded image patch sequence into the first image feature extraction module for feature extraction and the first image root mean square normalization layer for normalization processing to obtain first image normalized features; The first image normalized features are sequentially passed through the first temporal transformation layer to introduce temporal information, the second image feature extraction module for feature extraction, and the second image root mean square normalization layer for normalization processing to obtain second image normalized features; After the normalized features of the second image and the normalized features of the first image are weighted and fused by the first fuser, they are then successively passed through the second time conversion layer to introduce temporal information, the third image feature extraction module to extract features, and the third image root mean square normalization layer to perform normalization processing, resulting in the normalized features of the third image; Repeat the above steps of weighted fusion, introducing temporal information, feature extraction, and normalization processing until the normalized features of the sixth image are obtained; After the normalized features of the sixth image and the normalized features of the fifth image are weighted and fused by the fifth fuser, they are then passed through the sixth time conversion layer to introduce temporal information, and finally a series of video frame features are obtained.
3. The multi-modal emotion discrimination method based on human body postures according to claim 2, wherein, The first to sixth image feature extraction modules all include a linear layer, a first SiLU activation function, a forward one-dimensional convolutional layer, a forward state space module, a second SiLU activation function, a backward one-dimensional convolutional layer, a backward state space module, a linear layer, a third SiLU activation function, a first multiplier, a second multiplier, a sixth fuser, a fourth SiLU activation function, and a linear layer; The first image feature extraction module to the sixth image feature extraction module perform the following processing steps on the input image features: Input the input image features into the linear layer and the linear layer for transformation to generate a first input feature and a second input feature; The first input feature is successively processed by the first SiLU activation function and the forward one-dimensional convolutional layer, and then the forward output feature is obtained by capturing the dependency relationship through the forward state space module; moreover, the first input feature is successively processed by the second SiLU activation function and the backward one-dimensional convolutional layer, and then the backward output feature is obtained by capturing the dependency relationship through the backward state space module; Meanwhile, after the second input feature is non-linearly transformed by the third SiLU activation function, the forward output feature and the backward output feature are respectively gated by the first multiplier and the second multiplier to obtain the forward extraction feature and the backward extraction feature; Fuse the forward-extracted features and the backward-extracted features through the sixth fuser, and then successively pass through the fourth SiLU activation function and the linear layer for transformation to finally obtain the image-extracted features.
4. The multi-modal emotion discrimination method based on human body postures according to claim 1, characterized in that The video frame feature fusion module includes a second tagging layer and a Transformer encoding layer; The method of successively inputting a series of video frame features into the video frame feature fusion module for fusion operation and the first linear module for normalization processing to generate video fusion features includes: Using the second tagging layer to add a time tag before each video frame feature in a series of video frame features to obtain video frame features with time tags; Passing the video frame features with time tags through the Transformer encoding layer to perform feature fusion in the temporal dimension to obtain the fused video frame features; Passing the fused video frame features through the first linear module for linear mapping to finally obtain video fusion features.
5. The multi-modal emotion discrimination method based on human body postures according to claim 1, wherein The text encoding module includes a first text feature extraction module, a first text root mean square normalization layer, a second text feature extraction module, a second text root mean square normalization layer, a seventh fuser, a third text feature extraction module, a third text root mean square normalization layer, an eighth fuser, a fourth text feature extraction module, a fourth text root mean square normalization layer, a ninth fuser, a fifth text feature extraction module, a fifth text root mean square normalization layer, a tenth fuser, a sixth text feature extraction module, a sixth text root mean square normalization layer, and an eleventh fuser; The method of using the text encoding module to extract features from the text sequence data to generate text primary features includes: The text sequence data is sequentially input into the first text feature extraction module for feature extraction and the first text root mean square normalization layer for normalization processing to obtain the first text normalized feature; The first text normalized feature is sequentially passed through the second text feature extraction module for feature extraction and the second text root mean square normalization layer for normalization processing to obtain the second text normalized feature; After the second text normalized feature and the first text normalized feature are weighted and fused by the seventh fuser, they are sequentially passed through the third text feature extraction module for feature extraction and the third text root mean square normalization layer for normalization processing to obtain the third text normalized feature; Repeat the above steps of weighted fusion, feature extraction, and normalization processing until the sixth text normalized feature is obtained; The sixth text normalized feature and the fifth text normalized feature are weighted and fused by the eleventh fuser to finally obtain the text primary feature.
6. The multimodal emotion discrimination method based on human body postures according to claim 5, wherein, The first text feature extraction module to the sixth text feature extraction module all include an input linear layer, a convolutional layer, a fifth SiLU activation function, a state space module, a sixth SiLU activation function, a twelfth fuser, a seventh SiLU activation function, and an output linear layer; The first text feature extraction module to the sixth text feature extraction module perform the following processing steps on the input text features: The input text features are input into the input linear layer for conversion to obtain initial word vectors; The initial word vectors are sequentially passed through the convolutional layer for multi-dimensional feature fusion, the fifth SiLU activation function to introduce non-linear features, and the state space module to capture dependency relationships to obtain intermediate word vectors; After the initial word vectors are non-linearly transformed by the sixth SiLU activation function, they are residually connected with the intermediate word vectors through the twelfth fuser to obtain fused word vectors; The fused word vectors are sequentially passed through the seventh SiLU activation function and the output linear layer for dimensionality conversion to finally obtain the text extraction features.
7. The multimodal emotion discrimination method based on human body postures according to claim 1, wherein The text feature enhancement module includes a multi-head attention mechanism and a feed-forward neural network; The method for sequentially inputting the text primary feature and a series of video frame features into the text feature enhancement module for fusion operation and the second linear module for normalization processing to generate the text enhancement feature includes: Using the text primary feature as the query signal and a series of video frame features as the key and value, and sequentially processing them through the multi-head attention mechanism and the feed-forward neural network to obtain the text enhancement prompt; Among them, the expression of the text enhancement prompt is: , , In the formula, is the primary text feature; is a series of video frame features; is the multi-head attention mechanism; is the feed-forward neural network; is the text enhancement prompt; After the text enhancement prompt and the text primary feature are weighted and fused, they are linearly mapped through the second linear module to finally obtain the text enhancement feature.
8. The multi-modal emotion discrimination method based on human body postures according to claim 1, wherein The method for calculating the similarity between the video fusion feature and the text enhancement feature using the contrast learning module and outputting the emotion discrimination result includes: Performing contrast learning on the video fusion feature and the text enhancement feature in the mapped feature space; Defining a similarity matrix and calculating the cosine similarity between the video fusion feature and the text enhancement feature; Among them, the expression of the similarity matrix is: , In the formula, represents the cosine similarity between the -th text enhancement feature and the -th video fusion feature pair; is the -th text enhancement feature vector; is the -th video fusion feature vector; Output the emotion label corresponding to the maximum similarity, which is the emotion discrimination result corresponding to the human body posture video to be discriminated.
9. The multi-modal emotion discrimination method based on human body postures according to claim 1, characterized in that, The training method of the multi-modal emotion discrimination network model includes: Construct a human body posture emotion dataset; Divide the human body posture emotion dataset into a video training set and a video test set according to the ratio of 8:2; According to the emotion labels in the human body posture emotion dataset, extract posture description words and construct a text prompt template; Concatenate the text prompt template with various emotion labels to obtain text training data; Use the video training set and the text training data to train the pre-constructed multi-modal emotion discrimination network model, use the video test set and the text training data to test the trained multi-modal emotion discrimination network model, and use the multi-modal emotion discrimination network model with the optimal test result as the finally trained multi-modal emotion discrimination network model.
10. The multi-modal emotion discrimination method based on human body posture according to claim 9, characterized in that, The training method of the multi-modal emotion discrimination network model also includes: Construct a contrast loss function according to the losses of positive sample pairs and negative sample pairs; Optimize the multi-modal emotion discrimination network model with the goal of minimizing the value of the contrast loss function; Among them, the loss of the positive sample pair is: , , Wherein, represents the loss between text and video pairs belonging to the same emotion; is the Sigmoid activation function; represents the -th text enhancement feature and the -th cosine similarity between video fusion features; is the -th text enhancement feature vector; is the -th video fusion feature vector; The loss of the negative sample pair is: , wherein, represents the loss between text and video pairs belonging to different emotions; is a hyperparameter. When the cosine similarity is small, the loss of negative sample pairs is small. When the cosine similarity is large, the loss of negative sample pairs increases; The expression of the contrast loss function is: , Wherein, is the contrastive loss function; indicates that the -th text and video pair belongs to the positive sample pair, indicates that the -th text and video pair belongs to the negative sample pair.