Micro-posture recognition method based on multi-modal feature fusion and fine adjustment
Through the multimodal feature fusion and fine-tuning method, combined with video, skeleton and text data, the problem that traditional single-modal methods are difficult to capture micro-pose subtle information, achieving higher recognition accuracy and robustness.
Patent Information
- Application Number
- CN202510377218.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-28
AI Technical Summary
Traditional single-modal-based deep learning methods are difficult to effectively capture subtle action information in micro-poses, especially in unconscious actions and fuzzy features, lacking sufficient context and multi-dimensional support, resulting in limited model performance.
A micro-pose recognition method based on multimodal feature fusion and fine-tuning is proposed. Features are extracted through multimodal encoders of video, skeleton and text data, and feature fusion is realized using cross-modal fusion module. Comparative learning and cross-entropy loss training model are used to reduce the computational complexity.
It effectively makes up for the lack of single-modal information, improves the accuracy and robustness of micro-pose recognition, can capture the fine-grained features of micro-pose more deeply, and improves the system's ability to perceive subtle movements.
Smart Images

Figure FT_1 
Figure SMS_17 
Figure SMS_18
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and action recognition, and particularly relates to a micro gesture recognition method based on multi-modal feature fusion and fine-tuning. Background Art
[0002] Action recognition is a key research direction in the field of computer vision, whose goal is to automatically recognize and analyze human postures and actions through computer technology to infer the state and intention of the human body. By analyzing the overall posture of the human body, the positions and movements of body joints, more comprehensive information can be obtained, protecting people's privacy and providing a more comprehensive and accurate method for emotion analysis. Action recognition is commonly used in fields such as motion capture, human-computer interaction, and virtual reality recognition.
[0003] In the field of action recognition, most people are committed to the recognition of descriptive postures. "Descriptive postures" mainly refer to purposeful and more significant limb movements, such as drinking water and running, through which people can clearly express their emotions and attitudes. However, in some specific situations, such as interviews and competitions, people may deliberately hide or suppress their true feelings, which is not conducive to further analysis by computers. Therefore, micro gesture recognition is proposed. "Micro gestures" refer to spontaneous and unconscious subtle movements that are difficult to detect but can reflect an individual's true emotional state, especially tension, stress, etc., and can further detect abnormal mental states, such as Alzheimer's disease and autism. It has important significance and research prospects in psychology, behavioral analysis, and communication research.
[0004] Due to the unique complexity and subtlety of micro gesture data, traditional single-modal deep learning methods are unable to capture the detailed information of these tiny movements. Single-modal methods usually only rely on video frames or skeleton data and are difficult to comprehensively depict the emotional and semantic information hidden in micro gestures. Especially in the unconscious actions and fuzzy features contained in micro gestures, a single data source often lacks sufficient context and multi-dimensional support, resulting in limited performance of the model. In recent years, with the rapid development of vision-language models, multi-modal methods have gradually become an important means for studying complex behavior patterns. Vision-language models combine video data with text descriptions, not only making up for the deficiencies of single-modal methods but also endowing the model with stronger semantic understanding capabilities. Therefore, the present invention proposes a micro gesture recognition method based on multi-modal fusion and fine-tuning. On the basis of vision-text, skeleton information is further introduced to form a multi-modal fusion method of video-skeleton-text. Summary of the Invention
[0005] To solve the above problems, the present invention proposes a micro gesture recognition method based on multi-modal feature fusion and fine-tuning, and proposes a general cross-modal knowledge fusion framework. By fine-tuning the network to extract the features of videos, skeletons and texts respectively, and performing subsequent feature fusion and interaction, the information missing in the single modality is compensated, and finer-grained micro gesture details are captured.
[0006] The technical solution adopted by the present invention is as follows:
[0007] A micro gesture recognition method based on multi-modal feature fusion and fine-tuning, comprising the following steps:
[0008] Step 1: Sample and preprocess the multi-modal data.
[0009] Step 2: Construct a micro gesture recognition model, which includes a multi-modal encoder and a cross-modal fusion module.
[0010] Step 3: Construct a loss function and train the fine-tuning model.
[0011] Step 4: Output the current human micro gesture behavior obtained by the camera into the trained model to obtain the micro gesture category to which it belongs.
[0012] Further, the specific process of the step 1 is as follows:
[0013] For video data, uniformly sample the original video, and compress, crop and flip the image at a ratio of 1:1.
[0014] For skeleton data, obtain the skeleton joint point data of the original video through the OpenPose pose estimation algorithm, normalize and centralize the joint point coordinates, and perform alignment in the time dimension to ensure the same frame rate as the video. For text data, use the different category names of actions as the semantic text library, and set this library to the pre-input format of the CLIP text encoder.
[0015] Further, the specific process of the step 2 is as follows:
[0016] The multi-modal encoder includes a CLIP-based video encoder, a CLIP-based text encoder, and a CTR-GCN-based skeleton encoder.
[0017] The CLIP-based video encoder is a vision Transformer architecture with strong feature extraction capabilities. The CLIP-based text encoder is a Transformer model, and through the self-attention mechanism, the model can capture the relationships between words. CTR-GCN is a graph convolution model, which refines the modeling of the skeleton by performing feature extraction and aggregation on each channel through a convolutional network, and improves the recognition performance.
[0018] The transmembrane fusion module includes a video-skeleton fusion module and a text-skeleton fusion module.
[0019] Among them, the video-skeleton fusion module is based on a single-layer Transformer architecture. For the video features encoded by the video encoder, the global class token Cls∈R is extracted from them T×C×1 , where T represents time and C represents the number of channels output by the video encoder. For the skeleton features encoded by the skeleton encoder, the feature is obtained, where T represents time, C2 represents the number of channels output by the graph convolutional network, and N represents the dimension of the joint points. The skeleton features are linearly mapped in the channel dimension to unify the number of channels, and then the global class token Cls and the skeleton features F P are concatenated in the third dimension to obtain the fused feature F1∈R T×C×(1+N) , and this fused feature is fed into the Transformer encoder to encode the fused feature.
[0020] The text-skeleton fusion module is based on the K most likely classes in the skeleton classification results. For the skeleton features encoded by the skeleton encoder, the skeleton features are obtained. For the feature F P , average pooling is performed in the joint point dimension N to obtain The pooled feature is input into a fully connected layer and mapped to the class space to obtain , where N2 represents the number of action classes. The softmax function is applied to the output of the fully connected layer to obtain the probability distribution of each class, and K classes with the highest probabilities are selected as the several most likely classes. According to the predicted K classes, the corresponding class embedding features are loaded from the text modality and obtained through the K-class index These text features are obtained through a CLIP-based text encoder and can provide accurate semantic information related to the classes. Subsequently, F T is replicated in the time dimension and averaged along the K-class dimension to obtain the feature This operation helps to integrate the information of the text and skeleton modalities into a single feature space.
[0021] Next, the skeleton features F before pooling P and the text features F T are merged through a concatenation operation to form a fused feature The fused feature F2 further extracts the time series-related feature representation through a one-dimensional convolutional layer, and the enhanced skeleton features generate a classification probability distribution through a fully connected layer, which is weighted and fused with the original skeleton classification results to obtain the final prediction result.
[0022] Furthermore, the specific process of step 3 is as follows:
[0023] For the result F1 of the fusion of the video modality and the skeleton modality, contrastive learning is used to align the features of F1 and the text. First, the features of F1 and the text feature F T are normalized, and then the similarity matrix is calculated The features of F1 and the text are aligned through a symmetric contrastive loss function. For the result F2 of the fusion of the text modality and the skeleton modality, the predicted class is obtained through a fully connected layer, and cross-entropy loss is used for training. The total loss of the model is the weighted sum of the two losses.
[0024] During the training process of the model, the freeze-fine-tuning method is adopted. Specifically, for the video encoder and the text encoder, the parameters of the Transformer part are frozen, and an adapter module is added after each layer of attention. For the video encoder, the added adapter type is a spatio-temporal adapter. For the text encoder, since the text information does not have features of time and space dimensions, the added adapter is based on a multi-layer perceptron. The adapter parts in the video encoder and the text encoder are frozen, and fine-tuning training is performed with the skeleton encoder, and the rest of the Transformer part is frozen to reduce the computational complexity of the model.
[0025] Furthermore, the specific process of step 4 is as follows:
[0026] For the video data to be processed obtained by the monitoring camera, video data and skeleton data are obtained through preprocessing. The video data, skeleton data, and text data pass through the video encoder, skeleton encoder, and text encoder respectively to obtain video features, skeleton features, and text features. The video and skeleton features are fused through a fusion module to obtain fusion feature 1. The similarity between fusion feature 1 and the text features is calculated to obtain the micro-pose category with the highest similarity. The skeleton and text features are fused through a fusion module to obtain fusion feature 2. Fusion feature 2 is classified through a fully connected layer to obtain the predicted micro-pose category. The two prediction results are added to obtain the final prediction result. Brief Description of the Drawings
[0027] Figure 1 It is a schematic diagram of the model of the micro-pose recognition method based on multi-modal feature fusion and fine-tuning of the present invention. Detailed Embodiment
[0028] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0029] The present invention proposes a micro gesture recognition method based on multi-modal feature fusion and fine-tuning. Features are extracted through multi-modal encoders for video, skeleton, and text data, and a cross-modal fusion module is used to achieve feature fusion. The model is trained using contrastive learning and cross-entropy loss, and the computational complexity is reduced through a freeze-fine-tuning strategy and an adapter module. Finally, the trained model is used to classify the micro gesture behaviors of the human body obtained by the camera, and the micro gesture categories are output. This method improves the performance of micro gesture recognition through multi-modal feature fusion and a parameter-efficient fine-tuning strategy.
[0030] As Figure 1 shown, a micro gesture recognition method based on multi-modal feature fusion and fine-tuning specifically includes the following steps:
[0031] Step 1. Sample and preprocess the multi-modal data. The specific process is as follows:
[0032] Step 1.1. Select video frames from the original video data at fixed time intervals to ensure that all video samples have a consistent frame rate. Then, perform image compression on the sampled images, scale the images at a ratio of 1:1; image cropping, perform central cropping on the images to remove redundant background regions; image flipping, perform horizontal flipping on the images.
[0033] Step 1.2. Use the OpenPose pose estimation algorithm to detect skeleton key points for the original video data. First, input the video frames into a convolutional neural network to extract features. After being processed by the feed-forward convolutional network, the feature map is divided into two branches. Branch one predicts the joint point confidence map through the convolutional network post-processing, and each joint point has an independent confidence heat map; the other branch predicts the limb connection vector field through a set of two-dimensional vector fields, representing the connection relationship between joint points. Finally, through the greedy matching algorithm, the optimal joint point connection is calculated according to the limb connection vector field, and the skeleton information of the human body is parsed.
[0034] Step 1.3. Organize the micro gesture category text, convert the text into the format required for CLIP pre-training, such as "a person is {action name}", to enhance the generality and generalization ability of the text description.
[0035] Step 2. Build a micro gesture recognition model, which includes a multi-modal encoder and a cross-modal fusion module. The specific process is as follows:
[0036] Step 2.1: For the input video, input the video frames into a CLIP-based Vision Transformer with an adapter. The video frames are first sliced into multiple fixed-size image patches, each patch having a size of 16×16. Each small image patch is mapped to a high-dimensional feature space through a linear projection layer and positional encoding is added to preserve spatial information. Then, these small image patches are processed through 12 layers of Transformer, each layer containing multi-head attention and a feed-forward neural network to capture global context information. The encoded output includes multiple tokens, where the global class token Cls ∈ R T×C×1 represents the global feature of the entire frame, where T is the number of frames and C is the feature dimension.
[0037] Step 2.2: For the input text, input the text vector into a CLIP-based Transformer with an adapter. Specifically, the input text is first tokenized into a series of discrete tokens, then mapped to a 512-dimensional vector space through a word embedding layer. Next, the embedded text is processed by the Transformer, and to preserve sequence information, each token is added with positional encoding. After encoding, the output of the text sequence includes multiple token-level features, where the end token is used to extract the global semantic representation of the entire text, generating the final text-level feature vector where N2 is the number of classes and C2 is the feature dimension.
[0038] Step 2.3: For the input skeleton, input the skeleton data into the CTR-GCN network for encoding. CTR-GCN adaptively adjusts the connection relationship between joint points through a graph topology learning mechanism, and processes the features of each joint point through a channel attention mechanism, assigning different weights to different channels. The input skeleton is calculated through a multi-layer graph convolutional network to obtain the skeleton feature representation where T is the number of frames, C2 is the feature dimension, and N is the number of joint points.
[0039] Step 2.4: In the video-skeleton fusion module, the processed video feature is represented as Cls ∈ R T×C×1 , where T is the time frame and C is the video feature dimension. The processed skeleton feature is represented as where T is the number of frames, C2 is the feature dimension, and N is the number of joint points.
[0040] To fuse the skeleton feature and the video feature, channel alignment is performed on the skeleton feature. Through linear mapping, the skeleton feature is mapped to the same dimension C as the video feature:
[0041] F aligned = Linear(F P ) ∈ R T×C×N (1)
[0042] Concatenate the video feature Cls with the aligned skeleton feature F aligned in the third dimension to obtain the fused feature:
[0043] F1 = [Cls, F aligned ∈ R T×C×(N+1) (2)
[0044] Feed the fused feature F1 into a single-layer Transformer for further modeling, and capture the relationship between the video and the skeleton through the multi-head self-attention mechanism. Calculate the similarity between the modeled fused feature and the text feature to obtain the most likely micro-pose category.
[0045] Step 2.5. In the text-skeleton fusion module, select the most likely K categories according to the skeleton classification results, and fuse the text feature and the skeleton feature. For the encoded skeleton feature, first perform average pooling in the joint dimension:
[0046]
[0047] Then feed it into a fully connected layer and map it to the category space to obtain the category prediction result:
[0048] P = Softmax(FC(F pool )) (4)
[0049] Select the K categories with the highest probabilities as candidate categories. For the encoded text feature, obtain the embedding features of the K categories from the text modality to get and replicate them in the time dimension, and perform average pooling in the category dimension to obtain the fused text feature Combine the skeleton feature before pooling and the fused text feature through a concatenation operation:
[0050]
[0051] Feed the fused feature into a convolutional layer to extract temporal information, and finally generate the final classification probability distribution through a fully connected layer.
[0052] Finally, through a weighted fusion strategy, add the results predicted by the two fusion modules to improve the performance of micro-pose recognition.
[0053] Step 3. Construct a loss function and train and fine-tune the model. The specific process is as follows:
[0054] Step 3.1. For the video-skeleton fused feature and the text feature, construct a similarity loss through contrastive learning. First, normalize the fused feature F1 and the text feature F T Then calculate F1 and F TCosine similarity between:
[0055]
[0056] Then, the fused features and text features are aligned through contrastive learning, using a symmetric contrastive loss function:
[0057] Loss from video - skeleton fusion modality to text modality:
[0058]
[0059] Loss from text modality to video - skeleton fusion modality:
[0060]
[0061] where t is the temperature parameter and B is the batch size. Add them together:
[0062]
[0063] This loss function promotes the alignment between modalities by maximizing the similarity between the fused features and the text and minimizing the differences between them.
[0064] Step 3.2: For the result F2 of the fusion of the text modality and the skeleton modality, obtain the predicted class through a fully - connected layer and use cross - entropy loss for training:
[0065]
[0066] where p[i][labels[i]] is the predicted probability corresponding to the true class of the i - th sample.
[0067] The total loss of the model is:
[0068] L = L1+α×L2 (11)
[0069] where α is the weight coefficient of the loss function.
[0070] Step 3.3: During the model training process, a freeze - fine - tuning strategy is adopted to optimize the training efficiency and reduce the computational complexity. In the video encoder and the text encoder, all parameters of the Transformer part are frozen, keeping their pre - trained states unchanged to avoid updating these parameters during training. The role of freezing the Transformer is to maintain its powerful feature extraction ability. After the frozen Transformer part, an adapter module is added to enhance the model's ability. The purpose of the adapter module is to add a small number of parameters for fine - tuning while keeping most of the model parameters unchanged.
[0071] For the video encoder, a spatio-temporal adapter is added. First, the dimension of the input features is halved through a dimensionality reduction convolution; a temporal convolution is used to perform convolution operations in the time dimension, and a convolution kernel size of (3,1) is adopted to capture temporal information; finally, the dimension of the features is restored through a dimensionality increase convolution. The output of the adapter is added to the original features to obtain enhanced features.
[0072] Since text information does not have features in the time and space dimensions, the text adapter is constructed based on a multi-layer perceptron. First, the text features are input into the multi-layer perceptron to obtain enhanced features. Then, through the adapter module, including two linear layers, the dimension is first reduced to 1 / 4 of the original dimension, then the non-linear ability is enhanced through the GELU activation function, and finally it is restored to the original dimension.
[0073] During the training process, only the parameters of the adapter part are updated, and the parameters of the Transformer part are frozen. This approach can reduce the number of model parameters during the training process, reduce the computational complexity, and at the same time effectively fine-tune the adapter module to adapt to specific tasks.
[0074] Step 4: The current human micro-gesture behavior obtained through the camera is output to the trained model to obtain the micro-gesture category to which it belongs.
[0075] Due to the adoption of the technical solution based on multi-modal feature fusion and fine-tuning, the present invention makes full use of the multi-modal information of video, skeleton and text, extracts the features of each modality through a fine-tuning network respectively, and performs feature fusion and interaction in the subsequent stage. Therefore, it effectively makes up for the lack of single-modal information and improves the accuracy and robustness of micro-gesture recognition. In addition, the cross-modal knowledge fusion framework proposed by the present invention can capture the fine-grained features of micro-gestures more deeply, improve the system's perception ability of subtle actions, and still maintain high recognition performance in complex environments. The present invention has strong adaptability and generalization ability, is applicable to multiple application scenarios such as behavior recognition, security monitoring, and human-computer interaction, and has high practical application value.
Claims
1. A micro-gesture recognition method based on multimodal feature fusion and fine-tuning, comprising the following steps: Step 1: Sampling and preprocessing multimodal data; Step 2: construct a micro-gesture recognition model, which includes a multimodal encoder and a trans-membrane state fusion module; Step 3: Construct the loss function and train the fine-tuning model; Step 4: The current human micro-gesture behavior obtained by the camera is output to the trained model to obtain the corresponding micro-gesture category.
2. The micro-gesture recognition method based on multimodal feature fusion and fine-tuning according to claim 1 is characterized in that: The step 1 processes the data of the video, skeleton, and text modes respectively, samples the video data to ensure that all video samples have a consistent frame rate, compresses, crops, and flips the image, uses the OpenPose posture estimation algorithm to detect skeleton key points of the processed image, extracts features through a convolutional neural network, parses the human skeleton information, and formats the micro-posture category text.
3. The micro-gesture recognition method based on multimodal feature fusion and fine-tuning according to claim 1 is characterized in that: The specific process of step 2 is: Step 2.1, construct a multimodal encoder branch network; Step 2.2, construct the video-skeleton fusion module; Step 2.3: Build a text-skeleton fusion module.
4. The micro-gesture recognition method based on multimodal feature fusion and fine-tuning according to claim 1 is characterized in that: In step 2.1, a video encoder is constructed based on the visual Transformer of CLIP, a text encoder is constructed based on the Transformer of CLIP, and a skeleton encoder is constructed based on CTR-GCN.
5. The micro-gesture recognition method based on multimodal feature fusion and fine-tuning according to claim 1 is characterized in that: In the step 2.2, the video-skeleton fusion module adopts a single-layer Transformer architecture. For the video features obtained by the video encoder, the global class tokens are extracted, and the dimension of the token is (T, C), where T represents the time step, and C is the number of channels output by the video encoder. For the skeleton features processed by the skeleton encoder, the feature dimension obtained is (T, C', N), where T represents the time step, C' is the number of channels output by the graph convolutional network, and N is the dimension of the joint point. The skeleton features are adjusted in the channel dimension through linear mapping to ensure that the number of channels of the video features and the skeleton features are consistent. The global class tokens in the video features and the skeleton features after linear mapping are spliced along the third dimension to form fused features. The fused features are input into the Transformer encoder for further encoding to generate the final fused feature representation.
6. The micro-gesture recognition method based on multimodal feature fusion and fine-tuning according to claim 1 is characterized in that: In the step 2.3, for the skeleton features obtained by the skeleton encoder, average pooling is performed on the joint point dimension, the pooled features are mapped to the category space through the fully connected layer, the Softmax function is applied to the output of the fully connected layer, the probability distribution of each category is calculated, the K categories with the highest prediction probability are used as indexes, the text features of the text encoder are used as queries, and the text features of K categories are queried from the text features through the indexes. The text features obtained by the query are copied in the time dimension, and averaged along the dimensions of the K categories to obtain enhanced text feature representation, the skeleton features before pooling and the enhanced text features are merged through a splicing operation to form a fused feature, the fused feature generates a new classification probability distribution through a one-dimensional convolutional layer and a fully connected layer, and is weightedly fused with the original skeleton classification result to obtain the final prediction result.
7. The micro-gesture recognition method based on multimodal feature fusion and fine-tuning according to claim 1 is characterized in that: In the step 3, for the fusion feature F of the video modality and the skeleton modality, a contrastive learning method is used to align F with the encoded text feature, the similarity matrix between the normalized fusion feature F and the text feature is calculated, and a symmetric contrast loss function is constructed. The symmetric contrast loss function is composed of the contrast loss from the fusion feature F to the text feature, and the contrast loss from the text feature to the fusion feature F. The two are added to obtain the loss of contrastive learning. For the fusion feature F' of the text modality and the skeleton modality, a cross entropy classification loss is used to construct a loss function, and the contrast loss and the cross entropy loss are weighted summed to obtain the total loss in the training process.
8. The micro-gesture recognition method based on multimodal feature fusion and fine-tuning according to claim 1 is characterized in that: In the step 3, a freeze-fine-tune strategy is adopted. In the video encoder and the text encoder, all parameters of the Transformer part are frozen, and their pre-trained state is kept unchanged. An adapter module is added after the frozen Transformer part, and fine-tuning is performed through a small number of parameters. The video adapter is based on a spatiotemporal adapter. The dimension of the input feature is halved through dimensionality reduction convolution, the timing information is captured using a time domain convolution operation, and the feature dimension is restored through dimensionality increase convolution. The adapter output is added to the original video feature to obtain an enhanced video feature. The text adapter uses an adapter module based on a multi-layer perceptron. The feature dimension is reduced to 1 / 4 of the original dimension through two linear layers, and the nonlinear ability is enhanced through a GELU activation function. Finally, it is restored to the original dimension. During the training process, the model only fine-tunes the skeleton encoder, the video adapter, and the text adapter.
Citation Information
Patent Citations
Fine-grained action recognition method based on cross-modal knowledge alignment
CN118196888A
Skeleton action recognition method based on prompt type contrast learning and storage medium
CN118865490A
Emotion recognition method based on visual language pre-training and multi-modal collaborative fusion
CN119026071A
Zero-sample multi-mode first-view-angle behavior recognition method introduced based on visual language knowledge
CN119203019A
Multi-modal dialogue emotion recognition method
CN119293740A
Cited By
Human motion posture recognition method and system based on multi-modal data fusion
CN120804842A
Lightweight Transform video action recognition method based on video and text fusion
CN121353996A
Behavior identification and abnormity early warning method and system
CN121768084A
A behavior recognition and abnormality early warning method and system
CN121768084B
Multi-modal gait recognition method based on three-domain fusion and frequency domain guided learning
CN122049990A