Human body activity identification method based on bidirectional cross-modal attention mechanism

Through bidirectional cross-modal attention mechanism and adversarial contrast learning, combined with millimeter wave radar and RGB image data, the accuracy problem of multimodal human activity recognition in complex environments is solved, and high recognition accuracy and adaptability are achieved.

CN120260119APending Publication Date: 2025-07-04HUNAN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510308492.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing multimodal human activity recognition technology has low recognition accuracy in complex environments and is difficult to fully mine complementary information of multimodal data, especially in the face of environmental and user changes.

Method used

Using a method based on a two-way cross-modal attention mechanism, combining millimeter wave radar and RGB image data, pre-trained and fine-tuned models are constructed through unsupervised confrontational learning and bidirectional cross-modal attention mechanism to extract domain-independent consistent information and complementary information of multimodal data.

Benefits of technology

It achieves high recognition accuracy in multiple scenarios and multi-user conditions, adapts to complex environment changes, and improves the accuracy of human body activity recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260119A_ABST
    Figure CN120260119A_ABST
Patent Text Reader

Abstract

The invention relates to a human body activity identification method based on a bidirectional cross-modal attention mechanism, and the method comprises the following steps: firstly, collecting an activity data set (such as millimeter wave radar data, RGB image data and the like) of a user at the same time through employing a plurality of wireless devices, carrying out the preprocessing, and dividing the training data into unmarked data and marked data; training a pre-training model built based on an unsupervised confrontation contrast learning technology by adopting unmarked data; thirdly, finely adjusting the model based on the bidirectional transmembrane state attention mechanism by using the mark data to obtain a human body activity recognition model; and finally, inputting multi-modal data in a user or scene which is not used for training into the model to finish accurate recognition. The domain adaptivity of the model is improved through an adversarial contrast learning method, and a bidirectional cross-modal attention mechanism is designed to enhance information fusion between modals, so that the method is suitable for human body activity recognition tasks with only a small number of labeled samples in various complex environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of pervasive computing and multimodal sensing and recognition, and particularly relates to a human activity recognition method based on a bidirectional cross-modal attention mechanism. Background Art

[0002] In recent years, due to the performance leap of deep learning algorithms and the significant progress of sensor technology, important progress has been made in human activity recognition (HAR) technology. The HAR technology aims to collect data generated by human activities through various sensors, and through the analysis and processing of deep learning algorithms, accurately identify specific human activities. The HAR technology has been widely applied in fields such as smart furniture, healthcare, and digital sports, and is expected to become one of the basic functions of various emerging applications in the future. For example, in the field of digital sports, the HAR technology can detect the standard degree of the user's movements throughout the exercise, correct incorrect postures in a timely manner, and avoid sports injuries.

[0003] Currently, the single-modal HAR technology has been very mature. However, due to the high complexity and dynamics of human activities, it is difficult to completely capture the data information of these activities using only a single sensor, which will result in a lower recognition accuracy of the single-modal HAR when dealing with some complex activities. In addition, the data information collected by a single sensor has certain defects due to the limitations of the sensor itself design. For example, the gyroscope is more sensitive to angular velocity and azimuth changes and has better performance in recognizing dynamic activities than static activities; while the accelerometer has better performance in recognizing static activities than dynamic activities due to static force factors. Therefore, the concept of multimodal data fusion is proposed to improve the performance of activity detection by using the complementary relationship between various sensors, such as the data fusion of RGB cameras and millimeter-wave radars. Previous studies have shown that building HAR models with multimodal data is promising and can improve the performance of HAR.

[0004] Then, HAR in the real-world scenario still faces practical challenges and application gaps when using multi-modal data methods. On the one hand, in practical HAR applications, environmental (domain) information has an adverse effect on recognition performance. Domain information, such as user changes or background interference, may introduce biases during the feature extraction process, thereby reducing system performance. This problem is more severe in multi-modal HAR because multi-modal data collection requires the use of multiple sensors, which inevitably collect domain information. On the other hand, due to the different physical characteristics of different types of sensors, highly heterogeneous information about the same activity can be captured. For example, millimeter-wave radar presents data in the form of point clouds or range-Doppler maps, reflecting target distance, speed, etc.; RGB images describe the scene through visual dimensions such as color, shape, and texture. There are significant differences in their structures and contents. These heterogeneous information make it extremely difficult to mine complementary information between modalities. Existing multi-modal HAR systems only use self-attention mechanisms, focusing on intra-modal correlations and ignoring cross-modal interactions; or adopt single cross-modal attention, only relating modalities from a specific perspective and failing to fully mine complementary information. Summary of the Invention

[0005] In view of the above problems, the object of the present invention is to propose a human activity recognition method based on a bidirectional cross-modal attention mechanism, aiming to achieve the domain adaptation ability of the model in the face of various complex environments; at the same time, considering the small number of labeled samples in actual situations, a two-stage training strategy is proposed: a pre-training stage and a fine-tuning stage; and a bidirectional cross-modal attention mechanism is designed to fully mine the complementary information between multi-modalities in limited labeled samples, thereby improving the accuracy of human activity recognition.

[0006] In the first aspect of the present invention, a human activity recognition method based on a bidirectional cross-modal attention mechanism is provided. Taking millimeter-wave radar data and RGB image data as examples, the method includes the following steps:

[0007] Simultaneously collect the activity data set of the user (such as millimeter-wave radar data, RGB image data, etc.) using multiple wireless devices and perform preprocessing, and at the same time divide the training data into unlabeled data and labeled data;

[0008] Use the unlabeled data to train a pre-training model built based on unsupervised adversarial contrast learning technology;

[0009] Use the labeled data to fine-tune the model based on the bidirectional cross-modal attention mechanism to obtain a human activity recognition model;

[0010] Input the multi-modal data of users and scenarios not used for training into the model to complete accurate recognition;

[0011] Further, the multimodal data preprocessing method includes: acquiring the original radar signals scanned by the millimeter-wave radar, extracting the active distance information and velocity information based on the radar signals through Fourier transform (FFT) in the fast-time and slow-time dimensions, and then constructing an angular spatial spectrum using the Multiple Signal Classification (MUSIC) algorithm to obtain the active angle information; and performing multi-frame accumulation and normalization processing along the time dimension to generate complete three-frame radar data (RVA). Frame extraction is performed on the acquired RGB image data, and 16 frames of image data are extracted from the original image data at equal intervals. The radar data and image data obtained in this way are used as the training set.

[0012] Further, a pre-training model based on unsupervised adversarial contrastive learning technology is constructed. The contrastive learning network consists of a radar feature encoder module, an image feature encoder module, a projection neural network module, and a fusion enhancement module, which is used to extract the consistent information of multimodal data. Then, an adversarial learning network is formed by the domain discriminator module and the contrastive learning network, and an adversarial contrastive learning network is built in this way to force the feature encoder to extract domain-agnostic consistent information.

[0013] Further, a large amount of unlabeled multimodal data is used for training to obtain a pre-training model that can extract domain-agnostic consistent information.

[0014] Further, an activity recognition model based on a bidirectional cross-modal attention mechanism is constructed, which consists of a radar feature encoder module, an image feature encoder module, a bidirectional cross-modal attention module, and a classifier module. The radar feature encoder module and the image feature encoder module load the feature encoder parameters of the pre-training model, while the bidirectional cross-modal attention module fully exploits the complementary information between different modalities.

[0015] Further, a small amount of labeled multimodal data is used to fine-tune the model to obtain the final activity recognition model.

[0016] Further, data different from that of the users or scenarios in the training set is used as test data, and the test data is input into the trained model to calculate the category probability distribution of the target activity. The sum of the probabilities of all categories is 1; the activity category with the highest probability value is selected as the final recognition result.

[0017] In the second aspect, the present invention also provides a human activity recognition system based on a bidirectional cross-modal attention mechanism. Taking millimeter-wave radar data and RGB image data as examples, it includes:

[0018] A data acquisition module, which is used to acquire millimeter-wave radar data and RGB image data in multiple scenarios and multiple users, perform preprocessing, and divide the data into unlabeled data and labeled data;

[0019] A pre-training module, which is used to construct an unsupervised adversarial contrast learning network model, including a contrast learning network and an adversarial learning network architecture, and is trained using unlabeled data to obtain a pre-trained model;

[0020] A transfer learning module, which is used to construct a neural network model based on a bidirectional cross-modal attention mechanism, uses attention mechanisms both within the modality and between multiple modalities, and is trained using a small amount of labeled sample data, thereby obtaining an activity recognition model.

[0021] An activity recognition module, which inputs the activity information of users or scenarios not used for training into the activity recognition model, thereby achieving accurate recognition of activity categories.

[0022] Furthermore, the method for the data acquisition module to obtain activity data includes: obtaining original multi-modal data, dividing the training data into unlabeled data and labeled data; extracting the distance, speed, and angle features of activity actions based on the millimeter-wave radar data therein, and obtaining RVA through multi-frame accumulation and normalization; performing frame extraction on the RGB image data therein to obtain 16-frame image data; the preprocessed millimeter-wave radar data and RGB image data form a training set.

[0023] The structure of the pre-trained model includes: a radar feature encoder, an image feature encoder, a projection neural network module, a fusion enhancement module, and a domain discriminator module.

[0024] Furthermore, the pre-trained model is trained based on the unlabeled sample data in the training set to obtain a neural network model capable of extracting domain-independent multi-modal consistent information.

[0025] The structure of the neural network model for transfer learning includes: a radar feature encoder, an image feature encoder, a bidirectional attention module, and a classifier module.

[0026] Furthermore, the neural network model is trained based on the labeled sample data in the training set, and an activity recognition model is obtained through training.

[0027] Furthermore, the steps for the activity recognition module to recognize activity categories include: based on the trained activity recognition model, calculating the probabilities of all activity categories of the input activity action data, and selecting the activity category with the highest probability as the final recognition result.

[0028] The beneficial effects of the present invention:

[0029] Cross-domain adaptability: The activity recognition method of the present invention can adapt to activity recognition tasks in multiple scenarios and multiple users, and solves the problem of data differences caused by changes in the environment and user habits in existing methods.

[0030] Self-supervised learning: Introduce transfer learning methods into human activity recognition, use a large amount of unlabeled sample data to train a pre-trained model, and use partially labeled data to train an activity recognition model. This is in line with the actual situation in reality where data labeling is difficult and only a small amount of labeled data exists.

[0031] High recognition accuracy: By combining self-attention and cross-modal attention in deep learning, the present invention can enhance its own feature representation while capturing complementary information between different modalities, thereby achieving accurate activity recognition. Brief Description of the Drawings

[0032] Figure 1 It is a flowchart of a human activity recognition method based on a bidirectional cross-modal attention mechanism provided by an embodiment of the present invention;

[0033] Figure 2 It is a flowchart of collecting and preprocessing original multi-modal data provided by an embodiment of the present invention;

[0034] Figure 3 It is a research framework of human activity recognition based on a bidirectional cross-modal attention mechanism provided by an embodiment of the present invention;

[0035] Figure 4 It is a structural diagram of an adversarial contrast learning network provided by an embodiment of the present invention;

[0036] Figure 5 It is a structural diagram of a bidirectional cross-modal attention module provided by an embodiment of the present invention.

[0037] Figure 6 It is an example diagram of a confusion matrix tested according to an embodiment of the present invention. Detailed Embodiments

[0038] In order to make the purpose and advantages of the present invention clearer, the present invention will be specifically described below in conjunction with embodiments and with reference to the accompanying drawings. It should be understood that the following text description is only an example of the specific implementation manner of the present invention and does not strictly limit the scope of protection of the specific claims of the present invention.

[0039] Figure 1 It is a flowchart of a human activity recognition method based on a bidirectional cross-modal attention mechanism according to an embodiment of the present invention; as Figure 1 shown, it includes operations S101 - S104.

[0040] In operation S101, use a variety of wireless devices to simultaneously collect the user's activity dataset (such as millimeter-wave radar data, RGB image data, etc.) and perform preprocessing, and at the same time divide the training data into unlabeled data and labeled data;

[0041] In operation S102, a pre-training model built based on the unsupervised adversarial contrastive learning technique is trained using unlabeled data;

[0042] In operation S103, the model based on the bidirectional cross-modal attention mechanism is fine-tuned using labeled data to obtain a human activity recognition model;

[0043] In operation S104, the multi-modal data of users and scenarios not used for training are input into the model to complete accurate recognition.

[0044] Through the method provided above, the millimeter-wave radar device and the RGB camera are used to detect the activity actions of the target user, and the obtained multi-modal data passes through the constructed neural network model to achieve accurate activity recognition.

[0045] Figure 2 It is a flowchart for collecting and preprocessing the original multi-modal data provided by an embodiment of the present invention; as Figure 2 shown, collecting the millimeter-wave radar activity action information and processing it into RVA includes steps S201 - step S203, and the image data frame sampling operation is step S204. Taking 2.5 seconds as an observation time for a human activity action, where the millimeter-wave radar will collect 64 frames of radar signals, and the RGB camera records the human activity video at a data rate of 20 frames per second.

[0046] In operation S201, the millimeter-wave radar device and the RGB camera are used to collect multi-modal data of activity actions in multiple scenarios and multiple users, and the training data is divided into unlabeled data and partially labeled data;

[0047] In operation S202, the millimeter-wave radar data is subjected to Fourier transform in the time dimension and the multiple signal classification algorithm is used to construct the angular spatial spectrum to obtain distance, speed, and angle parameters;

[0048] In operation S203, the distance, speed, and angle parameters are subjected to multi-frame accumulation and normalization processing in the time dimension to obtain three-frame radar data RVA, which is used to reflect human activities;

[0049] In operation S204, compared with the 3-frame radar data, the number of frames of the image data is too large, which is not suitable for the dual-threaded operation of synchronously processing the image data and the radar data. Therefore, a frame sampling operation is performed on the image data, and 16 frames of image data are extracted to reflect human activities.

[0050] Through the method provided above, the RVA data can be extracted from the millimeter-wave radar data using the Doppler transform in the time dimension and the multiple signal classification algorithm, and 16 frames of image data are obtained from the image data using frame sampling. Finally, the frame-sampled image data and the RVA data form a multi-modal data set as the input of the model.

[0051] Figure 3 This is the research framework for human activity recognition based on a bidirectional cross-modal attention mechanism provided by an embodiment of the present invention. As Figure 3 shown, it includes a pre-trained model and an activity recognition model. The process of obtaining the activity recognition result will be introduced in detail below according to Figure 3 this.

[0052] For the pre-trained model, the model is trained using unlabeled multimodal data. First, a radar feature encoder (R) and an image feature encoder (C) are used to extract features from the multimodal data to obtain a feature representation (z); secondly, the feature representation is mapped to the same space through a projection neural network to obtain a feature vector (r); then, based on a fusion-based feature enhancement module, these feature representations are widely enhanced into a set of fused features (v); finally, we perform contrastive learning on these fused features to strengthen the consistency between modalities; at the same time, a domain classifier is trained to identify the domain to which the features belong, and the two form a dynamic balance through an adversarial manner to ensure that the features retain both consistency and domain independence.

[0053] For the activity recognition model, only partial labeled multimodal data is used to train the model. First, we load the feature encoder of the first stage and fine-tune it. Secondly, we design a novel bidirectional cross-modal attention module, which can capture complementary information between different modalities from the limited labeled multimodal data, while balancing the consistency between modalities and the unique features within modalities. For example, when monitoring the activity of "having lunch", the RGB image will show the sitting posture of the subject, while the millimeter-wave radar can capture the arm movements of the subject during eating. Then, the classifier processes the fused features and outputs the prediction result of the target classification.

[0054] Figure 4 This is the structural diagram of the adversarial contrastive learning network provided by an embodiment of the present invention. The adversarial contrastive learning network will be introduced in detail below according to Figure 4 this.

[0055] Based on the fusion-based feature enhancement module, the projected features can be used to generate a set of fused features by weighted linking or weighted summation.

[0056] These fused features are used for contrastive learning, enabling the feature encoder to generate features that are invariant to different fusion schemes in order to extract multimodal consistent information. Therefore, the system needs to push the enhanced features (positive samples) from the same original multimodal sample closer and move the enhanced features (negative samples) from different original multimodal samples farther away. For this purpose, we design the following contrastive loss function:

[0057]

[0058] where v sFor enhanced features. S represents any enhanced feature, · represents the inner product of vectors, s represents the anchor point, P(s) represents the enhanced feature (the positive sample set of the anchor point) from the same original multimodal sample, p represents the index of the positive sample, and a represents the index of the negative sample.

[0059] Meanwhile, contrastive learning is introduced into the adversarial framework, and a domain classifier is also designed in the framework, whose purpose is to identify the domain labels of the data. Then, by reducing the differences between data in different domains, the fused features are forced not to depend on a specific domain. The domain classifier consists of 3 fully connected layers (FC), where the last fully connected layer classifies the domain labels of the data. The loss function of the domain classifier is as follows:

[0060]

[0061] where X u represents the amount of data, D represents the number of environments (domains), Y id is the true domain label, and P id is the predicted output of the domain discriminator.

[0062] In order to make the extracted features be able to deceive the domain classifier while maximizing the ability of the feature encoder to extract consistent information. We need to maximize the loss function of the domain classifier and minimize the loss function of contrastive learning. Considering the above problems, we construct the loss function of the following pre-trained model:

[0063] L = Loss conf -γLoss d .

[0064] where γ is a coefficient. By minimizing the overall loss function L as the final goal, the feature encoder is forced to learn the domain-independent consistent information of all input data.

[0065] Figure 5 This is the structural diagram of the bidirectional cross-modal attention module provided by the embodiment of the present invention. The following will be based on Figure 5 to introduce in detail the processing process of generating fused features from the features of each modality.

[0066] First, we input the extracted features into the self-attention module (SA), calculate the attention for each modality respectively, so as to strengthen the representation within each modality. The formula of SA is as follows:

[0067]

[0068]

[0069]

[0070]

[0071] Among them represents the feature input of SA, and d represents the dimension of the input vector. represents different representations of the input, and U qkv is a learnable transformation matrix, and norm(·) represents normalization.

[0072] Secondly, a cross-modal attention module (CA) is introduced, and the enhanced features obtained are input into it. The formula of CA is as follows:

[0073]

[0074]

[0075]

[0076]

[0077] Among them, i represents different modalities. CA enhances the interaction and complementary information between modalities. It focuses on the irrelevant information between different modalities by querying from one modality and having keys and values from another modality, combined with the re-softmax activation function.

[0078] Finally, the feature vectors of each modality passing through CA are concatenated to obtain a fused feature vector, which is used as the data input for the activity classifier.

[0079] Figure 6 is an example diagram of the confusion matrix tested according to the embodiments of the present invention.

[0080] Combined with Figure 6 To further illustrate the effectiveness of the above system. The public multi-modal dataset UTD consists of 8 users and 27 activities. Each activity needs to be collected 4 times repeatedly, and there are a total of 864 samples. In this experiment, the public multi-modal dataset UTD is used. 6 users are used as training data, and the remaining 2 users are used as test data; and 5% of the training data is used as labeled data, and the remaining training data is used as unlabeled data. For the convenience of display, 8 activities are randomly selected to verify the accurate performance of the system under the condition of user change. The experimental results show that the average recognition accuracy of the system can reach 64.44% in the scenario of new users.

[0081] In summary, the present invention uses an adversarial contrastive learning network for unsupervised learning tasks, and through the collaborative optimization of the adversarial-contrast dual objectives, it realizes the unification of consistent feature representation and domain robustness; at the same time, a bidirectional cross-modal attention module is used in the supervised learning task, so that each modality not only focuses on its own features, but also can focus on the features of other modalities, thereby achieving a deeper level of fusion. For the activity recognition model trained by this method, when using the new user activity action data in the new scenario to input into the activity recognition model, the activity action category can be obtained quickly and accurately.

Claims

1. A human activity recognition method based on a bidirectional cross-modal attention mechanism, comprising the following steps: Simultaneously collect the activity data set of the user (such as millimeter-wave radar data, RGB image data, etc.) using multiple wireless devices and perform preprocessing, and at the same time divide the training data into unlabeled data and labeled data; Train a pre-training model built based on the unsupervised adversarial contrast learning technique using the unlabeled data; Fine-tune the model based on the bidirectional cross-modal attention mechanism using the labeled data to obtain a human activity recognition model; Input the multi-modal data of the user or scenario not used for training into the model to complete accurate recognition.

2. The human activity recognition method based on a bidirectional cross-modal attention mechanism according to claim 1, taking millimeter-wave radar data and RGB image data as an example, the steps of constructing the training set include: Use a millimeter-wave radar device and an RGB camera to collect multi-modal data of activity actions of multiple scenarios and multiple users, and divide the training data into unlabeled data and labeled data; Based on the original millimeter-wave radar data, obtain distance, speed, and angle information, perform multi-frame accumulation and normalization processing on the distance, speed, and angle parameters in the time dimension, and integrate to obtain radar data RVA (RTM, VTM, ATM); Based on the RGB image data, since the number of frames of the acquired image data is too large, 16 frames of image data are extracted at equal intervals from the total amount of collected image data.

3. The human activity recognition method based on a bidirectional cross-modal attention mechanism according to claim 1, taking millimeter-wave radar data and RGB image data as an example, construct a pre-training model built based on the unsupervised adversarial contrast learning technique. Step 3-1: The contrast network consists of a radar feature encoder module, an image feature encoder module, a projection neural network module, and a fusion enhancement module; its contrast loss function is designed as: where v s is an enhanced feature. S represents any enhanced feature, · represents the vector inner product, s represents the anchor point, P(s) represents the enhanced features (the positive sample set of the anchor point) from the same original multimodal sample, p represents the index of the positive sample, and a represents the index of the negative sample. Contrastive learning forces the network to extract features containing multimodal consistent information, which requires narrowing the distance between positive samples while moving negative samples away, thus minimizing the contrastive loss function. Step 3-2: Construct a domain discriminator module, the purpose is to identify the domain label of the data, and then force the fused features to be domain-independent by reducing the differences between domain data. Its domain loss function can be designed as: Where X u represents the data volume of the data, D represents the number of environments (domains), Y id is the true domain label, P id is the predicted output of the domain discriminator. Backpropagation and training are performed on the output results to blur the difference between domains and domain data, and it is necessary to maximize the domain loss function. Step 3-3: An adversarial network is formed between the contrast learning network and the domain discriminator module. Through adversarial learning, the network finds a balance between extracting multi-modal consistent information and domain-independent information, ensuring that the extracted features can deceive the domain discriminator while maximizing the results of contrast learning. For this purpose, the following loss function is constructed: L = Loss conf -γLoss d where γ is a coefficient. By minimizing the overall loss function L as the final goal, force the feature encoder to learn the domain-independent multi-modal consistent information of all input data.

4. The human activity recognition method based on a bidirectional cross-modal attention mechanism according to claim 1, taking millimeter-wave radar data and RGB image data as an example, construct a human activity recognition model based on the bidirectional cross-modal attention mechanism. It consists of a radar feature encoder module, an image encoder module, a bidirectional cross-modal attention module, and a classifier module. Among them, the radar feature encoder module and the image feature encoder module load the feature encoder parameters of the pre-training model, and the bidirectional cross-modal attention module is used to fully explore the complementary information between different modalities.

5. A human activity recognition method based on a bidirectional cross-modal attention mechanism according to claim 4, which focuses on a bidirectional cross-modal attention module for generating fused features, including two self-attention modules and two cross-modal attention modules. Step 5-1: The unimodal features extracted by the feature encoder pass through their respective self-attention modules to obtain enhanced feature representations. The self-attention weights can be expressed as: Step 5-2: The two enhanced features are jointly input into the two cross-modal attention modules, and finally an enhanced feature representation that fuses the complementary information of the two modalities is obtained. The cross-modal attention weights can be expressed as: Among them, i represents different modalities. Step 5-3: The finally obtained enhanced features are concatenated to obtain fused features.

6. A human activity recognition method based on a bidirectional cross-modal attention mechanism according to claim 1, wherein the steps of identifying human activity categories include: Step 6-1: Obtain the possible probabilities of each activity category according to the activity recognition model, and the sum of the activity probabilities of all categories is 1; Step 6-2: Select the activity category with the highest probability as the finally classified activity category.