Human body activity identification method based on multi-mode self-supervised network

By using a multimodal self-supervised network for four-way feature extraction and self-supervised learning, the problems of labeled data dependence and insufficient multimodal feature fusion in existing technologies are solved, thereby improving the accuracy and security of human activity recognition.

CN120995268APending Publication Date: 2025-11-21GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511083424.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing methods for human activity recognition rely heavily on labeled data, are costly, and lack sufficient fusion of multimodal features, resulting in low recognition accuracy, especially posing security risks in the recognition of low-frequency activities.

Method used

A multimodal self-supervised network is adopted, which combines time-domain and frequency-domain features from accelerometers and gyroscopes through four-way feature extraction, time-frequency domain fusion mechanism and three sets of self-supervised tasks. The Transformer module is used for feature fusion, and feature alignment is optimized through self-supervised learning. The self-supervised loss function is used for training.

Benefits of technology

It significantly improves the accuracy of human activity recognition, especially in low-frequency activity recognition, reduces costs, and enhances the robustness and security of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120995268A_ABST
    Figure CN120995268A_ABST
Patent Text Reader

Abstract

The invention relates to a human body activity identification method based on a multi-mode self-supervised network, and belongs to the technical field of artificial intelligence. According to the method, the data insight of the network is enhanced by fusing data of different sensors and time-frequency domain data of the sensors. In addition, three self-supervised learning tasks are utilized to assist feature extraction, so that data from different sensors can represent the same activity, an encoder is helped to extract more effective features, and the network performance is improved. And meanwhile, a Transform module is introduced to perform feature fusion, so that the classification performance is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence technology, specifically relating to a method and system for human activity recognition that integrates multimodal sensor data and self-supervised learning, applicable to smart home, health monitoring and other scenarios. Background Technology

[0002] The high dependence of existing technologies on labeled data is a major problem. Current mainstream human activity recognition methods, such as Convolutional Neural Networks (CNN), Long Short-Term Memory Networks (LSTM), and CNN-GRU, heavily rely on large-scale labeled datasets for training, leading to high costs in practical applications. For example, on the public dataset UCI-HAR, traditional CNN models require 50% labeled data to achieve an accuracy of 92.18%. Manually labeling inertial sensor data (such as accelerometer and gyroscope signals) requires professional knowledge, and in scenarios such as medical monitoring, clinical experts are also needed, further increasing the time and economic burden. Semi-supervised learning methods, such as pseudo-labeling techniques, attempt to reduce the need for labeling, but they have obvious drawbacks. For example, setting a fixed confidence threshold can easily discard a large number of unlabeled samples in the early stages of training, resulting in low data utilization. At the same time, the pseudo-label generation process often causes class imbalance, causing the model to favor high-frequency activity categories such as "walking" while ignoring low-frequency key activities such as "falling," resulting in recognition bias and safety risks.

[0003] Insufficient multimodal feature fusion is another key limitation. Existing methods fail to effectively exploit the complementary characteristics of sensor time-domain and frequency-domain information. At the feature engineering level, manually extracting statistical features such as mean or variance cannot capture dynamic time-series patterns, while frequency-domain processing often simply applies Fast Fourier Transform without combining amplitude features to optimize signal power distribution, which limits the representational ability of features. In deep learning model design, mainstream methods such as DeepConvLSTM only process single-modal time-domain signals, ignoring the discriminative value of frequency-domain features such as accelerometer and gyroscope frequency-domain features. Furthermore, feature fusion strategies use shallow concatenation or weighted averaging, which cannot achieve deep interaction between time-frequency features, resulting in the model failing to fully utilize the synergistic advantages of multimodal data. Summary of the Invention

[0004] The core of this invention lies in proposing a human activity recognition method based on a multimodal self-supervised network. Through innovative four-way feature extraction, time-frequency domain fusion mechanism and three sets of self-supervised tasks, the network performance is significantly improved.

[0005] The specific technical solution of this invention is as follows:

[0006] A human activity recognition method based on a multimodal self-supervised network includes the following steps: converting the raw time-domain signals from the accelerometer and gyroscope into frequency-domain signals using a fast Fourier transform; extracting time-domain features, frequency-domain features, and gyroscope features from the accelerometer and gyroscope respectively using four independent encoders; fusing the time-domain and frequency-domain features of the same sensor using a Transformer module; implementing three sets of self-supervised learning tasks to optimize feature alignment; and inputting the fused features into a classifier to output the activity category.

[0007] Furthermore, in the signal preprocessing stage, the original time-domain signals from the accelerometer and gyroscope are first converted into frequency-domain signals through fast Fourier transform, and the amplitude characteristics are obtained by taking the modulus of the frequency-domain transformation result.

[0008] ,

[0009] This amplitude characteristic effectively characterizes the signal power distribution;

[0010] Four independent encoders process four types of input signals: acceleration and timing domain signals. accelerometer frequency domain signal gyroscope time-domain signal gyroscope frequency domain signal , where the subscript i represents the i-th sample in the batch.

[0011] The encoder structure includes an input layer that receives T×C dimensional sensor signals (T is the time step and C is the number of channels), a 64-channel one-dimensional convolutional layer (kernel size 3) for local feature extraction, two-stage processing units, each containing a ReLU activation function and a size 2 max pooling layer for feature dimensionality reduction and receptive field expansion, and a global average pooling layer at the end to compress spatiotemporal features into a 64-dimensional feature vector. The entire structure achieves multi-scale feature extraction through cascaded convolution and pooling operations.

[0012] The time-frequency domain signal is encoded to obtain the time-domain characteristics. , Frequency domain characteristics , , where the subscript i represents the i-th sample in the batch.

[0013] The four encoders have the same structure but independent parameters, ensuring that there is no interference in the extraction of modal features.

[0014] Furthermore, the time-frequency characteristics of the same sensor are fused through a single-layer Transformer module.

[0015]

[0016] The Transformer configuration includes two independent attention heads, a fixed input feature dimension of 64, a feedforward neural network hidden layer dimension of 256, and uses Dropout regularization with a dropout rate of 0.1 to effectively prevent overfitting.

[0017] The entire module contains only a single-layer encoder structure, achieving an optimized balance between performance and computational efficiency.

[0018] Self-supervised learning includes three sets of comparison tasks: accelerometer internal alignment task, gyroscope internal alignment task, and cross-sensor alignment task.

[0019] Among them, the task of accelerometer internal alignment constructs time-domain features. and frequency domain features Comparative learning.

[0020] Gyroscope internal alignment task to construct temporal features and frequency domain features Comparative learning.

[0021] Cross-sensor alignment task to construct accelerometer fusion features Features integrated with gyroscope Comparative learning.

[0022] The self-supervised loss function is defined as a composite structure.

[0023]

[0024] in For the standard cross-entropy loss of the classification task, The contrast loss corresponding to cross-sensor alignment tasks, and These correspond to the time-frequency contrast loss of the accelerometer and gyroscope, respectively.

[0025] α is a dynamic equilibrium parameter, and its optimal value has been verified by experiments to be 0.3.

[0026] Furthermore, the classifier structure used includes an input layer of 128 neurons that receive fused feature vectors, a hidden layer of 256 neurons that sequentially perform batch normalization → ReLU activation → Dropout regularization with a probability of 0.5 during forward propagation, and an output layer of K-dimensional fully connected layers that generate probability distributions for each class through the Softmax function.

[0027] The present invention discloses a human activity recognition method based on a multimodal self-supervised network. By integrating frequency domain data while retaining time domain data from accelerometers and gyroscopes to enrich the classification basis, it also employs a Transformer module for functional fusion to enhance correlation and utilizes self-supervised learning (SSL) to assist in training and enhancing functional interactions, thereby significantly improving the accuracy of HAR.

[0028] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description

[0029] Figure 1 This is a multimodal self-supervised network diagram of a human activity recognition method based on a multimodal self-supervised network in this invention;

[0030] Figure 2 This is a structural diagram of the encoder and classifier in the multimodal self-supervised network of this invention; Detailed Implementation

[0031] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0032] This specific implementation details a human activity recognition method based on a multimodal self-supervised network. This method significantly improves recognition accuracy by integrating frequency domain data, fusing multimodal features, and optimizing through self-supervised learning. Figure 1 As shown, the SSMF model comprises five core modules: domain transformation, four-way feature extraction, Transformer fusion module, three self-supervised tasks, and a classifier.

[0033] The system processing flow is as follows: First, the original time-domain signal is converted into a frequency-domain signal through Fast Fourier Transform (FFT). Then, time-frequency features are extracted separately, and cross-modal fusion is performed through the Transformer module. The feature representation is optimized by combining a self-supervised task. Finally, the processed features are fed into the classifier to obtain the classification result.

[0034] The signal preprocessing stage employs Fast Fourier Transform (FFT) to achieve time-frequency domain conversion. The raw three-axis time-domain signals from the accelerometer and gyroscope are converted as follows:

[0035]

[0036] in This represents frequency domain data with a batch size of B, while This refers to the time domain. To preserve the signal power distribution characteristics, amplitude information is extracted:

[0037]

[0038] Where i ∈ {1, 2, ..., B}. This process will be applied to the frequency domain data of the accelerometer and gyroscope to generate the frequency domain information required for feature extraction. and Then, the data is fed into the encoder for feature extraction to obtain frequency domain features. and .

[0039] Feature extraction employs a four-channel independent encoder structure. For example... Figure 2 As shown in (a), each encoder includes: an input layer that receives T×C dimensional sensor signals (T is the time step and C is the number of channels), a 64-channel one-dimensional convolutional layer (kernel size 3) for local feature extraction, two-stage processing units, each containing a ReLU activation function and a size 2 max pooling layer for feature dimensionality reduction and receptive field expansion, and a global average pooling layer at the end to compress spatiotemporal features into a 64-dimensional feature vector. The entire structure achieves multi-scale feature extraction through cascaded convolution and pooling operations.

[0040] The time-frequency feature fusion uses a single-layer Transformer module. The configuration parameters include: a fixed input feature dimension of 64, a multi-head self-attention mechanism with two independent attention heads, a feedforward neural network hidden layer dimension of 256, and a Dropout regularization technique with a dropout rate of 0.1 to prevent overfitting. The entire module contains only a single-layer encoder structure to balance performance and computational efficiency.

[0041]

[0042] The output accelerometer features are then fused. Features integrated with gyroscope A 128-dimensional fused feature vector is formed by cascading along the feature dimensions. , The input is a multilayer perceptron classifier for end-to-end activity classification, and the output dimension of the classifier is equal to the number of activity categories K.

[0043] Self-supervised learning includes three sets of comparative tasks: accelerometer time-frequency alignment, gyroscope time-frequency alignment, and cross-sensor alignment.

[0044] Among them, acceleration timing frequency alignment is maximized. and Cosine similarity.

[0045] Among them, gyroscope time-frequency alignment, maximizing and Cosine similarity.

[0046] Among them, gyroscope time-frequency alignment, maximizing and Cosine similarity.

[0047] The total loss function used in this invention

[0048]

[0049] in and Let B represent the predicted label and the actual label of the i-th instance, respectively, and B represent the batch size.

[0050]

[0051] D ∈ {A, G} represents the features of the accelerometer (A) or gyroscope (G), and X ∈ {T, F, W} represents time-domain information (T), frequency-domain information (F), or fused information (W). For the self-supervised loss of the accelerometer, yes and For gyroscopes, it is and In the self-supervised loss of different sensors, for and .

[0052] Our experiments used an RTX 3090 GPU and PyTorch for deep learning. The Adam optimizer was employed with a weight decay of 1×10⁻⁷ and a momentum of 0.9. The learning rate started at 0.001 and halved every 20 epochs. Training was performed for over 200 epochs, and performance was evaluated on the final epoch. The batch size for each dataset was 64, and evaluation was conducted on the validation set. The initial temperature t for the three self-supervised learning (SSL) tasks was set to 0.07, with SSL loss weights of 0.3 for different sensors and 0.35 for the two tasks targeting the same sensor. The network was run on a remote server.

[0053] We evaluate model performance using accuracy and weighted F1 score. Accuracy measures the proportion of correct predictions and is suitable for balanced data; the weighted F1 score adjusts the F1 score according to the sample size of each class to handle class imbalance.

[0054] We compared our model with seven other models using the UCI-HAR, mHealth, and PAMAP2 datasets, and the quantification results are shown in Table 1.

[0055] Table 1 compares the method described in this invention (SSMF) with seven other methods (bold numbers indicate best performance).

[0056] This invention proposes a sensor-based HAR method that enhances the network's data insight by fusing data from different sensors and their time-frequency domain data. It also utilizes three self-supervised learning tasks to assist feature extraction, enabling data from different sensors to represent the same activity and helping the encoder extract more effective features, thereby improving network performance. Furthermore, a Transformer module is introduced for feature fusion to further improve classification performance.

[0057] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications can be made to the technical solutions of the present invention without departing from the spirit and scope of the present invention, and all such modifications should be covered within the scope of the claims of the present invention.

Claims

1. A method for human activity recognition based on a multimodal self-supervised network, characterized in that... include: The method includes the following steps: 1) Convert the raw time-domain signals from the accelerometer and gyroscope into frequency-domain signals using a fast Fourier transform; 2) Four independent encoders are used to extract the acceleration time domain features, accelerometer frequency domain features, gyroscope time domain features, and gyroscope frequency domain features from the time and frequency domain signal in 1). 3) The time-domain and frequency-domain characteristics of the same sensor from step 2) are fused using the Transformer module; 4) Implement three sets of self-supervised learning tasks to optimize feature alignment; 5) Input the fused features into the classifier to output the activity category.

2. The human activity recognition method based on a multimodal self-supervised network according to claim 1, characterized in that: Frequency domain signal processing includes: obtaining amplitude characteristics by taking the modulus of the Fast Fourier Transform result. ,in As the result of frequency domain transformation, the amplitude feature reflects the signal power distribution and serves as the input for frequency domain feature extraction.

3. The human activity recognition method based on a multimodal self-supervised network according to claim 1, characterized in that: Feature extraction involves four independent encoders processing four types of input signals in parallel: the first encoder processes acceleration and timing domain signals. Output time domain features The second encoder processes the accelerometer frequency domain signal. Output frequency domain characteristics The third encoder processes the time-domain signal from the gyroscope. Output time domain features The fourth encoder processes the frequency domain signals from the gyroscope. Output frequency domain characteristics , where the subscript i represents the i-th sample in the batch, and each encoder has the same structure but independent parameters.

4. The human activity recognition method based on a multimodal self-supervised network according to claim 1, characterized in that: The self-supervised task includes an accelerometer internal alignment task to construct temporal features. Frequency domain characteristics Contrastive learning is used to maximize the similarity of sample features, and the gyroscope intra-alignment task constructs temporal features. Frequency domain characteristics Contrastive learning optimizes the consistency of homologous features and constructs accelerometer fusion features across sensor alignment tasks. Features integrated with gyroscope Cross-modal contrastive learning to facilitate feature space alignment of heterogeneous sensors.

5. The human activity recognition method based on a multimodal self-supervised network according to claim 1, characterized in that: The self-supervised loss function is defined as follows: in For the standard cross-entropy loss of classification tasks, The contrast loss corresponding to cross-sensor alignment tasks, and The time-frequency contrast loss corresponds to that of the accelerometer and gyroscope, respectively. α is a dynamic balance parameter and its optimal value has been verified to be 0.3 through experiments.

6. The human activity recognition method based on a multimodal self-supervised network according to claim 1, characterized in that: Feature fusion includes fusing time-frequency features from the same sensor using a single-layer Transformer encoder.

7. Where TF(⋅) represents the Transformer fusion operation.

8. The human activity recognition method based on a multimodal self-supervised network according to claim 1, characterized in that... Classification decisions include fusing accelerometer features from the Transformer output. Features integrated with gyroscope A 128-dimensional fused feature vector is formed by cascading along the feature dimensions. , The input is a multilayer perceptron classifier for end-to-end activity classification, and the output dimension of the classifier is equal to the number of activity categories K.