Multi-modal human body action recognition method based on comparison feature fusion

By combining wavelet decomposition and a three-branch modal network with an adaptive attention point network, the problem of insufficient detection accuracy in human action recognition is solved, achieving higher recognition accuracy and computational efficiency.

CN121365281APending Publication Date: 2026-01-20CHANGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511509569.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-22
Publication Date
2026-01-20

AI Technical Summary

Technical Problem

Existing technologies suffer from insufficient detection accuracy in human motion recognition, especially in context awareness and fine-grained activity recognition. Traditional methods struggle to effectively capture global trends and the differences in multi-axis data.

Method used

Wavelet decomposition is used to decompose the time-domain signal into a frequency-domain signal, and a three-branch modal network is constructed. Specific frequency signals are highlighted or suppressed by adaptive weight coefficients. Global motion trends are obtained by using convolutional layers, and an adaptive attention network is constructed for feature fusion to dynamically adjust the importance of data on each axis.

Benefits of technology

It improves the accuracy of action recognition, enhances the interpretability and computational efficiency of the model, effectively makes up for the shortcomings of traditional methods in global trend analysis and multi-axis data processing, and achieves higher classification accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121365281A_ABST
    Figure CN121365281A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of action recognition, in particular to a multi-modal human body action recognition method based on comparison feature fusion, which comprises the following steps: acquiring a time domain signal of a human body action; performing frequency domain decomposition on the time domain signal to obtain a frequency domain signal; inputting a frequency domain signal into a three-branch modal network, wherein a first branch highlights or suppresses a certain frequency signal in a time sequence by using an adaptive weight coefficient; the second branch obtains a global motion trend by using a plurality of convolutional layers; the third branch reserves a frequency domain signal; and inputting the output feature vector of the three-branch modal network into the comparison fusion network. The problem that the detection precision of an existing method needs to be further improved is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of action recognition processing, and particularly relates to a multi-modal human action recognition method based on comparative feature fusion. BACKGROUND

[0002] In recent years, with the continuous progress of sensor technology, low-cost inertial measurement units are widely used in smart phones and watches to perceive the motion state of users and provide intelligent services for users.

[0003] Action recognition can be defined as a multi-variate time series classification problem, and conventional algorithms such as KNN and ensemble learning techniques are highly dependent on the production of manual features, which brings a great workload to engineers, and these methods perform well in the recognition of low-level activities such as standing, walking and sitting, but the performance decreases in the recognition of context awareness and fine-grained activities; deep convolutional neural networks often convert measurement data into image data for processing; two-dimensional images inevitably only capture local time series patterns, and although frequency images process frequency information, current research only analyzes each frequency graph, ignoring the influence of global trends on the final accuracy, and standard discrete convolution can only perceive the changes of sampling points at a predefined scale, which may not well cover the range of motion information, and the standard convolution selects multiple sensor axes equally, which reduces the accuracy. SUMMARY

[0004] In view of the deficiencies of the existing method, the present application solves the problem that the detection accuracy of the existing method needs to be further improved.

[0005] The technical scheme adopted by the present application is: a multi-modal human action recognition method based on comparative feature fusion comprises the following steps: Step one, collecting time domain signals of human actions; As a preferred embodiment of the present application, the human actions include running and turning over.

[0006] As a preferred embodiment of the present application, the time domain signals are collected by sensors.

[0007] Step two, frequency domain decomposition is performed on the time domain signals to obtain frequency domain signals; As a preferred embodiment of the present application, the wavelet decomposition method is used for frequency domain decomposition.

[0008] Step three, inputting the frequency domain signals into a three-branch modal network, the first branch uses adaptive weight coefficients to highlight or suppress certain frequency signals in the time series; the second branch uses several convolution layers to obtain global motion trends; and the third branch retains the frequency domain signals. As a preferred embodiment of the present application, the formula of the three-branch modal network is:

[0009] wherein, , , are convolutions with kernel size 1x1, 3x3, 3x3, respectively; is the number of channels, is the number of frequencies, is the length of the time series of the original time-domain signal, , are the weights and biases in the convolution operation, respectively.

[0010] Step four, input the output feature vector of the three-branch modal network into the comparative fusion network; As a preferred embodiment of the present application, the formula of the comparative fusion network is:

[0011] wherein, Flatten is the flattening operation; ; is the 3x3 convolution.

[0012] As a preferred embodiment of the present application, the accuracy and loss value are used to evaluate the model of the fused three-branch modal network and the comparative fusion network.

[0013] As a preferred embodiment of the present application, the multi-modal human action recognition system based on comparative feature fusion comprises: a memory for storing instructions executable by a processor; and the processor is used to execute the instructions to realize the multi-modal human action recognition method based on comparative feature fusion.

[0014] As a preferred embodiment of the present application, the computer readable medium storing computer program code realizes the multi-modal human action recognition method based on comparative feature fusion when the computer program code is executed by the processor.

[0015] The present application has the following advantages: 1. The original signal is decomposed into different frequency time series signals through wavelet transform; 2. The three-branch parallel strategy is used to realize dynamic extraction of frequency weight, capture of cross-sensor global trend and reservation of original time series information; 3. Construct an adaptive attention point network, give each data point an independent weight, dynamically adjust the importance of each axis data, break through the limitations of traditional convolution fixed receptive field and homogeneous processing of multi-axis data, realize cross-axis feature decoupling, at the same time reduce the calculation complexity and enhance the model interpretability through dynamic optimization weight, effectively make up for the defects of traditional method in local time sequence capture, global trend analysis and multi-axis data processing. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1 is a multi-modal human motion recognition method based on comparative feature fusion of the application; Figure 2 is a three-modal principle schematic diagram of the application; Figure 3 is a comparative fusion network schematic diagram of the application; Figure 4 is an accuracy performance diagram of the model of the application on the training set and the test set; Figure 5 is a loss performance diagram of the model of the application on the training set and the test set. DETAILED DESCRIPTION

[0017] The application will be further described below in conjunction with the drawings and examples, the drawings are simplified schematic diagrams, only schematically show the basic structure of the application, therefore only show the structure related to the application.

[0018] Currently, researchers widely adopt two types of conventional methods: one is to convert the time series information collected by sensors into two-dimensional images, and the other is to convert into frequency image information; in the process of converting sensor data into two-dimensional image information, due to the fixed form of two-dimensional images, it is inevitable to capture only partial time series patterns; taking human motion recognition applications as an example, when using inertial sensors to monitor running actions, after being converted into two-dimensional images, it may only present the local correlation of arm swing angle and leg movement amplitude at a certain moment, but it is difficult to reflect the time series continuity of the actions of each part of the body in the entire running period, making the model's understanding of the action one-sided; the other type of method converts data into frequency image information, which can focus on the frequency characteristics of the signal, but current research generally has limitations; researchers often only analyze a single frequency graph independently, ignoring the influence of global trends on the final classification accuracy; in addition, the standard discrete convolution operation can only perceive the changes of sampling points at a pre-defined scale when processing sensor data; this means that its weight range may not be able to comprehensively cover complex and variable motion information, and when processing sensor data containing a mixture of high-speed and low-speed movements, the pre-defined weight range may not be able to effectively capture the differences in motion information at different speeds, resulting in incomplete feature extraction; at the same time, the standard convolution treats multiple sensor axes equally when processing multi-axis sensor data, and fails to extract differentiated features according to the relevance of each axis data to the target activity; taking the monitoring of sleep state by a smart bracelet as an example, among the data collected by the three-axis accelerometer on the bracelet, the Z-axis data is more critical in determining the turning-over action, but the standard convolution treats the three-axis data equally, making it difficult for the model to accurately locate the key information, ultimately leading to a decrease in classification accuracy.

[0019] As shown in Figure 1 , a multi-modal human motion recognition method based on comparative feature fusion includes the following steps: Step 1, collecting time domain signals of human motion; Collecting time domain signals of arm swing angle and leg movement amplitude when running by using sensors; human motion includes but is not limited to running; Step 2, frequency domain decomposition of the time domain signal to obtain the frequency domain signal; Wavelet decomposition is used for frequency domain decomposition, which first expands the original time domain signal through wavelet transform; wavelet transform is one of the most useful mathematical tools in time domain analysis, and wavelet transform is a convolution operation of signal and function; the function is obtained by shifting and scaling a base function; the base function is called mother wavelet, and the function after shifting or scaling is called wavelet; The wavelet function is defined as , with a mean of 0 and existing in a limited duration, by scaling the wavelet function, we get (1) in, , This is called the scaling parameter. This is a transitional value; The wavelet transform formula for a continuous signal is: (2) in, It is a time-domain signal. It is a complex number of the conjugate mother wavelet; From equations 1 and 2, we obtain: (3) The wavelet decomposition operation in Equation 3 is used to transform the motion signals of different sensors into frequency domain signals of different frequencies. like Figure 2 Step 2: Construct a three-branch modal network; First, frequency weights are extracted, and the frequency domain features are globally weighted and averaged to obtain a feature map with important weights. Through learnable weights, the frequency band with high sensitivity is dynamically focused to suppress noise interference and enhance feature representation. The frequency domain signal after wavelet transform is used as the input sequence, and the definition is... For the number of channels, The number of frequencies The length of the original time-domain signal is given by the time series length. The product of the dimension (C, H, B) of the frequency-domain signal and the weight coefficient W1 forms the first branch. An adaptive weight coefficient W1 is created, W1 = (C, H, 1). W1 assigns an independent weight to the combination of "each channel, each height," and then applies this weight to all time steps corresponding to the C and H combinations. Essentially, it assigns a dedicated filtering coefficient to each channel C and each spatial location H. This coefficient acts on the entire time series (B time points) corresponding to that location, achieving overall scaling of the time signal at a specific spatial location, thereby highlighting or suppressing information at specific frequencies in time series processing. The formula for the first branch is as follows: (4) in, This is the output of the first branch. This is the input for the first branch.

[0020] The second branch passes through 1 1. Convolution compresses frequency domain signals from different channels into a single channel, enabling spatial mapping across sensors and allowing neural networks to capture the global motion trends of multi-sensor fusion. The formula for the second branch is: (5) wherein, , are weight and bias in convolution operation respectively; is a convolution operation with kernel size 1x1, responsible for extracting motion features of each sensor; is a convolution operation with kernel size 3x3, responsible for extracting C dimensional reduction to one dimension to obtain the overall trend feature map; is a convolution operation with kernel size 3x3, converting 1 dimension to C dimension, the purpose is to keep the matching with the output data , dimension; , , automatic padding is adopted; The third branch retains the frequency domain signal to ensure that the underlying timing details are not lost, .

[0021] Step three, build a contrast fusion network; The contrast fusion network is shown in Figure 3 ; the calculation formula of the remaining network is: First, define two flat weights as and ; , for example, when C =1, H =3, B =2, then flatten to 6, corresponding to flattenweight1=6; , is the output of the first branch and the second branch after the adaptive attention network, and the adaptive attention point network formula is formula 6, 7; and the output of the contrast fusion is shown in formula 8: (6) (7) (8) In order to avoid the fixed convolution kernel processing data in multiple sensor axes, the neural network cannot be differentiated according to the correlation between the axis data and the target activity, and the adaptive attention point network is created, each data point is given an independent weight after flattening the data, and the weight size can be dynamically adjusted according to the importance of each axis data to the target activity (such as Z-axis data when judging the turning over action), so as to enhance the feature expression of key axis data and suppress non-key axis noise interference; At the same time, its global weighting characteristic can directly model the importance difference between axes, realize feature solution, and the adaptive attention point overcomes the problem that the fixed convolution kernel mechanism of the traditional convolution network has an incomplete receptive field, and the global weighting mechanism can directly model the importance difference between axes, avoid the feature coupling problem caused by the local receptive field limitation of the fixed convolution kernel, and realize the decoupled feature extraction of different axis data; In addition, the weight can be dynamically optimized according to the task demand, compared with the shared weight mode of the fixed convolution kernel, the parameter utilization efficiency is higher, which can not only reduce the calculation complexity through sparse weight, but also improve the model interpretability through weight visualization, and has obvious advantages in key information capture and accurate classification of multi-axis sensor data.

[0022] The original signal is decomposed into different frequency time series signals by wavelet transform, the three-branch parallel strategy is adopted to realize dynamic extraction of frequency weight, capture of global trend between sensors and reservation of original time series information, and an adaptive attention point network is constructed, each data point is given an independent weight, the importance of each axis data is dynamically adjusted, the limitations of traditional convolution fixed receptive field and homogeneous processing of multi-axis data are broken through, cross-axis feature decoupling is realized, and the calculation complexity is reduced and the model interpretability is enhanced through dynamic optimization of weight, effectively making up for the defects of traditional methods in local time series capture, global trend analysis and multi-axis data processing.

[0023] In the embodiment, the three-branch + contrast fusion network is trained for 150 rounds, and the accuracy and loss function value obtained are as shown in Figure 4 It can be seen that good performance is shown on the training set and the test set; The application firstly performs wavelet transform on the original signal, decomposes the multi-sensor motion signal into time sequence signals of different frequencies, and then extracts features by adopting a three-branch parallel strategy: the first branch realizes frequency weight extraction by global weighted average, focuses on the high-sensitive frequency band, and reduces the interference of noise on the final accuracy; the second branch realizes spatial mapping between cross-sensors by using convolution operation, realizes spatial mapping between cross-sensors, and enables the neural network to capture global motion trend; the third branch retains the original data to maintain the underlying time sequence details. Finally, an adaptive attention point network architecture is constructed, which directly models the importance difference between cross-axes by giving each data point an independent weight and dynamically adjusting according to the correlation between data and target activities, breaks through the limitations of traditional convolution kernel, realizes feature decoupling, and finally completes the construction of the contrast fusion network, realizes the key information capture and classification of multi-axis sensor feature data.

[0024] Based on the above ideal embodiments according to the application, through the above description, relevant personnel can make various changes and modifications without deviating from the technical idea of the application. The technical scope of the application is not limited to the contents of the specification, and must be determined according to the scope of the claims.

Claims

1. A multimodal human action recognition method based on contrast feature fusion, characterized in that, Includes the following steps: Step 1: Acquire time-domain signals of human movements; Step 2: Perform frequency domain decomposition on the time domain signal to obtain the frequency domain signal; Step 3: Input the frequency domain signal into the three-branch modal network. The first branch uses adaptive weight coefficients to highlight or suppress a certain frequency signal in the time series; the second branch uses several convolutional layers to obtain the global motion trend. The third branch retains the frequency domain signal; Step 4: Input the output feature vector of the three-branch modal network into the comparison and fusion network.

2. The multimodal human action recognition method based on contrast feature fusion according to claim 1, characterized in that, The formula for a three-branch modal network is: in, , , The convolutions are 1x1, 3x3, and 3x3 respectively; For the number of channels, The number of frequencies The time series length of the original time-domain signal. , These are the weights and the biases, respectively.

3. The multimodal human action recognition method based on contrast feature fusion according to claim 2, characterized in that, The formula for comparing fusion networks is: in, Flatten () indicates the flattening operation; It is a 3x3 convolution; .

4. The multimodal human action recognition method based on contrast feature fusion according to claim 1, characterized in that, Frequency domain decomposition employs wavelet decomposition.

5. The multimodal human action recognition method based on contrast feature fusion according to claim 1, characterized in that, Human movements include: running and rolling over.

6. The multimodal human action recognition method based on contrast feature fusion according to claim 1, characterized in that, Time-domain signals are acquired using sensors.

7. The multimodal human action recognition method based on contrast feature fusion according to claim 1, characterized in that, The models of the fusion three-branch modal network and the contrastive fusion network are evaluated using accuracy and loss values.

8. A multimodal human motion recognition system based on contrast feature fusion, characterized in that, include: Memory is used to store instructions that can be executed by the processor; A processor for executing instructions to implement the multimodal human motion recognition method based on contrast feature fusion as described in any one of claims 1-7.

9. A computer-readable medium storing computer program code, characterized in that, The computer program code, when executed by a processor, implements the multimodal human action recognition method based on contrast feature fusion as described in any one of claims 1-7.