Human motion perception method and related equipment based on three-dimensional hybrid attention model
Through the human body motion perception method of the three-dimensional hybrid attention model, non-local attention modules and cross attention modules are used to extract features from wireless signal spectrum data, solving the problem of difficult deployment of traditional models on resource-limited devices, and achieving efficient human body motion perception, which is suitable for scenarios such as smart home and smart medical care.
Patent Information
- Application Number
- CN202510687548.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-05-27
AI Technical Summary
Traditional deep learning models require a large number of parameters when processing high-dimensional wireless signal data, making it difficult to effectively deploy on hardware devices with limited resources, especially on mobile devices and small sensor nodes, resulting in a huge consumption of memory and computing resources.
The human body movement perception method based on the three-dimensional hybrid attention model is adopted. Through the three-dimensional hybrid attention mechanism, including the non-local attention module and the cross-attention module, the attention weight is calculated from the three dimensions of time, frequency and sub-carrier, and the three-dimensional attention feature map is generated, and the cross-attention is fusion to output the human body movement perception features.
With limited hardware resources, the recognition efficiency of human motion perception is greatly improved, allowing human motion perception technology to be effectively deployed on small devices in scenarios such as smart homes and smart medical care, reducing costs and energy consumption, and improving recognition accuracy and adaptability.
Smart Images

Figure CN120216966B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of human motion perception, and in particular to a human motion perception method and related equipment based on a three-dimensional hybrid attention model. Background Art
[0002] In the field of intelligent perception, wireless signals from Wi-Fi or millimeter-wave radar, due to their high channel count and large data volumes, provide rich data support for accurate human motion perception. However, this advantage also presents a significant challenge: traditional deep learning models require a large number of parameters to analyze features when processing such complex data, making effective deployment difficult on resource-constrained hardware. While traditional models can extract key features from high-dimensional data, the sheer number of parameters leads to significant consumption of memory and computing power in specific applications, especially on resource-constrained hardware such as mobile devices and small sensor nodes.
[0003] To address this issue, many past studies have relied on highly parameterized deep learning models to extract features from high-dimensional wireless signals, but this results in significant inefficiencies in terms of model size and inference time. Therefore, improving the efficiency of human motion perception and recognition to enable real-time perception for resource-limited IoT devices has become a significant challenge. Summary of the Invention
[0004] The main purpose of this application is to provide a human motion perception method and related equipment based on a three-dimensional hybrid attention model, aiming to solve the technical problem of low efficiency of human motion perception and recognition in resource-limited IoT devices.
[0005] To achieve the above objectives, this application proposes a human motion perception method based on a three-dimensional hybrid attention model, the method comprising:
[0006] Obtain wireless signal spectrum data;
[0007] The wireless signal spectrum data is feature processed according to a preset three-dimensional hybrid attention model to determine the human motion perception features, wherein the three-dimensional hybrid attention model first performs feature compression on the wireless signal spectrum data, and then calculates attention weights from the three dimensions of time, frequency, and subcarrier based on a non-local attention mechanism to generate a three-dimensional attention feature map. Finally, the three-dimensional attention feature map is fused through cross-attention, and the human motion perception features are output after dimension adjustment;
[0008] A human motion perception recognition result is determined according to the human motion perception feature.
[0009] In one embodiment, the step of performing feature processing on the wireless signal spectrum data according to a preset three-dimensional hybrid attention model to determine the human motion perception feature includes:
[0010] Performing feature compression on the wireless signal spectrum data using a preset first separable convolution model to generate an initial compressed feature map;
[0011] Inputting the initial compressed feature map into a preset non-local attention model, performing embedding space mapping along the time dimension, frequency dimension, and subcarrier dimension respectively, and generating a time attention weight, a frequency attention weight, and a subcarrier attention weight based on the similarity calculation between the embedded features, and performing weighted fusion of the time attention weight, the frequency attention weight, and the subcarrier attention weight with the initial compressed feature map respectively to generate a three-dimensional attention feature map;
[0012] Performing multi-dimensional joint feature fusion on the three-dimensional attention feature map through a preset cross-attention model to generate deep joint features;
[0013] The preset second separable convolution model is used to adjust the dimension of the deep joint feature, restore it to the target feature dimension, and output the human motion perception feature.
[0014] In one embodiment, the step of performing feature compression on the wireless signal spectrum data using a preset first separable convolution model to generate an initial compressed feature map includes:
[0015] The local feature extraction of the spatial dimension of the input spectrum data is performed through deep convolution to obtain a local feature map;
[0016] Compressing the channel dimension of the local feature map by point convolution to reduce the number of redundant feature channels and obtain a compressed feature map;
[0017] The compressed feature map is enhanced by a nonlinear activation function to determine an initial compressed feature map.
[0018] In one embodiment, the step of inputting the initial compressed feature map into a preset non-local attention model, performing embedding space mapping along the time dimension, frequency dimension, and subcarrier dimension, and generating a time attention weight, a frequency attention weight, and a subcarrier attention weight based on the similarity calculation between the embedded features, and weightedly fusing the time attention weight, the frequency attention weight, and the subcarrier attention weight with the initial compressed feature map to generate a three-dimensional attention feature map includes:
[0019] By linearly projecting the initial compressed feature map along the time dimension, frequency dimension and subcarrier dimension respectively, the embedded features of each dimension are generated.
[0020] Calculate the similarity matrix between the embedded features of each dimension based on the embedded Gaussian function to generate the attention weight of each dimension;
[0021] The attention weights of each dimension are weighted summed with the initial compressed feature map, and then fused with the learnable gain parameter through residual connection to obtain a three-dimensional attention feature map. In one embodiment, the step of performing multi-dimensional joint feature fusion on the three-dimensional attention feature map through a preset cross-attention model to generate a deep joint feature includes:
[0022] The three-dimensional attention feature maps are spatially aligned and spliced to generate a joint feature matrix.
[0023] Performing a secondary non-local attention operation on the joint feature matrix to calculate the global association weight across dimensions;
[0024] The joint feature matrix is dynamically weighted by the global association weight, and the deep joint feature is determined after a nonlinear transformation.
[0025] In one embodiment, the step of compressing the wireless signal spectrum data using a preset first separable convolution model to generate an initial compressed feature map includes:
[0026] In the non-local attention model, channel reduction technology is used to compress the number of channels in the embedding space to a preset ratio of the original features to reduce the amount of computation;
[0027] Introducing subsampling operations in the cross-attention model, reducing the resolution of feature maps through maximum pooling or average pooling to reduce the scale of matrix operations;
[0028] During model training, an incremental supervised data augmentation loss function is used, combined with a dynamic weight adjustment strategy to optimize model parameter updates to suppress overfitting and improve generalization capabilities.
[0029] The bottleneck structure design limits the number of intermediate channels in the separable convolutional layer, reducing the number of model parameters and memory usage.
[0030] In addition, to achieve the above objectives, the present application also proposes a human motion perception device based on a three-dimensional hybrid attention model, the human motion perception device based on a three-dimensional hybrid attention model comprising:
[0031] An acquisition module, used to acquire wireless signal spectrum data;
[0032] a processing module for performing feature processing on the wireless signal spectrum data according to a preset three-dimensional hybrid attention model to determine human motion perception features, wherein the three-dimensional hybrid attention model first performs feature compression on the wireless signal spectrum data, then calculates attention weights from the three dimensions of time, frequency, and subcarrier based on a non-local attention mechanism to generate a three-dimensional attention feature map, and finally fuses the three-dimensional attention feature map through cross-attention, outputting human motion perception features after dimensional adjustment;
[0033] The determination module is used to determine the human motion perception recognition result according to the human motion perception feature.
[0034] In addition, to achieve the above-mentioned purpose, the present application also proposes a human motion perception device based on a three-dimensional hybrid attention model, the device comprising: a memory, a processor, and a computer program stored on the memory and runnable on the processor, the computer program being configured to implement the steps of the human motion perception method based on a three-dimensional hybrid attention model as described above.
[0035] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the human motion perception method based on the three-dimensional hybrid attention model as described above are implemented.
[0036] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the human motion perception method based on the three-dimensional hybrid attention model as described above.
[0037] One or more technical solutions proposed in this application have at least the following technical effects:
[0038] Compared with the related art, which relies on highly parameterized deep learning models to extract features from high-dimensional wireless signals, resulting in serious inefficiency in model size and inference time, the present application obtains wireless signal spectrum data; performs feature processing on the wireless signal spectrum data according to a preset three-dimensional hybrid attention model to determine the human motion perception features, wherein the three-dimensional hybrid attention model first performs feature compression on the wireless signal spectrum data, and then calculates attention weights from the three dimensions of time, frequency, and subcarrier based on the non-local attention mechanism to generate a three-dimensional attention feature map, and finally fuses the three-dimensional attention feature map through cross-attention, outputs human motion perception features after dimensional adjustment; and determines the human motion perception recognition results based on the human motion perception features. It can be understood that the present application adopts a three-dimensional hybrid attention mechanism, including a non-local attention module and a cross-attention module. The non-local attention module uses the three dimensions of time, frequency, and subcarrier to calculate attention weights, so as to better capture the multi-dimensional features of the wireless signal. The cross-attention module further fuses these three-dimensional attention maps to enhance the feature extraction capability of the model. Under limited hardware resources, this model significantly improves recognition efficiency by reducing the number of parameters, enabling the effective deployment of human motion perception technology based on wireless signals on small devices in scenarios such as smart homes and smart healthcare. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0040] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0041] Figure 1 A flowchart of the first embodiment of the method for human motion perception based on a three-dimensional hybrid attention model of this application is provided;
[0042] Figure 2 A flowchart of the second embodiment of the human motion perception method based on the three-dimensional hybrid attention model of this application is provided;
[0043] Figure 3 A schematic diagram of a simplified flow chart of a method for human motion perception based on a three-dimensional hybrid attention model provided in Example 2 of the present application;
[0044] Figure 4This is a schematic diagram of the module structure of a human motion perception device based on a three-dimensional hybrid attention model according to an embodiment of the present application;
[0045] Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the human motion perception method based on the three-dimensional hybrid attention model in the embodiment of the present application.
[0046] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0047] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.
[0048] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0049] The main solutions of the embodiments of this application are:
[0050] Obtain wireless signal spectrum data;
[0051] The wireless signal spectrum data is feature processed according to a preset three-dimensional hybrid attention model to determine the human motion perception features, wherein the three-dimensional hybrid attention model first performs feature compression on the wireless signal spectrum data, and then calculates attention weights from the three dimensions of time, frequency, and subcarrier based on a non-local attention mechanism to generate a three-dimensional attention feature map. Finally, the three-dimensional attention feature map is fused through cross-attention, and the human motion perception features are output after dimension adjustment;
[0052] A human motion perception recognition result is determined according to the human motion perception feature.
[0053] In this embodiment, the present application uses a human motion perception device based on a three-dimensional hybrid attention model as the execution subject. For the sake of convenience, it is specifically described below as "device".
[0054] Existing technologies for processing this type of complex data require a large number of parameters to analyze features, making effective deployment difficult on resource-constrained hardware. While traditional models can extract key features from high-dimensional data, the sheer number of parameters leads to significant consumption of memory and computing power in real-world applications, especially on resource-constrained hardware like mobile devices and small sensor nodes.
[0055] This application provides a solution by designing a lightweight 3D hybrid attention model to address the challenge of efficiently processing high-dimensional wireless signal data under limited hardware resources. Specifically, the three-dimensional hybrid attention mechanism of the model includes a non-local attention module and a cross-attention module. The non-local attention module uses the three dimensions of time, frequency, and subcarrier to calculate the attention weights, so that it can better capture the multi-dimensional characteristics of the wireless signal. The cross-attention module further integrates these three-dimensional attention maps to enhance the feature extraction capability of the model. Under the condition of limited hardware resources, the model greatly improves the recognition efficiency by reducing the number of parameters, so that the human motion perception technology based on wireless signals can be effectively deployed on small devices in scenarios such as smart homes and smart medical care.
[0056] The model first obtains the spectrum data of the wireless signal through short-time Fourier transform (STFT), identifies the features composed of three dimensions: time, frequency and subcarrier, then uses the non-local attention module to calculate the weights of each dimension, and fuses these weights through the cross-attention module, and finally outputs the key features to achieve efficient analysis and real-time perception of complex data.
[0057] Based on this, the embodiment of the present application provides a method for human motion perception based on a three-dimensional hybrid attention model. Figure 1 , Figure 1 This is a flow chart of the first embodiment of the human motion perception method based on the three-dimensional hybrid attention model of this application.
[0058] In this embodiment, the human motion perception method based on the three-dimensional hybrid attention model includes steps S10 to S30:
[0059] Step S10, obtaining wireless signal spectrum data;
[0060] It should be noted that wireless signal spectrum data refers to the amplitude, phase and other information of wireless signals during transmission, which will change with time, frequency and other factors. After processing the wireless signal through specific technical means (such as short-time Fourier transform, etc.), it can be represented as data in the form of a spectrum diagram, which contains information such as the energy distribution of the wireless signal at different time points, different frequencies and different subcarriers. This information can reflect the state of human movement and can be used for analysis and identification related to human motion perception.
[0061] Understandably, in a smart home scenario, wireless signals are transmitted by a home Wi-Fi router. People's activities indoors (such as walking and waving) can cause changes in the wireless signal. Signal receiving equipment deployed indoors acquires the channel state information (CSI) of the wireless signal. This CSI data is then processed through a short-time Fourier transform (SFT) to convert it into wireless signal spectrum data. This spectrum data is then fed into a trained three-dimensional hybrid attention model. After processing, the model derives human motion perception features. Ultimately, these features can be used to identify specific movements of people indoors, such as when someone walks from one location to another in the living room.
[0062] It is understandable that by obtaining wireless signal spectrum data, basic data support is provided for the subsequent use of three-dimensional hybrid attention models for human motion perception, so that human motion perception technology no longer relies on traditional wearable devices or invasive means such as cameras, expanding the application scenarios of human motion perception, reducing deployment costs and difficulty, and wireless signals can penetrate obstacles, and have better adaptability and universality for indoor human activity perception in complex environments, thereby improving the convenience and practicality of human motion perception.
[0063] Step S20: Feature processing is performed on the wireless signal spectrum data according to a preset three-dimensional hybrid attention model to determine the human motion perception features. The three-dimensional hybrid attention model first performs feature compression on the wireless signal spectrum data. Then, based on the non-local attention mechanism, attention weights are calculated from the three dimensions of time, frequency, and subcarrier to generate a three-dimensional attention feature map. Finally, the three-dimensional attention feature map is fused through cross-attention, and the dimensions are adjusted to output the human motion perception features.
[0064] It should be noted that feature processing refers to a series of operations on the acquired wireless signal spectrum data to extract feature information that is valuable for human motion perception, remove irrelevant or redundant information, and make the data more suitable for subsequent analysis and recognition processes. It includes steps such as feature compression, calculation of attention weights, feature fusion, and dimensionality adjustment.
[0065] Feature compression uses specific algorithms or model structures (such as separable convolution) to reduce the dimension and data volume of wireless signal spectrum data, remove redundant features, and retain key information in the data to reduce computational complexity, improve model operation efficiency, and highlight the main features in the data.
[0066] The non-local attention mechanism is a mechanism that captures long-range dependencies in data. It generates attention weights by calculating the similarity or correlation between different positions in the feature map, and weights the features according to the weights. This enables the model to focus on key areas in the data related to the current task, rather than just being limited to local neighborhoods, which helps to extract more discriminative features.
[0067] The three-dimensional attention feature map is generated by weighted fusion of the attention weights calculated from the three dimensions of time, frequency, and subcarrier with the initial compressed feature map. It contains the key feature information of the wireless signal spectrum data in three dimensions. The features of each dimension are highlighted and emphasized, providing a basis for subsequent feature fusion.
[0068] Cross-attention is a mechanism used to fuse three-dimensional attention feature maps. By rebalancing the importance of features in different dimensions and comprehensively considering information in the three dimensions of time, frequency, and subcarriers, the model can more comprehensively and accurately understand the human motion perception information contained in the wireless signal spectrum data, achieve the synergy of multi-dimensional features, and improve the expressive ability of human motion perception features.
[0069] Dimensionality adjustment is the final stage of feature processing, which adjusts the dimensions of the deep joint features after cross-attention fusion and restores them to the target feature dimensions to meet the requirements of subsequent processing (such as classification, recognition, etc.). It also helps to further reduce the complexity and computational complexity of the model and improve the efficiency and adaptability of the model.
[0070] It's understandable that by compressing the features of wireless signal spectrum data, the data volume and computational complexity are reduced, improving the model's operational efficiency and enabling real-time execution on resource-limited devices (such as small wireless signal transceivers), reducing costs and energy consumption. The non-local attention mechanism is able to focus on key feature areas in the data. The combination of three-dimensional attention feature maps and cross-attention fully exploits the multi-dimensional information related to human motion perception in wireless signal spectrum data, improving the accuracy and richness of human motion perception features, thereby enhancing the reliability and precision of human motion perception recognition results. This provides strong technical support for accurate human motion perception in smart homes, smart healthcare, and other fields.
[0071] In a feasible implementation, the step of performing feature processing on the wireless signal spectrum data according to a preset three-dimensional hybrid attention model to determine the human motion perception feature includes:
[0072] Performing feature compression on the wireless signal spectrum data using a preset first separable convolution model to generate an initial compressed feature map;
[0073] Inputting the initial compressed feature map into a preset non-local attention model, performing embedding space mapping along the time dimension, frequency dimension, and subcarrier dimension respectively, and generating a time attention weight, a frequency attention weight, and a subcarrier attention weight based on the similarity calculation between the embedded features, and performing weighted fusion of the time attention weight, the frequency attention weight, and the subcarrier attention weight with the initial compressed feature map respectively to generate a three-dimensional attention feature map;
[0074] Performing multi-dimensional joint feature fusion on the three-dimensional attention feature map through a preset cross-attention model to generate deep joint features;
[0075] The preset second separable convolution model is used to adjust the dimension of the deep joint feature, restore it to the target feature dimension, and output the human motion perception feature.
[0076] It should be noted that the first separable convolution model is a convolutional model composed of depthwise convolution and pointwise convolution, used to compress the features of wireless signal spectrum data. Depthwise convolution extracts local features from the spatial dimension of the input data, while pointwise convolution compresses the channel dimension of the local feature map, reducing the number of redundant feature channels. The combination of the two reduces computational costs while preserving key feature information.
[0077] The initial compressed feature map is the feature map obtained after processing by the first separable convolution model. Compared with the original wireless signal spectrum data, its dimension is reduced, redundant information is reduced, the main features in the data are highlighted, and a basis for further feature processing and analysis is provided.
[0078] The non-local attention model is based on a non-local mechanism. It calculates the similarity or correlation between different locations in the feature map and generates attention weights in three dimensions: time, frequency, and subcarrier. This model can capture long-range dependencies in the data and focus on key areas related to human perception.
[0079] Embedding space mapping maps the feature vectors in the initial compressed feature map into a new embedding space, allowing the similarity or correlation between features to be calculated in that space. In wireless signal processing, embedding space mapping can better capture the inherent connections between features across different dimensions, such as time and frequency, helping to highlight key feature information.
[0080] Similarity calculation is the process of calculating the similarity between different feature vectors in the embedding space, typically using methods such as the embedded Gaussian function. This similarity calculation can determine the degree of correlation between features, providing a basis for generating attention weights, allowing the model to focus on more important or representative features.
[0081] Weighted fusion combines the generated attention weights for time, frequency, and subcarrier with the initial compressed feature map, fusing the feature information from different dimensions to generate a three-dimensional attention feature map. This weighted fusion process can highlight key feature areas in the data based on the attention weights, enhancing the expressiveness of the features while suppressing unimportant feature information.
[0082] The three-dimensional attention feature map contains key characteristic information of wireless signal spectrum data in three dimensions: time, frequency, and subcarrier. The features of each dimension are highlighted and emphasized, providing a foundation for subsequent cross-attention fusion, helping the model to more comprehensively and accurately understand the human perception information contained in wireless signals.
[0083] The cross-attention model is used to perform multi-dimensional joint feature fusion on three-dimensional attention feature maps. By rebalancing the importance of features across different dimensions and comprehensively considering information from time, frequency, and subcarriers, the model can more comprehensively understand the human perception information contained in wireless signal spectrum data. This enables the synergy of multi-dimensional features, improving the expressiveness and recognition accuracy of human perception features.
[0084] Deep joint features are fused through a cross-attention model, integrating key feature information from three dimensions: time, frequency, and subcarriers, to form a deeper and more comprehensive feature representation. Deep joint features can more accurately reflect perceptual information such as human motion, providing stronger support for further dimensionality adjustment and human motion recognition.
[0085] The second separable convolutional model is similar to the first, but is used to re-dimensionalize the deep joint features. Through depthwise and pointwise convolutions, the deep joint features are restored to their target dimensions, meeting the requirements of subsequent human action recognition and other processing. This also helps further reduce the model's complexity and computational effort, improving its efficiency and adaptability.
[0086] The human motion perception feature is the final output feature, containing key information from the wireless signal spectrum data after feature compression, attention mechanism processing, feature fusion, and dimensionality adjustment. The human motion perception feature accurately reflects perceptual information such as human motion and can be used for tasks such as human motion recognition and behavior analysis. It has broad application prospects in smart homes, smart healthcare, and other fields.
[0087] It is understood that in a smart sports monitoring scenario, the user wears a smart wristband with integrated wireless signal transceiver capabilities. The smart wristband acquires real-time wireless signal spectrum data around the user. This data is first input into a preset first separable convolutional model for feature compression, generating an initial compressed feature map. Next, the initial compressed feature map is input into a preset non-local attention model, where it is embedded in spatial mapping along the time, frequency, and subcarrier dimensions. Based on the similarity between the embedded features, time attention weights, frequency attention weights, and subcarrier attention weights are generated. These attention weights are then weightedly fused with the initial compressed feature map to generate a three-dimensional attention feature map. Subsequently, the three-dimensional attention feature map is subjected to multi-dimensional joint feature fusion using a preset cross-attention model to generate deep joint features. Finally, the deep joint features are resized using a preset second separable convolutional model, restoring them to the target feature dimension and outputting human motion perception features. These features can be used to identify various user motions, such as running, jumping, and waving, enabling real-time monitoring and analysis of the user's motion status, providing more accurate motion data feedback and health recommendations.
[0088] It can be understood that by using the first separable convolution model to perform feature compression on the wireless signal spectrum data, the data volume and computational complexity are effectively reduced, the operating efficiency of the model is improved, and it can run in real time on resource-constrained devices such as smart bracelets. The non-local attention model can focus on key feature areas in the data. The combination of the three-dimensional attention feature map and the cross-attention model fully mines the multi-dimensional information related to human motion perception in the wireless signal spectrum data, and improves the accuracy and richness of human motion perception features. The dimensionality adjustment operation of the second separable convolution model ensures that the output features meet the requirements of subsequent processing, while further reducing the complexity of the model. Overall, this technical process can improve the accuracy and efficiency of human motion perception, provide strong technical support for applications such as intelligent motion monitoring, and enhance user experience and application practicality.
[0089] In a feasible implementation, the step of performing feature compression on the wireless signal spectrum data using a preset first separable convolution model to generate an initial compressed feature map includes:
[0090] The local feature extraction of the spatial dimension of the input spectrum data is performed through deep convolution to obtain a local feature map;
[0091] Compressing the channel dimension of the local feature map by point convolution to reduce the number of redundant feature channels and obtain a compressed feature map;
[0092] The compressed feature map is enhanced by a nonlinear activation function to determine an initial compressed feature map.
[0093] It should be noted that deep convolution is a convolution operation within a convolutional neural network that primarily extracts local features from the spatial dimensions of the input data. It performs convolution operations on each channel of the input data separately, capturing local spatial correlations in the data. For example, in wireless signal spectrum data, it can extract local feature patterns at different time points and frequencies.
[0094] Local feature extraction is the process of identifying and extracting representative local features from input spectrum data. Local features typically reflect the characteristics of data within a small range, such as the changing trend of wireless signals over a short period of time or the energy distribution around specific frequencies. These features are important for understanding the local behavior and variation patterns of signals.
[0095] The local feature map is the result of a deep convolution operation, expressed in the form of a matrix or tensor. It contains the local feature information of the input spectrum data in the spatial dimension. Each element or channel represents the local eigenvalue of a specific location or feature, providing a basis for subsequent feature processing and analysis.
[0096] Point convolution is a convolution operation performed on the channel dimension, primarily used to compress local feature maps. Point convolution fuses features from different channels through linear combinations, reducing the number of redundant feature channels and preserving key feature information while lowering computational costs.
[0097] Channel dimension compression is the process of reducing the number of channels in a feature map, aiming to remove redundant feature channels and highlight important feature information. Channel dimension compression can reduce model complexity and computational effort while allowing the model to focus more on features that are important to the task.
[0098] A compressed feature map is a feature map obtained by compressing the channel dimension of a local feature map through point convolution. While retaining the main features of the original feature map, it has fewer channels and a smaller data size, making it more conducive to subsequent feature processing and model calculation.
[0099] A nonlinear activation function is a mathematical function that introduces nonlinear factors and is commonly used in neural networks. It performs a nonlinear transformation on input feature data, enabling the model to learn complex patterns and nonlinear relationships within the data. Common nonlinear activation functions include ReLU, Sigmoid, and Tanh. During feature enhancement, nonlinear activation functions can highlight important information in a feature and suppress unimportant information, thereby improving the expressiveness and discriminability of the feature.
[0100] Feature enhancement involves processing compressed feature maps using methods such as nonlinear activation functions to enhance the expressiveness and separability of features. Feature enhancement allows the model to focus more on key feature information and increases its sensitivity to different features, thereby improving model performance and generalization.
[0101] For example, referring to Figure 2 , the separable convolutional network composed of depthwise convolution and pointwise convolution is an effective alternative to the well-known traditional convolutional neural network (CNN). The reduced dimension of the wireless signal spectrum graph is But if In order to better summarize different datasets, this application uses a separable convolutional network to transform the data dimension from Reduce to . Determines the overall computational cost of evaluating the model.
[0102] It is understandable that deep convolution can effectively extract local features from wireless signal spectrum data, capturing subtle changes and patterns in the data's spatial dimensions, laying the foundation for subsequent feature processing. Point convolution's channel dimension compression operation reduces the data volume and complexity of the feature map, lowering the model's computational resource consumption and enabling the model to run more efficiently on resource-constrained devices. The feature enhancement effect of nonlinear activation functions improves the expressiveness and separability of features, enabling the model to more accurately identify and distinguish different feature patterns, thereby improving the accuracy and reliability of human perception. Through this series of operations, the overall technical process enables efficient processing and feature extraction of wireless signal spectrum data, providing high-quality feature input for human perception tasks, enhancing the model's performance and adaptability in practical applications, and expanding its application scenarios, such as human activity monitoring in smart homes and patient behavior analysis in smart healthcare.
[0103] In a feasible implementation manner, the initial compressed feature map is input into a preset non-local attention model, and embedded in the spatial mapping along the time dimension, frequency dimension, and subcarrier dimension respectively, and based on the similarity calculation between the embedded features, the time attention weight, the frequency attention weight, and the subcarrier attention weight are generated. The time attention weight, the frequency attention weight, and the subcarrier attention weight are weightedly fused with the initial compressed feature map respectively to generate a three-dimensional attention feature map. The steps include:
[0104] By linearly projecting the initial compressed feature map along the time dimension, frequency dimension and subcarrier dimension respectively, the embedded features of each dimension are generated.
[0105] Calculate the similarity matrix between the embedded features of each dimension based on the embedded Gaussian function to generate the attention weight of each dimension;
[0106] The attention weights of each dimension are weighted and summed with the initial compressed feature map, and then fused with the learnable gain parameters through residual connection to obtain a three-dimensional attention feature map.
[0107] It should be noted that linear projection is a linear transformation technique that projects the initial compressed feature map into a new embedding space by linearly projecting it along the time, frequency, and subcarrier dimensions. Linear projection can change the representation of features, making them more suitable for subsequent similarity calculations and attention weight generation, while preserving the key information of the features.
[0108] Embedded features are feature vectors obtained after linear projection and are located in the new embedding space. Embedded features can better capture the key feature information of the initial compressed feature map in different dimensions, providing a basis for further similarity calculation and attention mechanism processing.
[0109] The embedded Gaussian function is used to calculate similarity. Based on the Gaussian function, it calculates the similarity matrix between embedded features in each dimension. The embedded Gaussian function measures the degree of similarity between features. The resulting similarity matrix reflects the correlation between different features in the embedding space, providing a basis for generating attention weights.
[0110] The similarity matrix is calculated using an embedded Gaussian function and represents the similarity between embedded features. Each element in the similarity matrix represents the similarity between two feature vectors, with larger values indicating greater similarity between the features. The similarity matrix is a key basis for generating attention weights, which can be used to determine the relative importance and correlation between features.
[0111] Attention weights for each dimension are generated based on the similarity matrix, corresponding to attention weights for time, frequency, and subcarrier. These weights reflect the importance of features along these dimensions, with features with higher weights being more critical to the human perception task in those dimensions. The attention weights are used to perform a weighted summation of the initial compressed feature map, highlighting important features and suppressing unimportant ones.
[0112] Weighted summation is the process of performing a weighted operation on the attention weights of each dimension and the initial compressed feature map. This operation multiplies the attention weight of each dimension by the corresponding eigenvalue in the initial compressed feature map and then sums the results. This weighted summation operation fuses features based on the attention weights, generating a feature representation that incorporates multidimensional information while highlighting important features and suppressing unimportant ones.
[0113] Residual connections are a connection method used to alleviate the vanishing gradient problem in deep neural networks. They connect the input directly to the output, allowing gradients to be propagated more directly. During the generation of three-dimensional attention feature maps, residual connections preserve the original feature information, preventing feature loss during multiple iterations and contributing to model stability and convergence speed.
[0114] The learnable gain parameter is a learnable parameter used to adjust feature weights during the fusion process. It automatically learns the optimal weights based on the training data, enabling the model to better adapt to different feature distributions and task requirements, improving the model's expressiveness and generalization capabilities.
[0115] It can be understood that the non-local mechanism captures long-range dependencies by calculating the response at a certain position as the weighted sum of the features at all positions in the feature map. This "non-local" concept was originally designed to calculate the weighted spatial features of pixels in image processing and semantic segmentation. Compared with static images with only three RGB channels, wireless signal spectrograms contain complex patterns of monitoring the human body in many subcarriers at specific time and frequency. The three-dimensional hybrid attention proposed in this application is designed specifically for wireless signal spectrum datasets. The principle is to extract the weighted time-frequency-subcarrier features of the wireless signal through three non-local operations, and re-weigh the importance of the three extracted dimensions in the deep feature space. This application proposes that the model can simultaneously focus on the three dimensions of wireless signal data and adaptively shift the importance of each dimension to capture the unique patterns of human movement.
[0116] For example, given a feature map extracted from the frequency dimension ,The non-local module first calculates the similarity between the two embedding spaces using the embedded Gaussian function, i.e. , and then use the similarity to generate the attention map As shown below:
[0117]
[0118]
[0119] Among them, 、 、 is a learnable weight matrix used to project features into the embedding space.
[0120] After applying the focus mask, we get the observed features Then, residual connections and learnable gain parameters are further added to the non-local operations, allowing multiple non-local blocks in one network:
[0121]
[0122] in, Is a learnable gain parameter used to control the observed features Compared with the original feature map The proportion of the sum when adding can be dynamically adjusted through learning, so that the network can adaptively fuse non-local information and original features according to the characteristics of the data.
[0123] Similarly, while rotating the feature map, weighted features are generated from the time dimension and subcarrier dimension respectively. , , and forms a three-dimensional attention mechanism as the first stage.
[0124] It can be understood that by linearly projecting the initial compressed feature map along three dimensions, the features can be mapped into an embedding space more suitable for similarity calculation and attention weight generation, better capturing the relationships and similarities between features. The similarity matrix calculated by the embedded Gaussian function provides a reliable basis for generating accurate attention weights for each dimension, allowing the attention weights to accurately reflect the importance of features in different dimensions. The weighted summation operation combined with residual connections and learnable gain parameters not only integrates multi-dimensional feature information and highlights key features, but also preserves the original feature information, improving the stability and convergence speed of the model. The resulting three-dimensional attention feature map integrates key feature information from the three dimensions of time, frequency, and subcarrier. It can more comprehensively and accurately reflect the human perception information contained in the wireless signal spectrum data, improve the accuracy and reliability of human perception, enhance the performance and adaptability of the model in human perception tasks, and provide higher-quality feature representation and more accurate recognition results for applications in smart homes, smart healthcare, and other fields.
[0125] In a feasible implementation, the step of performing multi-dimensional joint feature fusion on the three-dimensional attention feature map through a preset cross-attention model to generate deep joint features includes:
[0126] The three-dimensional attention feature maps are spatially aligned and spliced to generate a joint feature matrix.
[0127] Performing a secondary non-local attention operation on the joint feature matrix to calculate the global association weight across dimensions;
[0128] The joint feature matrix is dynamically weighted by the global association weight, and the deep joint feature is determined after a nonlinear transformation.
[0129] It should be noted that spatial alignment and splicing is a technique that spatially aligns and splices feature maps from different sources or dimensions. By spatially aligning and splicing 3D attention feature maps, multiple feature maps are integrated into a joint feature matrix. This allows features from different feature maps to be fused and processed within the same spatial framework, providing a unified feature representation for subsequent feature analysis and fusion.
[0130] The joint feature matrix is a matrix or tensor composed of spatially aligned and concatenated feature maps. It integrates feature information from different dimensions or sources, contains a more comprehensive feature description, can more completely reflect the intrinsic characteristics of the data, and provides a richer information foundation for subsequent feature processing and analysis.
[0131] The quadratic non-local attention operation is a process that performs secondary attention calculation and feature fusion on the joint feature matrix based on the non-local attention mechanism. This operation can further capture the long-range dependencies between different positions in the joint feature matrix, explore the deep connections between features, and enhance the expressiveness and discriminability of features.
[0132] The cross-dimensional global correlation weight is calculated using a quadratic non-local attention operation and reflects the weight of the global correlation between features of different dimensions in the joint feature matrix. This cross-dimensional global correlation weight measures the importance and interaction of features of different dimensions in the overall feature space, providing a basis for subsequent dynamic weighting, enabling the model to perform feature fusion and weight adjustment based on the global correlation between features.
[0133] Dynamic weighting is the process of weighting features in a joint feature matrix based on global cross-dimensional correlation weights. Dynamic weighting automatically adjusts feature weights based on the degree of correlation and importance between features, highlighting key features and suppressing unimportant ones. This allows the model to focus more on task-relevant feature information, improving model performance and generalization.
[0134] Nonlinear transformation is the process of applying nonlinear processing to features, typically implemented through nonlinear activation functions. Nonlinear transformations introduce nonlinear factors, enabling the model to learn complex patterns and nonlinear relationships in the data, enhancing the expressive power of features and the model's fitting capabilities. In the process of determining deep joint features, nonlinear transformations help map dynamically weighted features into a new feature space, further improving the representation and discriminability of features.
[0135] Deep joint features are derived through dynamic weighting and nonlinear transformation. They fuse feature information from different dimensions and perform deep feature mining and fusion in the feature space. They can more comprehensively and accurately reflect the essential characteristics and inherent laws of the data, providing higher-quality feature representation for subsequent tasks such as human perception.
[0136] For example, the motion patterns carried by the three dimensions of the wireless signal spectrum graph may be different. As the second stage, cross attention is to re-weigh the importance of each dimension in the deep feature space.
[0137] Specifically, the generated weighted feature map , Rotate back and concatenate, and the resulting dimension is 3 The non-local operation then calculates the weights of the previous 3 weighted feature maps as follows:
[0138]
[0139] Afterwards, a second separable convolution is used to combine , and change the dimension from 3 Restore to The attention structure of these two stages can focus on specific regions in three dimensions (i.e., time, frequency, and subcarrier) to achieve accurate classification.
[0140] It is understandable that the spatial alignment and splicing operation integrates the three-dimensional attention feature map into a joint feature matrix, providing a unified feature representation for subsequent feature fusion, allowing features of different dimensions to be analyzed and processed within the same spatial framework. The secondary non-local attention operation can further explore the deep correlations between features and calculate the global correlation weights across dimensions, providing an accurate basis for dynamic weighting. Dynamic weighting combined with nonlinear transformations can highlight key features, suppress unimportant features, and enhance the expressive power of features and the model's fitting capabilities. The resulting deep joint features integrate multi-dimensional feature information and perform deep mining and fusion in the feature space. They can more comprehensively and accurately reflect the human perception information contained in wireless signal spectrum data, improve the accuracy and reliability of human perception recognition, enhance the performance and adaptability of the model in human perception tasks, and provide higher-quality feature representations and more accurate recognition results for applications in smart homes, smart healthcare, and other fields.
[0141] In a feasible implementation manner, before the step of performing feature compression on the wireless signal spectrum data using a preset first separable convolution model to generate an initial compressed feature map, the step includes:
[0142] In the non-local attention model, channel reduction technology is used to compress the number of channels in the embedding space to a preset ratio of the original features to reduce the amount of computation;
[0143] Introducing subsampling operations in the cross-attention model, reducing the resolution of feature maps through maximum pooling or average pooling to reduce the scale of matrix operations;
[0144] During model training, an incremental supervised data augmentation loss function is used, combined with a dynamic weight adjustment strategy to optimize model parameter updates to suppress overfitting and improve generalization capabilities.
[0145] The bottleneck structure design limits the number of intermediate channels in the separable convolutional layer, reducing the number of model parameters and memory usage.
[0146] It should be noted that channel reduction is a technique used to reduce the computational complexity of a model. In non-local attention models, channel reduction compresses the number of channels in the embedding space to a preset ratio of the original features, reducing computational complexity while preserving key feature information.
[0147] Subsampling is a technique for reducing the resolution of feature maps. Introducing subsampling in the crisscross attention model reduces the resolution of feature maps through max pooling or average pooling to reduce the scale of matrix operations and reduce computing resource consumption.
[0148] Max pooling is a pooling operation that divides the feature map and takes the maximum value within each divided area as the output, reducing the spatial dimension of the feature map while retaining key feature information.
[0149] Average pooling is a pooling operation that divides the feature map and takes the average value of each divided area as the output, reducing the spatial dimension of the feature map and smoothing the feature information.
[0150] The incremental supervised data augmentation loss function combines data augmentation with supervised learning. During model training, by gradually increasing the intensity and diversity of data augmentation, the model can adapt to more data changes during learning, improving its generalization ability.
[0151] The dynamic weight adjustment strategy is a strategy for optimizing model parameter updates. During model training, the dynamic weight adjustment strategy automatically adjusts the magnitude and direction of weight updates based on model performance and data changes, suppressing overfitting and improving the model's generalization capabilities.
[0152] Bottleneck structure design is a model structure optimization technique that limits the number of intermediate channels in the separable convolutional layer, thereby reducing the number of model parameters and memory usage, lowering computational costs, and improving model efficiency.
[0153] For example, this application uses bottleneck and subsampling techniques to further reduce the computational complexity of the four attention operations. and The number of channels represented is set to , or 1 / 8 of the number of channels in the sample. The sampling method is to use the maximum pooling to reduce , or The size of and reduces the matrix multiplication computation by 1 / 4 in one attention operation.
[0154] It's understandable that channel reduction techniques and subsampling operations reduce the data volume and computational complexity of feature maps in the channel and spatial dimensions, respectively. Channel reduction compresses the number of channels in the embedding space to a preset ratio of the original features, while subsampling reduces the resolution of feature maps through max pooling or average pooling. This effectively reduces the computational effort of non-local attention and crisscross attention models, alleviating the computational burden and improving model efficiency. The incremental supervised data augmentation loss function, combined with a dynamic weight adjustment strategy, gradually increases the intensity and diversity of data augmentation during model training, enabling the model to adapt to more data variations. The dynamic adjustment of the magnitude and direction of weight updates suppresses overfitting, improves model generalization, and achieves more stable performance across diverse scenarios and data distributions. The bottleneck structure design reduces the number of model parameters and memory usage by limiting the number of intermediate channels in the separable convolutional layer. This enables the model to run more efficiently on resource-constrained devices, reducing hardware reliance and expanding its application scenarios, such as deploying human perception technology on resource-constrained hardware such as mobile devices and small sensor nodes.
[0155] Step S30: determining a human motion perception recognition result according to the human motion perception feature.
[0156] It's important to note that human motion sensing features are extracted from wireless signal spectrum data and reflect key information about human motion. Through multi-layer processing and feature extraction within the model, these features retain crucial information related to human motion, such as motion type, amplitude, and frequency, providing the basis for human motion recognition.
[0157] The human motion perception and recognition result is the specific human motion type or state, such as walking, running, falling, etc., determined by classification and recognition algorithms based on the human motion perception characteristics. It is the final output of the entire model and is used to realize the monitoring and analysis of human motion.
[0158] For example, in a smart nursing home, wireless signal monitoring equipment installed in each room collects real-time wireless signal spectrum data from the elderly. This data is processed using a three-dimensional hybrid attention model to extract human motion perception features. Based on these features, a pre-set classification model is used to determine human motion perception recognition results, such as whether the elderly person is walking normally or has fallen. Once a fall is detected, the system immediately triggers an alarm, notifying caregivers to promptly address the situation, thereby ensuring real-time safety monitoring of the elderly.
[0159] This embodiment provides a human motion perception method based on a three-dimensional hybrid attention model, which adopts a three-dimensional hybrid attention mechanism, including a non-local attention module and a cross-attention module. The non-local attention module uses the three dimensions of time, frequency, and subcarrier to calculate the attention weight, so that it can better capture the multi-dimensional characteristics of the wireless signal. The cross-attention module further integrates these three-dimensional attention maps to enhance the feature extraction capability of the model. In the case of limited hardware resources, the model greatly improves the recognition efficiency by reducing the number of parameters, so that the human motion perception technology based on wireless signals can be effectively deployed on small devices in scenarios such as smart homes and smart medical care.
[0160] For example, in order to help understand the implementation process of the human motion perception method based on the three-dimensional hybrid attention model obtained by combining this embodiment with the above embodiment 1, please refer to Figure 3 , Figure 3 This paper provides a brief flowchart of a human motion perception method based on a three-dimensional hybrid attention model. Specifically:
[0161] Wireless signal spectrum data is collected through corresponding sensor devices (such as WiFi devices, millimeter wave radars, etc.). This data contains information such as signal strength and frequency that changes over time, reflecting the impact of human movement on wireless signals.
[0162] The collected wireless signal spectrum data is fed into the pre-set first separable convolutional model. First, deep convolution is used to extract local features from the spatial dimensions of the data, generating a local feature map. Point convolution is then used to compress the channel dimensions of the local feature map, reducing the number of redundant feature channels. A nonlinear activation function is then applied to enhance the compressed feature map, resulting in an initial compressed feature map. At this point, the feature dimensions of the data have been reduced, preserving key feature information.
[0163] The initial compressed feature map is fed into a pre-set non-local attention model. Embedding space mapping is performed along the time, frequency, and subcarrier dimensions. This involves linearly projecting the feature map along each dimension to generate embedded features for each dimension. A similarity matrix between the embedded features in each dimension is calculated based on an embedded Gaussian function, which then generates temporal, frequency, and subcarrier attention weights. These weights are then weighted summed with the initial compressed feature map and fused with a learnable gain parameter via a residual connection. This ultimately yields a three-dimensional attention feature map, highlighting features important for motion perception along these three dimensions.
[0164] The three-dimensional attention feature map is fed into a pre-set cross-attention model. The feature maps are first spatially aligned and concatenated to form a joint feature matrix. A secondary non-local attention operation is then performed on this joint feature matrix to calculate global cross-dimensional correlation weights. These global correlation weights are then dynamically weighted on the joint feature matrix and, after a nonlinear transformation, are used to generate deep joint features. This achieves deep fusion of multi-dimensional features and further enhances feature discriminability.
[0165] A preset second separable convolutional model is used to resize the deep joint features. Through the corresponding convolution operation, the features are restored to the target feature dimension, making them meet the input requirements of the subsequent human action recognition model, thereby outputting human action perception features that can effectively represent different human action states.
[0166] The human motion perception features are input into a pre-trained recognition model (which can be a classifier, etc.). The model determines the human motion perception recognition results based on the perceived features, such as judging the specific action type such as walking, running, waving, etc., and finally completes the entire process of human motion perception.
[0167] It should be noted that the above examples are only used to understand this application and do not constitute a limitation on the human motion perception method based on the three-dimensional hybrid attention model of this application. More simple transformations based on this technical concept are all within the scope of protection of this application.
[0168] This application also provides a human motion perception device based on a three-dimensional hybrid attention model, please refer to Figure 4 , the human motion perception device based on the three-dimensional hybrid attention model includes:
[0169] An acquisition module 10 is used to acquire wireless signal spectrum data;
[0170] a processing module 20 for performing feature processing on the wireless signal spectrum data according to a preset three-dimensional hybrid attention model to determine human motion perception features, wherein the three-dimensional hybrid attention model first performs feature compression on the wireless signal spectrum data, then calculates attention weights from the three dimensions of time, frequency, and subcarrier based on a non-local attention mechanism to generate a three-dimensional attention feature map, and finally fuses the three-dimensional attention feature map through cross-attention, outputting human motion perception features after dimensional adjustment;
[0171] The determination module 30 is used to determine the human motion perception recognition result according to the human motion perception feature.
[0172] And / or, the processing module 20 includes:
[0173] A first compression module, configured to perform feature compression on the wireless signal spectrum data using a preset first separable convolution model to generate an initial compression feature map;
[0174] a first mapping module, configured to input the initial compressed feature map into a preset non-local attention model, perform embedding spatial mapping along the time dimension, frequency dimension, and subcarrier dimension, respectively, and generate a time attention weight, a frequency attention weight, and a subcarrier attention weight based on similarity calculation between embedded features, and perform weighted fusion of the time attention weight, the frequency attention weight, and the subcarrier attention weight with the initial compressed feature map to generate a three-dimensional attention feature map;
[0175] A first fusion module is used to perform multi-dimensional joint feature fusion on the three-dimensional attention feature map through a preset cross-attention model to generate a deep joint feature;
[0176] The first adjustment module is used to use a preset second separable convolution model to adjust the dimension of the deep joint feature, restore it to the target feature dimension, and output the human motion perception feature.
[0177] And / or, the first compression module includes:
[0178] A first extraction module is used to extract local features of the spatial dimension of the input spectrum data through deep convolution to obtain a local feature map;
[0179] A second compression module is used to compress the channel dimension of the local feature map by point convolution to reduce the number of redundant feature channels and obtain a compressed feature map;
[0180] The first enhancement module is used to perform feature enhancement on the compressed feature map through a nonlinear activation function to determine an initial compressed feature map.
[0181] And / or, the first mapping module includes:
[0182] The first projection module is configured to generate embedded features in each dimension by linearly projecting the initial compressed feature map along the time dimension, the frequency dimension, and the subcarrier dimension, respectively.
[0183] A first embedding module is used to calculate the similarity matrix between the embedded features of each dimension based on the embedded Gaussian function to generate the attention weight of each dimension;
[0184] The second fusion module is used to perform weighted summation of the attention weights of each dimension and the initial compressed feature map, and then fuse them with the learnable gain parameter through a residual connection to obtain a three-dimensional attention feature map.
[0185] And / or, the first fusion module includes:
[0186] The first splicing module is used to spatially align and splice the three-dimensional attention feature map to generate a joint feature matrix.
[0187] A first computing module is configured to perform a secondary non-local attention operation on the joint feature matrix to calculate a global correlation weight across dimensions;
[0188] The first transformation module is used to dynamically weight the joint feature matrix through the global association weight, and determine the deep joint feature after nonlinear transformation.
[0189] And / or, the human motion perception device based on the three-dimensional hybrid attention model includes:
[0190] The third compression module is used to use channel reduction technology in the non-local attention model to compress the number of channels in the embedding space to a preset ratio of the original features to reduce the amount of calculation;
[0191] The first reduction module is used to introduce subsampling operations in the cross-attention model and reduce the resolution of the feature map through maximum pooling or average pooling to reduce the scale of matrix operations;
[0192] The first update module is used to use incremental supervised data augmentation loss function during model training, combined with dynamic weight adjustment strategy to optimize model parameter updates to suppress overfitting and improve generalization ability;
[0193] The second reduction module is used to limit the number of intermediate channels of the separable convolutional layer through bottleneck structure design, thereby reducing the number of model parameters and memory usage.
[0194] The human motion perception device based on a three-dimensional hybrid attention model provided by the present application adopts the human motion perception method based on a three-dimensional hybrid attention model in the above-mentioned embodiment, which can solve the technical problem of low efficiency of human motion perception and recognition in resource-limited IoT devices. Compared with the prior art, the beneficial effects of the human motion perception device based on a three-dimensional hybrid attention model provided by the present application are the same as the beneficial effects of the human motion perception method based on a three-dimensional hybrid attention model provided by the above-mentioned embodiment, and the other technical features of the human motion perception device based on a three-dimensional hybrid attention model are the same as the features disclosed in the above-mentioned embodiment method, which will not be repeated here.
[0195] The present application provides a human motion perception device based on a three-dimensional hybrid attention model. The human motion perception device based on the three-dimensional hybrid attention model includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the human motion perception method based on the three-dimensional hybrid attention model in the above-mentioned embodiment one.
[0196] Reference below Figure 5 , which shows a schematic structural diagram of a human motion perception device based on a three-dimensional hybrid attention model suitable for implementing an embodiment of the present application. The human motion perception device based on a three-dimensional hybrid attention model in the embodiment of the present application may include, but is not limited to, mobile terminals such as mobile phones, tablet computers, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital televisions and desktop computers. Figure 5 The human motion perception device based on the three-dimensional hybrid attention model shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0197] like Figure 5As shown, the human motion perception device based on the three-dimensional hybrid attention model may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM: Read Only Memory) 1002 or the program loaded from the storage device 1003 to the random access memory (RAM: Random Access Memory) 1004. Various programs and data required for the operation of the human motion perception device based on the three-dimensional hybrid attention model are also stored in RAM1004. The processing device 1001, ROM1002 and RAM1004 are connected to each other through a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the human motion perception device based on the three-dimensional hybrid attention model to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows a human motion perception device based on the three-dimensional hybrid attention model with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have instead.
[0198] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0199] The human motion perception device based on a three-dimensional hybrid attention model provided by this application adopts the human motion perception method based on a three-dimensional hybrid attention model in the above-mentioned embodiment, which can solve the technical problem of low efficiency of human motion perception and recognition in resource-limited IoT devices. Compared with the prior art, the beneficial effects of the human motion perception device based on a three-dimensional hybrid attention model provided by this application are the same as the beneficial effects of the human motion perception method based on a three-dimensional hybrid attention model provided by the above-mentioned embodiment, and the other technical features of the human motion perception device based on a three-dimensional hybrid attention model are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.
[0200] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0201] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0202] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, and the computer-readable program instructions are used to execute the human motion perception method based on the three-dimensional hybrid attention model in the above-mentioned embodiment.
[0203] The computer-readable storage medium provided herein may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including, but not limited to, wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0204] The above-mentioned computer-readable storage medium can be included in the human motion perception device based on the three-dimensional hybrid attention model; or it can exist alone without being assembled into the human motion perception device based on the three-dimensional hybrid attention model.
[0205] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by a human motion sensing device based on a three-dimensional hybrid attention model, the human motion sensing device based on a three-dimensional hybrid attention model: obtains wireless signal spectrum data;
[0206] The wireless signal spectrum data is feature processed according to a preset three-dimensional hybrid attention model to determine the human motion perception features, wherein the three-dimensional hybrid attention model first performs feature compression on the wireless signal spectrum data, and then calculates attention weights from the three dimensions of time, frequency, and subcarrier based on a non-local attention mechanism to generate a three-dimensional attention feature map. Finally, the three-dimensional attention feature map is fused through cross-attention, and the human motion perception features are output after dimension adjustment;
[0207] A human motion perception recognition result is determined according to the human motion perception feature.
[0208] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0209] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0210] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0211] The computer-readable storage medium provided in this application stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned method for human motion perception based on a three-dimensional hybrid attention model. This computer-readable storage medium can address the technical issue of low human motion perception and recognition efficiency in resource-limited IoT devices. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the method for human motion perception based on a three-dimensional hybrid attention model provided in the aforementioned embodiments, and are not further elaborated here.
[0212] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the human motion perception method based on the three-dimensional hybrid attention model as described above.
[0213] The computer program product provided in this application can address the technical issue of low efficiency in human motion perception and recognition in resource-limited IoT devices. Compared to the prior art, the computer program product provided in this application offers the same beneficial effects as the human motion perception method based on a three-dimensional hybrid attention model provided in the aforementioned embodiments, and will not be further elaborated here.
[0214] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A human motion perception method based on a three-dimensional hybrid attention model, characterized in that: The method includes: Obtain wireless signal spectrum data; The wireless signal spectrum data is feature processed according to a preset three-dimensional hybrid attention model to determine the human motion perception features, wherein the three-dimensional hybrid attention model first performs feature compression on the wireless signal spectrum data, and then calculates attention weights from the three dimensions of time, frequency, and subcarrier based on a non-local attention mechanism to generate a three-dimensional attention feature map. Finally, the three-dimensional attention feature map is fused through cross-attention, and the human motion perception features are output after dimension adjustment; The step of performing feature processing on the wireless signal spectrum data according to a preset three-dimensional hybrid attention model to determine the human motion perception features includes: Performing feature compression on the wireless signal spectrum data using a preset first separable convolution model to generate an initial compressed feature map; Inputting the initial compressed feature map into a preset non-local attention model, performing embedding space mapping along the time dimension, frequency dimension, and subcarrier dimension respectively, and generating a time attention weight, a frequency attention weight, and a subcarrier attention weight based on the similarity calculation between the embedded features, and performing weighted fusion of the time attention weight, the frequency attention weight, and the subcarrier attention weight with the initial compressed feature map respectively to generate a three-dimensional attention feature map; Performing multi-dimensional joint feature fusion on the three-dimensional attention feature map through a preset cross-attention model to generate deep joint features; Using a preset second separable convolutional model to adjust the dimension of the deep joint feature, restore it to the target feature dimension, and output the human motion perception feature; A human motion perception recognition result is determined according to the human motion perception feature.
2. The method according to claim 1, wherein The step of compressing the wireless signal spectrum data by using a preset first separable convolution model to generate an initial compressed feature map includes: The local feature extraction of the spatial dimension of the input spectrum data is performed through deep convolution to obtain a local feature map; Compressing the channel dimension of the local feature map by point convolution to reduce the number of redundant feature channels and obtain a compressed feature map; The compressed feature map is enhanced by a nonlinear activation function to determine an initial compressed feature map.
3. The method according to claim 1, wherein The step of inputting the initial compressed feature map into a preset non-local attention model, performing embedding space mapping along the time dimension, frequency dimension and subcarrier dimension respectively, and generating a time attention weight, a frequency attention weight and a subcarrier attention weight based on the similarity calculation between the embedded features, and performing weighted fusion of the time attention weight, the frequency attention weight and the subcarrier attention weight with the initial compressed feature map respectively to generate a three-dimensional attention feature map includes: By linearly projecting the initial compressed feature map along the time dimension, frequency dimension and subcarrier dimension respectively, the embedded features of each dimension are generated. Calculate the similarity matrix between the embedded features of each dimension based on the embedded Gaussian function to generate the attention weight of each dimension; The attention weights of each dimension are weighted and summed with the initial compressed feature map, and then fused with the learnable gain parameters through residual connection to obtain a three-dimensional attention feature map.
4. The method according to claim 1, wherein The step of performing multi-dimensional joint feature fusion on the three-dimensional attention feature map through a preset cross attention model to generate deep joint features includes: The three-dimensional attention feature maps are spatially aligned and spliced to generate a joint feature matrix. Performing a secondary non-local attention operation on the joint feature matrix to calculate the global association weight across dimensions; The joint feature matrix is dynamically weighted by the global association weight, and the deep joint feature is determined after a nonlinear transformation.
5. The method according to claim 1, wherein The step of compressing the wireless signal spectrum data by using a preset first separable convolution model to generate an initial compressed feature map includes: In the non-local attention model, channel reduction technology is used to compress the number of channels in the embedding space to a preset ratio of the original features to reduce the amount of computation; Introducing subsampling operations in the cross-attention model, reducing the resolution of feature maps through maximum pooling or average pooling to reduce the scale of matrix operations; During model training, an incremental supervised data augmentation loss function is used, combined with a dynamic weight adjustment strategy to optimize model parameter updates to suppress overfitting and improve generalization capabilities. The bottleneck structure design limits the number of intermediate channels in the separable convolutional layer, reducing the number of model parameters and memory usage.
6. A human motion perception device based on a three-dimensional hybrid attention model, characterized in that: The device comprises: An acquisition module, used to acquire wireless signal spectrum data; a processing module for performing feature processing on the wireless signal spectrum data according to a preset three-dimensional hybrid attention model to determine human motion perception features, wherein the three-dimensional hybrid attention model first performs feature compression on the wireless signal spectrum data, then calculates attention weights from the three dimensions of time, frequency, and subcarrier based on a non-local attention mechanism to generate a three-dimensional attention feature map, and finally fuses the three-dimensional attention feature map through cross-attention, outputting human motion perception features after dimensional adjustment; The step of performing feature processing on the wireless signal spectrum data according to a preset three-dimensional hybrid attention model to determine the human motion perception features includes: Performing feature compression on the wireless signal spectrum data using a preset first separable convolution model to generate an initial compressed feature map; Inputting the initial compressed feature map into a preset non-local attention model, performing embedding space mapping along the time dimension, frequency dimension, and subcarrier dimension respectively, and generating a time attention weight, a frequency attention weight, and a subcarrier attention weight based on the similarity calculation between the embedded features, and performing weighted fusion of the time attention weight, the frequency attention weight, and the subcarrier attention weight with the initial compressed feature map respectively to generate a three-dimensional attention feature map; Performing multi-dimensional joint feature fusion on the three-dimensional attention feature map through a preset cross-attention model to generate deep joint features; Using a preset second separable convolutional model to adjust the dimension of the deep joint feature, restore it to the target feature dimension, and output the human motion perception feature; The determination module is used to determine the human motion perception recognition result according to the human motion perception feature.
7. A human motion perception device based on a three-dimensional hybrid attention model, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the human motion perception method based on a three-dimensional hybrid attention model as described in any one of claims 1 to 5.
8. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the human motion perception method based on a three-dimensional hybrid attention model as described in any one of claims 1 to 5 are implemented.
9. A computer program product, characterized in that The computer program product includes a computer program, which, when executed by a processor, implements the steps of the human motion perception method based on a three-dimensional hybrid attention model as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Human body action recognition method based on separable three-dimensional residual attention network
CN113065450A
Spatial data sensing method and device, equipment and storage medium
CN119760625A