Human body posture recognition method based on small sample radar data
By combining deep convolutional neural networks and self-attention mappers, improving loss functions and data enhancement techniques, the problems of accuracy and robustness in radar human posture recognition in small sample data environments were solved, achieving efficient posture recognition effects and expanding application scenarios.
Patent Information
- Application Number
- CN202510914329.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-10-17
AI Technical Summary
Existing radar human posture recognition technology lacks accuracy and robustness in environments with small sample data, and has low training efficiency, making it difficult to quickly adapt to complex environments and dynamic changes.
A deep convolutional neural network combined with a self-attention mapper is used to improve the loss functions in the pre-training and meta-training stages, combine data enhancement technology, optimize feature extraction and matching metrics, build a multi-layer feature extraction model, reduce overfitting, and improve the performance of the model in a small sample data environment.
It significantly improves the accuracy and robustness of radar human posture recognition, and enhances the recognition accuracy and training efficiency of the model under small sample conditions, making it suitable for fields such as health monitoring and intelligent transportation.
Smart Images

Figure CN120804878A_ABST
Abstract
Description
[0001] The technical field
[0002] The present application relates to the field of radar signal processing and artificial intelligence, in particular to a radar data processing method based on small sample learning, specifically including a method for human posture recognition using radar signals. More specifically, the present application proposes a small sample radar data human posture recognition method combined with deep learning technology by modifying the number of attention mappers and improving the loss function expression in the pre-training and meta-training stages. BACKGROUND
[0003] Radar technology has shown great potential in human posture recognition, especially in health monitoring, intelligent transportation and other applications. However, traditional radar signal processing methods have many challenges in complex environments and dynamic conditions, especially in small sample data environments, where model recognition accuracy and robustness are key issues.
[0004] Existing human posture recognition technologies mainly rely on visual sensors and wearable devices. These technologies perform well in specific conditions, but often perform poorly in changing light, occlusion and complex environments. Traditional vision-based posture recognition systems usually require a large amount of labeled data to train the model to ensure its recognition accuracy and generalization ability. However, obtaining a large amount of high-quality labeled data is both time-consuming and expensive.
[0005] In contrast, two-dimensional imaging through-wall radar technology has gradually attracted attention in human posture recognition. Through-wall radar technology can work stably under various lighting conditions and has the ability to penetrate, allowing it to detect subtle motion changes that visual systems cannot capture. However, existing radar posture recognition technologies still have significant shortcomings when faced with complex environments and dynamic changes.
[0006] Existing technologies mainly face the following problems: data scarcity, poor recognition accuracy and robustness, and low training efficiency. Traditional radar human posture recognition methods rely on a large amount of labeled data to train the model, but in practical applications, especially in new application scenarios, it is difficult to quickly obtain a large amount of labeled data. This results in poor generalization ability of the model in small sample data environments. Existing technologies are easily disturbed by environmental noise and dynamic changes when processing data obtained by radar systems in complex environments, affecting the accuracy and robustness of recognition. The characteristics and complexity of radar data increase the difficulty of model training, and existing methods have low training efficiency when processing small sample data, making it difficult to achieve fast response.
[0007] Invention purposes
[0008] The present application aims to solve some of the problems existing in the prior art in human posture recognition, especially to improve the recognition accuracy and robustness in small sample data environment. The existing radar human posture recognition technology relies on a large amount of labeled data, and when facing complex environment and dynamic changes, the recognition accuracy and robustness are insufficient, and the training efficiency is low. The task of the present application is to provide a two-dimensional imaging through-wall radar data human posture recognition method combining small sample learning and deep learning technology to significantly improve the recognition performance and adaptability of the model. SUMMARY
[0009] The present application proposes a human posture recognition method based on small sample radar data, which specifically includes the following steps:
[0010] First, feature extraction of radar signals is performed through a deep convolutional neural network, and data augmentation techniques are used to increase the diversity and quantity of training samples, ensuring that the model can learn features in different environments during training. The number and distribution of self-attention mappers are optimized to enable the model to pay more attention to important information during feature extraction. The self-attention mechanism dynamically adjusts the focus area according to the input data, reducing the impact of environmental noise and dynamic changes on the recognition result.
[0011] In the model training stage, the loss function expression in the pre-training and meta-training stage is improved, and a penalty term based on KL divergence is added to reduce overfitting, encourage the predicted probability distribution to be closer to the uniform distribution, and improve the performance of the model in small sample data environment. A two-stage training process is adopted, including pre-training and meta-training. The pre-training stage uses the optimized loss function for training; the meta-training stage uses the episodic training method, constructs the support set and query set, calculates the classification probability, and updates the network parameters.
[0012] Model evaluation is performed by analyzing the confusion matrix, average recognition accuracy, and monitoring training loss. The confusion matrix intuitively shows the performance of the model in the classification task, and understands the classification accuracy and error types of the model. The average accuracy represents the average performance of the model in multiple classification tasks, reflecting the stability and reliability of the model as a whole. Monitoring the training loss helps to understand the convergence of the model and the training effect.
[0013] Compared with the prior art, the present application has significant advantages. By optimizing the number and distribution of self-attention mappers, the model can pay more attention to important information during feature extraction, thereby improving recognition accuracy. The improved loss function expression adds a penalty term based on KL divergence, reducing overfitting and improving model performance in small sample data environments. Data augmentation techniques are used to increase the diversity and number of training samples, improving training efficiency. Combining deep convolutional neural networks and self-attention mappers for multi-scale feature extraction, the system can capture multiple aspects of image features from different scales and levels, building a more comprehensive feature set to improve classification accuracy and robustness.
[0014] Through preprocessing and enhancement of data, radar data can still maintain high recognition accuracy and robustness in complex environments. The method of the present application is not only suitable for human posture recognition, but also can be extended to other application scenarios that require efficient and accurate recognition, such as health monitoring, intelligent transportation, etc., showing a wide range of application prospects. In summary, the present application significantly improves the accuracy, robustness and training efficiency of small sample radar data human posture recognition by using improved self-attention mechanism, optimized loss function and multi-scale feature extraction, etc., and has wide application prospects and practical value. DETAILED DESCRIPTION
[0015] A human target posture recognition method based on small sample radar data is proposed. With the rapid development of radar technology, human posture recognition based on radar has shown great potential in health monitoring, intelligent transportation and other fields. Traditional radar signal processing methods face challenges in complex environments and dynamic changes, especially in small sample data environments, where the generalization ability and learning efficiency of the model are particularly critical. The present application explores the application of small sample learning technology in radar human posture recognition, aiming to improve the recognition accuracy and robustness of radar systems under the condition of a small number of training samples. We constructed a set-to-set matching measure model, optimized the traditional loss function, and verified its effectiveness through comparative experiments.
[0016] The human posture recognition method of the present application is shown in the accompanying Figure 1 as shown, mainly including steps 100-140.
[0017] Step 100, radar data acquisition and data preprocessing.
[0018] In the present application, we chose TK-TWR-2DI two-dimensional imaging radar as the main data acquisition tool, the main reason is that it has high penetration ability, high resolution, portability, ease of use and excellent clutter suppression ability, which can better simulate real scenarios.
[0019] To ensure data diversity and reduce individual bias, improve the reliability and practicality of the model, we selected three participants for the experiment. Each participant performed seven different actions, including walking, running, jumping, crawling, crawling, punching, and standing, to cover high and low posture changes.
[0020] During data collection, the TK-TWR-2DI two-dimensional imaging radar was placed in room A, and the data collection personnel performed action simulation in the adjacent room B to fully utilize the high penetration capability of the radar and accurately capture the action details. We collected data in a unified environment to ensure the reliability of the data quality. Before each data collection, we accurately calibrated the radar system, including positioning accuracy and signal strength, to ensure the consistency and accuracy of the data. Real-time monitoring of the collection process ensured standardized action execution and immediate adjustment of abnormal conditions for re-collection. Strict data collection procedures ensured high-quality and consistent data, laying a solid foundation for subsequent data processing and model training.
[0021] In the data preprocessing stage, we first use a high-pass filter with a transfer function of
[0022]
[0023] where f is the frequency, f c is the cutoff frequency of the filter. By applying this filter, we remove low-frequency components and only retain high-frequency parts, highlighting the signals of dynamic targets. The filtered signal better reflects the motion characteristics of the target.
[0024] Next, we divide the filtered radar data into 18-microsecond-long data segments, with each segment overlapping the previous one by 16 microseconds, ensuring that at least one complete action is included in each action cycle. This processing method not only ensures the continuity of the action but also increases the number of samples, improving the training effect of the model. Then, we use a 1024-point short-time Fourier transform (STFT) to perform time-frequency transformation on each data segment, generating time-frequency images that show the frequency distribution of radar signals over time and capture the dynamic characteristics of the target. The mathematical expression of STFT is:
[0025]
[0026] where x(τ) is the input signal, w(τ-t) is the moving window function, t and f represent time and frequency, respectively. By selecting an appropriate window function w(τ) and window length, we can obtain the distribution of the signal at different times and frequencies.
[0027] After generating the time-frequency images, we normalize all images to standardize the values within the range [0, 1], preventing extreme values from affecting model training. Normalization helps accelerate model convergence and improve training stability. Then, we resize the time-frequency images to a uniform size (120x120 pixels) to ensure consistency in input data. Finally, we divide the preprocessed data into training, validation, and test sets in a 50%, 25%, and 25% ratio for model training, tuning, and evaluation. Through these steps, we ensure high-quality and consistent radar data, providing a reliable foundation for subsequent human pose recognition model training and evaluation.
[0028] Step 110, feature extraction.
[0029] The feature extraction system of the present application includes a convolutional neural network (CNN) encoder and multiple self-attention mappers. Different numbers of self-attention mappers are embedded in different layers of the convolutional network to extract multi-scale features. Through this multi-level feature extraction mechanism, the system can capture multiple aspects of image features from different scales and levels, constructing a more comprehensive and detailed feature set.
[0030] The convolutional neural network is the basic structure of the feature extraction system, responsible for extracting initial features from input images. The CNN architecture of the present application includes multiple convolutional blocks, each composed of several convolutional layers. Specifically, the CNN architecture includes four main convolutional blocks, each responsible for extracting features of different scales.
[0031] The self-attention mapper is the core component of the present application, used to further process and extract features after each convolutional block. Each mapper module converts the features output by the convolutional block into multiple embedded feature vectors through a single-head self-attention mechanism. The specific workflow of these mappers is as follows:
[0032] First, the feature matrix output by the convolutional block is divided into multiple non-overlapping patches. As shown in the accompanying Figure 2 , assume that the input feature matrix is where P represents the number of patches, and D p represents the dimension of each patch. In actual operation, the patching of the feature matrix usually uses 1x1 patches, which means that each patch is actually a one-dimensional vector containing D p elements. This patching method can capture local information in the feature matrix in fine granularity, providing a basis for subsequent attention calculation.
[0033] After feature patching, as shown in the accompanying Figure 3 , the self-attention mapper calculates the query vector (Query) and the key vector (Key) dot product, and perform Softmax normalization to obtain the attention score β m The specific calculation formula of this step is as follows:
[0034]
[0035] wherein, denotes the query vector (Query) denotes the parameter set of the key vector (Key) denotes the parameter set of the key vector (Key) r k is a scaling factor to prevent the dot product value from being too large.
[0036] The attention score matrix β m reflects the correlation between different small blocks. A high score indicates that the similarity between the query vector and the key vector is high, i.e., there is a strong correlation between the corresponding small blocks. In this way, the self-attention mapper can dynamically adjust the weight of each small block, so that the model can focus on small blocks with important features.
[0037] After calculating the attention score matrix β m , the self-attention mapper uses the value vector (Value) and the attention score β to calculate the weighted output a m , as shown in equation (4):
[0038]
[0039] wherein, denotes the parameter set used to calculate the value vector (Value) , D is the weighted output, and D a is the dimension of the feature vector. The core of the attention weighting step is to use the attention score β m to weight and sum the feature vectors of each small block, thereby generating a new feature representation. In this way, the self-attention mapper can focus on the most representative small block features, thereby improving the quality and recognition of the feature representation.
[0040] After completing the above steps, the self-attention mapper generates the final feature vector g m through the mean pooling operation. The specific calculation formula is as follows:
[0041]
[0042] The purpose of mean pooling is to average all small block feature vectors to generate a global feature representation. This global feature representation not only retains the details of the local small block, but also integrates the features of different small blocks, thereby constructing a more recognizable image representation.
[0043] Because different convolutional blocks of the convolutional neural network extract different levels of features, placing the self-attention mapper after different convolutional blocks can capture these features and integrate global information, improve the recognition and transferability of the features, and construct a more comprehensive feature representation. The specific placement method is as follows:
[0044] Two self-attention mappers are embedded after the first convolutional block to process the edge and basic texture information of the image and enhance the representation ability of these features. Two self-attention mappers are embedded after the second convolutional block to process the middle-level features and better capture and integrate the structure and shape information of the image. Three self-attention mappers are embedded after the third convolutional block to process high-level features, extract rich semantic information, and improve the classification accuracy. Five self-attention mappers are embedded after the fourth convolutional block to process the highest level of features, represent global information, and further enhance these features to make the model perform better in the few-shot classification task.
[0045] Step 120, similarity measurement.
[0046] The present application calculates the minimum distance of the query sample and the support set sample on all self-attention mappers, and accumulates these minimum distances to obtain the final matching measure. The specific steps are as follows:
[0047] 1. Feature set extraction: For the query sample and the support set sample, respectively extract the feature set through the self-attention mapper. Assume that the feature set of the query sample is {g q1 ,g q2 ,...,g qm}, and the feature set of the support set sample is {g s1 ,g s2 ,...,g sn}, where m and n are the number of self-attention mappers.
[0048] 2. Distance calculation: Calculate the distance between the feature vectors of the query sample and the support set sample on each self-attention mapper. The calculation formula is as follows:
[0049] d(g qi, g sj )=-cos(g qi, g sj ) (6)
[0050] 3. Minimum distance selection: On each self-attention mapper, select the minimum distance g between the query sample and the support set samples qi . Specifically, for each query sample feature vector, find the one with the minimum distance among all support set sample feature vectors, and denote it as min j d(g qi, g sj ).
[0051] 4. Distance accumulation: Accumulate the minimum distances on each self-attention mapper to obtain the final matching metric. The calculation formula is as follows:
[0052]
[0053] Step 130, model training.
[0054] The present application adopts a two-stage training process to fully utilize the proposed feature set extractor and set matching metric. The training process includes a pre-training phase and a meta-training phase.
[0055] In the pre-training phase, a batch of instances is randomly sampled from the training set, denoted as X batch . For each instance x i ∈X batch , a feature set is extracted by the convolutional feature extractor. The feature set is further processed by a shallow self-attention mapper to generate multiple embedding feature vectors. In order to realize the classification task, a fully connected layer is attached to the feature vector output by each mapper to convert it into classification logits. The output of the fully connected layer is denoted as o m.c (g m,i ), where c represents the class. Using a specially designed cross-entropy loss function, the KL divergence between the current prediction distribution and the ideal uniform distribution is calculated and added to the cross-entropy loss to regularize the model and prevent overfitting. Each mapper is trained independently. The specific loss calculation formula is as follows
[0056]
[0057] where o m,c is the FC layer output of mapper m for class c, g m,i is the feature set of mapper m for instance x i , y i is the target output corresponding to instance x i . λ is a weight factor to balance the original loss and the KL divergence term. P is the predicted probability of each class, and Q is the probability of the uniform distribution.
[0058] During the meta-training phase, the model uses episodic training to simulate a few-shot learning scenario. N categories are randomly sampled from the training set, with K samples per category, to construct a support set. Furthermore, Q query samples are sampled from the same category to form a query set. For each sample in the support and query sets, a feature set is extracted using a convolutional feature extractor and a self-attention mapper. A set matching metric is used to calculate the distance between each query sample and the support set sample. The Softmax function is then used to calculate the probability that the query sample belongs to each category. The cross-entropy loss is then calculated, and the network parameters are updated through backpropagation.
[0059] l meta =-logp(y q |x q ,S) (9)
[0060] Among them, y q is the query sample x q The true category labels are obtained. Finally, after each epoch, the model performance is evaluated using the validation set. The model with the smallest validation loss is selected as the final model, and its parameters are recorded.
[0061] Step 140: Model evaluation.
[0062] This method primarily analyzes the confusion matrix, average recognition accuracy, and training loss. The confusion matrix provides a visual representation of the model's performance in classification tasks, allowing users to understand the model's classification accuracy and error types. The average accuracy, suitable for small-sample learning tasks, represents the average performance of the model across multiple classification tasks, reflecting the overall stability and reliability of the model. By monitoring changes in the training loss, users can understand the model's convergence and training effectiveness. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] To clearly explain the technical steps of the present invention, all the drawings used in the description of the present invention are briefly described below. It should be noted that the drawings described below are only some examples of the implementation of the present invention, and other persons skilled in the art can still obtain other drawings in different scenarios based on these drawings.
[0064] Attachment Figure 1 It is the implementation process of the present invention;
[0065] Attachment Figure 2 It is a feature extractor framework diagram of the present invention;
[0066] Attachment Figure 3 is a structural diagram of the attention mapper of the present invention;
[0067] AttachmentFigure 4 is a model training confusion matrix diagram of the present application;
[0068] attached Figure 5 is a model training accuracy and loss image of the present application;
[0069] Advantages
[0070] The present application proposes a human pose recognition method based on small sample radar data, aiming to improve the recognition accuracy and robustness under the condition of small training samples. The method improves the self-attention mapper and optimizes the loss function to build an effective recognition framework, and realizes efficient pose recognition by using various advanced technologies.
[0071] Specifically, this paper constructs a multi-layer feature extraction model containing convolutional neural network and self-attention mapper, uses short-time Fourier transform (STFT) to preprocess radar signals and extract dynamic target features. By designing the number and distribution of mappers and optimizing the traditional loss function, the learning ability and generalization ability of the model under small sample conditions are enhanced. In addition, according to the parameters of the actual application scene, a large number of experiments and comparative analysis are carried out to verify the effectiveness and superiority of the proposed scheme.
[0072] The improved model shows high accuracy in various pose recognition tasks. The experimental results are shown in Table 1, the 1-shot recognition accuracy is improved from 90.3% to 92.2%, and the 5-shot recognition accuracy is improved from 92.3% to 94.0%. Through small sample learning technology, the dependence on large-scale training data is reduced, and the model performance remains stable even if the sample number is reduced. The experimental results show that the recognition accuracy of the optimized model under the condition of 2 / 3 samples is almost the same as that under the condition of all samples, which verifies its effectiveness.
[0073] Table 1 Comparison of experimental results
[0074]
[0075] The recognition method of the present application not only has significant advantages in the field of human pose recognition, but also has wide application potential in the fields of intelligent transportation, health care, etc. In the future, the model structure will be further optimized, the recognition speed and real-time performance will be improved, the application scenarios will be expanded, and the radar technology will be promoted in more fields.
Claims
1. A human pose recognition method based on small-sample radar data. The method comprises the following steps: First, radar data processing is performed by modifying the number of attention mappers and improving the loss function expressions in the pre-training and meta-training phases, combined with deep learning techniques. Second, radar signal feature extraction is performed using a deep convolutional neural network, and data augmentation techniques are used to increase the diversity and number of training samples.
2. The method according to claim 1, wherein the recognition accuracy in the feature extraction process is improved by adjusting the number and distribution of self-attention mappers, and the self-attention mechanism dynamically adjusts the key areas of attention according to the different input data to reduce the impact of environmental noise and dynamic changes on the recognition results.
3. The method according to claim 1, wherein in the pre-training and meta-training stages, the traditional loss function expression is improved and a penalty term based on KL divergence is added to reduce overfitting. The penalty term based on KL divergence can encourage the predicted probability distribution to be closer to a uniform distribution, thereby improving the performance of the model in a small sample data environment.
4. The method according to claim 1, wherein the feature The extraction system includes a convolutional neural network encoder and multiple self-attention mappers. Different numbers of self-attention mappers are embedded in different layers of the convolutional network to extract multi-scale feature representations. The convolutional neural network contains multiple convolutional blocks, each of which consists of several convolutional layers for extracting feature representations at different scales.
5. The method according to claim 4, wherein each self-attention mapper obtains an attention score by calculating the dot product of the query vector and the key vector and performing Softmax normalization, calculates a weighted output using the value vector and the attention score, and generates a final feature vector through a mean pooling operation (Formula (5) in the specification).
6. The method according to claim 1, wherein in the data preprocessing stage, a high-pass filter is used to remove low-frequency components, a short-time Fourier transform (STFT) is used to generate a time-frequency image, and normalization is performed.
7. The method according to claim 1, wherein the loss function expression is improved and the optimized loss function is used for training in the pre-training stage, a penalty term based on KL divergence is added to encourage the predicted probability distribution to be closer to the uniform distribution to reduce overfitting, and an episodic training method is used to optimize the model in the meta-training stage.