Intelligent human body posture recognition method based on multi-mode micro-motion features and Transform network

By introducing multimodal micro-movement features and Transformer networks into millimeter-wave radar human posture recognition technology, the problems of low recognition accuracy and poor robustness are solved, and efficient and accurate human posture recognition in small sample scenarios are achieved.

CN120198955AActive Publication Date: 2025-06-24WUHU YABOSION ELECTRONIC TECH CO LTD

Patent Information

Application Number
CN202510259089.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-24
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

The recognition accuracy of human posture recognition technology based on micro-movement features under existing millimeter wave radar is low, especially in small sample scenarios, model performance has been greatly reduced and its robustness is poor.

Method used

The intelligent recognition method of human posture based on multimodal micro-modal features and Transformer network is adopted, and accurate recognition of human posture is achieved through radar echo model establishment, data preprocessing, Doppler image and spectrum sequence acquisition, Transformer network feature extraction and multimodal network classification output.

Benefits of technology

It improves the accuracy and robustness of human posture recognition, especially in small sample scenarios, which can still maintain a high recognition rate to adapt to the recognition needs in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198955A_ABST
    Figure CN120198955A_ABST
Patent Text Reader

Abstract

The invention discloses a human body posture intelligent identification method based on a multi-modal micro-motion feature and a Transform network, and the method comprises the steps: deducing micro-Doppler signal mathematical analysis expressions corresponding to different parts of a human body from human body target radar echo modeling; the method comprises the following steps: performing data preprocessing on target echo data acquired by a millimeter wave radar by adopting two-dimensional fast Fourier transform to obtain a one-dimensional spectrum sequence and a two-dimensional Doppler image corresponding to different postures, and respectively sending the one-dimensional spectrum sequence and the two-dimensional Doppler image into a one-dimensional Transform network and a two-dimensional Transform network for feature extraction; and carrying out feature splicing by adopting a multi-modal feature fusion model and outputting a classification result. According to the method, the problems of low recognition rate, poor robustness and the like in the existing human body posture intelligent recognition means can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of millimeter-wave radar signal processing technology and deep learning algorithm technology, and particularly relates to a human body posture intelligent recognition method based on multi-modal micro-motion features and a Transformer network. Background Art

[0002] The human body posture intelligent recognition technology based on multi-modal micro-motion features and a Transformer network relies on radar signal analysis and combines the powerful analysis ability of the Transformer network to accurately extract human body posture information. Compared with traditional optical methods and wearable devices, the radar data acquisition and analysis combined with a neural network have the advantages of non-contact, all-weather, and strong penetrability, and can overcome limitations such as light, occlusion, and user dependence. On this basis, the multi-modal fusion technology further improves the recognition accuracy and robustness, provides reliable support for human body behavior monitoring in complex scenarios, and is widely applied in fields such as security monitoring, intelligent interaction, and medical health. Therefore, it is of great research significance to carry out research on the human body posture intelligent recognition method based on multi-modal micro-motion features and a Transformer network.

[0003] Most of the existing human body posture recognition technologies based on micro-motion features under millimeter-wave radar use a single traditional neural network, not only with low recognition accuracy, especially in small-sample scenarios, where the model performance drops significantly and is extremely vulnerable to the surrounding environment. Therefore, there is an urgent need for a human body posture intelligent recognition method with high recognition accuracy and strong robustness. Summary of the Invention

[0004] The purpose of the present invention is to propose a human body posture intelligent recognition method based on multi-modal micro-motion features and a Transformer network, which can overcome problems such as low recognition rate and poor robustness existing in the existing human body posture intelligent recognition means.

[0005] To achieve the above technical purpose, the technical solution adopted by the present invention is as follows:

[0006] A human body posture intelligent recognition method based on multi-modal micro-motion features and a Transformer network, the method comprising the following steps:

[0007] S1, establishing a radar echo model of human body movement: modeling the micro-motions caused by the movements of different parts of the human body, including the upper limbs, lower limbs, and torso, and calculating the mathematical analytical expressions of the echo signals of different parts of the human body;

[0008] S2. Data preprocessing: For the target echo data collected by the millimeter-wave radar, the real and imaginary parts of the intermediate-frequency signal obtained by mixing the received signal and the transmitted signal are recombined into a complex form, then rearranged into a two-dimensional array, and finally the echo data after single-channel processing is extracted for superposition;

[0009] S3. Acquisition of Doppler image and spectrum sequence: Perform a two-dimensional fast Fourier transform on the preprocessed radar echo signal to obtain a two-dimensional Doppler time-frequency diagram. The maximum amplitude value of each frequency point of the two-dimensional Doppler time-frequency diagram is rearranged according to the time dimension to obtain a one-dimensional Doppler spectrum sequence corresponding to different postures;

[0010] S4. Feature extraction by Transformer network: Define a one-dimensional Transformer network and a two-dimensional Transformer network, and send the two-dimensional Doppler time-frequency diagram and the one-dimensional Doppler spectrum sequence corresponding to different postures into the defined two-dimensional Transformer network and one-dimensional Transformer network respectively for feature extraction;

[0011] S5. Classification output of multi-modal network: Construct a multi-modal feature fusion model, splice the features extracted by the one-dimensional Transformer network and the two-dimensional Transformer network, obtain the classification output through a fully connected layer, and identify the human body posture.

[0012] Step S1 further includes:

[0013] The human body is abstracted into a linear rigid body structure. The human body target moves along the radar line of sight. The distance between the human body target and the radar in the horizontal direction is R0. The plane determined by the radar and the human body target is the illumination plane, and the angle between the illumination plane and the human body target movement direction is

[0014] The human body movement is simplified to a linear rigid body movement. Calculate the instantaneous distance vector and instantaneous radial distance between the scatter points on different parts involved in the human body movement and the radar at time t. According to the instantaneous radial distance of the scatter points relative to the radar, obtain the radar echo of the scatter points; perform mixing and low-pass filtering on the radar echo to obtain a zero-intermediate-frequency echo signal, and further obtain the mathematical analytical expression of the linear structure target radar echo signal of different parts involved in the human body movement.

[0015] Further, when the upper limb of the human body moves, the process of modeling the radar echo of the upper limb movement includes the following steps:

[0016] Use the coordinate system (U, V, W) for irradiating the upper limb of the human body by the radar, where the radar is located at the origin Q(0, 0, 0) of the coordinate system, and the reference coordinate system (X, Y, Z) is a translation of the radar coordinate system, and the origin Q of this coordinate systemS The coordinates in the radar coordinate system are (x s , y s , z s ). Assume that the human target moves along the radar line of sight, and the distance between the human target and the radar in the horizontal direction is R0. The plane determined by the radar and the human target is the illumination plane, and the angle between the illumination plane and the moving direction of the human target is

[0017] Simplify the movement of the two arms into a pendulum movement in the front-back direction. At time t, the angle between the right arm of the human target and the negative direction of the Z-axis is θ S (t). Point A is one of the scattering points on the right arm, and 0 ≤ x A ≤ L armR , where L armR is the length of the right arm; the variation law of θ S (t) is θ s (t) = θ smax sin(ωt + ψ0), where θ smax is the angular amplitude when the arm swings, ω is the swing speed, ψ0 is the initial phase of the arm swing, the initial phase of the right arm is 0, and the initial phase of the left arm is π; the instantaneous distance vector between point A and the radar at time t is:

[0018]

[0019] where, is the initial rotation matrix, is the rotation matrix at time t, is 's unit vector, and its magnitude is β1 is 's elevation angle relative to the origin Q, is 's unit vector, and its magnitude is:

[0020]

[0021] where α A (t) and β A (t) are respectively 's azimuth angle and elevation angle relative to the starting point Q S , and are respectively:

[0022]

[0023] Combined with the instantaneous distance vector between point A and the radar at time t, let ω ref = 0, α0 = 0, k is the period extension variable, and the instantaneous radial distance between point A and the radar is calculated as:

[0024]

[0025] Scattering point A on the left arm L The instantaneous radial distance from the radar is:

[0026]

[0027] where 0 ≤ x AL ≤ L armL , L armL is the length of the left arm, θ sL (t) = θ s (t + T / 2) = θ smax sin(ωt + π) = -θ s (t);

[0028] According to the instantaneous radial distance of the scattering point relative to the radar, the radar echo of the scattering point is:

[0029] s t (t) = u1exp(j2πf0t)

[0030] In the formula, u1 is the amplitude of the radar transmitted signal. By mixing and low-pass filtering the radar echo, a zero-intermediate-frequency echo signal is obtained, which is expressed as:

[0031] s t (t) = ρ(x, y, z)exp{j2πf0(-2R(t) / c)} = ρ(x, y, z)exp{jΦ(t)}

[0032] where ρ(x, y, z) is the scattering coefficient of the scattering point, t0 ≤ t ≤ t s , t0 is the echo response time of the target at a distance R0 from the radar at time 0. When the moving speed of the target is much smaller than the speed of light, t0 ≈ 2R0 / c, t s is the radar signal sampling time; Φ(t) = 2πf0(-2R(t) / c) is the phase of the zero-intermediate-frequency echo signal; the radar echo of the linear structure target of upper limb movement is expressed as:

[0033] where L is the arm length.

[0034] Substituting the expression of the instantaneous radial distance of the arm into the expression of the zero-intermediate-frequency echo signal, the radar echo of A and A L can be obtained:

[0035]

[0036] The radar echoes of the right arm and the left arm during human upper limb movement are:

[0037]

[0038] The modeling and analysis of the radar echoes of the human lower limbs and torso are similar.

[0039] Step S2 further includes:

[0040] The millimeter-wave radar single channel transmits a chirp signal. The four receiving channels mix the echo reflected by the target with the transmitted signal to obtain an intermediate-frequency signal. The real and imaginary parts of the obtained intermediate-frequency signal are recombined into a complex form to generate echo data containing complex signals, and then rearranged into a two-dimensional array. The rows represent the number of sampling points of the receiving channels, and the columns represent the total number of pulses. Each row corresponds to the data of one receiving channel, and each column corresponds to a time series. The combined echo data is readjusted into a matrix format suitable for subsequent processing;

[0041] According to the number of 4 receiving antennas, the data is separated and stored according to the receiving channels, and then the echo data of all receiving antenna channels is added point by point to obtain the combined radar echo signal. By calculating the average value of the pulse sequence within each frame and subtracting this average value from the signal, the DC bias or low-frequency component in the signal is eliminated, and the echo signal matrix without DC component is obtained.

[0042] Step S3 further includes:

[0043] For the radar echo signal after data preprocessing, use two-dimensional fast Fourier transform to extract its image and sequence; specifically, perform FFT on the radar echo matrix in the distance dimension to obtain the distance-time relationship, and then perform FFT on the distance-time relationship frame by frame in the Doppler dimension to obtain the distance-Doppler relationship; accumulate the 2D FFT data of each frame in the Doppler dimension to obtain the micro-Doppler at one moment; splice the micro-Doppler data of each frame according to time to obtain a two-dimensional Doppler time-frequency map; extract the maximum value of each frequency point of the two-dimensional Doppler time-frequency map according to the time dimension and splice them one by one to obtain a one-dimensional Doppler spectrum sequence.

[0044] Further, in step S4, the one-dimensional Transformer network includes a spectrum input layer, an embedding layer, a first position encoding layer, a Transformer encoder block, an average pooling layer, and a first feature output layer;

[0045] The embedding layer is used to map the dimension of the input one-dimensional Doppler spectrum sequence of 255 frequency points to the embedding dimension, defined as 256, and the output matrix shape is (255, 256). At the same time, the Xavier function is used to initialize the weights to make the gradient distribution uniform;

[0046] The first position encoding layer is used to add position information to the output data of the embedding layer, enabling the model to perceive the relative order of each moment and spatial point in the data, and performing a transpose operation to obtain an output matrix with a shape of (256, 255) to meet the input requirements of the Transformer encoder layer;

[0047] The Transformer encoder block is used to extract sequence features from the input data, including a multi-head self-attention layer, a first residual connection and normalization layer, a feed-forward neural network, and a second residual connection and normalization layer;

[0048] The multi-head self-attention layer is used to capture global dependencies in the input sequence. Through multiple attention heads, each containing different query matrices, key matrices, and value matrices, it calculates the correlations between different parts in parallel, and then learns different patterns or features, enabling the network to focus on each part of the sequence from multiple perspectives.

[0049] The residual connection and normalization layer is used to solve the problems of gradient vanishing and gradient explosion, and ensure the stability of the data distribution. The Transformer structure uses residual connections to directly add the input of each layer to the output, which helps the gradient to be better transmitted during backpropagation and enables the model to learn deeper features. A normalization operation is performed after each layer to accelerate the training process, reduce the oscillation phenomenon during training, and improve the convergence speed of the model.

[0050] The feed-forward neural network layer includes two linear transformations and activation functions, and is used to further transform and non-linearly map the features extracted by the self-attention layer. Through the first linear transformation and non-linear activation function, the input features are mapped to a higher-dimensional space to filter out more important features, thereby enhancing the expressive ability of the features. Through the second linear transformation, the high-dimensional features are mapped back to the dimension of the original input to ensure compatibility with subsequent pooling operations.

[0051] The average pooling layer calculates the mean value along the sequence dimension to generate global features with a fixed dimension, which are output via the feature output layer.

[0052] Furthermore, in step S4, the two-dimensional Transformer network includes an image input layer, a packet embedding module, a class token module, a connection layer, a second position encoding layer, a dropout layer, a Transformer encoder layer, a first normalization layer, a class token segmentation layer, and a second feature output layer;

[0053] The packet embedding module includes a flattening layer and a two-dimensional convolutional layer connected to each other. Among them, the flattening layer divides the two-dimensional Doppler time-frequency map received by the image input layer into multiple patches, and then flattens each patch into a vector with a length of 196. The two-dimensional convolutional layer maps the vector output by the flattening layer to a high-dimensional space and outputs an embedding vector with a matrix shape of (196, 768). This part can extract local features and improve the model's ability to understand spatial information.

[0054] The class token module is used to add a class token of the class token at the front end of the embedding vector. Before the embedding vector enters the Transformer encoder layer, an additional learnable class token is added. This token can be used for global feature aggregation in the final stage to help the classification task learn global information.

[0055] The connection layer is used to splice the embedding vector and the class token to obtain a matrix with a shape of (197, 768).

[0056] The second position encoding layer is used to add position information. The position encoding and the image embedding vector are added element by element, and the matrix shape remains unchanged, still (197, 768).

[0057] The dropout layer adds Dropout to prevent overfitting and improve the generalization ability of the model.

[0058] The Transformer encoder layer includes 10 Transformer encoder blocks stacked in a loop. Each Transformer encoder block contains the multi-head self-attention layer, the first residual connection and normalization layer, the feed-forward neural network, and the second residual connection and normalization layer introduced above.

[0059] The first normalization layer is used for the output processing of the Transformer encoder layer to improve the training stability of the model.

[0060] The class token segmentation layer is used to extract the class token containing global image features.

[0061] The second feature output layer outputs the class token of the class token segmentation layer.

[0062] Furthermore, the multi-modal feature fusion model includes a linear layer, a second normalization layer, a feature fusion layer, and a fully connected layer connected in sequence.

[0063] The linear layer uniformly converts the extracted feature dimensions into multi-modal dimensions, defined as 512; the second normalization layer normalizes the features of different modalities to ensure that the features are within a similar numerical range, reducing the differences between modalities and improving the fusion effect; the feature fusion layer concatenates the image and spectrum features with unified dimensions and normalized, obtaining an output matrix with a shape of (1, 1024); the fully connected layer maps the concatenated feature dimensions to the category dimension, defined as 9-classification, and outputs the category labels.

[0064] The human body posture intelligent recognition method based on multi-modal micro-motion features and Transformer network of the present invention starts from the modeling of the radar echo of the human body target, combines the millimeter-wave radar signal system, deduces the mathematical analytical expressions of the micro-Doppler signals corresponding to different parts of the human body, and clarifies the micro-motion coupling mechanism; for the target echo data collected by the millimeter-wave radar, two-dimensional fast Fourier transform is used to preprocess the data, thereby obtaining the one-dimensional spectrum sequence and two-dimensional Doppler image corresponding to different postures; they are respectively fed into a custom one-dimensional Transformer network and two-dimensional Transformer network for feature extraction; enter the multi-modal network for feature splicing corresponding to one-dimensional and two-dimensional extracted by the Transformer network, and finally the network outputs the classification result, realizing the intelligent recognition of human body postures. Description of the Drawings

[0065] Figure 1 It is a flowchart of the human body posture intelligent recognition method based on multi-modal micro-motion features and Transformer network.

[0066] Figure 2 It is the radar observation coordinate system for upper limb movement.

[0067] Figure 3(a) is the spectrogram and time-frequency diagram of the posture of marching in place.

[0068] Figure 3(b) is the spectrogram and time-frequency diagram of the waving posture.

[0069] Figure 3(c) is the spectrogram and time-frequency diagram of the walking posture.

[0070] Figure 3(d) is the spectrogram and time-frequency diagram of the jogging posture.

[0071] Figure 3(e) is the spectrogram and time-frequency diagram of the posture of walking with a backpack.

[0072] Figure 3(f) is the spectrogram and time-frequency diagram of the posture of punching in place.

[0073] Figure 3(g) is the spectrogram and time-frequency diagram of the posture of marching and waving.

[0074] Figure 3(h) is the spectrogram and time-frequency diagram of the posture of walking and waving.

[0075] Figure 3(i) is the spectrogram and time-frequency diagram of the running and waving gesture.

[0076] Figure 4 It is the defined one-dimensional Transformer network structure diagram.

[0077] Figure 5 It is the defined two-dimensional Transformer network structure diagram.

[0078] Figure 6 It is the defined multimodal feature fusion model structure diagram.

[0079] Figure 7 It is the line graph of the recognition accuracy rate of the multimodal feature fusion model when the proportion of training samples is 10-80%.

[0080] Figure 8 It is the curve graph of the loss values during the training and validation processes of the multimodal feature fusion model when the proportion of training samples is 10%.

[0081] Figure 9 It is the confusion matrix graph of the multimodal feature fusion model when the proportion of training samples is 10%. Detailed implementation manners

[0082] The following further describes the embodiments of the present invention in detail with reference to the accompanying drawings.

[0083] A human body posture intelligent recognition method based on multimodal micro-motion features and Transformer network, the method includes the following steps:

[0084] S1, Establishment of the radar echo model of human body movement: Model the micro-motions caused by the movements of different parts of the human body, including the upper limbs, lower limbs, and torso, and calculate the mathematical analytical expressions of the echo signals of different parts of the human body;

[0085] S2, Data preprocessing: For the target echo data collected by the millimeter-wave radar, recombine the real part and the imaginary part of the intermediate-frequency signal obtained by mixing the received signal and the transmitted signal into a complex form, then rearrange it into a two-dimensional array, and finally extract the echo data after single-channel processing for superposition;

[0086] S3, Acquisition of Doppler images and spectral sequences: Perform two-dimensional fast Fourier transform on the preprocessed radar echo signal to obtain a two-dimensional Doppler time-frequency diagram, and rearrange the maximum amplitude value of each frequency point of the two-dimensional Doppler time-frequency diagram in the time dimension to obtain a one-dimensional Doppler spectral sequence corresponding to different postures;

[0087] S4, Transformer Network Feature Extraction: Define a one-dimensional Transformer network and a two-dimensional Transformer network, and send the two-dimensional Doppler time-frequency maps corresponding to different postures and the one-dimensional Doppler spectrum sequences into the defined two-dimensional Transformer network and one-dimensional Transformer network respectively for feature extraction;

[0088] S5, Multimodal Network Classification Output: Construct a multimodal feature fusion model, splice the features extracted by the one-dimensional Transformer network and the two-dimensional Transformer network, and obtain the classification output through a fully connected layer to identify the human posture.

[0089] Traditional human posture recognition technologies rely on optical sensors and image processing technologies. However, in the case of poor lighting conditions, such as extreme environments like dim or strong direct light, the performance of optical image-based technologies will decline rapidly. In contrast, radar technology is not restricted by weather and lighting conditions and is an ideal tool for human posture recognition. The key based on micro-motion characteristics is to analyze the motion signals of the human body through radar technology. These signals include the micro-motions of the human body. When the human body moves, different parts of the body will reflect radar signals at different frequencies and amplitudes. These reflected signals carry the time and frequency characteristics of the motion. Therefore, different postures can be identified and classified using a multimodal Transformer network based on the differences in the reflected signal expressions.

[0090] The effects of the human body on radar signals are mainly manifested as refraction, reflection, and absorption. These phenomena are closely related to the complex structure of the human body. Since the human body is a non-rigid body and belongs to an inhomogeneous medium, it is very difficult to model the human body. In the present invention, the human body is abstracted into a simple linear rigid body structure, and the radar echo of the human body taking a step in place is modeled. Use the coordinate system (U, V, W) for irradiating the upper limb of the human body with radar, where the radar is located at the origin position Q(0, 0, 0) of the coordinate system. The reference coordinate system (X, Y, Z) is a translation of the radar coordinate system, and the origin Q of this coordinate system S has the coordinates (x s , y s , z s ) in the radar coordinate system. Assume that the human target moves along the radar line of sight direction, and the distance between the human target and the radar in the horizontal direction is R0. In the present invention, the plane determined by the radar and the human target is the irradiation plane, and the angle between the irradiation plane and the human target's motion direction is φ. After simplifying the human motion into a linear rigid body motion, the motion of the two arms can be simplified into a pendulum motion moving in the front-back direction. Assume that at time t, the right arm of the human target swings to the position shown in the radar observation coordinate system for upper limb motion. At this time, the angle between this arm and the negative direction of the Z axis is θ S(t), assume that A is a certain point on the right arm, and 0 ≤ x A ≤ L armR , L armR is the length of the right arm. The variation law of θ S (t) is θ s (t) = θ smax sin(ωt + ψ0), where θ smax is the angular amplitude when the arm swings, ω is the swing speed, ψ0 is the initial phase of the arm swing, the initial phase of the right arm is 0, and the initial phase of the left arm is π; The instantaneous distance vector between point A and the radar at time t is:

[0091]

[0092] where, is the initial rotation matrix, is the rotation matrix at time t, is 's unit vector, and its magnitude is β1 is 's elevation angle relative to the origin Q, is 's unit vector, and its magnitude is:

[0093]

[0094] where α A (t) and β A (t) are respectively 's azimuth angle and elevation angle relative to the starting point Q S , and are respectively:

[0095]

[0096] Combined with the instantaneous distance vector between point A and the radar at time t, let ω ref = 0, α0 = 0, k is the periodic extension variable, and the instantaneous radial distance between point A and the radar is calculated as:

[0097]

[0098] The instantaneous radial distance between the scattering point A on the left arm L and the radar is:

[0099]

[0100] where, 0 ≤ x AL ≤ L armL , L armL is the length of the left arm, θ sLψ(t) = θ s ψ(t + T / 2) = θ smax sin(ωt + π) = -θ s ψ(t);

[0101] The radar echo of the scattering point is obtained according to the instantaneous radial distance of the scattering point relative to the radar as follows:

[0102] s t (t) = u1exp(j2πf0t)

[0103] Where u1 is the amplitude of the radar transmitted signal. By mixing and low-pass filtering the radar echo, the zero-intermediate-frequency echo signal is obtained, which is expressed as:

[0104] s t (t) = ρ(x, y, z)exp{j2πf0(-2R(t) / c)} = ρ(x, y, z)exp{jΦ(t)}

[0105] Where ρ(x, y, z) is the scattering coefficient of the scattering point, t0 ≤ t ≤ t s , t0 is the echo response time of the target at a distance R0 from the radar at time 0. When the moving speed of the target is much smaller than the speed of light, t0 ≈ 2R0 / c, t s is the radar signal sampling time; Φ(t) = 2πf0(-2R(t) / c) is the phase of the zero-intermediate-frequency echo signal; The radar echo of the linear structure target of upper limb movement is expressed as:

[0106] s target = ∫0 L s(t)dx

[0107] Where L is the arm length.

[0108] Substituting the expression of the instantaneous radial distance of the arm into the expression of the zero-intermediate-frequency echo signal, the radar echoes of A and A L can be obtained:

[0109]

[0110] The radar echoes of the right arm and the left arm during human upper limb movement are:

[0111]

[0112] The modeling and analysis of the radar echoes of human lower limbs and torso movements are similar to this.

[0113] The millimeter-wave radar's single channel transmits a chirp signal. After the four receiving channels mix the echo reflected by the target with the transmitted signal to obtain an intermediate-frequency signal, the real and imaginary parts of the obtained intermediate-frequency signal are recombined into a complex form to generate echo data containing complex signals. Then it is rearranged into a two-dimensional array, where the rows represent the number of sampling points of the receiving channels and the columns represent the total number of pulses. In this way, each row corresponds to the data of one receiving channel, and each column corresponds to a time series. Finally, the combined echo data is readjusted into a matrix format suitable for subsequent processing. According to the number of 4 receiving antennas, the data is separated and stored according to the receiving channels, and then the echo data of all receiving antenna channels is added point by point to obtain the combined radar echo signal. By calculating the average value of the pulse sequence within each frame and subtracting this average value from the signal, the DC bias or low-frequency components in the signal are eliminated, and finally an echo signal matrix without DC components is generated.

[0114] For the radar echo signal after data preprocessing, two-dimensional fast Fourier transform (2D FFT) is used to extract its image and sequence. First, FFT is performed on the radar echo matrix in the range dimension to obtain the range-time relationship. Then, FFT in the Doppler dimension is performed frame by frame on the range-time relationship to obtain the range-Doppler relationship. Next, the 2D FFT data of each frame is accumulated in the Doppler dimension to obtain the micro-Doppler at a certain moment. In other words, each time a frame of data is taken out, the abscissa of the data represents Doppler and the ordinate represents range. Since only Doppler information is needed, the range information can be ignored, and the data of each row is accumulated correspondingly. What is obtained is a one-dimensional Doppler data. Finally, the micro-Doppler data of each frame is spliced according to time to obtain the time-frequency diagram, that is, the Doppler image. The maximum value of each frequency point of the Doppler image is extracted in the time dimension and then spliced one by one to obtain the Doppler sequence.

[0115] For the obtained Doppler images and sequences, a two-dimensional Transformer and a one-dimensional Transformer network are respectively constructed to extract their features. According to the basic structure of the Transformer, a one-dimensional Transformer network is defined as follows: First, the initialization method is defined. An embedding layer is created to map the input spectrum sequence dimension (1) to the embedding dimension (512). The weights are initialized using the Xavier function to ensure uniform gradient distribution. Then, positional information is added through class instantiation of the positional encoder. Next, the encoder block is class instantiated, which contains a multi-head self-attention layer, a feed-forward neural network, two residual connections, and a normalization layer. Stacking the encoder block 10 times in a loop is defined as a Transformer encoder layer. Finally, the forward propagation method is defined. First, it is checked whether the input dimension is 3D (batch, 255, 1), and then feature mapping is performed. The input is mapped from 1D to 256D through a fully connected layer, resulting in a shape of (batch, 255, 256). Then, positional encoding is added. After that, a transpose operation is performed to adapt to the input of the Transformer encoder layer, and then sequence features are extracted. Finally, the mean value of the sequence dimension is obtained through an average pooling layer to generate a global feature output with a fixed dimension and a shape of (batch, 256), where batch is the training batch size.

[0116] Based on the one-dimensional Transformer network, an extension is made to define a two-dimensional Transformer network: First, a packet embedding class is defined. The class method includes splitting the image into multiple patches. Then, a flattening layer is defined to flatten the split patches into a vector with a length of 196. Finally, a two-dimensional convolutional layer is defined to map this vector to a high-dimensional space (768), and finally, an embedding vector with a matrix shape of (196, 768) is obtained. A connection layer is defined, and a class token category marker is added at the front end of the image, and the shape becomes (197, 768). Then, positional information is added through class instantiation of the positional encoder. The positional encoding and the image embedding vector are added element by element to obtain a matrix with a shape of (197, 768). Before entering the Transformer encoder block, a dropout class is defined to improve the generalization ability of the model. Then, the encoder block of the Transformer network is defined, and its internal structure is the same as that of the one-dimensional Transformer network. After the stacked output of the encoder block passes through a normalization layer, the optimization parameters are updated to improve the model stability and accelerate the training process. Finally, a class marker extraction layer is defined to extract the class token containing the global image features to obtain a feature output with a shape of (batch, 768).

[0117] For the features of the spectrum sequence and time-frequency image extracted by the one-dimensional and two-dimensional Transformer networks, a multi-modal feature fusion model is defined for feature splicing and class output: First, a linear layer is defined to uniformly convert the extracted feature dimensions into the multi-modal dimension (512); then a normalization layer is defined to ensure that the feature values are in the range of [0,1]; a feature fusion class is defined to splice the normalized image and spectrum dimensions with the unified dimension, and the shape is (batch, 2*512); finally, a fully connected layer is defined to map the spliced feature dimensions to the class dimension (9) to achieve class output. Change the proportion of different training samples (10-80%) and repeat this step.

[0118] Example

[0119] To verify the effectiveness of the proposed solution of the present invention, the following test experiments are carried out. The system parameters are defined as follows: For the radar signal processing, the number of transmitting antenna channels is set to 1, the number of receiving antenna channels is 4, the sampling rate is 6 MHz, the number of sampling points per pulse is 256, the number of pulses per frame is 255, and the number of data frames is 400; for the neural network, the embedding dimension of the one-dimensional Transformer network is set to 256, the maximum input sequence length is 255, the number of encoder blocks is 4, the number of heads in the self-attention mechanism layer is 8, and the dimension of the feed-forward neural network is 256; the embedding dimension of the two-dimensional Transformer network is set to 768, the number of encoder blocks is 10, the number of heads in the self-attention mechanism layer is 12, the dimension of the feed-forward neural network is 2048, and the dropout rate is 0.1; the fusion dimension of the multi-modal feature fusion model is set to 512, and the number of classes is 9.

[0120] Figure 2 is the radar observation coordinate system for upper limb movement Figures 3(a) to 3(i) The spectrogram and time-frequency diagram corresponding to different human postures are shown. It can be seen that the spectrogram is actually the time-frequency diagram rearranged by the maximum amplitude in the time dimension at each frequency point. Figure 4 is the structural diagram of the defined one-dimensional Transformer network Figure 5 is the structural diagram of the defined two-dimensional Transformer network Figure 6 is the structural diagram of the defined multi-modal feature fusion model Figure 7 is the line chart of the recognition accuracy of the multi-modal feature fusion model under the training sample ratio of 10-80%. It can be seen that the recognition accuracy of the multi-modal network increases continuously with the increase of the training sample ratio, and the recognition rate remains above 95% in the small sample scenario. Figure 8 is the curve chart of the loss values during the training and validation process of the multi-modal feature fusion model under the training sample ratio of 10%. It can be seen that the training and validation process of the multi-modal feature fusion model converges stably and there is no obvious overfitting phenomenon.Figure 9 It is the confusion matrix diagram of the multi-modal feature fusion model when the proportion of training samples is 10%. It can be seen that the classification quantities are significantly concentrated on the anti-diagonal line, indicating excellent model performance.

[0121] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present application can be implemented using various computer languages. For example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript.

[0122] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions run by the processors of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0123] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0124] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are run on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions running on the computer or other programmable device provide steps for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0125] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present application.

[0126] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

Claims

1. A human body posture intelligent recognition method based on multimodal micro-motion features and Transformer network, characterized in that: The method comprises the following steps: S1, Establishment of radar echo model of human motion: Model the micro-motion caused by the motion of different parts of the human body, including upper limbs, lower limbs and trunk, and calculate the mathematical analytical expressions of echo signals of different parts of the human body; S2, data preprocessing: for the target echo data collected by the millimeter wave radar, the real and imaginary parts of the intermediate frequency signal obtained by mixing the received signal and the transmitted signal are recombined into a complex form, and then rearranged into a two-dimensional array, and finally the single-channel processed echo data is extracted for superposition; S3, Doppler image and spectrum sequence acquisition: perform two-dimensional fast Fourier transform on the preprocessed radar echo signal to obtain a two-dimensional Doppler time-frequency diagram, and rearrange each frequency point of the two-dimensional Doppler time-frequency diagram according to the maximum amplitude value in the time dimension to obtain a one-dimensional Doppler spectrum sequence corresponding to different postures; S4, Transformer network feature extraction: define a one-dimensional Transformer network and a two-dimensional Transformer network, and send the two-dimensional Doppler time-frequency diagram and the one-dimensional Doppler spectrum sequence corresponding to different postures to the defined two-dimensional Transformer network and one-dimensional Transformer network for feature extraction; S5, multimodal network classification output: Build a multimodal feature fusion model, concatenate the features extracted by the one-dimensional Transformer network and the two-dimensional Transformer network, obtain the classification output through a fully connected layer, and recognize the human posture.

2. The method for intelligent human posture recognition based on multimodal micro-motion features and Transformer network according to claim 1 is characterized in that: Step S1 further comprises: The human body is abstracted into a linear rigid body structure. The human target moves along the radar line of sight. The horizontal distance between the human target and the radar is R0. The plane determined by the radar and the human target is the illumination plane. The angle between the illumination plane and the movement direction of the human target is The human body motion is simplified into linear rigid body motion, and the instantaneous distance vector and instantaneous radial distance between the scattering points on different parts involved in the human body motion and the radar at time t are calculated. The radar echo of the scattering point is obtained according to the instantaneous radial distance of the scattering point relative to the radar; the radar echo is mixed and low-pass filtered to obtain a zero intermediate frequency echo signal, and then the mathematical analytical expression of the linear structure target radar echo signal of different parts involved in the human body motion is obtained.

3. The method for intelligent human posture recognition based on multimodal micro-motion features and Transformer network according to claim 1 is characterized in that: When the upper limbs of a person move, the process of modeling the radar echo of the upper limb movement includes the following steps: The coordinate system (U, V, W) of the human upper limbs is illuminated by radar, where the radar is located at the origin of the coordinate system Q (0, 0, 0). The reference coordinate system (X, Y, Z) is the translation of the radar coordinate system. The origin of the coordinate system Q S The coordinates in the radar coordinate system are (x s ,y s ,z s ), assuming that the human target moves along the radar line of sight, and the horizontal distance between the human target and the radar is R0, the plane determined by the radar and the human target is the illumination plane, and the angle between the illumination plane and the moving direction of the human target is Simplify the movement of both arms into a pendulum motion moving in the forward and backward directions. At time t, the angle between the right arm of the human target and the negative direction of the Z axis is θ S (t), point A is one of the scattering points on the right arm, And 0≤x A ≤L armR , L armR is the length of the right arm; θ S The variation law of (t) is θ s (t) = θ smax sin(ωt+ψ0), where θ smax is the angular amplitude of the arm when it swings, ω is the swing speed, ψ0 is the initial phase of the arm swing, the initial phase of the right arm is 0, and the initial phase of the left arm is π; the instantaneous distance vector between point A and the radar at time t is: in, is the initial rotation matrix, is the rotation matrix at time t, yes The unit vector of β1 is The elevation angle relative to the origin Q, yes The unit vector of is: where α A (t) and β A (t) are Relative to the starting point Q S The azimuth and elevation angles are: Combined with the instantaneous distance vector between point A and the radar at time t, let ω ref =0, α0=0, k is the period extension variable, and the instantaneous radial distance between point A and the radar is calculated as: Scattering point A on the left arm L The instantaneous radial distance from the radar is: in, 0≤x AL ≤L armL , L armL is the length of the left arm, θ sL (t) = θ s (t+T / 2)=θ smax sin(ωt+π)=-θ s (t); The radar echo of the scattering point is obtained according to the instantaneous radial distance of the scattering point relative to the radar: s t (t)=u1 exp(j2πf0t) Where u1 is the amplitude of the radar transmission signal. By mixing and low-pass filtering the radar echo, a zero intermediate frequency echo signal is obtained, which is expressed as: s t (t)=ρ(x,y,z)exp{j2πf0(-2R(t) / c}=ρ(x,y,z)exp{jΦ(t)} where ρ(x,y,z) is the scattering coefficient of the scattering point, t0≤t≤t s , t0 is the echo response time of the target at the radar R0 at time 0. When the target's moving speed is much smaller than the speed of light, t0≈2R0 / c, t s is the radar signal sampling time; Φ(t) = 2πf0 (-2R(t) / c) is the phase of the zero intermediate frequency echo signal; the linear structure target radar echo of the upper limb movement is expressed as: Where L is the arm length; Substituting the arm instantaneous radial distance expression into the zero intermediate frequency echo signal expression, we can get A and A L The radar echo is: When the upper limbs of the human body move, the radar echoes of the right and left arms are:

4. The method for intelligent human posture recognition based on multimodal micro-motion features and Transformer network according to claim 1 is characterized in that: Step S2 further comprises: The millimeter-wave radar sends a linear frequency modulation signal on a single channel. The four receiving channels mix the echo reflected by the target with the transmitted signal to obtain an intermediate frequency signal. The real and imaginary parts of the intermediate frequency signal are recombined into a complex form to generate echo data containing complex signals. The data is then rearranged into a two-dimensional array. The rows represent the number of sampling points of the receiving channel, and the columns represent the total number of pulses. Each row corresponds to the data of a receiving channel, and each column corresponds to a time series. The merged echo data is re-adjusted into a matrix format suitable for subsequent processing. According to the number of the four receiving antennas, the data is separated according to the receiving channels and stored separately. Then the echo data of all the receiving antenna channels are added point by point to obtain the combined radar echo signal. By calculating the average value of the pulse sequence in each frame and subtracting the average value from the signal, the DC bias or low-frequency component in the signal is eliminated to obtain the echo signal matrix with the DC component removed.

5. The method for intelligent human posture recognition based on multimodal micro-motion features and Transformer network according to claim 1 is characterized in that: Step S3 further comprises: The radar echo signal after data preprocessing is subjected to image and sequence extraction using two-dimensional fast Fourier transform. Specifically, the radar echo matrix is ​​subjected to FFT in the distance dimension to obtain the distance-time relationship, and then the distance-time relationship is subjected to FFT in the Doppler dimension frame by frame to obtain the distance-Doppler relationship. The 2D FFT data of each frame is accumulated in the Doppler dimension to obtain the micro-Doppler at one moment. The micro-Doppler data of each frame is spliced ​​in time to obtain a two-dimensional Doppler time-frequency diagram. The maximum value of each frequency point in the two-dimensional Doppler time-frequency diagram is extracted in the time dimension, and they are spliced ​​one by one to obtain a one-dimensional Doppler spectrum sequence.

6. The method for intelligent human posture recognition based on multimodal micro-motion features and Transformer network according to claim 1 is characterized in that: In step S4, the one-dimensional Transformer network includes an embedding layer, a first position encoding layer, a Transformer encoder block, an average pooling layer and a first feature output layer; The embedding layer is used to map the dimension of the input one-dimensional 255-frequency Doppler spectrum sequence to an embedding dimension, which is defined as 256, and the output matrix shape is (255, 256). At the same time, the weights are initialized using the Xavier function to make the gradient distribution uniform; The first position encoding layer is used to add position information to the output data of the embedding layer so that the model can perceive the relative order of each moment and spatial point in the data, and perform a transposition operation to obtain an output matrix shape of (256, 255) to adapt it to the input requirements of the Transformer encoder layer; The Transformer encoder block is used to extract sequence features from input data, including a multi-head self-attention layer, a first residual connection and normalization layer, a feedforward neural network, and a second residual connection and normalization layer; The multi-head self-attention layer is used to capture the global dependencies in the input sequence. Through multiple attention heads, each of which contains a different query matrix, key matrix and value matrix, the correlation between different parts is calculated in parallel, and different patterns and features are learned, so that the network can focus on various parts in the sequence from multiple perspectives. The residual connection and normalization layer are used to solve the gradient disappearance and gradient explosion problems and stabilize the data distribution. The Transformer structure uses residual connections to add the input of each layer directly to the output to learn deeper features; normalization is performed after each layer; The feedforward neural network layer includes two linear transformations and activation functions, which are used to further transform and nonlinearly map the features extracted by the self-attention layer. Through the first linear transformation and nonlinear activation function, the input features are mapped to a higher-dimensional space, and through the second linear transformation, the high-dimensional features are mapped back to the dimensions of the original input. The average pooling layer averages the sequence dimensions to generate global features of fixed dimensions, which are output via the feature output layer.

7. The method for intelligent human posture recognition based on multimodal micro-motion features and Transformer network according to claim 1 is characterized in that: In step S4, the two-dimensional Transformer network includes an image input layer, a packet embedding module, a class labeling module, a connection layer, a second position encoding layer, a random inactivation layer, a Transformer encoder layer, a first normalization layer, a class labeling segmentation layer and a second feature output layer; The packet embedding module includes a flattening layer and a two-dimensional convolutional layer connected to each other, wherein the flattening layer divides the two-dimensional Doppler time-frequency image received by the image input layer into multiple patches, and flattens each patch into a vector with a length of 196; the two-dimensional convolutional layer maps the vector output by the flattening layer to a high-dimensional space, and outputs an embedding vector with a matrix shape of (196,768) to extract local features; The class tag module is used to add a class token category tag to the front end of the embedding vector. Before the embedding vector enters the Transformer encoder layer, an additional learnable class tag token is added to the embedding vector. The token is used for global feature aggregation and learning global information in the final stage. The connection layer is used to concatenate the embedding vector and the class label to obtain a matrix with a shape of (197,768); The second position encoding layer is used to add position information, the position encoding and the image embedding vector are added element by element, and the matrix shape remains unchanged; The random inactivation layer is added with a Dropout layer; The Transformer encoder layer includes Transformer encoder blocks stacked 10 times in a loop, each Transformer encoder block including a multi-head self-attention layer, a first residual connection and normalization layer, a feedforward neural network, and a second residual connection and normalization layer; The first normalization layer is used for output processing of the Transformer encoder layer; The class tag segmentation layer is used to extract the class token containing the global image features; The second feature output layer outputs the class token of the class label segmentation layer.

8. The method for intelligent human posture recognition based on multimodal micro-motion features and Transformer network according to claim 1 is characterized in that: The multimodal feature fusion model includes a linear layer, a second normalization layer, a feature fusion layer and a fully connected layer connected in sequence; The linear layer uniformly converts the extracted feature dimensions into multimodal dimensions, which is defined as 512; the second normalization layer normalizes the features of different modalities so that the features are within a similar numerical range; the feature fusion layer concatenates the image and spectrum features after unifying the dimensions and normalizing them to obtain an output matrix shape of (1, 1024); the fully connected layer maps the concatenated feature dimensions to the category dimension, which is defined as 9 categories, and outputs the category label.

Citation Information

Patent Citations

  • Gesture recognition detection method based on time-frequency spectrum and distance Doppler spectrum

    CN110309690A

  • Human body identity recognition method based on dual-channel convolutional neural network

    CN112433207A

  • Indoor human body posture recognition method based on millimeter wave radar

    CN116561700A

  • Molding apparatus having clamp

    KR1020210148679A

Cited By

  • Tracking method and system applied to rehabilitation process of orthopedic patient

    CN121506508A

  • Radar micro-motion feature extraction method based on deep learning

    CN121541170A