Multi-modal millimeter wave gait recognition method based on multi-scale attention

Through the combination of multi-scale attention mechanism and bidirectional long and short-term memory network, the problem of modal singleness and insufficient feature extraction in the existing technology is solved, and efficient multi-modal millimeter wave gait recognition is achieved, which improves the recognition accuracy and robustness.

CN120299085APending Publication Date: 2025-07-11SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510424927.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing millimeter-wave radar gait recognition methods mainly rely on single mode data, making it difficult to effectively extract local motion details and relative spatial position relationships of human body, resulting in limited recognition accuracy, and multimodal methods do not fully utilize the complementarity and dynamic changes of different mode features.

Method used

The multi-scale attention mechanism is adopted, and multi-scale feature extraction is performed through point cloud branches and micro-Doppler branches, and the attention weighting mechanism is used to adaptively fuse feature information of different scales, and timing modeling is performed in combination with a bidirectional long and short-term memory network, and finally gait recognition is achieved through a classifier.

Benefits of technology

It improves the accuracy and generalization ability of millimeter-wave radar gait recognition, can adaptively fuse feature information of different scales and modes, and enhances recognition ability and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299085A_ABST
    Figure CN120299085A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal millimeter wave gait recognition method based on multi-scale attention, and the method comprises the steps: obtaining three-dimensional point cloud data and micro-Doppler spectrum data of a pedestrian at the same time through a millimeter wave radar, and carrying out the preprocessing operation; performing multi-scale feature extraction on the preprocessed three-dimensional point cloud data by using point cloud branches, and performing micro-Doppler multi-scale feature extraction on the micro-Doppler spectrum data by using micro-Doppler branches; features output by the point cloud branches and the micro-Doppler branches at different scales are combined through linear projection and attention weighting, and multi-frame gait sequence features are obtained; and inputting the multi-frame gait sequence features into a bidirectional long-short-term memory network to perform modeling of a time dimension, paying attention to key moment action features through a time attention mechanism, and outputting a pedestrian gait recognition result by adopting a classifier. The method has the advantages of strong privacy protection, high recognition precision and strong robustness, and is suitable for the fields of security monitoring, smart home, medical health and protection and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of millimeter-wave radar gait recognition, and particularly to a multi-modal millimeter-wave gait recognition method based on multi-scale attention. Background Art

[0002] In recent years, with the rapid development of deep learning technology, artificial intelligence has made significant breakthroughs in the field of biometric recognition, promoting the deep integration of intelligent perception technology with fields such as security and healthcare. As an important biometric recognition method, gait recognition technology based on millimeter-wave radar has received extensive attention in the industrial and academic communities. This technology uses the point cloud data and micro-Doppler spectrograms collected by the radar, extracts human motion features through a deep neural network, and realizes identity recognition and behavior analysis, playing an important role in scenarios such as smart homes and health monitoring.

[0003] Existing millimeter-wave radar gait recognition methods mainly rely on single-modal data, such as point cloud data or micro-Doppler spectrograms. Among them, point cloud-based recognition methods usually focus on capturing the three-dimensional morphological features of the target. However, since point clouds are essentially sparse data, conventional methods are difficult to effectively extract local motion details of the human body, such as the swinging of the arms and legs. At the same time, micro-Doppler spectrogram-based methods are good at reflecting the velocity information of the target, but have limitations in describing the relative spatial position relationships of various parts of the human body. This single-modal feature extraction method limits the expressive ability of the model, thus affecting the accuracy of gait recognition.

[0004] In recent years, multi-modal data fusion methods have gradually emerged, aiming to improve the performance of gait recognition by combining point clouds and micro-Doppler spectrograms. However, existing multi-modal methods mostly adopt simple feature splicing strategies and do not fully consider the complementarity and correlation of different modal features. In particular, the importance of different modal features varies dynamically at different scales, and direct splicing alone cannot fully express these complex relationships. In addition, traditional methods often lack a comprehensive modeling of time series features and cannot capture long-term dependencies and periodic change patterns in gait sequences, thus limiting their applicability in complex scenarios.

[0005] The current millimeter-wave radar gait recognition algorithms mainly have the following limitations. Existing feature extraction methods only focus on coarse-grained global features and do not fully utilize the local spatial structure in point cloud data and the fine-grained velocity features in micro-Doppler spectrograms, resulting in limited expressive ability of the model. Current point cloud-based methods are difficult to capture local motion details such as the arms and legs, and micro-Doppler spectrogram-based methods cannot accurately characterize the relative position relationships of various parts of the human body. Summary of the Invention

[0006] The object of the present invention is to overcome the disadvantages and deficiencies of the prior art, and provide a multi-modal millimeter-wave gait recognition method based on multi-scale attention to achieve accurate and highly private pedestrian gait recognition.

[0007] To achieve the above object, the technical solution provided by the present invention is: A multi-modal millimeter-wave gait recognition method based on multi-scale attention, comprising the following steps:

[0008] Step 1, data acquisition and preprocessing: The millimeter-wave radar simultaneously acquires two types of modal data, namely three-dimensional point cloud data and micro-Doppler spectrogram data of pedestrians, and performs preprocessing operations on the two types of modal data respectively to form a standardized input for subsequent feature extraction.

[0009] Step 2, multi-modal feature extraction: Perform multi-scale feature extraction on the preprocessed three-dimensional point cloud data using a point cloud branch, and at the same time perform micro-Doppler multi-scale feature extraction on the micro-Doppler spectrogram data using a micro-Doppler branch; wherein, the point cloud branch gradually extracts and aggregates local geometric features through four layers of Set Abstraction modules, and this Set Abstraction module is abbreviated as the SA module; the micro-Doppler branch is processed by a four-layer two-dimensional convolutional network, and a pooling operation is inserted between the convolutional layers to finally obtain a multi-scale micro-Doppler spectrogram feature representation.

[0010] Step 3, multi-scale attention multi-modal fusion: Combine the features output by the point cloud branch and the micro-Doppler branch at different scales through linear projection and attention weighting, adaptively allocate the weights of each modal feature in the fusion vector, and perform weighting or splicing on the fusion results at different scales to obtain the finally fused multi-frame gait sequence features.

[0011] Step 4, temporal modeling and classification: Input the fused multi-frame gait sequence features into a bidirectional long short-term memory network for temporal dimension modeling, and pay attention to the action features at critical moments through a temporal attention mechanism, and finally use a classifier to output the recognition result of the pedestrian gait.

[0012] Further, in Step 1, the three-dimensional point cloud data and the micro-Doppler spectrogram data collected by the millimeter-wave radar are synchronously acquired according to time frames \(t = 1, 2, \cdots, T\); perform denoising, downsampling, and coordinate normalization processing on the three-dimensional point cloud data to obtain frame-level point cloud where is the set of real numbers, \(N\) is the number of sampling points per frame, \(D\) is the feature dimension of each point, including three-dimensional coordinates \((x, y, z)\) and Doppler velocity \(v\), signal-to-noise ratio \(s\);

[0013] Perform amplitude normalization and logarithmic amplitude conversion on the micro-Doppler spectrogram data to obtain frame-level micro-Doppler spectrogram Where H and W represent the height and width of the spectrogram, respectively.

[0014] Furthermore, in step 1, the acquired 3D point cloud data and the micro-Doppler spectrogram data are collected using the same millimeter-wave radar and are synchronously aligned in the time dimension.

[0015] Furthermore, in step 2, the specific configuration of the four-layer SA module adopted by the point cloud branch is as follows:

[0016] The first-layer SA module: the number of sampled points n1 = 512, the neighborhood radius r1 = 0.2, the number of neighborhood points k1 = 32, and the MLP channel configuration is [D, 64, 64, 128];

[0017] The second-layer SA module: the number of sampled points n2 = 256, the neighborhood radius r2 = 0.4, the number of neighborhood points k2 = 32, and the MLP channel configuration is [128, 128, 128, 256];

[0018] The third-layer SA module: the number of sampled points n3 = 128, the neighborhood radius r3 = 0.8, the number of neighborhood points k3 = 32, and the MLP channel configuration is [256, 256, 256, 512];

[0019] The fourth-layer SA module: the number of sampled points n4 = 256, the neighborhood radius r4 = 1.6, the number of neighborhood points k4 = 64, and the MLP channel configuration is [512, 512, 512, 1024];

[0020] The above configurations respectively control the farthest point sampling number, the spherical query neighborhood radius, the upper limit of the number of points selected in the neighborhood, and the output dimension after the multi-layer perceptron mapping for each layer to perform point cloud processing;

[0021] The four-layer SA module is used to perform layer-by-layer feature abstraction on the frame-level point cloud PC(t), and the point cloud features output by each layer of the SA module are denoted as

[0022]

[0023] In the formula, represents the point cloud feature of the i-th layer of the SA module;

[0024] The i-th layer of the SA module performs the following operations: ① Select the number of sampled points n i through farthest point sampling FPS to ensure that the sampled points are evenly distributed in space; ② Perform a spherical neighborhood query centered on the sampled points, define the neighborhood radius r i , and select the number of neighborhood points k i; ③ Use a multi-layer perceptron (MLP) and max pooling (Max Pooling) to obtain local geometric features from neighboring points, fuse them with the output features of the previous layer to form the output of this layer, and finally obtain 4-layer point cloud features

[0025] Furthermore, in step 2, the micro-Doppler branch includes a four-layer two-dimensional convolutional network based on a residual structure, denoted as:

[0026] {Conv1, Conv2, Conv3, Conv4}

[0027] The specific parameters are as follows:

[0028] Conv1: The convolution kernel size is 7×7, the stride is 2, and the number of output channels is 64;

[0029] Conv2: The convolution kernel size is 5×5, the stride is 2, and the number of output channels is 128;

[0030] Conv3: The convolution kernel size is 3×3, the stride is 2, and the number of output channels is 256;

[0031] Conv4: The convolution kernel size is 3×3, the stride is 2, and the number of output channels is 512;

[0032] Let the output of the i-th layer convolution operation be:

[0033]

[0034] In the formula, represents the micro-Doppler spectrogram feature of the i-th layer convolution operation Conv i , represents the micro-Doppler spectrogram feature of the (i - 1)-th layer convolution operation Conv i . The network is followed by a ReLU activation function after a specific layer; in each residual block, the convolution path includes a two-dimensional convolution operation, followed by batch normalization and a non-linear activation ReLU function after each convolution; another residual connection directly connects the input and output, and is fused through an element-wise addition operation to finally obtain 4-layer micro-Doppler spectrogram features

[0035] Furthermore, in step 3, the steps of the multi-scale attention multi-modal fusion include:

[0036] Feature mapping: Map the point cloud feature and the micro-Doppler spectrogram feature to the same dimension respectively, and obtain the point cloud feature and the micro-Doppler spectrogram feature

[0037]

[0038] In the formula, and are learnable weight matrices, and are bias terms;

[0039] Attention weight calculation: For the i-th scale, the point cloud feature after feature mapping and the micro-Doppler spectrogram feature calculate the correlation weight α between the two through attention weight calculation i :

[0040]

[0041] In the formula, and are learnable weight matrices, is the bias term, and σ is the Sigmoid activation function;

[0042] Weighted fusion:

[0043] In the formula, represents the multi-scale fusion feature, that is, the fused multi-frame gait sequence feature;

[0044] The above operations are performed at four scales of i = 1, 2, 3, 4 respectively to obtain the multi-scale fusion feature set:

[0045] Furthermore, after the multi-scale attention multi-modal fusion, for the obtained multi-scale fusion feature set: Further process to obtain the global multi-modal feature F mm :

[0046] F mm = ReLU(W fusion F + b fusion )

[0047] In the formula, the intermediate variable Concat represents the concatenation operation in the feature dimension, W fusion is the learnable weight matrix, b fusion is the bias term, and ReLU is the non-linear activation function;

[0048] Finally, F mm is used as the final global multi-modal feature in the temporal modeling step. This feature integrates multi-scale information and enhances the discriminative ability of the feature.

[0049] Furthermore, in step 4, perform temporal modeling on the multi-frame gait sequence features obtained in step 3, specifically including:

[0050] Feature sequence construction: Organize the global multi-modal feature F obtained in step 3 mm into a sequence over T consecutive time frames where represents the input feature at time step t, and t = 1, 2,..., T is the number of time frames collected by the millimeter-wave radar;

[0051] Bidirectional long short-term memory network Bi-LSTM processing: Perform recurrent calculations on this sequence separately from the forward and backward directions to obtain the hidden state and finally concatenate them into a bidirectional hidden state In the bidirectional long short-term memory network Bi-LSTM, for the input feature at each time step t and the corresponding output bidirectional hidden state h t a residual connection is established between them to obtain a new bidirectional hidden state h t ':

[0052]

[0053] where W r is a linear transformation matrix used to adjust the feature dimension for addition operations;

[0054] Temporal attention calculation: To highlight the time steps or action segments that have a more significant impact on gait recognition, assign attention weights e t to the bidirectional hidden states h t ' at each moment:

[0055] e t = tanh(W e h t '+ b e )

[0056] where W h and b e are learnable weight matrices and bias vectors. Using the attention weights, perform a weighted sum on the hidden states to obtain the global temporal feature H seq :

[0057]

[0058] After the temporal attention calculation, add a skip connection to the global temporal feature H seq to obtain a new global temporal feature H seq ':

[0059]

[0060] where denotes the sequence Average value in the time dimension, W s is a learnable weight matrix;

[0061] Classification prediction: The classifier passes through the fully connected layer MLP and uses the Softmax function to classify and predict the global temporal feature H seq ' for classification prediction:

[0062]

[0063] where W cls and b cls are the weight matrix and bias vector of the classifier, is the predicted class probability distribution, representing the recognition result of the pedestrian gait.

[0064] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0065] 1. The present invention proposes a multi-scale attention mechanism that can adaptively fuse feature information of different scales and different modalities. By designing independent attention for each feature scale, it can make full use of the spatial structure information of the point cloud and the dynamic velocity information of the Doppler map, improving the gait feature expression ability of the millimeter-wave radar.

[0066] 2. The present invention adopts hierarchical feature extraction, and gradually extracts multi-scale features through the SA module of the point cloud branch and the two-dimensional convolutional network of the micro-Doppler branch. This hierarchical feature extraction method can capture local details and global semantic information simultaneously, enhancing the recognition ability.

[0067] 3. The multi-scale attention multi-modal fusion of the present invention adopts a fusion method guided by feature projection and attention, avoiding the interference between modalities that may be brought by simple feature splicing. By learning adaptive fusion weights, it can highlight important features and suppress redundant information, improving the effect of feature fusion.

[0068] 4. The method of the present invention has good generalization ability and robustness. Through the complementarity of multi-scale features and the self-adaptability of the attention mechanism, it can adapt to gait recognition tasks in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 is a schematic flow chart of the method of the present invention.

[0070] Figure 2 is a structural diagram of the micro-Doppler branch of the present invention.

[0071] Figure 3 is a schematic flow chart of the multi-scale attention multi-modal fusion of the present invention.

[0072] Figure 4This is the architecture diagram of the method of the present invention. Detailed implementation manners

[0073] The present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings, but the implementation manners of the present invention are not limited thereto.

[0074] As Figure 1 shown, this embodiment discloses a multi-modal millimeter-wave gait recognition method based on multi-scale attention. The specific implementation manner focuses on the data processing of two modalities of the three-dimensional point cloud data and the micro-Doppler spectrogram data of pedestrians. Multi-modal fusion is achieved through a multi-scale attention mechanism, and temporal modeling is performed using a bidirectional long short-term memory network (Bi-LSTM) combined with a temporal attention mechanism in the time dimension. Finally, gait recognition is achieved through classifier training. The whole method includes four main steps: Step 1 Data acquisition and preprocessing, Step 2 Multi-modal feature extraction (point cloud multi-scale feature extraction and micro-Doppler multi-scale feature extraction), Step 3 Multi-scale attention multi-modal fusion, and Step 4 Temporal modeling and classification.

[0075] As Figure 4 shown, it details the entire method architecture, and the implementation details of each core part will be introduced in turn below.

[0076] 1) Data acquisition and preprocessing

[0077] The three-dimensional point cloud data and the micro-Doppler spectrogram data are collected using the same millimeter-wave radar. The three-dimensional point cloud data and the micro-Doppler spectrogram data collected by the millimeter-wave radar are synchronously obtained according to the time frames t = 1, 2,..., T; the three-dimensional point cloud data is subjected to denoising, downsampling, and coordinate normalization processing to obtain frame-level point cloud where is the set of real numbers, N is the number of sampling points per frame, D is the feature dimension of each point, PC(t) contains multiple time frames, each frame is composed of a large number of points, and each point contains five feature components: including three-dimensional coordinates (x, y, z) and Doppler velocity v, signal-to-noise ratio s; the normalization method is to subtract the mean value μ x , and then divide by the standard deviation σ x :

[0078]

[0079] where, x norm represents the normalized point cloud feature x represents any one of the feature components in x, y, z, v, s.

[0080] The micro-Doppler spectrogram data is preprocessed. The micro-Doppler spectrogram data is subjected to amplitude normalization and logarithmic amplitude conversion to obtain frame-level micro-Doppler spectrogram Where H and W represent the height and width of the spectrogram respectively. The micro-Doppler spectrogram MD(t) is a time-frequency image obtained by performing a Short-Time Fourier Transform (STFT) on the millimeter-wave radar echo signal. The horizontal axis represents time, and the vertical axis represents the Doppler frequency. To meet the input requirements of the neural network, the image is first normalized to normalize the pixel values to the range [0, 1], improving the training effect. For the first normalization:

[0081]

[0082] Among them, MD(t) norm represents the normalized micro-Doppler spectrogram, MD(t) mean is the mean of the micro-Doppler spectrogram, and MD(t) std is the standard deviation of the micro-Doppler spectrogram.

[0083] 2) Point cloud multi-scale feature extraction

[0084] After completing the above point cloud preprocessing, the point cloud is hierarchically and scale-by-scale feature extracted through 4 Set Abstraction modules (abbreviated as SA modules) in the point cloud branch. Each SA module includes three stages: Sampling, Grouping, and Feature Extraction, which can effectively capture the local details and global structure of the point cloud at different spatial scales.

[0085] To select representative key points from the point cloud data, in the i-th SA module of this embodiment, the Farthest Point Sampling (FPS) algorithm is used to select several sampling points Centers i-1 from the point set PC output by the previous layer i . FPS iteratively selects the point with the maximum distance from the currently existing sampling points, ensuring that the sampling points are evenly distributed in space and preserving the overall structure of the point cloud.

[0086] For the i-th sampling point, the Ball Query method is used to search for the neighborhood points Neighborhoods i within the search radius r i . To obtain multi-scale features, {r1, r2, r3, r4} from small to large are set at different layers of the network, such that the previous layer focuses on details and the subsequent layer pays attention to a larger range of spatial structures. After neighborhood retrieval, these neighborhood points are organized together for subsequent feature aggregation operations.

[0087] For each point in the neighborhood, its coordinates are converted into local coordinates relative to the sampling point:

[0088] Δx = x neighbor -x center

[0089] And concatenate with original features such as velocity and signal-to-noise ratio to form the input of neighborhood points. Subsequently, use a multi-layer perceptron (MLP) with shared parameters to perform non-linear mapping on the features of neighborhood points, enhancing the representation ability of local geometry and motion information. Finally, adopt the max pooling operation to aggregate the point features within the neighborhood, generating the feature vector of the sampling point at this layer.

[0090] By stacking 4 SA modules, gradually expand the receptive field and reduce the number of sampling points layer by layer to achieve feature abstraction from fine-grained to coarse-grained. At some levels, the Feature Propagation (FP) module can also be combined. Through interpolation and skip connections, fuse the high-level abstract information with the low-level details to improve the resolution. To further enhance the ability to capture the spatio-temporal features of point clouds, in this embodiment, two-dimensional convolution (Conv2D) operations, batch normalization BN, and ReLU activation functions are added in several layers to learn more complex local patterns.

[0091] At the 4 SA layers of the network, respectively extract the point cloud feature representations i = 1, 2, 3, 4 corresponding to the point cloud features of different scales and different abstraction levels, forming a point cloud feature set These multi-scale point cloud features will interact and fuse with the corresponding scale features extracted by the micro-Doppler branch.

[0092] 3) Micro-Doppler multi-scale feature extraction

[0093] As Figure 2 shown, the frame-level spectrogram MD(t) reflects the change of target velocity information over time and contains rich motion pattern information. The present invention adopts an architecture composed of a four-layer two-dimensional convolutional network based on the residual structure to achieve feature extraction of the micro-Doppler spectrogram.

[0094] The first-layer convolution operation Conv1 processes the input micro-Doppler spectrogram MD(t). The convolution kernel size is 7×7, the stride is 2, and the number of output channels is 64. After the convolution operation, batch normalization and ReLU activation functions are connected to form features The first-layer convolution focuses on extracting the local texture and edge features of the spectrogram and capturing the basic patterns of velocity distribution.

[0095] The second-layer convolutional network Conv2 processes the feature map output by the first layer The convolution kernel size is 5×5, the stride is 2, and the number of output channels is 128, obtaining The second convolutional layer further increases the number of channels and reduces the size of the feature map, capturing more abstract texture combinations and medium-scale patterns. This layer introduces a residual connection structure that adds the input features to the convolutional output to form where W r is a linear transformation matrix used to adjust the feature dimensions for the addition operation.

[0096] The third convolutional network Conv3 processes with a convolutional kernel size of 3×3, a stride of 2, and an output channel number of 256, resulting in This layer further reduces the size of the feature map while significantly increasing the number of channels, capturing higher-level time-frequency features.

[0097] The fourth convolutional network Conv4 processes with a convolutional kernel size of 3×3, a stride of 2, and an output channel number of 512, resulting in After the fourth convolution, through global average pooling operation, the feature map is converted into a feature vector with a fixed dimension

[0098] Through the above four-layer convolutional structure, the micro-Doppler branch captures time-frequency features in the micro-Doppler spectrogram from different scales and abstraction levels, forming a feature set These multi-scale features cover multi-level information from local details to global patterns, providing a rich dynamic feature representation for subsequent multi-modal fusion.

[0099] In this embodiment, batch normalization operations are added between convolutional layers to reduce the internal covariate shift phenomenon, accelerate network convergence, and improve generalization ability. The introduction of residual connections enables gradients to directly propagate from deep layers to shallow layers, solving the vanishing gradient problem in deep network training.

[0100] The stride settings between convolutional layers cause the size of the feature map to decrease layer by layer while the number of channels increases layer by layer, forming a pyramid-shaped feature extraction structure. This structural design enables the network to balance computational efficiency and expressive power, making it suitable for processing micro-Doppler spectrogram data containing multi-scale time-frequency information.

[0101] 4) Multi-scale attention multi-modal fusion

[0102] As Figure 3 shown, after obtaining the multi-scale point cloud features of the point cloud branch and the multi-scale micro-Doppler spectrogram features of the micro-Doppler branch this embodiment realizes the effective combination of corresponding scale features through a multi-scale attention fusion module. This module first aligns the dimensions of different modal features, and then adaptively assigns feature weights through an attention mechanism, finally generating a fused representation.

[0103] First, perform a linear mapping on the point cloud features and the micro-Doppler spectrogram features to ensure they have the same feature dimension. Assume that at the i-th scale, after simple feature mapping, we obtain and The mapping operation can be expressed as:

[0104]

[0105] where and are learnable weight matrices, and are bias terms.

[0106] Calculate the attention weight α in the following way i :

[0107]

[0108] where and are learnable weight matrices, is a bias term, and σ is the Sigmoid activation function, where σ maps the output to [0, 1]. Ensure that the sum of the weights of the two modalities at the same scale is 1.

[0109] Use the normalized weights to perform a weighted sum of the point cloud and micro-Doppler spectrogram features to obtain the fused feature at this scale:

[0110]

[0111] To further integrate the fused features of 4 different scales, define the multi-scale fusion result set:

[0112]

[0113] Since different scales may have different importance in describing different levels of target features, for the obtained multi-scale fusion feature set: Perform further processing to obtain the global multi-modal feature F mm :

[0114] F mm = ReLU(W fusion F + b fusion )

[0115] where the intermediate variable Concat represents the concatenation operation in the feature dimension, W fusion is a learnable weight matrix, b fusionis the bias term, and ReLU is the non-linear activation function;

[0116] Through the above multi-scale attention fusion, F is obtained mm As the final global multi-modal feature, the above steps can adaptively learn the importance of different-level features, give full play to the complementarity of the point cloud spatial feature and the micro-Doppler time-frequency feature, and provide rich multi-modal representations for subsequent temporal modeling.

[0117] 5) Temporal Modeling and Classification

[0118] During the continuous acquisition process of the millimeter-wave radar, a series of 3D point cloud data and micro-Doppler spectrogram data at moments t = 1, 2,..., T are obtained. After the above multi-modal fusion, a fused feature vector can be output at each time step t Arrange them in chronological order to form a gait feature sequence:

[0119]

[0120] Pedestrian gait has temporal characteristics such as periodicity and long-distance dependence. To capture the forward and backward context information of the sequence simultaneously, this embodiment uses a bidirectional long short-term memory network, namely bidirectional LSTM (Bi-LSTM), to perform modeling. For each time step t, the forward LSTM calculates the hidden state from t = 1 to t = T in the forward direction, and the backward LSTM calculates from t = T to t = 1 in the reverse direction. After concatenating the two, the bidirectional hidden state is obtained:

[0121]

[0122] In the bidirectional long short-term memory network Bi-LSTM, a residual connection is established between the input feature at each time step t and the corresponding output bidirectional hidden state h t to obtain a new bidirectional hidden state h t ':

[0123]

[0124] In the formula, W r is the linear transformation matrix, which is used to adjust the feature dimension for the addition operation;

[0125] In a long time series, not all moments are equally important for gait discrimination. Some moments contain key actions or pose changes, while other moments may be more redundant. For this reason, this embodiment further introduces temporal attention based on the hidden state output by Bi-LSTM, and calculates the attention weight e t at each time step, which can be specifically written as:

[0126] e t = tanh(W e h t '+ b e )

[0127] Perform a weighted sum on the hidden state h t ' to obtain the global temporal feature:

[0128]

[0129] After the temporal attention calculation, add a skip connection to the global temporal feature H seq to obtain the new global temporal feature H seq ':

[0130]

[0131] In the formula, represents the average value of the sequence in the time dimension, and W s is a learnable weight matrix;

[0132] The introduction of the above residual connection and skip connection enhances the ability to capture long-sequence information on the one hand, improving the expression ability for complex gait patterns; on the other hand, it improves the stability of network training, effectively alleviating the vanishing gradient problem, especially having obvious advantages when dealing with longer gait sequences.

[0133] Finally, the classifier performs classification prediction on H seq ' through the fully connected layer MLP and using the Softmax function:

[0134]

[0135] In the formula, W cls and b cls are the weight matrix and bias vector of the classifier, is the predicted class probability distribution, representing the recognition result of the pedestrian gait.

[0136] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A multi-modal millimeter-wave gait recognition method based on multi-scale attention, characterized in that, It includes the following steps: Step 1, Data acquisition and preprocessing: The millimeter-wave radar simultaneously acquires two types of modal data, namely the three-dimensional point cloud data and the micro-Doppler spectrogram data of pedestrians, and performs preprocessing operations on the two types of modal data respectively to form a standardized input for subsequent feature extraction; Step 2, Multimodal feature extraction: Perform multi-scale feature extraction on the preprocessed three-dimensional point cloud data using a point cloud branch, and at the same time perform micro-Doppler multi-scale feature extraction on the micro-Doppler spectrogram data using a micro-Doppler branch; among them, the point cloud branch gradually extracts and aggregates local geometric features through four layers of Set Abstraction modules, and this Set Abstraction module is abbreviated as the SA module; the micro-Doppler branch is processed through a four-layer two-dimensional convolutional network, and pooling operations are inserted between the convolutional layers to finally obtain a multi-scale micro-Doppler spectrogram feature representation; Step 3, Multi-scale attention multimodal fusion: Combine the features output by the point cloud branch and the micro-Doppler branch at different scales through linear projection and attention weighting, adaptively allocate the weights of each modal feature in the fusion vector, and perform weighting or splicing on the fusion results at different scales to obtain the finally fused multi-frame gait sequence features; Step 4, Temporal modeling and classification: Input the fused multi-frame gait sequence features into a bidirectional long short-term memory network for temporal dimension modeling, and pay attention to the action features at critical moments through a temporal attention mechanism, and finally use a classifier to output the recognition result of the pedestrian gait.

2. The multi-modal millimeter-wave gait recognition method based on multi-scale attention according to claim 1, characterized in that, In step 1, the three-dimensional point cloud data and micro-Doppler spectrogram data collected by the millimeter-wave radar are synchronously acquired according to time frames \(t = 1, 2, \cdots, T\); the three-dimensional point cloud data is subjected to denoising, downsampling, and coordinate normalization processing to obtain frame-level point clouds where is the set of real numbers, \(N\) is the number of sampling points per frame, and \(D\) is the feature dimension of each point, including three-dimensional coordinates \((x, y, z)\), Doppler velocity \(v\), and signal-to-noise ratio \(s\); Perform amplitude normalization and log amplitude transformation on the micro-Doppler spectrogram data to obtain the frame-level micro-Doppler spectrogram where H and W represent the height and width of the spectrogram, respectively.

3. A multi-modal millimeter-wave gait recognition method based on multi-scale attention according to claim 2, characterized in that, In Step 1, the obtained three-dimensional point cloud data and micro-Doppler spectrogram data are collected using the same millimeter-wave radar and are synchronously aligned in the time dimension.

4. A multi-modal millimeter-wave gait recognition method based on multi-scale attention according to claim 3, characterized in that, In Step 2, the specific configurations of the four-layer SA modules adopted by the point cloud branch are as follows: The first-layer SA module: The number of sampling points n1 = 512, the neighborhood radius r1 = 0.2, the number of neighborhood points k1 = 32, and the MLP channel configuration is [D, 64, 64, 128]; The second-layer SA module: The number of sampling points n2 = 256, the neighborhood radius r2 = 0.4, the number of neighborhood points k2 = 32, and the MLP channel configuration is [128, 128, 128, 256]; The third-layer SA module: The number of sampling points n3 = 128, the neighborhood radius r3 = 0.8, the number of neighborhood points k3 = 32, and the MLP channel configuration is [256, 256, 256, 512]; The fourth-layer SA module: The number of sampling points n4 = 256, the neighborhood radius r4 = 1.6, the number of neighborhood points k4 = 64, and the MLP channel configuration is [512, 512, 512, 1024]; The above configurations respectively control the farthest point sampling number, the ball query neighborhood radius, the upper limit of the number of points selected in the neighborhood, and the output dimension after the multi-layer perceptron mapping for each layer to perform on the point cloud; The four-layer SA module is used to perform layer-by-layer feature abstraction on the frame-level point cloud PC(t), and the point cloud features output by each layer of the SA module are denoted as i = 1, 2, 3, 4; In the formula, represents the point cloud feature of the i-th layer SA module; The i-th layer SA module performs the following operations: ① Select the number of sampling points n through farthest point sampling (FPS) i , to ensure that the sampling points are evenly distributed in space; ② Perform spherical neighborhood query with the sampling points as the center, and define the neighborhood radius r i , and select the number of neighborhood points k within the neighborhood of each center point i ; ③ Use a multi-layer perceptron (MLP) and max pooling to obtain local geometric features from the neighborhood points, and fuse them with the output features of the previous layer to form the output of this layer, and finally obtain 4-layer point cloud features 5. A multi-modal millimeter-wave gait recognition method based on multi-scale attention according to claim 4, characterized in that, In Step 2, the micro-Doppler branch includes a four-layer two-dimensional convolutional network based on a residual structure, denoted as: [Conv1, Conv2, Conv3, Conv4} The specific parameters are as follows: Conv1: The convolution kernel size is 7×7, the stride is 2, and the number of output channels is 64; Conv2: The convolutional kernel size is 5×5, the stride is 2, and the number of output channels is 128; Conv3: The convolutional kernel size is 3×3, the stride is 2, and the number of output channels is 256; Conv4: The convolutional kernel size is 3×3, the stride is 2, and the number of output channels is 512; Let the output of the i-th layer convolutional operation be: In the formula, represents the micro-Doppler spectrogram feature of the i-th layer convolutional operation Conv i ; represents the micro-Doppler spectrogram feature of the (i-1)-th layer convolutional operation Conv i ; the network is followed by a ReLU activation function after a specific layer; in each residual block, the convolutional path includes two-dimensional convolutional operations, followed by batch normalization and the non-linear activation ReLU function after each convolution; additionally, a residual connection is directly connected to the input and output, and fused through an element-wise addition operation to finally obtain the micro-Doppler spectrogram features of 4 layers 6. A multi-modal millimeter-wave gait recognition method based on multi-scale attention according to claim 5, characterized in that, In step 3, the steps of the multi-scale attention multi-modal fusion include: Feature mapping: Map the point cloud features and the micro-Doppler spectrogram features to the same dimension respectively, and obtain the point cloud features and the micro-Doppler spectrogram features after feature mapping respectively: wherein, and are learnable weight matrices, and are bias terms; Attention weight calculation: For the i-th scale, the point cloud features after feature mapping and the micro-Doppler spectrogram features Calculate the correlation weight α between the two through attention weight i : wherein, and are learnable weight matrices, is a bias term, and σ is the Sigmoid activation function; Weighted fusion: In the formula, represents the multi-scale fusion feature, that is, the fused multi-frame gait sequence feature; The above operations are carried out at four scales of i = 1, 2, 3, and 4 respectively to obtain a multi-scale fusion feature set:

7. A multi-modal millimeter-wave gait recognition method based on multi-scale attention according to claim 6, characterized in that, After the multi-scale attention multi-modal fusion, for the obtained multi-scale fusion feature set: Further process to obtain the global multi-modal feature F mm : F mm = ReLU(W fusion F + b fusion ) In the formula, the intermediate variable Concat represents the concatenation operation in the feature dimension, and W fusion is a learnable weight matrix, b fusion is the bias term, and ReLU is a non-linear activation function; Finally, F mm As the final global multi-modal feature, it is used in the temporal modeling step. This feature integrates multi-scale information and enhances the discriminative ability of the feature.

8. A multi-modal millimeter-wave gait recognition method based on multi-scale attention according to claim 7, characterized in that In step 4, temporal modeling is performed on the multi-frame gait sequence features obtained in step 3, specifically including: Feature sequence construction: Organize the global multi-modal feature F obtained in step 3 mm into a sequence over consecutive T time frames where represents the input feature at time step t, where t = 1, 2, ..., T is the number of time frames collected by the millimeter-wave radar; Bidirectional Long Short-Term Memory Network Bi-LSTM Processing: Calculate the sequence cyclically from the forward and backward directions respectively to obtain the hidden state and finally splice it into a bidirectional hidden state In the bidirectional long short-term memory network Bi-LSTM, a residual connection is established between the input feature at each time step t and the corresponding output bidirectional hidden state h t to obtain a new bidirectional hidden state h t ': Where, W r is a linear transformation matrix for adjusting the feature dimension for addition operation; Temporal attention calculation: To highlight the time steps or action segments that have a more significant impact on gait recognition, the bidirectional hidden states h at each moment are used to assign attention weights e t ' t : t ' t : e t = tanh(W e h t '+ b e ) where, W h and b e are learnable weight matrices and bias vectors, and the hidden states are weighted and summed using the attention weights to obtain the global temporal feature H seq : After calculating the temporal attention, for the global temporal feature H seq Add skip connections to obtain a new global temporal feature H seq ′: wherein, represents the average value of the sequence in the time dimension, and W s is a learnable weight matrix; Classification prediction: The classifier classifies and predicts the global temporal feature H through the fully connected layer MLP and uses the Softmax function seq ' for classification prediction: where W cls and b cls are the weight matrix and bias vector of the classifier, is the predicted class probability distribution, representing the recognition result of the pedestrian gait.

Citation Information

Cited By

  • Twin model construction method, system and equipment for storage tank decommissioning and medium

    CN121093719A