Real-time dynamic trajectory tracking method and system for millimeter-wave radar gesture recognition

By combining frequency-modulated continuous wave millimeter-wave radar signal processing and a multimodal motion model library with a Poisson-Dobernoli hybrid filtering framework, the problem of insufficient real-time performance and anti-interference capability of existing millimeter-wave radar gesture recognition methods under complex gestures is solved, achieving highly robust gesture trajectory tracking and meeting the low latency requirements of human-computer interaction scenarios.

CN120802203BActive Publication Date: 2026-03-06SHENZHEN YUNENG WIRELESS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2026-03-06

AI Technical Summary

Technical Problem

Existing millimeter-wave radar gesture recognition methods still have room for improvement in terms of real-time performance and anti-interference capabilities for complex gestures, and it is difficult to achieve highly robust gesture trajectory tracking under low latency conditions.

Method used

High-precision dynamic trajectory tracking of gesture point clouds is achieved by employing frequency-modulated continuous wave millimeter-wave radar signal processing, constant false alarm rate detection, multimodal motion model library, Poisson-Dobernoli hybrid filtering framework and multi-hypothesis tracking strategy, combined with Kalman filtering smoothing processing.

Benefits of technology

High-precision, low-latency gesture motion tracking was achieved in complex scenarios, significantly improving the system's anti-interference capability and recognition accuracy, and meeting the low-latency requirements of human-computer interaction scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120802203B_ABST
    Figure CN120802203B_ABST
Patent Text Reader

Abstract

This invention discloses a real-time dynamic trajectory tracking method and system for millimeter-wave radar gesture recognition, belonging to the field of gesture recognition and tracking technology. The method includes: receiving echo signals reflected from gestures, extracting potential target point clouds, and clustering to generate gesture point cloud sequences; establishing a multimodal motion model library, dynamically selecting the optimal motion model using graph matching, and generating predicted states based on the gesture point cloud sequences; dynamically adjusting observation weights and optimizing the observation point clouds based on the signal-to-noise ratio and spatial distribution of the current gesture point cloud sequences using a Poisson-Doubouli hybrid filtering framework, and performing optimal association using Mahalanobis distance combined with dynamic time warping; maintaining trajectory hypotheses using a multi-hypothesis tracking strategy, selecting the optimal trajectory through a trajectory scoring mechanism, and smoothing the optimal trajectory using Kalman filtering. This invention achieves high-precision, low-latency tracking of gesture movements, adapting to different users' gesture habits and complex environmental interference while maintaining high trajectory accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gesture recognition technology, and more specifically, to a real-time dynamic trajectory tracking method and system for millimeter-wave radar gesture recognition. Background Technology

[0002] With the rapid development of Human-Computer Interaction (HCI) technology, gesture recognition, as a natural and intuitive interaction method, has been widely applied in smart homes, in-vehicle systems, virtual reality (VR), and augmented reality (AR). Traditional camera-based gesture recognition technology is affected by lighting conditions, occlusion issues, and privacy restrictions, making it difficult to operate stably in all scenarios. Millimeter-wave radar, due to its high resolution, strong penetration, immunity to ambient light interference, and good privacy protection characteristics, has become a research hotspot in the field of gesture recognition. Millimeter-wave radar operates in the 30 GHz to 300 GHz frequency band and can extract the distance, speed, and angle information of a target by transmitting frequency-modulated continuous wave (FMCW) or pulse signals and receiving the echo reflected from the target. Compared to optical sensors, millimeter-wave radar can more accurately capture minute movements and achieve real-time dynamic detection of gestures. However, millimeter-wave radar gesture recognition still faces some technical challenges, such as dynamic target interference caused by the Doppler effect, trajectory modeling of complex gestures, and the high computational complexity of real-time processing.

[0003] Existing millimeter-wave radar gesture recognition methods are mainly divided into two categories: static feature classification and dynamic trajectory tracking. Static feature classification methods extract micro-Doppler features or point cloud distributions of gestures for pattern matching, but their generalization ability is limited when processing continuous dynamic gestures. Dynamic trajectory tracking methods, on the other hand, capture the motion path of gestures in real time and combine Kalman filtering, particle filtering, or deep learning models for trajectory prediction and optimization, thereby achieving high-precision continuous gesture recognition. However, existing methods still have room for improvement in terms of real-time performance and anti-interference capabilities for complex gestures. Therefore, how to combine efficient signal processing algorithms and machine learning models to achieve highly robust gesture trajectory tracking under low-latency conditions is an urgent problem to be solved. Summary of the Invention

[0004] To address the aforementioned technical issues, this invention proposes a real-time dynamic trajectory tracking method and system for millimeter-wave radar gesture recognition. This method achieves highly robust gesture trajectory tracking under low latency conditions, adapts to different users' gesture habits and complex environmental interference, and maintains high trajectory accuracy and recognition accuracy.

[0005] The first aspect of this invention provides a real-time dynamic trajectory tracking method for millimeter-wave radar gesture recognition, comprising the following steps:

[0006] The system transmits millimeter-wave signals using a frequency-modulated continuous wave millimeter-wave radar and receives echo signals reflected from hand gestures. The echo signals are preprocessed, noise is filtered using constant false alarm rate detection, potential target point clouds are extracted, and the potential target point clouds are clustered to generate a hand gesture point cloud sequence.

[0007] A multimodal motion model library including uniform motion model, acceleration model and arc motion model is established. The optimal motion model is dynamically selected by graph matching. The next state of the gesture is predicted by combining the gesture point cloud sequence and the predicted state is generated.

[0008] Based on the Poisson-Bernoulli hybrid filtering framework, the observation weights are dynamically adjusted according to the signal-to-noise ratio and spatial distribution of the current gesture point cloud sequence to optimize the observation point cloud. The Mahalanobis distance is used to calculate the matching degree between the observation point cloud and the predicted state, and dynamic time warping is combined to achieve optimal association.

[0009] A multi-hypothesis tracking strategy is adopted to maintain trajectory hypotheses, and the optimal trajectory is selected through a trajectory scoring mechanism. The optimal trajectory is then smoothed by Kalman filtering to output the dynamic trajectory of the gesture.

[0010] In this scheme, the echo signal is preprocessed, and constant false alarm rate detection is used to filter noise and extract potential target point clouds, specifically as follows:

[0011] The received echo signal is mixed with the current transmitted signal to obtain an intermediate frequency signal. The intermediate frequency signal is then sampled by analog-to-digital conversion to obtain a discrete time series. The discrete time series is organized based on the radar frame structure of the frequency-modulated continuous wave millimeter-wave radar to generate a range dimension data matrix and a velocity dimension data matrix.

[0012] The distance dimension data matrix is ​​processed using Fast Fourier Transform to convert the time domain signal into the frequency domain. The distance to the gesture target is obtained based on the frequency corresponding to the peak value in the frequency domain. The velocity dimension data matrix is ​​then subjected to a second-dimensional Fast Fourier Transform to extract the Doppler frequency shift of the gesture target and obtain the velocity of the gesture target.

[0013] A distance-velocity matrix is ​​formed based on the distance and velocity of the gesture target. In the distance-velocity matrix, an ordered statistical constant false alarm rate detection method is used to calculate the local noise level through a sliding window to distinguish between real gesture targets and noise.

[0014] Set a dynamic threshold, retain units larger than the dynamic threshold, suppress warnings, output binarized detection results, mark potential target point clouds, map the potential target point clouds to three-dimensional space, and normalize data of different dimensions.

[0015] In this scheme, the potential target point cloud is clustered to generate a gesture point cloud sequence, specifically as follows:

[0016] The neighborhood radius is defined based on the physical size of the gesture and the preset radar resolution. All potential target point clouds are traversed. When the number of points within the neighborhood radius of a potential target point cloud is greater than the preset minimum number of point clouds, the point cloud is marked as the core point cloud, and the core point cloud is used to represent the main part of the gesture.

[0017] Starting from the core point cloud, recursively merge all reachable potential target point clouds in its neighborhood to form independent clusters, and assign boundary points to the nearest independent cluster. Isolated point clouds that are not assigned to any independent cluster are marked as noise and removed.

[0018] Motion consistency is checked on the independent clusters generated by clustering. The centroid coordinates and bounding box size of each independent cluster are calculated, and the velocity distribution of the point cloud within the cluster is extracted. Cluster features are generated to obtain gesture labels, and gesture point cloud sequences with gesture labels are output.

[0019] This scheme establishes a multimodal motion model library including uniform motion, acceleration, and arc motion models, and uses graph matching to dynamically select the optimal motion model, specifically:

[0020] Construct a uniform motion model suitable for steady linear motion, an acceleration model suitable for variable speed motion, and an arc motion model suitable for rotational or arc trajectory; and establish a multimodal motion model library based on the uniform motion model, acceleration model, and arc motion model.

[0021] A time series graph structure is constructed using a sliding window approach. Each node contains a motion feature vector, a motion model prediction state, and an observation matching degree, wherein the observation matching degree is the Mahalanobis distance score between the motion model prediction trajectory and the measured point cloud.

[0022] The time sequence graph structure includes time sequence edges, model competition edges, and cross-frame similar edges. Time sequence edges are constructed by forcibly connecting adjacent frame nodes and edge weights are calculated based on motion continuity. Model competition edges are constructed by connecting different model nodes in the same frame and edge weights are calculated based on KL divergence between models. Cross-frame similar edges are constructed by connecting nodes in non-continuous frames but with similar motion patterns and edge weights are calculated based on DTW distance of trajectory morphology.

[0023] After obtaining new frame data, a new node containing three model prediction states is generated, and the graph structure is updated. A weighted bipartite graph is constructed based on the updated graph structure. The left node represents the candidate motion model of a specific frame, and the right node corresponds to the observation data of a specific frame. The edge connections and edge weights of the weighted bipartite graph are generated according to the three edge structures and edge weights in the time series graph structure.

[0024] The Hungarian algorithm is improved by introducing motion smoothness constraints as a penalty term. The maximum weight matching is used to select the optimal motion model. Based on the optimal motion model, the next state of the gesture is predicted according to the gesture point cloud sequence, and the predicted state is generated.

[0025] In this scheme, the probability distribution is managed based on a Poisson-Bernoulli hybrid filtering framework. The observation weights are dynamically adjusted and the observation point cloud is optimized based on the signal-to-noise ratio and spatial distribution of the current gesture point cloud sequence. Specifically:

[0026] The newly generated hand gesture target is described by the Poisson point process. The spatial probability density of the newly generated hand gesture target is established within the radar field of view. The existence probability and state distribution of the known hand gesture target are characterized by the multi-Bernoulli mixture model. Each known hand gesture target maintains a Bernoulli tuple, which includes the existence probability and the state distribution described by the Gaussian mixture model.

[0027] The signal-to-noise ratio (SNR) of the observed point cloud in the gesture point cloud sequence is calculated, and the normalized entropy value of the observed point cloud is calculated to represent the spatial distribution. The observation weights are dynamically optimized by combining the SNR and spatial distribution, and hierarchical data association is performed through the optimized observation weights.

[0028] In the association of new gesture targets, when the observation weight corresponding to the unmatched observation point cloud is greater than the preset weight threshold and the distance to the known gesture target is greater than the preset distance threshold, a new Bernoulli tuple is generated.

[0029] In the known gesture target association, a noise matrix with fused observation weights is constructed based on the predicted state and the observed point cloud combined with the observation noise. The matching degree between the predicted state and the observed point cloud is evaluated by weighted Mahalanobis distance. The n observed point clouds with the smallest Mahalanobis distance to the predicted position are selected as candidate observations. The association probability of the candidate observations is calculated, and the existence probability and state distribution are updated.

[0030] The optimized observation point cloud is obtained by reconstructing the point cloud based on Bernoulli tuples.

[0031] In this scheme, Mahalanobis distance is used to calculate the matching degree between the observed point cloud and the predicted state, and dynamic time warping is combined to achieve optimal association, specifically:

[0032] For the optimized observation point cloud, Mahalanobis distance is used to calculate the single-point matching degree with the predicted state. For trajectory segments of a preset length, the matching degree matrix is ​​calculated based on the single-point matching degree.

[0033] Optimal association is achieved using dynamic time warping. A warped path is defined based on the observed point cloud and the predicted state. A cost matrix is ​​constructed by combining the matching degree matrix and kinematic constraints. Dynamic programming is used to solve for the minimum cumulative cost. A preset number of optimal association pairs are obtained by backtracking, and a preset number of hypothetical trajectories are generated.

[0034] In this scheme, a multi-hypothesis tracking strategy is used to maintain trajectory hypotheses, and the optimal trajectory is selected through a trajectory scoring mechanism. The optimal trajectory is then smoothed using Kalman filtering to output the dynamic gesture trajectory, specifically:

[0035] A dynamic trajectory hypothesis library is constructed and maintained through a multi-hypothesis tracking strategy. Each hypothesis node contains trajectory status, existence probability and historical observation matching sequence. A multi-dimensional scoring function is constructed using matching score, smoothing score and continuous score. The multi-dimensional scoring function is used to score the hypothetical trajectory.

[0036] The hypothetical trajectory with the highest score is selected as the optimal trajectory. Kalman smoothing is applied to the optimal trajectory, and the process noise is adjusted according to the dynamic characteristics of the gesture target to output the dynamic trajectory of the gesture.

[0037] The second aspect of this invention provides a real-time dynamic trajectory tracking system for millimeter-wave radar gesture recognition, the system comprising a signal acquisition and preprocessing module, a point cloud optimization and clustering module, a motion model matching module, a multimodal tracking module, and a trajectory smoothing module;

[0038] The signal acquisition and preprocessing module transmits millimeter-wave signals through a frequency-modulated continuous wave millimeter-wave radar and receives echo signals reflected from hand gestures, and preprocesses the echo signals.

[0039] The point cloud optimization and clustering module uses constant false alarm rate detection to filter noise, extracts potential target point clouds, clusters the potential target point clouds, and generates a gesture point cloud sequence.

[0040] The motion model matching module establishes a multimodal motion model library including uniform motion model, acceleration model and arc motion model, uses graph matching to dynamically select the optimal motion model, and combines the gesture point cloud sequence to predict the next moment state of the gesture to generate the predicted state.

[0041] The multimodal tracking module is based on a Poisson-Bernoulli hybrid filtering framework. It dynamically adjusts the observation weights and optimizes the observation point cloud according to the signal-to-noise ratio and spatial distribution of the current gesture point cloud sequence. It uses Mahalanobis distance to calculate the matching degree between the observation point cloud and the predicted state, and combines dynamic time warping to perform optimal association. It adopts a multi-hypothesis tracking strategy to maintain trajectory hypotheses and selects the optimal trajectory through a trajectory scoring mechanism.

[0042] The trajectory smoothing module performs Kalman filtering on the optimal trajectory to output a dynamic gesture trajectory.

[0043] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0044] The real-time dynamic trajectory tracking method of this invention achieves high-precision, low-latency tracking of gesture movements through multi-module collaborative processing, demonstrating significant technical advantages in complex scenarios.

[0045] During the signal acquisition and preprocessing stage, the system acquires gesture reflection signals using high-resolution millimeter-wave radar. Combined with adaptive filtering and detection algorithms, it effectively extracts potential target point clouds, ensuring the reliability and stability of the original data. The point cloud optimization module employs dynamic clustering and motion consistency verification to significantly suppress environmental noise and static interference, improving the detection rate of valid gesture targets while avoiding false tracking caused by clutter.

[0046] The multimodal tracking module, based on the PMBM filtering framework and combined with a multi-hypothesis tracking strategy, can robustly manage the probability distribution of emerging and known targets. By dynamically adjusting observation weights and optimizing data association, the system can maintain stable trajectory continuity even under complex conditions such as target occlusion, intersection, or brief disappearance, significantly reducing the frequency of identity switching. The motion prediction and smoothing module utilizes multi-model Kalman filtering to adaptively select the optimal motion mode and eliminates trajectory jitter through a forward-backward smoothing algorithm, outputting a smooth gesture motion path that conforms to physical laws.

[0047] The gesture recognition module integrates spatiotemporal and geometric features, and uses a lightweight neural network for real-time classification, mapping trajectories into semantically clear gesture commands. The system as a whole adopts a pipelined architecture and resource optimization strategy, achieving efficient real-time processing capabilities on an embedded platform, meeting the requirements of low latency and high response speed in human-computer interaction scenarios. Attached Figure Description

[0048] To more clearly illustrate the technical solutions in the embodiments or examples of the present invention, the drawings used in the embodiments or examples will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained according to these drawings without creative effort.

[0049] Figure 1 A flowchart illustrating a real-time dynamic trajectory tracking method for millimeter-wave radar gesture recognition is shown.

[0050] Figure 2 The flowchart illustrates the dynamic selection of the optimal motion model using graph matching.

[0051] Figure 3 The flowchart of optimizing the observed point cloud based on the Poisson-Dobernoli hybrid filter is shown;

[0052] Figure 4 A block diagram of a real-time dynamic trajectory tracking system for millimeter-wave radar gesture recognition is shown. Detailed Implementation

[0053] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.

[0054] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.

[0055] Figure 1 A flowchart of a real-time dynamic trajectory tracking method for millimeter-wave radar gesture recognition is shown.

[0056] like Figure 1 As shown, this embodiment provides a real-time dynamic trajectory tracking method for millimeter-wave radar gesture recognition, including:

[0057] S102, a millimeter-wave signal is transmitted by a frequency-modulated continuous wave millimeter-wave radar and the echo signal reflected by the gesture is received. The echo signal is preprocessed, noise is filtered by constant false alarm rate detection, potential target point cloud is extracted, the potential target point cloud is clustered, and a gesture point cloud sequence is generated.

[0058] S104, establish a multimodal motion model library including uniform motion model, acceleration model and arc motion model, use graph matching to dynamically select the optimal motion model, combine the gesture point cloud sequence to predict the next moment state of the gesture, and generate the predicted state.

[0059] S106, based on the Poisson-Bernoulli hybrid filtering framework, dynamically adjusts the observation weights and optimizes the observation point cloud according to the signal-to-noise ratio and spatial distribution of the current gesture point cloud sequence. It uses Mahalanobis distance to calculate the matching degree between the observation point cloud and the predicted state, and combines dynamic time warping to perform optimal association.

[0060] S108, a multi-hypothesis tracking strategy is adopted to maintain trajectory hypotheses, and the optimal trajectory is selected through a trajectory scoring mechanism. The optimal trajectory is then smoothed by Kalman filtering, and the dynamic gesture trajectory is output.

[0061] It should be noted that the radar transmitter generates a linear frequency modulated continuous wave (FMCW) signal, typically in the 60 GHz or 77 GHz band, which has high resolution characteristics; the electromagnetic waves reflected by the gesture target are captured by the radar receiving antenna and form an echo signal. Due to the Doppler frequency shift introduced by the gesture movement, the echo signal carries the target's distance, speed and angle information. The received echo signal is mixed with the current transmitted signal to obtain an intermediate frequency (IF) signal. This IF signal contains target range and velocity information, and its frequency is proportional to the target range. The IF signal is then sampled using analog-to-digital conversion to obtain a discrete-time sequence. This discrete-time sequence is organized based on the radar frame structure of a frequency-modulated continuous-wave millimeter-wave radar, generating a range-dimensional data matrix and a velocity-dimensional data matrix. The range-dimensional data matrix is ​​processed using a Fast Fourier Transform (FFT) to convert the time-domain signal to the frequency domain. The distance to the gesture target is obtained based on the frequency corresponding to the frequency domain peak. A second-dimensional FFT is performed on the velocity-dimensional data matrix to extract the Doppler frequency shift of the gesture target, obtaining the gesture target velocity. The magnitude of the frequency shift is related to the target velocity. A range-velocity matrix is ​​formed based on the gesture target range and velocity. In the range-velocity matrix, ordered statistical constant false alarm rate (CFAR) detection is used. Local noise levels are calculated using a sliding window to distinguish between real gesture targets and noise. A dynamic threshold is set, retaining cells larger than the dynamic threshold to suppress warnings. A binarized detection result is output, and potential target point clouds are marked.

[0062] The DBSCAN clustering algorithm is used to merge neighboring points to form gesture-related point cloud clusters. The potential target point clouds are mapped to three-dimensional space, and the azimuth information of each point is calculated using radar antenna array configuration and beamforming technology. Data from different dimensions is normalized to ensure clustering calculations are performed under the same units. A neighborhood radius is defined based on the physical size of the gesture and a preset radar resolution. All potential target point clouds are traversed; if the number of points within the neighborhood radius of a potential target point cloud is greater than a preset minimum number of points, that point cloud is marked as a core point cloud, representing the main part of the gesture. Starting from the core point cloud, all reachable potential target point clouds within its neighborhood are recursively merged to form independent clusters. Boundary points are assigned to the nearest independent cluster, and isolated point clouds not assigned to any independent cluster are marked as noise and removed. Motion consistency checks are performed on the clustered independent clusters, including Doppler velocity filtering. The average Doppler velocity of points within the same cluster is calculated, and outliers with excessively large velocity differences are removed. The spatiotemporal continuity of the current cluster is verified by combining the clustering results from the previous frame. Calculate the centroid coordinates and bounding box size of each independent cluster, extract the velocity distribution of the point cloud within the cluster to distinguish different gestures, generate cluster features to obtain gesture labels, and output a gesture point cloud sequence with gesture labels.

[0063] Figure 2 The flowchart illustrates the dynamic selection of the optimal motion model using graph matching.

[0064] According to an embodiment of the present invention, a multimodal motion model library including uniform motion model, acceleration model and arc motion model is established, and the optimal motion model is dynamically selected using graph matching, specifically as follows:

[0065] S202, Construct a uniform motion model suitable for steady linear motion, an acceleration model suitable for variable speed motion, and an arc motion model suitable for rotational or arc trajectory, and establish a multimodal motion model library based on the uniform motion model, acceleration model, and arc motion model;

[0066] S204, a time series graph structure is constructed in a sliding window manner. Each node contains a motion feature vector, a motion model prediction state, and an observation matching degree, wherein the observation matching degree uses the Mahalanobis distance score between the motion model prediction trajectory and the measured point cloud.

[0067] S206, The time sequence graph structure includes time sequence edges, model competition edges, and cross-frame similar edges. Time sequence edges are constructed by forcibly connecting adjacent frame nodes and edge weights are calculated based on motion continuity. Model competition edges are constructed by connecting different model nodes in the same frame and edge weights are calculated based on KL divergence between models. Cross-frame similar edges are constructed by connecting nodes in non-continuous frames but with similar motion patterns and edge weights are calculated based on trajectory morphology DTW distance.

[0068] S208. After obtaining new frame data, a new node containing three model prediction states is generated. The graph structure is updated. A weighted bipartite graph is constructed based on the updated graph structure. The left node represents the candidate motion model of a specific frame, and the right node corresponds to the observation data of a specific frame. The edge connections and edge weights of the weighted bipartite graph are generated according to the three edge structures and edge weights in the time series graph structure.

[0069] S210, the Hungarian algorithm is improved by introducing motion smoothness constraints as a penalty term, solving the maximum weight matching problem to select the optimal motion model, and predicting the next state of the gesture based on the gesture point cloud sequence according to the optimal motion model, and generating the predicted state.

[0070] It should be noted that the uniform motion model uses a linear dynamics model as its state equation, and the state vector includes position and velocity; the acceleration model extends the state vector by adding an acceleration component; and the arc motion model uses a curvature parameterization model, introducing angular velocity as a state variable. A time-series graph structure is constructed using a sliding window approach. Each node contains a motion feature vector, the predicted state of the motion model, and the observation matching degree. The motion feature vector includes the rate of change of displacement, curvature, velocity variance, and acceleration magnitude. In the weighted bipartite graph construction, the left node set is the candidate motion model sequence, containing all candidate model nodes for each frame, and the right node set is the observation feature sequence, containing the observation data nodes for each frame, including point cloud spatial distribution, kinematic features, and signal-to-noise ratio. The edge weights integrate the model-observation matching degree and temporal consistency. In the bipartite graph model, the weights of temporal edges are encoded as model transition penalty terms, restricting the Hungarian algorithm to selecting only model combinations with higher temporal edge weights. The KL divergence of competing model edges is transformed into the intrinsic weights of nodes on the left side of the bipartite graph, significantly distinguishing different models under similar observations. Cross-frame similar edges are mapped to virtual connections on the right side of the bipartite graph, and a cross-frame consistency term is added to the Hungarian algorithm. In the graph structure, model-observation nodes within the same frame must be connected, while nodes of different models across frames are not connected.

[0071] An improvement to the Hungarian algorithm is introduced by incorporating motion smoothness constraints as a penalty term. If adjacent frames have different model types, the total score is penalized. The top-two weighted edges are retained for each frame to generate a coarse selection map. The Hungarian algorithm is then run on this map, employing row and column reduction, finding independent zero elements, and cover line checks to execute the algorithm. The output model sequence is then used to obtain the normalized total weight, and the optimal motion model is selected. A user gesture habit database is established, and high-frequency model selection sequences for specific users are statistically analyzed. Corresponding models are preloaded at the start of the same gesture.

[0072] Figure 3 A flowchart is shown for optimizing the observed point cloud based on a Poisson-Dobernoli hybrid filter.

[0073] According to an embodiment of the present invention, a probability distribution is managed based on a Poisson-Bernoulli hybrid filtering framework. The observation weights are dynamically adjusted and the observation point cloud is optimized based on the signal-to-noise ratio and spatial distribution of the current gesture point cloud sequence. Specifically,

[0074] S302, describes the new gesture target through the Poisson point process, establishes the spatial probability density of the new gesture target within the radar field of view, and uses multiple Bernoulli mixtures to characterize the existence probability and state distribution of known gesture targets. Each known gesture target maintains a Bernoulli tuple, which includes the existence probability and the state distribution described by the Gaussian mixture model.

[0075] S304: Calculate the signal-to-noise ratio (SNR) of the observed point cloud in the gesture point cloud sequence, and calculate the normalized entropy value of the observed point cloud to represent the spatial distribution. Combine the SNR and spatial distribution to dynamically optimize the observation weights, and perform hierarchical data association through the optimized observation weights.

[0076] S306, In the association of new gesture targets, when the observation weight corresponding to the unmatched observation point cloud is greater than the preset weight threshold and the distance to the known gesture target is greater than the preset distance threshold, a new Bernoulli tuple is generated.

[0077] S308, In the known gesture target association, a noise matrix with fused observation weights is constructed based on the predicted state and the observation point cloud combined with the observation noise. The degree of matching between the predicted state and the observation point cloud is evaluated by weighted Mahalanobis distance. The n observation point clouds with the smallest Mahalanobis distance to the predicted position are selected as candidate observations. The association probability of the candidate observations is calculated, and the existence probability and state distribution are updated.

[0078] S310, reconstruct the point cloud based on Bernoulli tuples to obtain the optimized observation point cloud.

[0079] It should be noted that for all point clouds in the current frame, two core metrics are calculated: signal-to-noise ratio (SNR) and spatial distribution density. SNR is obtained by comparing the difference between the target point amplitude and the background noise level; high SNR observations typically correspond to realistic hand gesture reflections. Spatial distribution density counts the number of point clouds per unit volume; excessively sparse distributions may indicate sensor misses, while overly dense areas may indicate multipath interference. Point clouds are categorized into three types: high-confidence observations (high SNR and spatial continuity), edge observations (medium SNR or locally sparse), and suspected clutter (low SNR or isolated points).

[0080] The Poisson point process specifically handles newly emerging targets. When a high signal-to-noise ratio (SNR) observation point cannot be associated with an existing trajectory, its weight is automatically increased to ensure that newly emerging gesture targets are not filtered out. Simultaneously, the enhancement area is limited to avoid over-response to distant noise. For existing Bernoulli components, the quality variation trend of their historical matching observations is analyzed. If the SNR of three consecutive matching observations decreases, the matching threshold for new observations is gradually reduced to increase tolerance for low-quality observations; conversely, the standard is tightened to enhance trajectory purity. The spatial distribution characteristics of the detection point cloud directly affect the weights. For sparse observations caused by gesture occlusion, a neighborhood compensation algorithm is used. Centered on the trajectory prediction location, low-weight points within a 15cm radius are progressively weighted to compensate for missed detections. For abnormally dense clusters, competitive weight decay is implemented, retaining only the core points with the highest spatial consistency. The adjusted observation weights directly affect the two key probabilities of the PMBM: detection probability and clutter density. The detection probability corresponding to high-weighted observations is adjusted upwards, making it easier to correlate with existing trajectories; simultaneously, the spatial clutter density parameter is dynamically adjusted to appropriately increase the expected clutter quantity in low-weighted observation regions. This two-way adjustment allows the filter to effectively suppress false alarms while maintaining high sensitivity.

[0081] It should be noted that high-confidence target point clouds are extracted from the PMBM filter update step as optimized observation point clouds. Mahalanobis distance is used to calculate the single-point matching degree with the predicted state. For trajectory segments of a preset length, a matching degree matrix is ​​calculated based on the single-point matching degree. Dynamic time warping is used for optimal association. A warped path is defined based on the observation point cloud and the predicted state. A cost matrix is ​​constructed by combining the matching degree matrix and kinematic constraints. Represented as:

[0082]

[0083] in This is the distance weighting coefficient. Normalized distance to Ma City Let be the predicted velocity vector of the i-th trajectory. Let j be the equivalent velocity vector of the j-th observed point cloud, and let spatial distance be the term. It reflects the spatial fit between the predicted trajectory location and the observed point cloud, with the velocity difference term penalizing the mismatch in motion direction.

[0084] Dynamic programming is used to solve for the minimum cumulative cost, backtracking to obtain a predetermined number of optimal association pairs and generating a predetermined number of hypothetical trajectories. A dynamic trajectory hypothesis library is constructed and maintained through a multi-hypothesis tracking strategy. Each hypothesis node includes the trajectory state, existence probability, and historical observation matching sequence. If the current trajectory successfully matches an observation, the hypothesis continues; if an unmatched observation generates a new trajectory, a new hypothesis is created; a trajectory without a match for three consecutive frames is eliminated, and the hypothesis is terminated using the matching score. Smooth scoring and sustained scoring Construct a multi-dimensional scoring function, where The total number of frames within the current evaluation window. For time frame indexing, For the first t The weighted squared Mahalanobis distance of the frames represents the deviation between the observations and the model predictions. The first t Frame, First t-1 The acceleration vector of the frame. This represents the number of frames the current trajectory has survived. The maximum lifespan reference value is preset. The exponential term of the matching score maps the Mahalanobis distance to the matching confidence level in the [0,1] interval. The smoothing score compresses the acceleration changes in the infinite domain to the [0,1] interval through the hyperbolic tangent function, penalizing drastic acceleration changes and encouraging smooth motion. The continuous score quantifies the continuous tracking stability of the trajectory, avoiding the misidentification of transient noise as a valid target. Low weights are applied to newly initialized trajectories, and high weights are given to stable trajectories. The hypothetical trajectory is scored using the multi-dimensional scoring function. The hypothetical trajectory with the highest score is selected as the optimal trajectory. Kalman smoothing is applied to the optimal trajectory, and the process noise is adjusted according to the dynamic characteristics of the gesture target to output the gesture dynamic trajectory.

[0085] A second aspect of the present invention provides a real-time dynamic trajectory tracking system for millimeter-wave radar gesture recognition.

[0086] Figure 4 The architecture diagram of a real-time dynamic trajectory tracking system for millimeter-wave radar gesture recognition is shown.

[0087] The second embodiment of the present invention provides a real-time dynamic trajectory tracking system 4 for millimeter-wave radar gesture recognition. The system includes: a signal acquisition and preprocessing module 401, a point cloud optimization and clustering module 402, a motion model matching module 403, a multimodal tracking module 404, and a trajectory smoothing module 405.

[0088] The signal acquisition and preprocessing module transmits millimeter-wave signals through a frequency-modulated continuous wave millimeter-wave radar and receives echo signals reflected from hand gestures, and preprocesses the echo signals.

[0089] The point cloud optimization and clustering module uses constant false alarm rate detection to filter noise, extracts potential target point clouds, clusters the potential target point clouds, and generates a gesture point cloud sequence.

[0090] The motion model matching module establishes a multimodal motion model library including uniform motion model, acceleration model and arc motion model, uses graph matching to dynamically select the optimal motion model, and combines the gesture point cloud sequence to predict the next moment state of the gesture to generate the predicted state.

[0091] The multimodal tracking module is based on a Poisson-Bernoulli hybrid filtering framework. It dynamically adjusts the observation weights and optimizes the observation point cloud according to the signal-to-noise ratio and spatial distribution of the current gesture point cloud sequence. It uses Mahalanobis distance to calculate the matching degree between the observation point cloud and the predicted state, and combines dynamic time warping to perform optimal association. It adopts a multi-hypothesis tracking strategy to maintain trajectory hypotheses and selects the optimal trajectory through a trajectory scoring mechanism.

[0092] The trajectory smoothing module performs Kalman filtering on the optimal trajectory to output a dynamic gesture trajectory.

[0093] The third embodiment of the present invention provides a computer-readable storage medium, which includes a real-time dynamic trajectory tracking method program for millimeter-wave radar gesture recognition. When the real-time dynamic trajectory tracking method program for millimeter-wave radar gesture recognition is executed by a processor, it implements the steps of a real-time dynamic trajectory tracking method for millimeter-wave radar gesture recognition.

[0094] In the several embodiments provided in this application, it should be understood that the disclosed methods and systems can be implemented in other ways. The system embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, and can be electrical, mechanical, or other forms. Furthermore, in the various embodiments of the present invention, all functional units can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0095] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0096] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A real-time dynamic trajectory tracking method for millimeter wave radar gesture recognition, characterized in that, The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: A multi-modal motion model library comprising a uniform speed model, an acceleration model and an arc motion model is established, and an optimal motion model is dynamically selected by graph matching, and a next time state of the gesture is predicted based on the gesture point cloud sequence to generate a predicted state; Based on a Poisson multi-Bernoulli mixed filtering framework, the observation weight is dynamically adjusted according to the signal-to-noise ratio and spatial distribution of the current gesture point cloud sequence, the observation point cloud is optimized, the matching degree of the observation point cloud and the predicted state is calculated by using the Mahalanobis distance, and the optimal association is performed by combining the dynamic time warping; A multi-hypothesis tracking strategy is used to maintain a trajectory hypothesis, and the optimal trajectory is selected by a trajectory scoring mechanism, and the optimal trajectory is smoothed by Kalman filtering to output a gesture dynamic trajectory; A multi-modal motion model library comprising a uniform speed model, an acceleration model and an arc motion model is established, and an optimal motion model is dynamically selected by graph matching, and a next time state of the gesture is predicted based on the gesture point cloud sequence to generate a predicted state; A uniform speed model suitable for smooth linear motion, an acceleration model suitable for variable speed motion and an arc motion model suitable for rotating or arc trajectory are constructed, and a multi-modal motion model library is established based on the uniform speed model, the acceleration model and the arc motion model; A sliding window is used to construct a time sequence graph structure, each node comprises a motion feature vector, a motion model predicted state and an observation matching degree, and the observation matching degree is obtained by using the Mahalanobis distance score of the motion model predicted trajectory and the measured point cloud; The time sequence graph structure comprises a time sequence edge, a model competition edge and a cross-frame similar edge, the time sequence edge is constructed by forcibly connecting adjacent frame nodes, the edge weight is calculated based on motion continuity, the model competition edge is constructed by connecting nodes of different models in the same frame, the edge weight is calculated based on the KL divergence between models, the cross-frame similar edge is constructed by connecting nodes with similar motion modes in non-continuous frames, and the edge weight is calculated based on the DTW distance of the trajectory shape; After obtaining new frame data, a new node comprising three model predicted states is generated, the graph structure is updated, a weighted bipartite graph is constructed based on the updated graph structure, the left side node represents a candidate motion model of a specific frame, the right side node corresponds to observation data of a specific frame, and the edge connection and edge weight of the weighted bipartite graph are generated according to the three edge structures and edge weights in the time sequence graph structure; 2. The real-time dynamic trajectory tracking method of mm-wave radar gesture recognition according to claim 1, characterized in that, The motion smoothness constraint is introduced as a penalty term to improve the Hungarian algorithm, the optimal motion model is selected by solving the maximum weight matching, and the next time state of the gesture is predicted based on the gesture point cloud sequence according to the optimal motion model to generate a predicted state. The method comprises the following steps: The received echo signal is mixed with the current transmission signal to obtain an intermediate frequency signal, the intermediate frequency signal is sampled by analog-to-digital conversion to obtain a discrete time sequence, the discrete time sequence is organized based on the radar frame structure of the frequency-modulated continuous wave millimeter wave radar to generate a distance dimension data matrix and a speed dimension data matrix; The distance dimension data matrix is processed using fast Fourier transform to convert the time domain signal into a frequency domain, a gesture target distance is obtained according to a frequency corresponding to a frequency domain peak value, a second dimension fast Fourier transform is performed on the speed dimension data matrix, and a Doppler frequency shift of the gesture target is extracted to obtain a gesture target speed; A distance-speed matrix is formed according to the gesture target distance and the gesture target speed, and a local noise level is calculated by using a sliding window in the distance-speed matrix to distinguish a real gesture target from noise; A dynamic threshold is set, units greater than the dynamic threshold are retained, a pre-warning is suppressed, a binary detection result is output, a potential target point cloud is marked, the potential target point cloud is mapped to a three-dimensional space, and data in different dimensions are normalized.

3. The real-time dynamic trajectory tracking method of mm-wave radar gesture recognition according to claim 2, characterized in that, The potential target point cloud is clustered to generate a gesture point cloud sequence, specifically: A neighborhood radius is defined according to a physical size of the gesture and a preset radar resolution, all potential target point clouds are traversed, when a point number in the neighborhood radius of a potential target point cloud is greater than a preset minimum point cloud number, the point cloud is marked as a core point cloud, and the core point cloud is used to represent a main part of the gesture; Starting from the core point cloud, all reachable potential target point clouds in the neighborhood thereof are recursively merged to form independent clusters, and boundary points are attributed to the nearest independent cluster, and isolated point clouds not attributed to any independent cluster are marked as noise and removed; Motion consistency of the independent clusters generated by clustering is tested, a centroid coordinate and a bounding box size of each independent cluster are calculated, a speed distribution of a point cloud in the cluster is extracted, a cluster feature is generated, a gesture label is obtained, and a gesture point cloud sequence with the gesture label is output.

4. The real-time dynamic trajectory tracking method of mm-wave radar gesture recognition according to claim 1, characterized in that, Based on a Poisson multi-Bernoulli mixture filtering framework, an observation weight is dynamically adjusted according to a signal-to-noise ratio and a spatial distribution of a current gesture point cloud sequence, and an observed point cloud is optimized, specifically A new gesture target is described by a Poisson point process, a spatial probability density of the new gesture target is established in a radar field of view, an existence probability and a state distribution of a known gesture target are described by a multi-Bernoulli mixture, each known gesture target maintains a Bernoulli tuple including the existence probability and the state distribution described by a Gaussian mixture model; A signal-to-noise ratio of the observed point cloud in the gesture point cloud sequence is calculated, a normalized entropy value of the observed point cloud is calculated to represent the spatial distribution, the observation weight is dynamically optimized in combination with the signal-to-noise ratio and the spatial distribution, hierarchical data association is performed by using the optimized observation weight; In new gesture target association, when an observation weight corresponding to an unmatched observed point cloud is greater than a preset weight threshold and a distance from the known gesture target is greater than a preset distance threshold, a new Bernoulli tuple is generated; In known gesture target association, a noise matrix with a fusion observation weight is constructed according to a predicted state and the observed point cloud in combination with observation noise, a weighted Mahalanobis distance is used to evaluate a matching degree of the predicted state and the observed point cloud, n observed points with the smallest Mahalanobis distance from the predicted position are selected as candidate observations, an association probability of the candidate observations is calculated, and the existence probability and the state distribution are updated; The point cloud is reconstructed according to the Bernoulli tuple to obtain an optimized observed point cloud.

5. The real-time dynamic trajectory tracking method of mm-wave radar gesture recognition according to claim 4, characterized in that, The matching degree of the observation point cloud and the predicted state is calculated by using the Mahalanobis distance, and optimal association is performed by using dynamic time warping, specifically as follows: For the optimized observation point cloud, the single-point matching degree with the predicted state is calculated by using the Mahalanobis distance, and a matching degree matrix is calculated for a preset length of a track segment according to the single-point matching degree; Optimal association is performed by using dynamic time warping, a warping path is defined according to the observation point cloud and the predicted state, a cost matrix is constructed by combining the matching degree matrix and kinematic constraints, the minimum cumulative cost is solved by using dynamic programming, a set of optimal association pairs of a preset number is obtained by backtracking, and a preset number of assumed tracks are generated.

6. The real-time dynamic trajectory tracking method of mm-wave radar gesture recognition according to claim 1, characterized in that, The track hypotheses are maintained by using a multi-hypothesis tracking strategy, and the optimal track is selected by using a track scoring mechanism, and the optimal track is smoothed by using Kalman filtering, and a dynamic track of the hand gesture is output, specifically as follows: A dynamic track hypothesis library is constructed and maintained by using a multi-hypothesis tracking strategy, each hypothesis node includes a track state, a probability of existence and a historical observation matching sequence, a multi-dimensional scoring function is constructed by using a matching score, a smoothing score and a persistence score, and the multi-dimensional scoring function is used for scoring the assumed track; The hypothesis track with the highest score is selected as the optimal track, the optimal track is smoothed by using Kalman filtering, and the process noise is adjusted according to the dynamic characteristics of the hand gesture target, and a dynamic track of the hand gesture is output.

7. A real-time dynamic trajectory tracking system for millimeter wave radar gesture recognition, characterized in that, The system is used for implementing the real-time dynamic track tracking method for millimeter wave radar gesture recognition according to any one of claims 1-6, and the system includes a signal acquisition and preprocessing module, a point cloud optimization and clustering module, a motion model matching module, a multi-modal tracking module and a track smoothing module; The signal acquisition and preprocessing module transmits a millimeter wave signal by using a frequency-modulated continuous wave millimeter wave radar and receives a return signal of a hand gesture reflection, and the return signal is preprocessed; The point cloud optimization and clustering module filters noise by using a constant false alarm rate, extracts a potential target point cloud, clusters the potential target point cloud, and generates a hand gesture point cloud sequence; The motion model matching module establishes a multi-modal motion model library including a uniform speed model, an acceleration model and an arc motion model, dynamically selects an optimal motion model by using graph matching, predicts a next time state of the hand gesture according to the hand gesture point cloud sequence, and generates a predicted state; The multi-modal tracking module is based on a Poisson multi-Bernoulli hybrid filtering framework, dynamically adjusts an observation weight according to a signal-to-noise ratio and a spatial distribution of a current hand gesture point cloud sequence, optimizes an observation point cloud, calculates a matching degree of the observation point cloud and the predicted state by using the Mahalanobis distance, and performs optimal association by using dynamic time warping; and the track hypotheses are maintained by using a multi-hypothesis tracking strategy, and the optimal track is selected by using a track scoring mechanism; The track smoothing module smoothes the optimal track by using Kalman filtering, and outputs a dynamic track of the hand gesture.

Citation Information

Patent Citations

  • Multi-extended-target PMBM tracking method of Gaussian process regression model

    CN117784115A

  • In-cabin gesture recognition method and system based on non-facing scene point cloud

    CN117789295A