UWB radar multi-target sensing and tracking method and system
By constructing a multimodal dataset using UWB radar combined with an optical motion capture system, and employing a three-level cascaded processing and graph neural network model, the problems of high hardware cost, poor real-time performance, and insufficient adaptability of traditional radar multi-target perception and tracking methods are solved, achieving high-precision multi-target detection and tracking in complex environments.
Patent Information
- Application Number
- CN202511136325.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-11-21
AI Technical Summary
Existing radar multi-target perception and tracking methods rely on multimodal sensors, which have high hardware costs, insufficient real-time processing capabilities, poor adaptability of tracking algorithms, and limited application scenarios, making it difficult to achieve high-precision and stable multi-target detection and tracking in complex environments.
UWB radar is used to collect echo signals from multiple scenes. A multimodal dataset is constructed by combining it with an optical motion capture system. Features are extracted through a three-level cascaded processing flow and a hybrid encoder. A detection and tracking model is constructed using graph neural networks and dynamic filtering theory. The model is optimized using a multi-task joint loss function and deployed to an edge computing device for real-time processing.
Maintaining centimeter-level ranging accuracy in complex environments, reducing noise interference, improving trajectory continuity and target detection accuracy, reducing the probability of false tracking and missed tracking, and achieving real-time multi-target perception and tracking.
Smart Images

Figure CN120993394A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wireless communication, and in particular to a UWB radar multi-target sensing and tracking method and system. Background Technology
[0002] Ultra-wideband (UWB) radar, as a high-precision and highly interference-resistant sensing technology, possesses unique advantages in multi-target detection and tracking in complex scenarios. With the rapid development of autonomous driving, intelligent security, and industrial robots, the demand for high-precision, real-time environmental perception technologies is increasingly urgent. However, traditional radar or optical sensors often suffer from insufficient resolution, high false detection rates, and poor tracking stability under conditions of dense targets, dynamic obstruction, or adverse weather. For example, in complex urban roads, multiple targets such as vehicles, pedestrians, and bicycles interact frequently, making it difficult for traditional millimeter-wave radar to distinguish close-range targets, while camera performance deteriorates sharply at night or in rainy or foggy weather. UWB radar, with its nanosecond-level pulse signals and wide bandwidth characteristics, can achieve centimeter-level ranging accuracy and strong penetration capabilities. However, in multi-target scenarios, solving problems such as signal aliasing and dynamic correlation remains a current technological bottleneck. Therefore, developing a multi-target perception and tracking system based on UWB radar is of great significance for overcoming perception limitations in complex environments and improving the reliability and security of intelligent systems.
[0003] Existing radar multi-target sensing and tracking methods are as follows:
[0004] Chinese invention patent application CN202210126527.3 discloses a target tracking method and apparatus based on millimeter-wave radar and lidar. The method includes: acquiring a first tracking result obtained by the millimeter-wave radar during a previous detection cycle of a monitored scene, where the monitored scene includes at least one target to be tracked; adjusting the detection range of the lidar in the current detection cycle based on the first tracking result to obtain a first detection range; detecting the target to be tracked within the first detection range using the lidar to obtain first detection information; and updating the first tracking result based on the first detection information to obtain a second tracking result for the target to be tracked within the current detection cycle. Through this method, the present application can improve tracking accuracy and detection efficiency.
[0005] Chinese invention patent application number CN202410767647.0 discloses a roadside perception method and system for video target fusion and LiDAR target fusion, relating to the field of intelligent traffic detection. It solves the problems of excessive computation and low efficiency in existing roadside perception fusion technologies for scenarios with multiple LiDARs and cameras. By independently fusing and recognizing LiDAR point cloud data and independently recognizing and fusing camera data, the fusion results of the two sensors are finally fused at a bird's-eye view level to obtain the intersection BEV (Bird's Eye View) image. Through the above steps, multi-LiDAR and multi-camera data fusion can effectively avoid blind spot detection and improve data accuracy. Simultaneously, by combining the respective advantages of LiDAR, cameras, and the BEV level, it achieves efficient, accurate, and intuitive roadside perception.
[0006] However, the above method has the following limitations:
[0007] 1) Strong dependence on multimodal sensors. Existing methods typically require a combination of millimeter-wave radar and lidar or vision sensors to achieve effective target tracking. While this multi-sensor fusion approach improves perception accuracy, it also significantly increases hardware costs and system complexity.
[0008] 2) Insufficient real-time processing capability. Existing methods require computationally intensive operations such as complex point cloud segmentation, clustering, and feature extraction when processing LiDAR point cloud data. While these processes can provide relatively accurate target detection results, they introduce significant computational latency, making it difficult to meet the needs of applications with high real-time requirements.
[0009] 3) Insufficient adaptability of tracking algorithms. Existing multi-target association methods are mainly based on traditional methods such as the Hungarian algorithm. When faced with complex scenarios such as dense targets, intersecting trajectories, and frequent occlusion, these methods often suffer from problems such as trajectory confusion and target loss.
[0010] 4) Limited application scenarios. Existing methods are often optimized for specific hardware platforms or sensor configurations, and this inflexible architectural design limits the widespread application of the technology.
[0011] Therefore, it is necessary to propose a solution to improve one or more problems existing in the above-mentioned related technical solutions.
[0012] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0013] The purpose of this disclosure is to provide a UWB radar multi-target perception and tracking method and system, thereby overcoming, at least to some extent, one or more problems caused by the limitations and defects of related technologies.
[0014] The first aspect of this disclosure provides a UWB radar multi-target sensing and tracking method, including the following steps:
[0015] The original echo signals from multiple scenes are collected by UWB radar, and the real position information of the target is obtained by synchronous optical motion capture system. The data diversity is enhanced by noise injection and time-frequency transformation, and a multimodal dataset containing time domain, frequency domain and spatial domain features is constructed.
[0016] A three-level cascaded processing flow is used to denoise and enhance the signal of the original echo. The basic features in the time and frequency domain are extracted by wavelet transform and short-time Fourier transform. A high-dimensional feature tensor is output by a hybrid encoder containing CNN and Transformer.
[0017] Based on graph neural network and dynamic filtering theory, a model including an improved PointPillars detection head and a GAT tracking head is constructed to perform target position estimation and trajectory correlation.
[0018] The model is trained using a multi-task joint loss function, and optimized by adaptive loss balancing, knowledge distillation, and cosine annealing learning rate scheduling.
[0019] The optimized model is deployed to edge computing devices to build a complete processing link from radar signal input to trajectory output, outputting target trajectory and motion status and visualization, thereby realizing real-time multi-target perception and tracking.
[0020] In an exemplary embodiment of this application, the step of acquiring raw echo signals from multiple scenes using UWB radar, simultaneously obtaining the target's true location information using an optical motion capture system, enhancing data diversity through noise injection and time-frequency transformation, and constructing a multimodal dataset containing time-domain, frequency-domain, and spatial-domain features includes:
[0021] Collect static target grid distribution data and dynamic target preset motion pattern data in multiple indoor and outdoor scenarios;
[0022] The number of targets, their 3D positions, velocity vectors, and trajectory IDs are labeled, and data augmentation is performed through noise injection, temporal pruning, and frequency domain filtering.
[0023] Spectral and spatial features are extracted using short-time Fourier transform and beamforming algorithms.
[0024] In an exemplary embodiment of this application, the three-level cascaded processing flow includes:
[0025] The input raw echo signal undergoes adaptive median filtering to remove impulse noise;
[0026] An LMS adaptive filter is used to eliminate environmental multipath interference;
[0027] The z-score method is used for normalization to eliminate the dimensional differences in signal amplitude;
[0028] Principal component analysis was performed using scikit-learn to retain the principal components corresponding to energy and remove redundant features in noisy or highly correlated frequency bands.
[0029] In one exemplary embodiment of this application, the hybrid encoder comprising CNN and Transformer includes:
[0030] CNN uses a variant of ResNet18 to extract local scattering features;
[0031] A multi-head self-attention mechanism is used to model long-range dependencies across frames and output a spatiotemporal feature tensor.
[0032] In an exemplary embodiment of this application, the construction of the model comprising the improved PointPillars detection head and the GAT tracking head includes:
[0033] The UWB radar point cloud is projected onto the BEV plane, and a non-uniform mesh is generated by dynamic voxelization.
[0034] PointNet++ is used to extract local geometric features, and the SSD detection head is used to fuse multi-scale receptive fields;
[0035] The sparse self-attention module calculates the weights of the relationships between targets.
[0036] In an exemplary embodiment of this application, the construction of the model including the improved PointPillars detection head and the GAT tracking head further includes:
[0037] A motion consistency and appearance similarity cost matrix is constructed based on an improved Hungarian algorithm;
[0038] The extended Kalman filter uses a constant turning rate model to predict the target state;
[0039] A two-layer graph attention network aggregates single-hop neighbor features and group interaction patterns.
[0040] In an exemplary embodiment of this application, the step of training the model using a multi-task joint loss function and combining adaptive loss balancing, knowledge distillation, and cosine annealing learning rate scheduling to optimize the model includes:
[0041] A multi-task joint loss function is constructed based on the PyTorch framework, including object detection loss function, trajectory regression loss function and interaction consistency loss function;
[0042] The AdamW optimizer is used for parameter updates, and the learning rate scheduling adopts a combination of linear preheating and cosine annealing strategies.
[0043] In one exemplary embodiment of this application, the target detection loss function employs an improved Focal Loss: Where, p i Let represent the confidence score of the i-th predicted box, γ be the focus factor (set to 2.0), and α be the focus factor. i Balance the weights for each category;
[0044] The trajectory regression loss function uses adaptive IoU loss: L reg =1-IoU(B pred B gt )+λ||Δv||;where, B pred and B gt These represent the predicted bounding box and the ground truth bounding box, respectively; Δv represents the velocity vector error; and λ is the balance coefficient.
[0045] The interaction consistency loss function is calculated using a graph attention network: L int =∑ i,j ||GAT(ft i )-GAT(ft j )||2; where GAT(·) represents the graph attention layer output, ft i and ft j These are target features in adjacent frames.
[0046] In an exemplary embodiment of this application, the expression for the learning rate scheduling employing a combination strategy of linear preheating and cosine annealing is as follows:
[0047]
[0048] Where, η v Let η be the learning rate at step v. base The base learning rate is T, where T is the total number of parameter updates during the entire model training process. warm η is the number of preheating steps. max and η min These represent the maximum and minimum learning rates.
[0049] A second aspect of this disclosure provides a UWB radar multi-target sensing and tracking system, the UWB radar multi-target sensing and tracking system comprising:
[0050] The data acquisition module collects raw echo signals from multiple scenes using UWB radar, and simultaneously obtains the target's real location information using an optical motion capture system. After noise injection and time-frequency transformation to enhance data diversity, a multimodal dataset containing time-domain, frequency-domain, and spatial-domain features is constructed.
[0051] The signal processing module uses a three-level cascaded processing flow to denoise and enhance the original echo. It extracts the basic features in the time and frequency domain through wavelet transform and short-time Fourier transform, and outputs a high-dimensional feature tensor using a hybrid encoder that includes CNN and Transformer.
[0052] The detection and tracking module, based on graph neural network and dynamic filtering theory, constructs a model that includes an improved PointPillars detection head and a GAT tracking head to perform target position estimation and trajectory correlation.
[0053] The training optimization module uses a multi-task joint loss function to train the model, and combines adaptive loss balancing, knowledge distillation and cosine annealing learning rate scheduling optimization model.
[0054] The edge deployment module deploys the optimized model to edge computing devices, builds a complete processing link from radar signal input to trajectory output, outputs target trajectory and motion status and visualizes it, and realizes real-time multi-target perception and tracking.
[0055] The present invention has the following beneficial effects:
[0056] (1) By collecting echoes from multiple scenes through UWB radar and constructing a multimodal dataset containing time-domain, frequency-domain, and spatial-domain features, the hardware cost bottleneck of traditional multi-sensor fusion is overcome. It can still maintain stable centimeter-level ranging accuracy in complex electromagnetic environments or harsh weather conditions, and overcome the defects of visual / laser sensors that are affected by environmental interference.
[0057] (2) The three-level cascaded processing effectively removes impulse noise and multipath interference, and combines wavelet transform and short-time Fourier transform to extract time-frequency features, providing high signal-to-noise ratio input data for subsequent models, thereby reducing the average absolute error of static target detection.
[0058] (3) The CNN and Transformer hybrid encoder extracts local scattering features and combines multi-head self-attention mechanism to model cross-frame long-range dependence, which improves the trajectory integrity rate in multi-person interaction scenarios and effectively solves the trajectory confusion problem when targets are densely intersecting.
[0059] (4) The improved PointPillars detection head and GAT tracking head, constructed based on graph neural network and dynamic filtering theory, improve the tracking time in target occlusion scenarios, enhance the trajectory continuity in complex scenarios, can intelligently predict the motion trajectory of occluded targets, and accurately resume tracking when the target reappears, greatly reducing the probability of mistracking and missed tracking; the multi-task joint loss function combined with cosine annealing learning rate scheduling enables efficient model training. Attached Figure Description
[0060] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0061] Figure 1 This diagram illustrates the steps of a UWB radar multi-target perception and tracking method in an exemplary embodiment of this application.
[0062] Figure 2 This diagram illustrates a flowchart of a UWB radar multi-target perception and tracking method in an exemplary embodiment of this application.
[0063] Figure 3 This illustration shows a schematic diagram of the model structure including the improved PointPillars detection head and GAT tracking head in an exemplary embodiment of this application;
[0064] Figure 4 This illustration shows a flowchart of the process of deploying the optimization model to an edge computing device in an exemplary embodiment of this application;
[0065] Figure 5 This diagram illustrates a UWB radar multi-target sensing and tracking system in an exemplary embodiment of this application. Detailed Implementation
[0066] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0067] Furthermore, the accompanying drawings are merely illustrative of this application and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0068] This example implementation provides a UWB radar multi-target sensing and tracking method, such as... Figures 1-2 As shown, the following steps may be included:
[0069] Step S101: Collect raw echo signals from multiple scenes using UWB radar, obtain the target's real location information using a synchronous optical motion capture system, enhance data diversity through noise injection and time-frequency transformation, and construct a multimodal dataset containing time-domain, frequency-domain, and spatial-domain features.
[0070] Step S102: The original echo is denoised and signal enhanced using a three-level cascaded processing flow. The basic features in the time and frequency domain are extracted by wavelet transform and short-time Fourier transform. A high-dimensional feature tensor is output using a hybrid encoder containing CNN and Transformer.
[0071] Step S103: Based on graph neural network and dynamic filtering theory, construct a model that includes an improved PointPillars detection head and a GAT tracking head to perform target position estimation and trajectory association;
[0072] Step S104: Train the model using a multi-task joint loss function, and optimize the model by combining adaptive loss balancing, knowledge distillation, and cosine annealing learning rate scheduling.
[0073] Step S105: Deploy the optimized model to an edge computing device to build a complete processing link from radar signal input to trajectory output, output the target trajectory and motion status and visualize it to achieve real-time multi-target perception and tracking.
[0074] This application proposes a UWB radar multi-target perception and tracking method. Firstly, it collects echoes from multiple scenes using UWB radar and constructs a multimodal dataset containing time, frequency, and spatial features, overcoming the hardware cost bottleneck of traditional multi-sensor fusion. This method maintains stable centimeter-level ranging accuracy even in complex electromagnetic environments or adverse weather conditions, overcoming the environmental interference limitations of visual / laser sensors. Secondly, a three-stage cascaded processing flow effectively removes impulse noise and multipath interference, and combines wavelet transform and short-time Fourier transform to extract time-frequency features, providing high signal-to-noise ratio input data for subsequent models and reducing the mean absolute error of static target detection. Thirdly, a hybrid CNN and Transformer encoder extracts local scattering features, and a multi-head self-attention mechanism models long-range dependencies across frames, improving trajectory integrity in multi-person interactive scenarios and effectively solving the trajectory confusion problem when targets are densely intersecting. Finally, an improved PointPillars detection head and GAT (Graph Attention) are constructed based on graph neural networks and dynamic filtering theory. The Graph Attention Network (GAN) tracking head improves tracking time in target occlusion scenarios and enhances trajectory continuity in complex scenes. It can intelligently predict the motion trajectory of occluded targets and accurately resume tracking when the target reappears, greatly reducing the probability of mistracking and missed tracking. The multi-task joint loss function combined with cosine annealing learning rate scheduling enables efficient model training.
[0075] Below, as Figures 1-4 As shown, a more detailed explanation of the UWB radar multi-target perception and tracking method proposed in this example embodiment will be provided.
[0076] Step S101: Collect raw echo signals from multiple scenes using UWB radar, obtain the target's real location information using a synchronous optical motion capture system, enhance data diversity through noise injection and time-frequency transformation, and construct a multimodal dataset containing time-domain, frequency-domain, and spatial-domain features.
[0077] Specifically, UWB radar equipment is used to collect raw echo signals of static and dynamic targets in indoor, outdoor, and obstructed environments. A synchronous optical motion capture system acquires real-world location information as a labeling benchmark. The collected data is labeled with the number of targets, their location coordinates, movement speed, and complete trajectory information. Enhancement methods such as noise injection and time-frequency transformation are used to expand data diversity. The labeled radar echo signals are then subjected to short-time Fourier transform to extract time-frequency features, constructing a multimodal dataset containing time-domain, frequency-domain, and spatial-domain features. This dataset is then divided into training and validation sets according to a predetermined ratio.
[0078] Step S101 specifically includes:
[0079] Step S1001: Collect gridded distribution data of static targets and preset motion mode data of dynamic targets in multiple indoor and outdoor scenarios.
[0080] To construct a radar multi-target perception dataset with broad scenario coverage, a UWB radar development kit was used as the core data acquisition device. Data was collected in various typical environments, including indoor open spaces, outdoor open areas, and complex scenarios with metal obstacles. Static target data was acquired through a gridded distribution to ensure coverage of radar echo characteristics at different distances and angles. Dynamic target data was collected by simulating target movement behavior in real-world scenarios through various preset motion modes, such as uniform speed, variable speed, change of direction, and multi-person interaction. An optical motion capture system was used as a reference, and strict time synchronization and spatial calibration ensured precise alignment between radar data and real target positions. Environmental noise and multipath effects were also recorded to improve the dataset's realism and generalization ability.
[0081] Step S1002: Label the number of targets, their 3D positions, velocity vectors, and trajectory IDs, and perform data augmentation through noise injection, temporal pruning, and frequency filtering.
[0082] The dataset annotation employs specialized tools to perform refined processing on each frame of radar echo, including complete information such as target quantity, 3D position, velocity vector, acceleration, radar scattering characteristics, and target trajectory ID association. For dynamic targets, additional recording of their motion state changes ensures the data can be used for continuous tracking tasks. To enhance data diversity, various data augmentation strategies are employed, including noise injection to simulate different signal-to-noise ratio conditions, temporal domain pruning to adjust signal length, and frequency domain filtering to simulate attenuation characteristics across different frequency bands. Each basic scenario data undergoes multiple rounds of augmentation to ensure the final dataset covers a wide range of practical applications.
[0083] Step S1003: Extract spectral and spatial features using short-time Fourier transform and beamforming algorithms.
[0084] Feature extraction employs a multimodal fusion strategy. First, the spectral features of the target are extracted using short-time Fourier transform and Mel-Cepstral coefficients, and the echo characteristics of targets of different materials are distinguished by combining time-frequency features such as zero-crossing rate and spectral centroid. Second, the target's angle of arrival is calculated using radar array data through beamforming algorithms, and spatial coherence analysis is combined to enhance multi-target resolution. Finally, motion features such as target velocity, acceleration, and trajectory curvature are calculated based on continuous frame data to improve the accuracy of dynamic target tracking. The final feature set comprehensively covers the target's physical attributes, motion state, and spatial distribution information, providing highly discriminative input features for subsequent algorithms.
[0085] Step S102: The original echo is denoised and signal enhanced using a three-level cascaded processing flow. The basic features in the time and frequency domain are extracted by wavelet transform and short-time Fourier transform. A high-dimensional feature tensor is output using a hybrid encoder containing CNN and Transformer.
[0086] Specifically, to address the issues of noise interference and insufficient feature representation in UWB signals, a two-level feature extraction architecture is constructed based on digital signal processing and deep learning. Adaptive filters and principal component analysis are used to denoise and enhance the original echo, while wavelet transform and short-time Fourier transform are employed to extract fundamental time-frequency features. The preprocessed signal is input into a hybrid encoder containing CNN and Transformer, where convolutional layers capture local scattering features, and a self-attention mechanism is used to model long-range dependencies, ultimately outputting a high-dimensional feature tensor with spatiotemporal correlation.
[0087] Step S102 employs a three-level cascaded processing flow, including:
[0088] An adaptive LMS filter is used for dynamic noise suppression, followed by wavelet packet transform for signal denoising. Finally, sliding window PCA is used to extract principal component features to improve the signal-to-noise ratio of the original signal.
[0089] Input raw radar echo signal x raw (t) First, impulse noise is removed using adaptive median filtering:
[0090] x filtered (t)=medfilt(x raw (t), wf=5)
[0091] Where t is the time variable, x filtered (t) is the signal after adaptive median filtering, x raw (t) represents the original radar echo signal, medfilt(·) is the median filter function, and wf represents the size of the filter window.
[0092] An LMS (Least Mean Square) adaptive filter is used to eliminate environmental multipath interference.
[0093] e(n)=d(n)-fw T (n)x(n)fw(n+1)=fw(n)+μe(n)x(n)
[0094] Where d(n) is the desired signal, fw(n) is the filter coefficient vector, μ is the step size factor (taken as 0.01), fw(n+1) is the updated filter coefficient vector, the error e(n) can reflect the difference between the current filter output and the ideal signal, and x(n) is the input signal vector at time n.
[0095] The normalization process uses the z-score method to eliminate the dimensional differences in signal amplitude, making data from different radar devices or scenarios comparable.
[0096]
[0097] Where, x norm (t) represents the normalized signal, μ x and σ x These are the signal mean and standard deviation, respectively.
[0098] Principal Component Analysis (PCA) is implemented using scikit-learn, retaining principal components corresponding to 95% of the energy, removing redundant features (such as noise or highly correlated frequency bands), retaining key time-frequency patterns of the target echo, reducing the computational complexity of subsequent models, and preserving the main signal features.
[0099] Step S102, which uses a hybrid encoder containing CNN and Transformer to output a high-dimensional feature tensor, includes:
[0100] A hybrid architecture consisting of convolutional layers and a Transformer encoder is constructed. Local scattering patterns are extracted through convolutional networks, and cross-frame spatiotemporal correlations are modeled using a self-attention mechanism. Finally, a high-dimensional spatiotemporal fusion feature tensor is output.
[0101] First, the time-frequency analysis uses Morlet wavelet transform:
[0102]
[0103] Where ψ is the mother wavelet function, x norm (t) is the normalized input signal, t is the time variable, a is the scale parameter, b is the translation parameter, and the output MW(a,b) is a two-dimensional matrix (scale × time) used for subsequent target detection.
[0104] Deep feature extraction uses a variant of ResNet18, and the first layer convolutional kernel is changed to a 1D form to adapt to radar signals, which can enhance the model's robustness to noise and occlusion.
[0105] y l =σ(LW l *y l-1 )
[0106] Among them, y l Let LW be the output feature vector of the l-th layer. l σ represents the learnable weights, * denotes a one-dimensional convolution operation, and σ is the ReLU activation function.
[0107] Multi-head attention calculation for Transformer encoders:
[0108]
[0109] Where Q, K, and V are the query, key, and value matrices, respectively, and d k This is the feature dimension. Long-term time series modeling can capture the long-term dependencies of target motion.
[0110] Step S103: Based on graph neural network and dynamic filtering theory, construct a model that includes an improved PointPillars detection head and a GAT tracking head to perform target position estimation and trajectory association.
[0111] Specifically, to address the issues of false detection and trajectory fragmentation in dense targets, a joint detection and tracking framework is constructed based on graph neural networks and dynamic filtering theory. The detection module uses an improved PointPillars network structure to process radar point clouds, focusing on effective scattering points through a spatial attention mechanism to output target position and velocity estimates. The tracking module inputs continuous frame detection results into a Hungarian algorithm for data association, combines Kalman filtering to predict motion states, and employs a GAT graph attention network for interactive modeling.
[0112] like Figure 3 As shown, step S103, which involves constructing a model that includes the improved PointPillars detection head and the GAT tracking head, includes:
[0113] An improved PointPillars detection head was constructed. A feature encoding network based on columnar voxels was built, projecting the 3D point cloud onto the BEV plane and generating a non-uniform mesh using a dynamic voxelization strategy. Local geometric features were extracted within each mesh using the PointNet++ architecture. The improved SSD detection head design incorporates a multi-scale receptive field fusion module, deploying anchor box prediction branches on three feature maps at different resolutions. The self-attention module employs a sparsity design to calculate the weights of relationships between targets.
[0114] The UWB radar point cloud is projected onto the BEV plane, and a non-uniform mesh is generated through dynamic voxelization. An improved PointPillars network structure is used to divide the radar point cloud into an xy-plane mesh, with each mesh generating a multi-dimensional feature vector (containing x, y, z coordinates, RCS value, and velocity components). The feature extraction network includes:
[0115]
[0116] PC represents point cloud features. MLP(·) represents feature concatenation, MLP(·) represents multilayer perceptron, which performs a nonlinear transformation on the features of each point, and MaxPool(·) represents taking the maximum value of the features of all points in the grid.
[0117] PointNet++ is used to extract local geometric features, and an SSD detection head is used to fuse multi-scale receptive fields. The detection head adopts an SSD architecture, and the loss function is:
[0118] L det =αL cls +βL reg
[0119] Among them, L cls For focal loss, classification loss, L reg The loss is for smooth L1 regression, where α and β are weighting coefficients, α = 1.0 and β = 2.0.
[0120] The sparse self-attention module calculates the weights of the relationships between targets, and the self-attention module calculates the weights of the relationships between targets, thus modeling the relationships between targets:
[0121]
[0122] in, Let k represent the query vector for the i-th target. j Let A represent the key vector of the j-th target, and output A. ij Here, N represents the attention weight matrix, where N represents the number of all targets in the current computational scenario.
[0123] A GAT tracking head is constructed, and a bipartite graph matching cost matrix is built during the data association stage, integrating motion consistency (Mahathano distance metric) and appearance similarity (cosine distance metric). The state equation of the extended Kalman filter is modeled as a constant steering rate model, and the process noise covariance matrix is adaptively estimated. The graph attention network is designed as a two-layer structure: the first layer aggregates the motion features of single-hop neighbors, and the second layer captures group interaction patterns. Edge features encode the relative distance and velocity angle.
[0124] A motion consistency and appearance similarity cost matrix is constructed based on an improved Hungarian algorithm; data association is performed using the improved Hungarian algorithm, and the cost matrix includes motion consistency and appearance similarity terms:
[0125]
[0126] in, This represents the position of the i-th predicted target. Let represent the position of the j-th measurement target, λ be the weighting coefficient (taken as 0.7), p represent the position, f be the appearance feature vector, and cos(·) represent the cosine similarity. i,f j They represent the i-th and j-th prediction targets, respectively.
[0127] The extended Kalman filter (EPF) uses a constant steering rate model to predict the target state; motion prediction also uses the EPF, with the prediction phase based on the motion model to predict the state at the next moment, and the update phase using measured values to correct the prediction.
[0128]
[0129] in, Let f(·) represent the target state vector (position, velocity, etc.), and let f(·) represent the nonlinear state transition function. pn represents the state transition Jacobian matrix. k Q represents process noise. k P represents the process noise covariance. k Let P represent the state covariance matrix. k-1 Let represent the state covariance matrix at time k-1.
[0130] A two-layer graph attention network aggregates single-hop neighbor features and group interaction patterns, and uses the graph attention network GAT for interaction modeling.
[0131]
[0132] Where, h′ i h is the processed output feature vector. j Let N(i) represent the features of node j, N(i) be the summation of all neighboring nodes j of node i, W be the learnable weight matrix, and α be the weight matrix. ij This represents the attention coefficient.
[0133] Step S104: Train the model using a multi-task joint loss function, and optimize the model by combining adaptive loss balancing, knowledge distillation, and cosine annealing learning rate scheduling.
[0134] Specifically, to address the gradient conflict problem in multi-task learning, a joint optimization strategy based on adaptive loss balancing and knowledge distillation is constructed. A multi-objective function incorporating detection loss, regression loss, and interaction loss is designed, and dynamic weight adjustment is used to balance the gradients of each task. A teacher-student architecture is introduced, and knowledge from the multimodal teacher network is transferred to the pure radar student network through feature distillation. A cosine annealing learning rate scheduler with preheating is used to optimize the training process, combined with stochastic depth regularization to prevent overfitting. The data pipeline constructed in S101-S103 is integrated into the training framework along with the model architecture, and network parameters are updated through end-to-end backpropagation.
[0135] A multi-task joint loss function is designed, and a hierarchical loss function system is constructed: the detection branch adopts an improved FocalLoss, and the parameters are dynamically adjusted according to the target density; the regression loss introduces the Smooth L1 loss with directional decomposition, and the horizontal / vertical errors are weighted separately; the interaction consistency loss is achieved through graph comparison learning, where positive samples are temporally continuous trajectory node pairs and negative samples are spatially close non-associated targets.
[0136] A joint loss function is constructed based on the PyTorch framework, including the object detection loss L. det Trajectory regression loss L reg and interaction consistency loss L int Three parts. The target detection loss uses an improved Focal Loss: Wherein, p i This represents the confidence score of the i-th predicted box, γ is the focus factor (set to 2.0), and α... i Balance the weights for each category.
[0137] The trajectory regression loss uses adaptive IoU loss: L reg =1-IoU(B pred B gt )+λ||Δv||;where, where B pred and B gt Let represent the predicted bounding box and the ground truth bounding box, respectively; Δv represents the velocity vector error; and λ is the balance coefficient.
[0138] Interaction consistency loss is calculated using a graph attention network: L int =∑ i,j ||GAT(ft i )-GAT(ft j )||2; where GAT(·) represents the graph attention layer output, ft i and ft j These are target features in adjacent frames.
[0139] The AdamW optimizer (implemented via torch.optim.AdamW) is used for parameter updates. To balance convergence speed and final accuracy, a combination of linear warm-up and cosine annealing strategies is employed for learning rate scheduling.
[0140]
[0141] Where, η v Let η be the learning rate at step v. base The base learning rate is T, where T is the total number of parameter updates during the entire model training process. warm η is the number of preheating steps. max and η min This represents the maximum and minimum learning rates.
[0142] The model regularization employs the Stochastic Depth technique, with a probability p. l Randomly skip the l-th layer of the network:
[0143]
[0144] Where, x l+1 x represents the output feature after processing at layer l. l Let l be the input feature tensor of the l-th layer. This represents the residual transformation of the l-th layer.
[0145] Step S105: Deploy the optimized model to an edge computing device to build a complete processing link from radar signal input to trajectory output, output the target trajectory and motion status and visualize it to achieve real-time multi-target perception and tracking.
[0146] Specifically, such as Figure 4 As shown, the trained model is converted into a specific hardware inference format, and a multi-threaded processing pipeline is designed to achieve parallel execution of data acquisition, feature extraction, and model inference. A visual interactive interface is developed to display the target trajectory and motion status in real time, and a network interface is provided to support data exchange with autonomous driving or security systems. The optimized model from step S104 is deployed to an edge computing device, constructing a complete processing chain from radar signal input to trajectory output.
[0147] A complete radar perception system is built based on an embedded AI computing platform, performing quantization compression and hardware adaptation optimization on the trained deep learning model. A model conversion tool is used to convert the PyTorch model into an inference format supported by the specific hardware platform, achieving computation graph optimization and operator fusion. To address the real-time requirements of UWB radar data streams, a dedicated memory management mechanism is employed to ensure efficient pipeline operation of data acquisition, preprocessing, and model inference. A multi-threaded scheduling module is developed to rationally allocate computing resources. The data acquisition thread receives and buffers the raw radar signal, the preprocessing thread performs signal filtering and feature extraction, the inference thread executes the neural network forward computation, and the post-processing thread performs target association and trajectory prediction. A system runtime monitoring module continuously tracks the latency and resource usage of each stage, dynamically adjusting task priorities to ensure the real-time performance of the critical path.
[0148] A complete radar sensing data processing pipeline is constructed, starting from receiving raw radar echo signals at the hardware interface layer. Through multiple functional modules including signal analysis, noise suppression, target detection, data association, and state prediction, it ultimately outputs stable and reliable target trajectory information. The system employs a time-window-based batch processing mechanism to balance real-time performance and computational efficiency, maximizing target tracking continuity while ensuring minimal latency. A dedicated motion state prediction module is designed, fusing current observation data and historical trajectory information, and compensating for positional jitter caused by radar measurement noise through a kinematic model. To address target occlusion issues in complex scenarios, a prediction and compensation algorithm based on motion inertia is introduced to maintain trajectory continuity even when the target is briefly lost. The system output includes complete state information for each target, such as a unique identifier, spatial coordinates, velocity, and orientation angle, along with a tracking confidence level.
[0149] A layered rendering technique is employed to visualize radar sensing results. The core interface displays a real-time spatial distribution map within the radar scan range, using different colors and shapes to distinguish various targets, and dynamically drawing target trajectories and predicted paths. The side panel provides detailed information on the target list, supporting interactive viewing of specific target attribute data and historical trajectory playback. Multiple visualization modes are designed, including raw point cloud display, target clustering view, and heatmap, to meet the observation needs of different application scenarios. The system supports multi-view collaborative display, enabling spatiotemporal alignment and overlay display of radar sensing results with other sensor data (such as camera footage). A recording and playback function has been developed, supporting the loading and analysis of historical data and the reproduction of typical scenarios, facilitating system debugging and performance evaluation.
[0150] It provides multiple interface protocols for interfacing with upper-layer application systems. A dedicated middleware layer is developed to encapsulate the core functions of the radar perception system and provide a unified API interface. It supports remote calls based on network protocols, enabling real-time push of perception data and dynamic configuration of system parameters. For autonomous driving applications, message formats and communication protocols conforming to the requirements of the autonomous driving framework are designed to ensure seamless integration with the planning and control module. An alarm event triggering mechanism is developed for security monitoring systems, proactively pushing alarm information when abnormal motion patterns or specific types of targets are detected. The system supports flexible expansion of functional modules, allowing the integration of additional analysis algorithms and business logic through a plug-in mechanism. Comprehensive development documentation and sample code are provided to reduce the integration difficulty of third-party systems.
[0151] This example implementation, in its second aspect, provides a UWB radar multi-target sensing and tracking system. For example... Figure 5 As shown, the UWB radar multi-target sensing and tracking system includes:
[0152] The data acquisition module collects raw echo signals from multiple scenes using UWB radar, and simultaneously obtains the target's real location information using an optical motion capture system. After noise injection and time-frequency transformation to enhance data diversity, a multimodal dataset containing time-domain, frequency-domain, and spatial-domain features is constructed.
[0153] The signal processing module uses a three-level cascaded processing flow to denoise and enhance the original echo. It extracts the basic features in the time and frequency domain through wavelet transform and short-time Fourier transform, and outputs a high-dimensional feature tensor using a hybrid encoder that includes CNN and Transformer.
[0154] The detection and tracking module, based on graph neural network and dynamic filtering theory, constructs a model that includes an improved PointPillars detection head and a GAT tracking head to perform target position estimation and trajectory correlation.
[0155] The training optimization module uses a multi-task joint loss function to train the model, and combines adaptive loss balancing, knowledge distillation and cosine annealing learning rate scheduling optimization model.
[0156] The edge deployment module deploys the optimized model to edge computing devices, builds a complete processing link from radar signal input to trajectory output, outputs target trajectory and motion status and visualizes it, and realizes real-time multi-target perception and tracking.
[0157] Existing solutions are typically designed for specific scenarios and struggle to adapt to diverse application needs. The innovation of this invention lies in constructing a highly configurable, general-purpose framework that supports rapid adaptation to different application scenarios through modular design. It possesses online learning capabilities, allowing for continuous optimization of model parameters to adapt to new environments and target features, demonstrating excellent generalization performance.
[0158] To evaluate the effectiveness of the proposed method, this experiment collected dynamic target data for 10 consecutive days in complex indoor and outdoor scenarios using a deployed UWB radar multi-target perception system. These scenarios included single-person walking, multi-person interaction, and static obstacle occlusion. Data acquisition employed UWB radar at a sampling frequency of 100Hz, and the OptiTrack optical motion capture system (accuracy ±0.5mm) was used simultaneously to acquire the target's true location information as a baseline label. The experimental dataset was divided into training and test sets in an 8:2 ratio.
[0159] In terms of model performance evaluation, verification was conducted from three dimensions: target detection accuracy, multi-target tracking stability, and real-time performance. For the target detection task, the following evaluation metrics were used for analysis:
[0160] Mean Absolute Error (MAE): Measures the absolute deviation between the predicted target position and the actual position, reflecting the positioning accuracy;
[0161] Detection recall: Evaluates the model's ability to detect true targets;
[0162] False Alarm Rate (FAR): Reflects the likelihood of the model misidentifying false targets.
[0163] Table 1 Experimental Results of Target Detection Performance
[0164] Target type MAE(m) Recall (%) FAR (%) Static target 0.08 98.2 1.5 Dynamic goals 0.12 96.7 2.1
[0165] For multi-target tracking tasks, the following metrics are used to evaluate tracking stability:
[0166] Trajectory Completeness (TC): Measures the continuity and completeness of a target trajectory;
[0167] Maximum Track Duration (MTD): Evaluates the model's ability to continuously track a target.
[0168] Table 2 Experimental results of multi-target tracking performance
[0169] Scene complexity TC (%) MTD(s) Single-player scene 99.1 60.5 Multi-person interaction 90.3 45.7
[0170] Experimental results show that the proposed method achieves high accuracy (MAE < 0.12 m) and low false detection rate (FAR < 2.1%) in target detection tasks, exhibits good stability (TC > 90%) in multi-target tracking tasks, and meets real-time processing requirements. The comprehensive performance of the system in complex scenarios verifies its effectiveness and reliability in practical applications.
[0171] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0172] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. In addition, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.
[0173] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.
Claims
1. A UWB radar multi-target sensing and tracking method, characterized in that, Includes the following steps: The original echo signals from multiple scenes are collected by UWB radar, and the real position information of the target is obtained by synchronous optical motion capture system. The data diversity is enhanced by noise injection and time-frequency transformation, and a multimodal dataset containing time domain, frequency domain and spatial domain features is constructed. A three-level cascaded processing flow is used to denoise and enhance the signal of the original echo. The basic features in the time and frequency domain are extracted by wavelet transform and short-time Fourier transform. A high-dimensional feature tensor is output by a hybrid encoder containing CNN and Transformer. Based on graph neural network and dynamic filtering theory, a model including an improved PointPillars detection head and a GAT tracking head is constructed to perform target position estimation and trajectory correlation. The model is trained using a multi-task joint loss function, and optimized by adaptive loss balancing, knowledge distillation, and cosine annealing learning rate scheduling. The optimized model is deployed to edge computing devices to build a complete processing link from radar signal input to trajectory output, outputting target trajectory and motion status and visualization, thereby realizing real-time multi-target perception and tracking.
2. The UWB radar multi-target perception and tracking method according to claim 1, characterized in that, The steps of acquiring raw echo signals from multiple scenes using UWB radar, simultaneously obtaining the target's true location information using an optical motion capture system, enhancing data diversity through noise injection and time-frequency transformation, and constructing a multimodal dataset containing time-domain, frequency-domain, and spatial-domain features include: Collect static target grid distribution data and dynamic target preset motion pattern data in multiple indoor and outdoor scenarios; The number of targets, their 3D positions, velocity vectors, and trajectory IDs are labeled, and data augmentation is performed through noise injection, temporal pruning, and frequency domain filtering. Spectral and spatial features are extracted using short-time Fourier transform and beamforming algorithms.
3. The UWB radar multi-target perception and tracking method according to claim 2, characterized in that, The three-level cascaded processing flow includes: The input raw echo signal undergoes adaptive median filtering to remove impulse noise; An LMS adaptive filter is used to eliminate environmental multipath interference; The z-score method is used for normalization to eliminate the dimensional differences in signal amplitude; Principal component analysis was performed using scikit-learn to retain the principal components corresponding to energy and remove redundant features in noisy or highly correlated frequency bands.
4. The UWB radar multi-target perception and tracking method according to claim 3, characterized in that, The hybrid encoder comprising CNN and Transformer includes: CNN uses a variant of ResNet18 to extract local scattering features; A multi-head self-attention mechanism is used to model long-range dependencies across frames and output a spatiotemporal feature tensor.
5. The UWB radar multi-target perception and tracking method according to claim 4, characterized in that, The model constructed, which includes an improved PointPillars detection head and a GAT tracking head, includes: The UWB radar point cloud is projected onto the BEV plane, and a non-uniform mesh is generated by dynamic voxelization. PointNet++ is used to extract local geometric features, and the SSD detection head is used to fuse multi-scale receptive fields; The sparse self-attention module calculates the weights of the relationships between targets.
6. The UWB radar multi-target perception and tracking method according to claim 5, characterized in that, The model construction, which includes the improved PointPillars detection head and GAT tracking head, also includes: A motion consistency and appearance similarity cost matrix is constructed based on an improved Hungarian algorithm; The extended Kalman filter uses a constant turning rate model to predict the target state; A two-layer graph attention network aggregates single-hop neighbor features and group interaction patterns.
7. The UWB radar multi-target perception and tracking method according to claim 6, characterized in that, The steps of training the model using a multi-task joint loss function, and optimizing the model by combining adaptive loss balancing, knowledge distillation, and cosine annealing learning rate scheduling include: A multi-task joint loss function is constructed based on the PyTorch framework, including object detection loss function, trajectory regression loss function and interaction consistency loss function; The AdamW optimizer is used for parameter updates, and the learning rate scheduling adopts a combination of linear preheating and cosine annealing strategies.
8. The UWB radar multi-target perception and tracking method according to claim 7, characterized in that, The target detection loss function employs an improved Focal Loss: Where, p i Let represent the confidence score of the i-th predicted box, γ be the focus factor (set to 2.0), and α be the focus factor. i Balance the weights for each category; The trajectory regression loss function uses adaptive IoU loss: L reg =1-IoU(B pred B gt )+λ||Δv||;where, B pred and B gt These represent the predicted bounding box and the ground truth bounding box, respectively; Δv represents the velocity vector error; and λ is the balance coefficient. The interaction consistency loss function is calculated using a graph attention network: L int =∑ i,j ||GAT(ft i )-GAT(ft j )||2; where GAT(·) represents the graph attention layer output, ft i and ft j These are target features in adjacent frames.
9. The UWB radar multi-target perception and tracking method according to claim 8, characterized in that, The expression for the learning rate scheduling strategy employing a combination of linear preheating and cosine annealing is as follows: Where, η v Let η be the learning rate at step v. base The base learning rate is T, where T is the total number of parameter updates during the entire model training process. warm η is the number of preheating steps. max and η min These represent the maximum and minimum learning rates.
10. A UWB radar multi-target sensing and tracking system, characterized in that, The system is used to perform the method as described in any one of claims 1 to 9, wherein the UWB radar multi-target sensing and tracking system comprises: The data acquisition module collects raw echo signals from multiple scenes using UWB radar, and simultaneously obtains the target's real location information using an optical motion capture system. After noise injection and time-frequency transformation to enhance data diversity, a multimodal dataset containing time-domain, frequency-domain, and spatial-domain features is constructed. The signal processing module uses a three-level cascaded processing flow to denoise and enhance the original echo. It extracts the basic features in the time and frequency domain through wavelet transform and short-time Fourier transform, and outputs a high-dimensional feature tensor using a hybrid encoder that includes CNN and Transformer. The detection and tracking module, based on graph neural network and dynamic filtering theory, constructs a model that includes an improved PointPillars detection head and a GAT tracking head to perform target position estimation and trajectory correlation. The training optimization module uses a multi-task joint loss function to train the model, and combines adaptive loss balancing, knowledge distillation and cosine annealing learning rate scheduling optimization model. The edge deployment module deploys the optimized model to edge computing devices, builds a complete processing link from radar signal input to trajectory output, outputs target trajectory and motion status and visualizes it, and realizes real-time multi-target perception and tracking.
Citation Information
Patent Citations
Target tracking method and device based on millimeter wave radar and laser radar
CN114690174A
Roadside sensing method and system with video target fused with laser radar target
CN118736506A
Cited By
Tunnel millimeter wave radar target detection method based on beam forming
CN121613423A