Non-vision-field target detection and tracking method and system based on millimeter wave radar
By employing a non-line-of-sight target detection and tracking method based on millimeter-wave radar, and utilizing the specular reflection principle of virtual detection points and feature enhancement modules, the real-time performance and accuracy issues of non-line-of-sight detection and tracking in existing technologies are resolved. This achieves low-cost and efficient target identification and localization, making it suitable for autonomous driving and national defense counter-terrorism.
Patent Information
- Application Number
- CN202510946567.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-10-17
AI Technical Summary
Existing non-line-of-sight detection or tracking methods have shortcomings in real-time and accuracy, especially in outdoor environments where they are difficult to adapt to large-scale dynamic scenes. In addition, the equipment is expensive, the scanning speed is slow, and the imaging cycle is long.
A non-line-of-sight target detection and tracking method based on millimeter-wave radar is adopted. By emitting electromagnetic waves and receiving reflected signals, virtual detection points are determined by combining relay wall information and radar position to generate pseudo-images. Feature enhancement modules such as Swin Transformer network and CDPA network are used to construct a non-line-of-sight detection model to achieve multi-scale feature extraction, fusion and classification.
It achieves low-cost, real-time, and widely applicable non-line-of-sight target detection and tracking, enabling rapid and accurate target location and identification in large-scale dynamic outdoor environments, thus improving the safety of autonomous driving and national defense counter-terrorism.
Smart Images

Figure CN120802260A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of computer vision, and particularly relates to a non-visual field target detection and tracking method and system based on a millimeter wave radar. BACKGROUND
[0002] Traditional perception methods rely on devices such as cameras, lidar, and ultrasonic sensors, and their detection capabilities are essentially limited to the visual range or direct line-of-sight range of the sensors, resulting in existing computer vision methods that can only detect and track targets within the direct line of sight and cannot detect objects obscured by scene obstacles, which may pose a safety hazard in driving safety systems and security monitoring scenarios. Non-visual field (NLOS) methods restore the information of these obscured objects by analyzing the reflected or projected light or electromagnetic signals to the visible relay surface, enabling the detection and tracking of objects that are otherwise invisible. In the field of counter-terrorism, non-visual field methods can enable counter-terrorism personnel to detect hidden threats in advance, and in the field of intelligent driving, non-visual field methods can enhance the proactive safety system of the vehicle to detect "ghost probes" in advance, enhancing navigation safety and avoiding collisions.
[0003] Currently, a common non-visual field method uses ultrafast pulsed lasers and time-resolved photon detectors to collect light transient signals and measure their time-resolved echoes for non-visual field target detection. However, the device is expensive, the scanning speed is slow, and the imaging cycle is long, making it difficult to apply to real-time scenarios. Moreover, the intensity decreases by a quarter with the distance to the visible relay wall, and the imaging distance is limited to a meter-level scene, making this method limited to controlled laboratory environments. Its accuracy and real-time performance in non-visual field detection or tracking in outdoor environments are poor, and it cannot adapt to large-scale dynamic outdoor environments. SUMMARY
[0004] The present application aims to provide a non-visual field target detection and tracking method and system based on a millimeter wave radar to solve the problem of poor real-time performance and accuracy of existing non-visual field detection or tracking methods.
[0005] The application provides a non-visual field target detection method based on a millimeter wave radar to solve the above technical problems, comprising: using the millimeter wave radar to emit electromagnetic waves to a scene where a non-visual field target is located and receive reflected signals; extracting radar point cloud data, wherein the radar point cloud data comprises position, speed and amplitude; combining relay wall information and millimeter wave radar position to determine whether each point cloud is a virtual detection point, and if it is a virtual detection point, the point cloud data of a real detection point corresponding to the virtual detection point is calculated by using a mirror reflection principle; generating a pseudo image based on the point cloud data of each real detection point, inputting the pseudo image into a non-visual field detection model to obtain target category and target positioning information; the non-visual field detection model comprises a feature extraction module, a feature enhancement module, a feature fusion module and a detection head; the feature extraction module is used for performing at least two scale feature extractions on the pseudo image; the feature enhancement module is used for performing feature enhancement on at least one scale feature; the feature fusion module is used for fusing all scale features after feature enhancement and the remaining scale features except the scale after feature enhancement; and the detection head is used for processing the fused features to output target category and target positioning information.
[0006] Further, the speed and amplitude of the real detection point are the speed and amplitude of the corresponding virtual detection point, and the position of the real detection point is calculated according to the following formula:
[0007]
[0008] Wherein, x is the position of the real detection point corresponding to the virtual detection point; x' is the position of the virtual detection point; c is the position of the millimeter wave radar receiver; n w is a relay wall normal vector, pointing away from the millimeter wave radar receiver; w is the intersection position of the virtual detection point and the relay wall.
[0009] Further, the intersection position of the virtual detection point and the relay wall is calculated according to the following formula:
[0010]
[0011] Wherein, p1 and p2 are respectively two end points corresponding to the maximum visible range of the millimeter wave radar on the relay wall; p = p2-p1, indicating that p1 points to the vector of p2.
[0012] Further, the type of the feature enhancement module is a Swin Transformer network, a CDPA network or a CAM network; wherein the CDPA network comprises a spatial attention module and a channel attention module, the input corresponding scale feature is processed by the spatial attention module to obtain a first feature, the input corresponding scale feature is processed by the channel attention module to obtain a second feature, and the first feature and the second feature are multiplied with the input corresponding scale feature element by element to obtain the output value of the CDPA network, that is, the feature after feature enhancement; the CAM network comprises a Swin Transformer network and a CDPA network, the input corresponding scale feature is processed by the Swin Transformer network to obtain a first output value, the input corresponding scale feature is processed by the CDPA network to obtain a second output value, and the first output value and the second output value are added element by element to obtain the output value of the CAM network, that is, the feature after feature enhancement.
[0013] Further, in the non-visual field detection model, each scale feature extracted by the feature extraction module is subjected to feature enhancement, and the feature enhancement module used for feature enhancement is of the same type or different types.
[0014] Further, the detection head is an SSD, each anchor frame of the SSD comprises a target category and a boundary regression frame, and the boundary regression frame is used to represent target positioning information, including the center position, size, direction and speed of the frame.
[0015] To solve the above technical problems, the application further provides a non-visual field target detection system based on a millimeter wave radar, comprising a millimeter wave radar and a processor, and the system is used to realize the non-visual field target detection method based on the millimeter wave radar.
[0016] The beneficial effects of the above technical solution are as follows: This invention is a pioneering creation that utilizes a frequency-modulated continuous-wave millimeter-wave radar with a multi-input, multi-output array structure for non-line-of-sight sensing. The millimeter-wave radar signal is processed to generate a large amount of radar point cloud data. This point cloud data includes virtual detection points corresponding to the target's third reflection on a relay wall. Therefore, the point cloud data is first judged by combining relay wall information and the millimeter-wave radar position, and the point cloud data of the real detection points corresponding to the virtual detection points are calculated. The point cloud data corresponding to these real detection points are then converted into pseudo-images. A non-line-of-sight detection model is then constructed, including multi-scale feature extraction, feature enhancement, feature fusion, and classification detection. This model can integrate feature information at different scales, enhance the representation of key targets, suppress background noise interference, enable the model to more accurately locate key information, and quickly and accurately obtain target classification and positioning information. Furthermore, the detection cost is low, it is not heavily dependent on specific environmental conditions (such as light source angle and surface material), and it is adaptable to large-scale dynamic outdoor environments. This achieves low-cost, real-time, widely applicable, and highly accurate NLOS target detection, which is crucial for promoting the development of fields such as autonomous driving, national defense counterterrorism, and road safety.
[0017] In order to solve the above technical problems, the present invention also provides a non-line-of-sight target tracking method based on millimeter-wave radar, comprising: using a millimeter-wave radar to transmit electromagnetic waves to the scene where the non-line-of-sight target is located and receiving reflected signals; extracting radar point cloud data of the current frame and the previous n frames, the radar point cloud data including position, speed and amplitude, n>1; combining the relay wall information and the millimeter-wave radar position to determine whether each point cloud is a virtual detection point, and if it is a virtual detection point, using the mirror reflection principle to calculate the point cloud data of the real detection point corresponding to the virtual detection point; based on the point cloud data of the real detection point of the current frame and the previous n frames, generating corresponding pseudo images of the current frame and the previous n frames; inputting the generated n+1 frames of pseudo images into the non-line-of-sight tracking model to obtain the target category, target positioning information of the current frame and the unrealized target location information. The target positioning information of the incoming frame; the non-viewing area tracking model includes a feature extraction module, a feature connection module, a feature enhancement module, a feature fusion module and a detection head; there are n+1 feature extraction modules, each of which is used to extract at least two scale features from the pseudo image of each frame; the feature connection module is used to splice n+1 features of the same scale along the feature channel to obtain spliced features of each scale; the feature enhancement module is used to enhance the features of at least one scale spliced feature; the feature fusion module is used to fuse all the spliced features of each scale after feature enhancement and the spliced features of the remaining scales except the scale for feature enhancement; the detection head is used to process the fused features and output the target category, the target positioning information of the current frame and the target positioning information of the future frame.
[0018] Further, the type of the feature enhancement module is a Swin Transformer network, a CDPA network or a CAM network; the CDPA network comprises a spatial attention module and a channel attention module, the input corresponding scale feature is processed by the spatial attention module to obtain a first feature, the input corresponding scale feature is processed by the channel attention module to obtain a second feature, and the first feature and the second feature are multiplied element by element with the input corresponding scale feature to obtain the output value of the CDPA network, that is, the feature after feature enhancement; the CAM network comprises a Swin Transformer network and a CDPA network, the input corresponding scale feature is processed by the Swin Transformer network to obtain a first output value, the input corresponding scale feature is processed by the CDPA network to obtain a second output value, and the first output value and the second output value are added element by element to obtain the output value of the CAM network, that is, the feature after feature enhancement.
[0019] To solve the above technical problems, the application further provides a non-line-of-sight target tracking system based on a millimeter wave radar, comprising a millimeter wave radar and a processor, and the system is used to realize the non-line-of-sight target tracking method based on the millimeter wave radar.
[0020] The beneficial effects of the above technical solution are as follows: the application is an opening-type invention, the millimeter wave radar signal is processed to obtain point cloud data comprising a virtual detection point corresponding to the third reflection of the target on the relay wall, the point cloud data is judged in combination with the relay wall information and the millimeter wave radar position, and the point cloud data of the real detection point corresponding to the virtual detection point is calculated, the point cloud data of the current frame and the previous n frames is converted into a pseudo image corresponding to the frame. A non-line-of-sight tracking model comprising multi-scale feature extraction of each frame pseudo image, splicing of all frame extracted features, feature enhancement, feature fusion and classification detection is constructed. Thus, the time sequence information of multiple frames of data in multiple scales is aggregated, the detail information and motion features in the frame are captured, the background noise interference is suppressed, the representation of the key target is enhanced, the target category, the current frame and the positioning information of the future frame can be quickly and accurately obtained, the tracking cost is low, the specific environmental conditions (such as light source angle and surface material) are not seriously dependent, the large-scale dynamic outdoor environment can be adapted, the NLOS target tracking with low cost, strong real-time performance, wide application range and high accuracy is realized, and the development of the fields of automatic driving, national defense and anti-terrorism and road safety is crucial. BRIEF DESCRIPTION OF DRAWINGS
[0021] Figure 1 is a non-line-of-sight target detection flowchart of an embodiment of the non-line-of-sight target detection method of the application;
[0022] Figure 2is a radar non-visual observation model schematic diagram of an embodiment of the non-visual target detection method of the present application;
[0023] Figure 3 is a virtual detection point judgment principle schematic diagram of an embodiment of the non-visual target detection method of the present application;
[0024] Figure 4 is a non-visual target detection model diagram of an embodiment of the non-visual target detection method of the present application;
[0025] Figure 5 is a second non-visual target detection model diagram of an embodiment of the non-visual target detection method of the present application;
[0026] Figure 6 is a third non-visual target detection model diagram of an embodiment of the non-visual target detection method of the present application;
[0027] Figure 7 is a non-visual target tracking model diagram of an embodiment of the non-visual target tracking method of the present application;
[0028] Figure 8 is another non-visual target tracking model diagram of an embodiment of the non-visual target tracking method of the present application. DETAILED DESCRIPTION
[0029] In order to make the objects, technical solutions and advantages of the present application more clear and more apparent, the specific embodiments of the present application will be further described below with reference to the drawings.
[0030] The present application uses millimeter wave radar to emit electromagnetic waves and receive echo information to obtain point cloud data, reconstructs the point cloud data of hidden targets to generate pseudo images, constructs a network model including feature extraction, feature enhancement and detection heads to realize detection and tracking of hidden targets in large-scale dynamic scenes.
[0031] Non-visual target detection method embodiment
[0032] A non-visual target detection method based on millimeter wave radar of the present application, as shown in Figure 1 , includes the following steps:
[0033] 1. Use millimeter wave radar to emit electromagnetic waves to the scene where the non-visual target is located and receive reflected signals.
[0034] For millimeter wave radar, the electromagnetic waves have undergone multiple reflections in the transmission process, resulting in the appearance of virtual detection targets. When the radar signal is reflected from the relay wall to the hidden object in the non-visual area, part of the signal will be scattered and reflected back to the observable relay wall by the hidden object, and then the scattered signal is reflected back to the receiver from the relay wall, the whole process is as shown in Figure 2 . Figure 2In the middle, the radar transmits a signal from the transmitter to the relay wall, reflects to the real target through the relay wall as the first reflection, from the relay wall to the real target, reflects to the relay wall through the real target as the second reflection, and reflects to the radar receiving end through the relay wall from the real target as the third reflection. C represents the position of the radar receiver, x' is the position of the virtual detection point, x is the position of the real detection point corresponding to the virtual detection point, w is the intersection position of the virtual detection point and the relay wall, n is the normal vector of the real detection point, v is the moving speed, v r The millimeter wave radar used in the present application is a frequency modulation continuous wave (FMCW) radar with a multiple-input multiple-output (MIMO) array configuration. The radar transmitter transmits electromagnetic waves by linearly scanning from the carrier frequency as the starting point across the entire frequency band, and the radar receiver receives the reflected signals to obtain the original radar data.
[0035] 2. Extract the radar point cloud data.
[0036] The radar transmits and receives multiple frames of data within one second, processes the radar data, extracts the point cloud data of the radar, and represents u as a d-dimensional point in the radar point cloud. The point cloud data includes position, speed, and amplitude. For example, in one embodiment, the point cloud data includes X-axis coordinates, Y-axis coordinates, speed (i.e., radial Doppler velocity), and amplitude. At this time, d = 4.
[0037] 3. Determine whether each point cloud is a virtual detection point by combining the relay wall information and the position of the millimeter wave radar, and if it is a virtual detection point, calculate the point cloud data of the real detection point corresponding to the virtual detection point.
[0038] 1) Relay wall estimation.
[0039] The scene is scanned using a laser radar, and the laser radar point cloud data obtained after scanning is projected onto a two-dimensional plane to generate a bird's eye view (BEV). Based on the BEV, a scene model composed of line segments is generated, and line segments with a length less than 1 meter are filtered out, only long line segments that may represent a wall surface are retained, thereby removing noise and irrelevant short line segments, and ensuring that the detected line segments indeed represent a relay wall. Endpoints p1 and p2 are extracted from the detected line segments, p1 and p2 are respectively two endpoints corresponding to the maximum visible range of the millimeter wave radar on the relay wall, which define the geometric shape of the relay wall. This process is used to restore the geometric shape of the relay wall, providing geometric information for subsequent third reflection estimation and target reconstruction, ensuring the efficiency and accuracy of the target detection method in complex environments, and improving the target positioning accuracy. As an embodiment, the laser radar can use a solid-state laser radar. Preferably, a non-repetitive scanning vehicle-grade hybrid solid-state laser radar can be selected, which has the advantages of low cost and miniaturization, and can be deployed in single-person equipment, directly based on vehicle-grade laser radar integration extension.
[0040] 2) Third reflection estimation, determining whether each point cloud is a virtual detection point.
[0041] In combination with the relay wall information and the millimeter wave radar position, it is determined whether each point cloud is a virtual detection point. Specifically, assuming that a certain point cloud is a virtual detection point, the intersection position w of the virtual detection point and the relay wall is first calculated:
[0042]
[0043] where p = p2 - p1, indicating a vector from p1 to p2; x represents a two-dimensional cross product.
[0044] As shown in Figure 3 , φ is the incidence angle of the radar signal to the relay wall, γ w is the radar emission range bisector (the angle between the radar emission antenna pointing direction and the n w direction); γx' is the angle between the third reflection signal and the radar emission bisector. For the virtual detection point of the third reflection, it must satisfy two criteria: first, the virtual detection point and the millimeter wave radar receiver must be on the opposite side of the relay wall; second, the intersection position w of the virtual detection point and the relay wall must be between p1 and p2. If the following judgment formula is true, it indicates that the point is a virtual detection point of the third reflection:
[0045] n w ·(x′-p1)≥0∧||w-p1||≤||p||∧||w-p2||≤||p||
[0046] where n wis the normal vector of the relay wall, pointing away from the millimeter wave radar receiver, and |||| represents the vector norm.
[0047] 3) If it is a virtual detection point, the point cloud data of the real detection point corresponding to the virtual detection point is calculated.
[0048] The speed and amplitude of the real detection point are the speed and amplitude of the corresponding virtual detection point, and the position of the real detection point is calculated according to the following formula:
[0049]
[0050] where c is the position of the millimeter wave radar receiver, and x is the position of the real detection point corresponding to the virtual detection point.
[0051] 4. Sparse pseudo-image generation.
[0052] Based on the point cloud data of each real detection point, a pseudo-image is generated. Specifically, the point cloud of the real detection point is discretized into a grid uniformly distributed on a plane, and the amplitude and speed are respectively assigned to the first channel and the second channel of the grid to generate a pseudo-image with a size of (d-2, H, W), where H and W represent the height and width of the grid, respectively, and d-2 represents the dimension after removing the speed and amplitude. By discretizing the point cloud into a grid, the data volume can be effectively reduced while retaining key spatial information.
[0053] In other embodiments, in order to enhance the visibility of weak signals and effectively extract target information in a low signal-to-noise ratio environment, the logarithm of the amplitude is used to obtain the intensity measurement value, and the point cloud data includes X-axis coordinates, Y-axis coordinates, speed, and intensity measurement value. The intensity measurement value and the speed are respectively assigned to the first channel and the second channel of the grid to generate a pseudo-image with a size of (d-2, H, W). For pixels into which multiple points fall after discretization, the speed values of all points within the pixel can be averaged, and the intensity values of all points within the pixel can be summed.
[0054] 5. Constructing a non-visual domain detection model.
[0055] The non-visual domain detection model includes a feature extraction module, a feature enhancement module, a feature fusion module, and a detection head; the feature extraction module is used to perform at least two scale feature extractions on the pseudo-image; the feature enhancement module is used to perform feature enhancement on at least one scale feature; the feature fusion module is used to fuse all scale features after feature enhancement and the remaining scale features except for the scale of feature enhancement; and the detection head is used to process the fused features to output target category and target positioning information. The feature extraction module can use a deep learning model for extracting different scale feature information, such as a CNN model, a ResNet model, or a Transformer model, without limitation.
[0056] The feature extraction module is configured to extract different feature scales, for example, 2, 3 or 4 units for extracting different feature scales, etc. Each unit is configured with at least one feature extraction layer, and 1, 2, 3 or 4 convolution layers can be configured in each feature extraction module according to requirements. The feature fusion module is configured with a corresponding scale transpose convolution, which up-samples the features of different scales to the same size and then connects them to form the final output.
[0057] The feature enhancement module is used to optimize feature expression, and existing CNN networks, CBAM networks, ViT networks, SE networks, ResNet networks or ASFF networks, etc. can be used, which are not limited here. As long as the network can realize the optimization of feature expression, it belongs to the protection scope of the present application. Preferably, the feature enhancement module can use Swin Transformer network, CDPA network or CAM network; wherein the CDPA network includes a spatial attention module and a channel attention module, the input corresponding scale feature is processed by the spatial attention module to obtain the first feature, the input corresponding scale feature is processed by the channel attention module to obtain the second feature, and the first feature and the second feature are multiplied with the input corresponding scale feature to obtain the output value of the CDPA network; the CAM network includes a Swin Transformer network and a CDPA network, the input corresponding scale feature is processed by the Swin Transformer network to obtain the first feature value, and the input corresponding scale feature is processed by the CDPA network to obtain the second output value, and the first output value and the second output value are added element by element to obtain the output value of the CAM network. The network structure of the spatial attention module and the channel attention module is the same as that of the CBAM.
[0058] In one embodiment, as shown in Figure 4 In order to efficiently encode the high-level representation of the hidden detection, the feature extraction module in this embodiment is configured to extract a pyramid network of two feature scales, which contains two consecutive units. Each unit uses three 2D convolution layers with a kernel size of 3x3 to down-sample its input feature map by two to obtain the first scale feature and the second scale feature. The first scale feature is input to the feature enhancement module for feature enhancement.
[0059] In this embodiment, the feature enhancement module adopts a Swin Transformer network, including Window Partition, linear embedding and Swin Transformer Block. The Window Partition sets the window size to 30*40, so the feature map can be divided into 100 windows. The linear embedding converts the original pixel value of the input image into a higher-dimensional feature representation, which can be more effectively processed by the subsequent Swin Transformer Block. The Swin Transformer Block includes W-MSA (Window Multi-Head Self-Attention) and SW-MSA two parts, which are connected in series, W-MSA is a window-based attention calculation, which calculates the self-attention score of each window, and limits the attention calculation within each window, thereby reducing the amount of calculation. And SW-MSA is to recalculate the attention after the window slides, which realizes the cross-window communication through the shifted window design. Work with the W-MSA module to achieve effective self-attention calculation within the local window and across windows. By performing self-attention calculation within each window, the SW-MSA module significantly reduces the amount of calculation. In consecutive SW-MSA layers, the window partition is shifted so that each window partially overlaps with its neighboring windows. This design allows information to be transmitted between adjacent windows, enhancing the model's ability to capture global features. The SW-MSA module considers the relative position relationship between the tokens (local features in the image) within the window, and further enhances the model's ability to perceive spatial relationships by introducing a relative position bias. The SW-MSA and W-MSA modules are used alternately, where W-MSA performs self-attention on windows without offset, and SW-MSA performs self-attention on windows with offset, to effectively fuse local features and global features.
[0060] Next, the transposed convolution in the feature fusion module respectively up-samples the first scale feature after feature enhancement and the second scale feature without feature enhancement, and then splices them after the same size. The two transposed convolutions are transposed 2D convolution layers with different strides and kernel sizes of 3*3. The transposed 2D convolution layer is interleaved with BatchNorm and ReLU. Through the combination of the pyramid network and the amplification network, multi-scale features can be effectively extracted, and the model's detection ability for targets of different sizes is enhanced.
[0061] In another embodiment, the extracted first scale feature and second scale feature are both sent to the feature enhancement module for feature enhancement, and the feature enhancement module is a CDPA network, as shown in Figure 5As shown in the figure, the channel attention mechanism in the CDPA network focuses on the most discriminative channels in the feature map. By dynamically adjusting the contribution of each channel, the model can learn the importance of each channel, making the weight of important channels close to 1 and the weight of unimportant channels close to 0, thereby enhancing the feature representation of important channels, highlighting useful information, suppressing redundant information, and enabling the model to focus on key features. The spatial attention mechanism in the CDPA network focuses on information-intensive areas in the feature map. By selectively enhancing or weakening specific areas, the model can learn the importance of each spatial position. The weight of important positions is close to 1, and the weight of unimportant positions is close to 0, thereby enhancing the feature representation of key positions, suppressing background noise interference, and enabling the model to more accurately locate key information. The fusion of these two mechanisms can significantly improve the representation and generalization capabilities of the model. By cascading attention modules after the convolution blocks in the two stages, the model can more effectively learn and utilize key features at multiple scales.
[0062] In other embodiments, different types of feature enhancement modules are used to enhance the features of each scale extracted by the feature extraction module. That is, the two feature enhancement modules are CDPA network and CAM network, such as Figure 6 shown.
[0063] Detection head ( Figures 4-6 (not shown) can adopt one-stage detectors (One-Stage Detectors) or two-stage detectors (Two-Stage Detectors), two-stage detectors can adopt R-CNN, FAST R-CNN, FASTER R-CNN or Cascade R-CNN, etc., and one-stage detectors can adopt YOLO series or SSD (Single Shot MultiBoxDetector), etc., and customized settings are made according to user needs, which are not limited here. As a preferred embodiment, the detection head adopts SSD. When performing bounding box regression and classification tasks based on the SSD detection head, SSD performs two-dimensional target detection on the input feature data. Each predicted anchor box includes the target category (such as background or hostile target) and the bounding regression box. The bounding regression box is used to represent the target positioning information, including the 6-dimensional vector of the center position, size, direction and speed of the box. The detection head performs classification and positioning tasks simultaneously to ensure that the target category can be accurately identified and its position can be determined. The position and size of the bounding box are predicted by regression, and combined with the speed information, the precise positioning of the target is achieved.
[0064] A separate positioning system is used to capture training labels and training data to train the constructed non-line-of-sight detection model. The detection loss function used in training the non-line-of-sight detection model includes detection positioning loss and detection classification loss. The detection loss function is as follows:
[0065] L' = a'L' + b'L' loc + b'L' cls
[0066] wherein L' represents a detection loss function, a' and b' are corresponding weights, L' loc represents a detection positioning loss, L' cls represents a detection classification loss.
[0067] Non-visual target tracking method implementation
[0068] A non-visual target tracking method based on a millimeter wave radar according to the present application comprises the following steps: using a millimeter wave radar to emit electromagnetic waves to a scene where a non-visual target is located and receiving reflected signals; extracting radar point cloud data of a current frame and n previous frames, the radar point cloud data including position, speed and amplitude, n > 1; combining relay wall information and millimeter wave radar position to determine whether each point cloud is a virtual detection point, and if it is a virtual detection point, calculating point cloud data of a real detection point corresponding to the virtual detection point by using a mirror reflection principle. The implementation of the above steps is similar to the non-visual target tracking method implementation, and thus will not be described here again. Based on the point cloud data of the real detection points of the current frame and the n previous frames, pseudo images corresponding to the current frame and the n previous frames are generated; the n+1 generated pseudo images are input into a non-visual tracking model to obtain target categories, target positioning information of the current frame and target positioning information of future frames.
[0069] The non-visual tracking model comprises a feature extraction module, a feature connection module, a feature enhancement module, a feature fusion module and a detection head; the feature extraction module has n+1, each feature extraction module is used for performing at least two scale feature extractions on pseudo images of each frame; the feature connection module is used for splicing n+1 same scale features along a feature channel to obtain each scale splicing feature; the feature enhancement module is used for performing feature enhancement on at least one scale splicing feature; the feature fusion module is used for fusing all scale splicing features after feature enhancement and the remaining scale splicing features except for the scale of feature enhancement; and the detection head is used for processing the fused features to output target categories, target positioning information of the current frame and target positioning information of future frames.
[0070] In each feature extraction module, units for extracting different feature scales are arranged, for example, 2, 3 or 4 units for extracting different feature scales are arranged, etc. At least one feature extraction layer is arranged in each unit, and 1, 2, 3 or 4 convolution layers can be arranged in each feature extraction module according to requirements. The feature fusion module is provided with a corresponding scale transposed convolution, which up-samples features of different scales to the same size and then connects them together to form a final output.
[0071] The feature enhancement module is used to optimize feature expression. Existing CNN network, CBAM network, ViT network, SE network, ResNet network or ASFF network can be used, and the present application does not limit the network as long as it can realize the optimization of feature expression. Preferably, the feature enhancement module can use Swin Transformer network, CDPA network or CAM network. The CDPA network includes a spatial attention module and a channel attention module. The input corresponding scale feature is processed by the spatial attention module to obtain a first feature. The input corresponding scale feature is processed by the channel attention module to obtain a second feature. The first feature and the second feature are multiplied with the input corresponding scale feature element by element to obtain the output value of the CDPA network. The CAM network includes a Swin Transformer network and a CDPA network. The input corresponding scale feature is processed by the Swin Transformer network to obtain a first output value. The input corresponding scale feature is processed by the CDPA network to obtain a second output value. The first output value and the second output value are added element by element to obtain the output value of the CAM network. The network structure of the spatial attention module and the channel attention module is the same as that of the CBAM.
[0072] In one embodiment, as Figure 7As shown, in order to efficiently encode the high-level representation of concealment detection, each feature extraction module adopts a pyramid network, which contains two consecutive units to generate features with smaller and smaller spatial resolutions. That is, the unit for extracting different scales uses three 2D convolution blocks with a kernel size of 3x3 to down-sample the input feature map by a factor of two, so that the pseudo image corresponding to each frame obtains features at the first scale and the second scale. Subsequently, in the feature connection module, the features of all frames at the first scale and the second scale are respectively spliced in the channel dimension to obtain two feature maps with sizes of ((n+1)C1, H / 2, W / 2) and ((n+1)C2, H / 4, W / 4), C1 and C2 represent the feature dimensions after the first stage convolution block and the two consecutive stage convolution blocks, respectively. The feature map obtained by splicing at the first scale is input into the feature enhancement module for processing. In this embodiment, the feature enhancement module is a CDPA network. The features processed by the CDPA network are respectively sent to two up-sampling modules in the zoom-in network for up-sampling, and the two up-sampling modules perform transposed two-dimensional convolution with different strides, so as to up-sample the input features to the same size. All transposed convolutions use a kernel with a size of 3, and are interleaved with BatchNorm and ReLU. Then the features obtained after up-sampling are spliced in the feature fusion module. By fusing the features extracted from multiple frames of data at different scales, the model can aggregate the temporal information of multiple frames of data at multiple scales and input it into the attention module, so that the model can effectively learn and use the key features at different scales, while capturing the intra-frame detail information and motion features, thereby enhancing the representation of the key target.
[0073] In another embodiment, the feature maps spliced at each scale are input into the feature enhancement module, and the feature enhancement module is a CDPA network, as shown in Figure 8
[0074] The detection head Figure 7 and Figure 8 The detection head (not shown in the figure) can adopt One-Stage Detectors or Two-Stage Detectors, the Two-Stage Detectors can adopt R-CNN, FAST R-CNN, FASTER R-CNN or Cascade R-CNN, etc., the One-Stage Detectors can adopt YOLO series or SSD (Single Shot MultiBox Detector), etc., and the user can customize the settings according to the requirements, which are not limited here. As a preferred embodiment, the detection head adopts SSD, when the SSD detection head is used to perform the bounding box regression and classification tasks, the SSD performs two-dimensional target prediction on the input feature data, each anchor box of the prediction includes a target category (such as background or enemy target) and a bounding box regression, the bounding box regression is used to represent the target positioning information, including a 6-dimensional vector of the center position, size, direction and speed of the box. The detection head simultaneously performs the classification and positioning tasks, which ensures accurate identification of the target category and determination of the position. Through the regression prediction of the position and size of the bounding box, combined with the speed information, not only the accurate positioning of the target is realized, but also the motion trajectory of the target is predicted, which is suitable for target tracking in dynamic scenes.
[0075] The constructed non-visual field tracking model is trained by using separate positioning systems to capture training labels and training data pairs. The tracking loss function used when training the tracking model includes a tracking positioning loss and a tracking classification loss, and the tracking loss function is as follows:
[0076] L = aL loc + bL cls
[0077] Wherein, L represents the tracking loss function, a and b are corresponding weights, L loc represents the tracking positioning loss, and L cls represents the tracking classification loss.
[0078] The tracking positioning loss is the sum of the local positioning losses of the current frame T and the future n frames:
[0079]
[0080] Wherein, represents the local positioning loss of the t-th frame, a represents a coefficient, x a , y a , w a , l, and q represent the X-axis coordinate, Y-axis coordinate, width, height, and angle of the bounding box, respectively, u' represents the bounding box, and Au' represents the difference between the true value and the predicted value.
[0081] Non-visual field target detection system implementation
[0082] A non-line-of-sight target detection system based on millimeter wave radar according to the present application comprises a millimeter wave radar and a processor, and is used to implement a non-line-of-sight target detection method based on millimeter wave radar. The method has been described in detail in the non-line-of-sight target detection method embodiment, and will not be repeated here.
[0083] The millimeter wave radar is used to emit electromagnetic waves and receive reflected signals reflected from the scene. The processor comprises a radar processing unit and an NLOS detection network. The radar processing unit is used to estimate the position and speed of the occluded target from the millimeter wave radar signal, comprising a distance measurement module, a Doppler speed estimation module, an incident angle estimation module and a real detection point reconstruction module. The distance measurement module is used to estimate the distance according to the beat frequency of the received signal, and to determine the target position according to the distance. The Doppler speed estimation module is used to estimate the speed according to the phase change of the received signal. The incident angle estimation module is used to estimate the incident angle according to the phase difference of the signals received by the antenna array. The real detection point reconstruction module is used to determine whether each point cloud is a virtual detection point in combination with the relay wall information and the millimeter wave radar position, and to calculate the point cloud data of the real detection point corresponding to the virtual detection point if it is a virtual detection point.
[0084] The NLOS detection network is used to detect the information extracted by fusion over time to detect the occluded target. It comprises an input parameterization module and a non-line-of-sight detection model. The input parameterization module is used to convert the point cloud data of the real detection points corresponding to each frame into a sparse pseudo image. The non-line-of-sight detection model is used to obtain the target category and target positioning information.
[0085] Non-line-of-sight target tracking system embodiment
[0086] A non-line-of-sight target tracking system based on millimeter wave radar according to the present application comprises a millimeter wave radar and a processor, and is used to implement a non-line-of-sight target tracking method based on millimeter wave radar. The method has been described in detail in the non-line-of-sight target tracking method embodiment, and will not be repeated here.
[0087] The millimeter wave radar is used for transmitting electromagnetic waves and receiving reflected signals reflected from a scene. The processor includes a radar processing unit and an NLOS tracking network. The radar processing unit is used for estimating the position and speed of the occluded target from the millimeter wave radar signals, including a distance measurement module, a Doppler speed estimation module, an incident angle estimation module, and a real detection point reconstruction module. The distance measurement module is used for estimating the distance according to the beat frequency of the received signal, determining the target position according to the distance, the Doppler speed estimation module is used for estimating the speed according to the phase change of the received signal, and the incident angle estimation module is used for estimating the incident angle according to the phase difference of the signals received by the antenna array. The real detection point reconstruction module is used to determine whether each point cloud is a virtual detection point in combination with the relay wall information and the millimeter wave radar position, and if it is a virtual detection point, the point cloud data of the real detection point corresponding to the virtual detection point is calculated.
[0088] The NLOS tracking network is used for tracking hidden targets from radar data. It includes an input parameterization module and a non-line-of-sight tracking model. The input parameterization module is used to convert the point cloud data of the real detection points corresponding to each frame into sparse pseudo images. The non-line-of-sight detection model is used to process the pseudo images of the current frame and the previous n frames, and to predict the target category and target positioning information of the current frame plus the next F (F≤n) frames.
[0089] The present application combines frequency-modulated continuous wave radar with machine learning algorithms to detect occluded moving targets using reflective surfaces such as buildings or vehicles, without relying on optical reflections. Instead, it uses the specular reflection characteristics of millimeter wave radar for NLOS detection and tracking, overcoming the distance and reflection intensity limitations of optical NLOS methods and enabling longer-range NLOS detection and tracking. The combination of FMCW radar and machine learning algorithms enables real-time processing of radar data in dynamic scenes, adapting to fast-moving counter-terrorism environments. Through data-driven methods, the system can process radar reflection signals in real time, enabling real-time non-line-of-sight target detection in counter-terrorism scenarios, improving individual combat capabilities, reducing personnel casualties, and enhancing battlefield situation awareness. By fusing the estimated and measured speed information of NLOS targets during tracking, stable tracking of targets in time series can be achieved. Through the fusion of time information from multiple frames of radar data, the system can maintain efficient target tracking performance even in conditions with high environmental noise, improving the robustness of the system.
[0090] The effects of the non-line-of-sight target detection and tracking method of the present application are described below.
[0091] Dataset: 100 sequences were collected in 21 different outdoor scenarios, i.e. multiple repetitions of the same non-line-of-sight trajectory in different scenarios. The relay walls in this dataset include plastered walls of residential and industrial buildings, marble garden walls, fences, several parked cars, garages, warehouse walls, and concrete kerbs. The dataset is evenly distributed between hidden pedestrians and cyclists, adding up to more than 32 million radar points. The dataset was split into non-overlapping training and validation sets, where the validation set consists of 20 sequences from 4 scenarios, amounting to 3063 frames.
[0092] For models including velocity, the input parameterized pseudo-image size is (2, H, W); for models without velocity, the input parameterized pseudo-image size is (1, H, W). For pixels where multiple points fall into after discretization, the average value of the velocity values of all points in the pixel is calculated, and the sum of the intensity values of all points in the pixel is calculated. The present application also conducts experiments on averaging the intensity values, but finds that the results are not improved. For a large area of 60m x 80m, the axis is discretized into a 600 x 800 grid using a resolution of 0.1m. Each ground truth box is assigned to the highest overlapping prediction box for training. The average precision (AP) and the average distance between the center of the prediction box and the ground truth are used to evaluate the hidden classification and positioning performance, respectively. It is verified that the present application can still obtain a high positioning accuracy of 0.1m under the condition of detecting clutter and small scattering cross section of hidden targets. For challenging NLOS data, when the number of multiple object tracking accuracy (MOTA) increases, the model can still accurately locate most of the multiple object tracking precision (MOTP). These results verify the effectiveness of the proposed joint NLOS detection scheme in the anti-collision application.
Claims
1. A non-line-of-sight target detection method based on millimeter-wave radar, characterized in that: include: Use millimeter-wave radar to transmit electromagnetic waves to the scene where the non-line-of-sight target is located and receive the reflected signal; Extract radar point cloud data, which includes position, velocity and amplitude; Combine the relay wall information and the millimeter-wave radar position to determine whether each point cloud is a virtual detection point. If it is a virtual detection point, the point cloud data of the real detection point corresponding to the virtual detection point is calculated using the principle of mirror reflection; A pseudo image is generated based on the point cloud data of each real detection point, and the pseudo image is input into the non-line-of-sight detection model to obtain target category and target positioning information; the non-line-of-sight detection model includes a feature extraction module, a feature enhancement module, a feature fusion module and a detection head; the feature extraction module is used to extract at least two scale features from the pseudo image; the feature enhancement module is used to enhance at least one scale feature; The feature fusion module is used to fuse all scale features after feature enhancement and the remaining scale features except the scale for feature enhancement; the detection head is used to process the fused features and output the target category and target positioning information.
2. The non-line-of-sight target detection method based on millimeter-wave radar according to claim 1, characterized in that: The velocity and amplitude of the real detection point are the velocity and amplitude of the corresponding virtual detection point, and the position of the real detection point is calculated according to the following formula: Where x is the position of the real detection point corresponding to the virtual detection point; x′ is the position of the virtual detection point; c is the position of the millimeter wave radar receiver; n w is the normal vector of the relay wall, pointing in the direction away from the millimeter-wave radar receiver; w is the intersection of the virtual detection point and the relay wall.
3. The non-line-of-sight target detection method based on millimeter-wave radar according to claim 2, characterized in that: The intersection position of the virtual detection point and the relay wall is calculated according to the following formula: Among them, p1 and p2 are the two endpoints corresponding to the maximum visible range of the millimeter wave radar on the relay wall; p=p2-p1 represents the vector from p1 to p2.
4. The non-line-of-sight target detection method based on millimeter-wave radar according to claim 1, characterized in that: The type of the feature enhancement module is a Swin Transformer network, a CDPA network or a CAM network; wherein the CDPA network includes a spatial attention module and a channel attention module, the corresponding scale feature of the input is processed by the spatial attention module to obtain a first feature, the corresponding scale feature of the input is processed by the channel attention module to obtain a second feature, the first feature and the second feature are multiplied element-by-element with the corresponding scale feature of the input as the output value of the CDPA network, that is, the feature after feature enhancement; the CAM network includes a Swin Transformer network and a CDPA network, the corresponding scale feature of the input is processed by the Swin Transformer network to obtain a first output value, the corresponding scale feature of the input is processed by the CDPA network to obtain a second output value, the first output value and the second output value are added element-by-element as the output value of the CAM network, that is, the feature after feature enhancement.
5. The non-line-of-sight target detection method based on millimeter-wave radar according to claim 4, characterized in that: In the non-viewing area detection model, feature enhancement is performed on each scale feature extracted by the feature extraction module, and the feature enhancement modules used in the feature enhancement are of the same type or different types.
6. The non-line-of-sight target detection method based on millimeter-wave radar according to claim 1, characterized in that: The detection head is an SSD, and each anchor box of the SSD includes a target category and a bounding regression box. The bounding regression box is used to represent the target positioning information, including the center position, size, direction and speed of the box.
7. A non-line-of-sight target tracking method based on millimeter-wave radar, characterized in that: include: Use millimeter-wave radar to transmit electromagnetic waves to the scene where the non-line-of-sight target is located and receive the reflected signal; Extract the radar point cloud data of the current frame and the previous n frames. The radar point cloud data includes position, velocity and amplitude, n>1; Combine the relay wall information and the millimeter-wave radar position to determine whether each point cloud is a virtual detection point. If it is a virtual detection point, the point cloud data of the real detection point corresponding to the virtual detection point is calculated using the principle of mirror reflection; Generate corresponding pseudo images of the current frame and the previous n frames based on the point cloud data of the real detection points of the current frame and the previous n frames; The generated n+1 frames of pseudo images are input into the non-line-of-sight tracking model to obtain the target category, target positioning information of the current frame, and target positioning information of the future frames. The non-line-of-sight tracking model includes a feature extraction module, a feature connection module, a feature enhancement module, a feature fusion module, and a detection head. There are n+1 feature extraction modules, each of which is used to extract at least two scale features from the pseudo image of each frame. The feature connection module is used to splice n+1 features of the same scale along the feature channel to obtain spliced features of each scale. The feature enhancement module is used to enhance the features of at least one scale splicing feature; the feature fusion module is used to fuse all the scale splicing features after feature enhancement and the splicing features of the remaining scales except the scale for feature enhancement; The detection head is used to process the fused features and output the target category, target positioning information of the current frame, and target positioning information of future frames.
8. The non-line-of-sight target tracking method based on millimeter-wave radar according to claim 7, characterized in that: The type of the feature enhancement module is a Swin Transformer network, a CDPA network or a CAM network; wherein the CDPA network includes a spatial attention module and a channel attention module, the corresponding scale feature of the input is processed by the spatial attention module to obtain a first feature, the corresponding scale feature of the input is processed by the channel attention module to obtain a second feature, the first feature and the second feature are multiplied element-by-element with the corresponding scale feature of the input as the output value of the CDPA network, that is, the feature after feature enhancement; the CAM network includes a Swin Transformer network and a CDPA network, the corresponding scale feature of the input is processed by the Swin Transformer network to obtain a first output value, the corresponding scale feature of the input is processed by the CDPA network to obtain a second output value, the first output value and the second output value are added element-by-element as the output value of the CAM network, that is, the feature after feature enhancement.
9. A non-line-of-sight target detection system based on millimeter-wave radar, comprising a millimeter-wave radar and a processor, characterized in that: The system is used to implement the non-line-of-sight target detection method based on millimeter-wave radar as described in any one of claims 1 to 6.
10. A non-line-of-sight target tracking system based on millimeter-wave radar, comprising a millimeter-wave radar and a processor, characterized in that: The system is used to implement the non-line-of-sight target tracking method based on millimeter-wave radar as described in claim 7 or 8.