Semantic perception method and system based on multi-radar cooperation

Through the multi-radar collaboration method, local coordinate system data is converted to the global coordinate system, and spatial and temporal consistency matching and feature fusion are carried out, which solves the insufficient identification of traditional single radar systems under self-occlusion and multipath effects, and achieves more accurate semantic perception.

CN120490998APending Publication Date: 2025-08-15PEKING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510500779.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Traditional single-radar semantic perception systems have problems with insufficient semantic recognition accuracy and false targets in self-occlusion of target objects and complex indoor environments.

Method used

The multi-radar collaboration method is adopted to convert the local coordinate system data of each radar to a unified global coordinate system. Through space-time consistency matching and multi-view feature fusion, a multi-radar collaborative semantic perception network is designed to eliminate false targets and extract complete features.

Benefits of technology

Accurate identification of goals in complex environments is achieved, the accuracy and completeness of semantic perception is improved, and false goals are eliminated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120490998A_ABST
    Figure CN120490998A_ABST
Patent Text Reader

Abstract

The invention discloses a semantic perception method and system based on multi-radar cooperation, and belongs to the technical field of radar detection, and the method comprises the steps: converting target data observed by each radar in a local coordinate system into a unified global coordinate system, and obtaining a distance azimuth thermodynamic diagram of each radar at a t moment; time domain and space domain matching is carried out on the distance azimuth thermodynamic diagrams of all radars at the t moment, and a t-moment distance azimuth global thermodynamic diagram with enhanced time-space consistency is obtained; and based on the time-space consistency enhanced t-moment distance azimuth angle global thermodynamic diagram, obtaining t-moment semantic perception results based on all radars. According to the invention, false targets can be eliminated, and accurate identification of the targets is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of radar detection technology, and in particular to a semantic perception method and system based on multi-radar collaboration. Background Art

[0002] Millimeter-wave radar, with its high range resolution, low cost, and excellent privacy protection, demonstrates significant potential for indoor perception. Semantic perception with millimeter-wave radar involves analyzing environmental information acquired by millimeter-wave radar, such as point clouds, heat maps, and micro-Doppler maps, to extract target location information and semantic attributes, enabling precise positioning and identification of multiple targets. This makes it highly practical in areas such as security monitoring, smart homes, and industrial automation.

[0003] Traditional millimeter-wave radar semantic perception primarily relies on a single radar system. The basic process is as follows: first, raw data from a single radar is acquired, and spatial information such as point clouds or range-angle heatmaps is obtained through signal processing. Then, deep neural networks are used to perform semantic recognition. Specifically, the current common single-radar semantic perception solution includes the following steps: 1) The radar transmits a chirp signal and receives the echo; 2) The received signal is subjected to an FFT transform in the range, velocity, and angle dimensions; 3) Target features are extracted using algorithms such as CFAR to generate a point cloud or heatmap; and 4) The target's semantic attributes are identified through a neural network. However, this single-radar solution has two major limitations: First, due to the limitation of a single viewing angle, the system cannot fully capture the object's feature information when the target is self-occluded, affecting the accuracy of semantic recognition. Second, in complex indoor environments, radar signals are susceptible to multipath effects, generating false targets and reducing the recognition accuracy of the semantic perception system. Summary of the Invention

[0004] The present invention proposes a semantic perception method and system based on multi-radar collaboration, which can eliminate false targets and achieve accurate target recognition.

[0005] To achieve the above-mentioned purpose, the technical solution of the present invention includes the following contents.

[0006] A semantic perception method based on multi-radar collaboration, the method comprising:

[0007] The target data observed by each radar in the local coordinate system is converted to a unified global coordinate system to obtain the range and azimuth heat map of each radar at time t;

[0008] By matching the range-azimuth heat maps of each radar at time t in the time domain and the space domain, a global range-azimuth heat map at time t with enhanced temporal and spatial consistency is obtained.

[0009] Based on the global heat map of distance and azimuth at time t with enhanced spatiotemporal consistency, the semantic perception results at time t based on all radars are obtained.

[0010] Furthermore, the data observed in the local coordinate system of each radar is converted to a unified global coordinate system to obtain the range and azimuth heat map of each radar at time t, including:

[0011] Get the observation result of the i-th radar on a target in the local coordinate system Observation direction θ i And the relative horizontal distance Δr of the radar relative to the reference radar i and the relative vertical distance Δh i ; Among them, r i Indicates the distance from the radar to the target center in the local coordinate system, Indicates the azimuth of the radar relative to the target center in the local coordinate system;

[0012] According to the observation results The observation direction θ i , the relative horizontal distance Δr i and the relative vertical distance Δh i , calculate the radar's observation results of the target in the global coordinate system Among them, r i,global Indicates the distance from the radar to the target center in the global coordinate system, Indicates the azimuth of the radar relative to the target center in the global coordinate system;

[0013] Comprehensive observation results of all targets by the radar in the global coordinate system Get the range and azimuth heat map of the radar at time t Among them, r represents the distance set from the radar to the center of each target in the global coordinate system, It represents the set of azimuth angles of the radar relative to the center of each target in the global coordinate system.

[0014] Furthermore, by matching the range-azimuth heatmaps of each radar at time t in the time domain and the space domain, a global range-azimuth heatmap at time t with enhanced temporal and spatial consistency is obtained, including:

[0015] Combine the range-azimuth heatmaps of M frames before and after time t, average the range-azimuth heatmap at time t, and obtain the averaged range-azimuth heatmap at time t;

[0016] Based on the averaged distance-azimuth heat map at time t, time domain matching is performed to obtain a global distance-azimuth heat map at time t with enhanced time domain consistency.

[0017] The averaged time t distance and azimuth heat map is combined with the spatial domain matching of the time t distance and azimuth global heat map after temporal consistency enhancement to obtain the time t distance and azimuth global heat map with temporal and spatial consistency enhancement.

[0018] Furthermore, time domain matching is performed based on the averaged distance-azimuth heat map at time t to obtain a global distance-azimuth heat map at time t with enhanced time domain consistency, including:

[0019] Calculate the temporal consistency index using the averaged distance-azimuth heat map at time t of M frames before and after time t;

[0020] Based on the averaged distance-azimuth heat map at time t and the temporal consistency index, a global distance-azimuth heat map at time t with enhanced temporal consistency is obtained.

[0021] Furthermore, the averaged time-t distance-azimuth heat map is combined with the time-t distance-azimuth global heat map after temporal consistency enhancement to perform spatial matching, thereby obtaining the time-t distance-azimuth global heat map with temporal and spatial consistency enhancement, including:

[0022] Calculate the spatial consistency index using the averaged distance-azimuth heat map at time t of M frames before and after time t;

[0023] Based on the global heat map of distance and azimuth at time t after the temporal consistency is enhanced and the spatial consistency index, a global heat map of distance and azimuth at time t with enhanced temporal and spatial consistency is obtained.

[0024] Furthermore, the time-t range-azimuth global heat map with enhanced spatiotemporal consistency is input into a multi-radar collaborative semantic perception network to obtain a semantic perception result at time t based on all radars; wherein the multi-radar collaborative semantic perception network includes:

[0025] Distributed feature extraction module is used to extract the target semantic feature F of the i-th radar based on the global heat map of distance and azimuth at time t enhanced by spatiotemporal consistency i ;

[0026] Multi-view feature fusion module is used to combine the target semantic features F of all radars i Fusion is performed to obtain the fused target semantic feature F fused ;

[0027] Semantic recognition module, used to identify the fused target semantic features F fused Decoding is performed to obtain the semantic perception results at time t based on all radars.

[0028] Furthermore, based on the global heat map of distance and azimuth at time t with enhanced spatiotemporal consistency, the target semantic feature F of the i-th radar is extracted. i ,include:

[0029] Compressing the temporal and spatial consistency enhanced distance and azimuth global heat map at time t into a low-dimensional potential feature space through multi-layer convolution operations;

[0030] The features in the low-dimensional potential feature space are mapped to the semantic space through the deconvolution layer to obtain the target semantic feature F of the i-th radar i .

[0031] Furthermore, for all radar target semantic features F i Fusion is performed to obtain the fused target semantic feature F fused ,include:

[0032] Get the learned weight parameter w corresponding to each radar i ;

[0033] Based on the learned weight parameter w i For all radar target semantic features F i Fusion is performed to obtain the fused target semantic feature F fused .

[0034] Furthermore, a cross entropy loss function is used to train the multi-radar collaborative semantic perception network.

[0035] A semantic perception system based on multi-radar collaboration, the system comprising:

[0036] The coordinate conversion module is used to convert the target data observed by each radar in the local coordinate system into a unified global coordinate system to obtain the range and azimuth heat map of each radar at time t;

[0037] The spatiotemporal consistency enhancement module is used to match the range-azimuth heat map of each radar at time t in the time domain and the space domain to obtain a global range-azimuth heat map at time t with enhanced spatiotemporal consistency;

[0038] The multi-radar collaborative semantic perception module is used to obtain the semantic perception results at time t based on all radars based on the global heat map of distance and azimuth at time t enhanced by spatiotemporal consistency.

[0039] Compared with the prior art, the present invention has at least the following beneficial effects.

[0040] 1. Traditional single-radar semantic perception systems are limited to a single perspective and cannot obtain complete target feature information due to target self-occlusion. The multi-radar collaborative semantic perception network proposed in this paper utilizes a distributed deployment of multiple radars to simultaneously observe targets from different perspectives. By fusing multi-perspective features, it achieves more comprehensive and complete target feature extraction.

[0041] 2. Traditional solutions are susceptible to multipath effects and can generate false targets in complex indoor environments. This new approach, based on spatiotemporal consistency analysis, identifies and eliminates false targets by analyzing the consistency of targets in both time and space, improving perception accuracy in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 Framework diagram of the multi-radar collaborative semantic perception system.

[0043] Figure 2 Target data observed by the radar in the local coordinate system. DETAILED DESCRIPTION

[0044] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below with reference to the accompanying drawings and embodiments.

[0045] like Figure 1 As shown in the figure, the multi-radar collaborative semantic perception system proposed in this invention consists of two main components: false target elimination based on spatiotemporal consistency and multi-radar collaborative semantic perception. In the false target elimination component, the present invention performs coordinate transformation on data from N radars and identifies and filters false targets caused by multipath effects by analyzing the target's consistency characteristics in both temporal and spatial dimensions. In the semantic perception component, the present invention designs a semantic perception network based on multi-view feature fusion. By fusing the range-azimuth heat maps of multiple radars, accurate target identification is achieved.

[0046] 1. Elimination of false targets based on spatiotemporal consistency.

[0047] In complex multipath environments, radar signals generate multipath reflections, causing the system to detect non-existent false targets. A key difference between these false targets and real targets is that they cannot remain consistent across consecutive time frames and multiple radar observations.

[0048] Based on this feature, the present invention proposes a time-space domain matching method, which distinguishes real targets from false targets by jointly analyzing the consistency of targets in time and space. Specifically, it includes three steps: first, coordinate conversion is performed to unify the observation data of multiple radars into a global coordinate system; then, time domain matching is performed to analyze the time consistency of the target using continuous multi-frame data of a single radar; finally, spatial domain matching is performed to analyze the spatial consistency of the target using the observation results of multiple radars at the same time. Based on the consistency indicators of the time domain and spatial domain, the radar data is enhanced to highlight the real targets and suppress false targets. The specific implementation methods of these three steps will be introduced in detail below:

[0049] 1. Coordinate conversion. The present invention deploys N millimeter-wave radars in the scene for distributed perception. In order to achieve collaborative processing of multi-radar data, it is first necessary to convert the data observed by the local coordinate system of each radar into a unified global coordinate system. Figure 2 As shown, the observation result of the target by the i-th radar is where r i Indicates the distance from the radar to the target center, The observation direction of the i-th radar is θ i , the relative horizontal and vertical distances relative to the reference radar (defined from the N millimeter-wave radars mentioned above) are Δr i and Δh i The transformation of each radar's observation data from the local coordinate system to the global coordinate system can be expressed as:

[0050]

[0051] Where R(θ i ) is the rotation matrix of radar i, T(Δr i ,Δh i ) is the translation vector, and the central processing unit is responsible for receiving and processing the converted global coordinate data.

[0052] Measurement results of different targets r i,global , The collection of constitutes the range-azimuth heat map of the i-th radar at time t

[0053] 2. Time domain matching. It is known that the range and azimuth heat map of the i-th radar at time t is By combining the heatmaps of the previous and next M frames, the target signal is averaged in the time dimension to eliminate local anomalies caused by random noise while retaining the real target signal that appears continuously:

[0054]

[0055] Represents the range and azimuth heat map of the i-th radar at frame t+m.

[0056] In order to further distinguish real targets from false targets, the present invention introduces the time domain consistency index Indicates the consistency of the target in M consecutive frames:

[0057]

[0058] Where I is the indicator function, which takes the value of 1 when the condition is met and 0 otherwise. ∈ is the time consistency confidence threshold, which is used to filter high confidence target points. If is 1, it indicates that the target remains stable in multiple frames and can be judged as a real target; if If it is 0, it indicates that the target may be a false target. Averaging results after coordinate transformation Perform time consistency enhancement:

[0059]

[0060] in This is the global heat map after temporal consistency enhancement.

[0061] 3. Spatial matching. Similar to time domain matching, spatial matching introduces spatial consistency indicators. In the global coordinate system, the semantic perception results of N radars are combined to eliminate false targets and enhance real targets:

[0062]

[0063] where γ is the spatial consistency confidence threshold.

[0064] Using airspace consistency indicators Perform spatial domain consistency enhancement on the global heat map after temporal domain consistency enhancement:

[0065]

[0066] in It is the input of the multi-radar collaborative semantic perception network.

[0067] 2. Multi-radar collaborative semantic perception network.

[0068] The present invention also designs a multi-radar collaborative semantic perception network that extracts and fuses feature information from multiple heat maps to achieve collaborative semantic perception of targets. The network can be divided into three parts: distributed feature extraction, multi-view feature fusion, and semantic recognition.

[0069] 1. Distributed feature extraction. The input of the multi-radar cooperative semantic perception network is the range and azimuth heat map after the spatiotemporal consistency enhancement of N radars. i=1,2,…,N, each heat map corresponds to an independent CNN Auto-Encoder feature extractor. The Auto-Encoder consists of two parts: the encoder and the decoder: the encoder compresses the input heat map into a low-dimensional latent feature space through multi-layer convolution operations, and the decoder maps the features to the semantic space through the deconvolution layer to extract the semantic features of the target. i=1,2,…,N。

[0070] 2. Multi-view feature fusion. The multi-view feature fusion module receives feature maps from N radars. Each feature map contains the observation information of the target under that view. In order to effectively integrate the feature information from different viewpoints, this paper designs an adaptive feature fusion mechanism by introducing learnable weight parameters W=w1,w2,…,w N The features from different perspectives are weighted. The fused features can be expressed as: i=1,2,…,N。

[0071] 3. Semantic recognition. Finally, a CNN Auto-Encoder is used to identify F fused Processing is performed to obtain the millimeter wave semantic segmentation results c∈C, C is the number of object types.

[0072] The neural network provides semantic labels for the network through the 2.5D depth map and RGB image provided by the RGB-D camera. Perform fully supervised training. The network is optimized using the cross entropy loss function, which is as follows:

[0073]

[0074] In summary, existing millimeter-wave semantic perception systems typically consist of a single millimeter-wave radar. Due to the self-occlusion of target objects, millimeter-wave semantic perception systems from a single perspective cannot fully capture the object's feature information, affecting the accuracy of semantic recognition. This invention uses multiple radars to detect objects from different perspectives, achieving more comprehensive semantic feature extraction of the target object.

[0075] Existing millimeter-wave semantic perception systems are susceptible to multipath effects in complex indoor environments, generating false targets and reducing the system's semantic perception accuracy. This paper proposes a false target elimination mechanism based on multi-radar collaboration. This mechanism leverages the temporal and spatial consistency of multiple radars to eliminate false targets caused by multipath effects, effectively improving the recognition accuracy of millimeter-wave semantic perception systems in complex multipath environments.

[0076] The above embodiments are provided for the purpose of describing the present invention only and are not intended to limit the scope of the present invention. The scope of the present invention is defined by the appended claims. Various equivalent substitutions and modifications made without departing from the spirit and principles of the present invention are intended to be within the scope of the present invention.

Claims

1. A semantic perception method based on multi-radar collaboration, characterized in that: The method comprises: The target data observed by each radar in the local coordinate system is converted to a unified global coordinate system to obtain the range and azimuth heat map of each radar at time t; By matching the range-azimuth heat maps of each radar at time t in the time domain and the space domain, a global range-azimuth heat map at time t with enhanced temporal and spatial consistency is obtained. Based on the global heat map of distance and azimuth at time t with enhanced spatiotemporal consistency, the semantic perception results at time t based on all radars are obtained.

2. The method according to claim 1, characterized in that The data observed in the local coordinate system of each radar is converted to a unified global coordinate system to obtain the range and azimuth heat map of each radar at time t, including: Get the observation result of the i-th radar on a target in the local coordinate system Observation direction θ i And the relative horizontal distance Δr of the radar relative to the reference radar i and the relative vertical distance Δh i ; Among them, r i Indicates the distance from the radar to the target center in the local coordinate system, Indicates the azimuth of the radar relative to the target center in the local coordinate system; According to the observation results The observation direction θ i , the relative horizontal distance Δr i and the relative vertical distance Δh i , calculate the radar's observation results of the target in the global coordinate system Among them, r i,global Indicates the distance from the radar to the target center in the global coordinate system, Indicates the azimuth of the radar relative to the target center in the global coordinate system; Comprehensive observation results of all targets by the radar in the global coordinate system Get the range and azimuth heat map of the radar at time t Among them, r represents the distance set from the radar to the center of each target in the global coordinate system, It represents the set of azimuth angles of the radar relative to the center of each target in the global coordinate system.

3. The method according to claim 1, characterized in that By matching the range-azimuth heat maps of each radar at time t in the time domain and the space domain, a global range-azimuth heat map at time t with enhanced temporal and spatial consistency is obtained, including: Combine the range-azimuth heatmaps of M frames before and after time t, average the range-azimuth heatmap at time t, and obtain the averaged range-azimuth heatmap at time t; Based on the averaged distance-azimuth heat map at time t, time domain matching is performed to obtain a global distance-azimuth heat map at time t with enhanced time domain consistency. The averaged time t distance and azimuth heat map is combined with the spatial domain matching of the time t distance and azimuth global heat map after temporal consistency enhancement to obtain the time t distance and azimuth global heat map with temporal and spatial consistency enhancement.

4. The method according to claim 3, characterized in that Based on the averaged distance-azimuth heat map at time t, time domain matching is performed to obtain a global distance-azimuth heat map at time t with enhanced time domain consistency, including: Calculate the temporal consistency index using the averaged distance-azimuth heat map at time t of M frames before and after time t; Based on the averaged distance-azimuth heat map at time t and the temporal consistency index, a global distance-azimuth heat map at time t with enhanced temporal consistency is obtained.

5. The method according to claim 3, characterized in that The averaged time-t distance-azimuth heat map is combined with the spatial domain matching of the time-t distance-azimuth global heat map after temporal consistency enhancement to obtain the time-t distance-azimuth global heat map with temporal and spatial consistency enhancement, including: Calculate the spatial consistency index using the averaged distance-azimuth heat map at time t of M frames before and after time t; Based on the global heat map of distance and azimuth at time t after the temporal consistency is enhanced and the spatial consistency index, a global heat map of distance and azimuth at time t with enhanced temporal and spatial consistency is obtained.

6. The method according to claim 1, characterized in that The time-t range-azimuth global heat map with enhanced spatiotemporal consistency is input into a multi-radar collaborative semantic perception network to obtain a semantic perception result at time t based on all radars; wherein the multi-radar collaborative semantic perception network includes: Distributed feature extraction module is used to extract the target semantic feature F of the i-th radar based on the global heat map of distance and azimuth at time t enhanced by spatiotemporal consistency i ; Multi-view feature fusion module is used to combine the target semantic features F of all radars i Fusion is performed to obtain the fused target semantic feature F fused ; Semantic recognition module, used to identify the fused target semantic features F fused Decoding is performed to obtain the semantic perception results at time t based on all radars.

7. The method according to claim 6, characterized in that Based on the global heat map of distance and azimuth at time t with enhanced spatiotemporal consistency, the target semantic feature F of the i-th radar is extracted. i ,include: Compressing the temporal and spatial consistency enhanced distance and azimuth global heat map at time t into a low-dimensional potential feature space through multi-layer convolution operations; The features in the low-dimensional potential feature space are mapped to the semantic space through the deconvolution layer to obtain the target semantic feature F of the i-th radar i .

8. The method according to claim 6, characterized in that For all radar target semantic features F i Fusion is performed to obtain the fused target semantic feature F fused ,include: Get the learned weight parameter w corresponding to each radar i ; Based on the learned weight parameter w i For all radar target semantic features F i Fusion is performed to obtain the fused target semantic feature F fused .

9. The method according to claim 6, characterized in that The multi-radar cooperative semantic perception network is trained using a cross-entropy loss function.

10. A semantic perception system based on multi-radar collaboration, characterized in that: The system comprises: The coordinate conversion module is used to convert the target data observed by each radar in the local coordinate system into a unified global coordinate system to obtain the range and azimuth heat map of each radar at time t; The spatiotemporal consistency enhancement module is used to match the range-azimuth heat map of each radar at time t in the time domain and the space domain to obtain a global range-azimuth heat map at time t with enhanced spatiotemporal consistency; The multi-radar collaborative semantic perception module is used to obtain the semantic perception results at time t based on all radars based on the global heat map of distance and azimuth at time t enhanced by spatiotemporal consistency.