An edge-assisted radsar fusion unmanned aerial vehicle target detection method and system

CN117706506BActive Publication Date: 2026-09-25BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211080281.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-05
Publication Date
2026-09-25
Estimated Expiration
2042-09-05

AI Technical Summary

Technical Problem

[0005]本发明的目的是提供一种边缘辅助的雷视融合的无人机目标检测方法及系统,解决恶劣环境下在无人机上融合毫米波雷达与相机进行目标检测所面临的低鲁棒性和高时延的问题

Benefits of technology

[0037]本发明提出的边缘辅助的雷视融合的无人机目标检测方法及系统,对无人机采集的毫米波雷达点云和视频数据进行预处理,利用相机分支协助毫米波雷达分支进行多帧合成来解决无人机移动下毫米波雷达点云的稀疏性和噪声加剧的问题;利用多帧合成方法输出的雷达点云帧辅助相机分支进行目标显著区域提取与编码,以此减少传输的数据量和相应的卸载延迟;将毫米波雷达和相机数据的编码、传输、解码和推理并行化,以进一步减少卸载延迟;同时,在推理过程中采用边界框增强方案来确保准确率-延迟之间的平衡,最终达到实时、鲁棒的目标检测目的。本发明可以应用到各种需要无人机进行目标检测的任务中,可以高效率完成恶劣条件下对相关区域的巡检。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117706506B_ABST
    Figure CN117706506B_ABST
Patent Text Reader

Abstract

The application discloses an edge-assisted radar and vision fusion unmanned aerial vehicle target detection method and system, preprocesses millimeter wave radar point cloud and video data collected by an unmanned aerial vehicle, utilizes a camera branch to assist a millimeter wave radar branch to perform multi-frame synthesis to solve the problem of sparseness of millimeter wave radar point cloud and intensified noise under unmanned aerial vehicle movement, utilizes radar point cloud frames output by the multi-frame synthesis method to assist the camera branch to perform target salient region extraction and coding, so that the amount of data transmission and corresponding unloading delay are reduced, the coding, transmission, decoding and reasoning of millimeter wave radar and camera data are parallelized, so that unloading delay is further reduced, meanwhile, a bounding box enhancement scheme is adopted in the reasoning process to ensure the balance between accuracy and delay, and finally, the purpose of real-time and robust target detection is achieved. The application can be applied to various tasks requiring unmanned aerial vehicle target detection, and can efficiently complete the inspection of related regions under harsh conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned aerial vehicle (UAV) application technology, and in particular to an edge-assisted radar-visual fusion UAV target detection method and system. Background Technology

[0002] Due to their advantages such as small size, low cost, high maneuverability, high autonomy, ability to carry a wide range of sensors, communication and processing capabilities, and a broad "bird's-eye view," drones are becoming an important sensing platform. Their target detection capabilities, in particular, support many important applications, such as search and rescue, detection of farmland thieves, and detection of illegal vehicles.

[0003] Currently, with the success of deep learning, vision-based target detection from the perspective of drones has achieved good results. However, their performance degrades significantly under adverse conditions such as fog and insufficient lighting. Millimeter-wave radar, as an emerging autonomous perception mode, can overcome these adverse conditions and has been widely used in autonomous driving and intelligent transportation. Due to the sparse and noisy characteristics of millimeter-wave radar point clouds, fusing millimeter-wave radar with cameras has become a trend for robust target detection under adverse conditions.

[0004] This invention focuses on achieving target detection using millimeter-wave radar and camera fusion on unmanned aerial vehicles (UAVs). However, applying millimeter-wave radar and camera fusion technology to target detection on UAVs faces the following difficulties: (1) UAVs typically have the characteristics of "mobility + oblique view," which exacerbates the sparsity and noise of radar point clouds; (2) Existing fusion methods either involve a large number of sensors or the equipment is bulky and expensive, hindering the widespread application of UAVs; (3) Some advanced fusion methods require intensive computation and are difficult to run in real time on UAVs. Therefore, existing technologies struggle to achieve low-cost, real-time, and robust target detection on mobile UAVs. Summary of the Invention

[0005] The purpose of this invention is to provide an edge-assisted radar-camera fusion method and system for UAV target detection, which solves the problems of low robustness and high latency faced when fusing millimeter-wave radar and camera for target detection on UAVs in harsh environments.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] On the one hand, the present invention provides an edge-assisted radar-visual fusion method for UAV target detection, comprising the following steps:

[0008] S1. Multi-frame synthesis: The camera video stream is obtained using a camera, and the camera video stream is transformed to obtain camera frames; the radar data stream is obtained using millimeter-wave radar, and the camera video stream is processed by a displacement feature extractor to obtain the inter-frame displacement vector of the camera. The inter-frame displacement vector of the camera and the radar data stream are processed by a frame generator to generate radar point cloud frames.

[0009] S2. Salient Region Extraction and Encoding: The radar point cloud frame is processed into a radar frame by a salient region helper, and the camera frame and the processed radar frame are encoded into a camera frame by a mapped region cutter.

[0010] S3, Parallel Transmission and Inference: The encoding, transmission, decoding and inference of the processed radar frames and encoded camera frames are performed in parallel, while the bounding box enhancement method is executed during the inference process.

[0011] Furthermore, the displacement feature extraction method in step S1 is as follows: First, a lightweight optical flow method is used to extract corresponding interest point pairs (p) in the background of two adjacent frames i and i+1 in the camera branch. i ,p i+1 Then calculate each pair of interest points (p) separately. i,j ,p i+1,j The displacement vector d between ) i,j Then, clustering is performed on the displacement vectors, and the average value of the displacement vectors in the densest cluster is selected as the displacement vector d′ between the two frames. i Finally, a coordinate system transformation is performed to obtain the camera displacement vector d in the world coordinate system between frame i and frame i+1. i .

[0012] Furthermore, the radar point cloud frame generation method in step S1 is as follows: after obtaining the inter-frame displacement vector d... i Then, from the corresponding radar point cloud frame f i Subtract the displacement vector d from the middle i Generate the corresponding frame f′ after offset. i Subsequently, for 5 consecutive frames (f′) i-4 ,…,f′ i The radar point cloud frame F′ is generated by overlaying the points together. i Finally, the generated F′ i Clustering is performed to generate the final radar point cloud frame F. i .

[0013] Furthermore, in step S2, the image branches corresponding to the salient regions are compressed with low quality, while other regions of the image branches are not compressed. At the same time, during transmission, the radar point cloud is transmitted in the form of a set of points.

[0014] Furthermore, step S2 involves the following method for extracting salient regions:

[0015] First, based on the synthesized point cloud frame F output by the multi-frame synthesis module i Based on its clustering information, determine the boundary range {(x)} of the k-th cluster. k,min ,y k,min ),(x k,max ,y k,max The specific calculation formula is as follows:

[0016] x k,min =min{x1,x2,…,x n},[x1,x2,…,x n ]∈F i,k

[0017] y k,min =min{y1,y2,…,y n},[y1,y2,…,y n ]∈F i,k

[0018] x k,max =max{x1,x2,…,x n},[x1,x2,…,x n ]∈F i,k

[0019] y k,max =max{y1,y2,…,y n},[y1,y2,…,y n ]∈F i,k

[0020] Among them, (x n ,y n ) is the synthesized point cloud frame F i The coordinates of the nth point in the kth cluster;

[0021] Then, based on this boundary range, the salient region S corresponding to the k-th cluster is drawn. k The formula is as follows:

[0022] S k =[(x k,min ,y k,min ),(x k,max ,y k,max )],

[0023] Finally, the composite point cloud frame F is generated. i The corresponding set of salient regions {S k}

[0024] Furthermore, the method for encoding salient regions in step S2 is as follows:

[0025] First, the salient region is expanded using the following formula:

[0026] S′ k =[(x k,min -x padding ,y k,min -y padding ),(x k,max +x padding ,y k,max +y padding )]

[0027] Then, based on actual data, the padding values ​​m in the x and y directions are heuristically determined. The distance d between two salient regions is then compared with the padding value m. When d > 2m, the two salient regions are expanded and extracted separately; otherwise, they are merged into one and then expanded and extracted. This expansion operation is performed on all salient regions in the current frame to generate an expanded set of salient regions {S′}. k};

[0028] Finally, by encoding the salient regions and leaving the other parts uncompressed, an encoded image frame is formed.

[0029] Further, the bounding box enhancement method in step S3 is as follows: a detection head sub-network is used to obtain the detection results and confidence scores from the image branch and the millimeter-wave radar branch, and the detection results and confidence scores output by the two detection branches are fused using the perception modes of the two sensors, and the outputs are vectors with the same dimension. A deep neural network is used to integrate the two vectors and perform bounding box enhancement.

[0030] Furthermore, the detection head network first flattens the feature map, then uses a 512-dimensional fully connected layer, followed by two sibling fully connected layers for bounding box regression and classification, respectively. The output is a four-dimensional vector and a C+1-dimensional vector, where C is the number of classes.

[0031] Furthermore, the first fully connected layer of the deep neural network merges information from the two inputs of each class, the second fully connected layer captures global correlations between classes, and the final activation function layer outputs a 2D vector while setting a threshold to determine whether to preserve the bounding box.

[0032] On the other hand, the present invention also provides an edge-assisted radar-visual fusion drone target detection system for implementing any of the above methods, comprising the following modules:

[0033] The multi-frame synthesis module is used to convert the camera video stream obtained by the camera and the radar data stream obtained by the millimeter-wave radar into camera frames and radar point cloud frames. The multi-frame synthesis module includes a displacement feature extractor and a frame generator. The displacement feature extractor is used to obtain the inter-frame displacement vector of the camera, and the frame generator is used to generate radar point cloud frames by combining the inter-frame displacement vector of the camera and the radar data stream.

[0034] The salient region extraction and encoding module includes a salient region helper and a mapped region cutter. The radar point cloud frame is processed into a radar frame by the salient region helper, and the camera frame and the processed radar frame are encoded into a camera frame by the mapped region cutter.

[0035] The parallel transmission and inference module executes the encoding, transmission, decoding, and inference of processed radar frames and encoded camera frames in parallel, while simultaneously performing bounding box enhancement methods during inference. The parallel transmission and inference module includes a detection head sub-network, a fusion module, and an integration module. The detection head sub-network first flattens the feature maps, then uses a 512-dimensional fully connected layer, followed by two sibling fully connected layers for bounding box regression and classification. The fusion module uses intermediate convolutional layers to obtain confidence scores, then adds the confidence scores from the two branches and sends them to the activation function layer to obtain the fusion score. The integration module is based on a deep neural network; the first fully connected layer merges information from the two inputs of each class, the second fully connected layer captures global correlations between classes, and the final activation function layer outputs a 2D vector.

[0036] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0037] This invention proposes an edge-assisted radar-visual fusion UAV target detection method and system. It preprocesses millimeter-wave radar point cloud and video data collected by the UAV, utilizing a camera branch to assist the millimeter-wave radar branch in multi-frame synthesis to address the sparsity and increased noise issues of the millimeter-wave radar point cloud during UAV movement. The radar point cloud frames output by the multi-frame synthesis method assist the camera branch in extracting and encoding salient target regions, thereby reducing the amount of data transmitted and the corresponding offloading latency. The encoding, transmission, decoding, and inference of millimeter-wave radar and camera data are parallelized to further reduce offloading latency. Simultaneously, a bounding box enhancement scheme is employed during inference to ensure a balance between accuracy and latency, ultimately achieving real-time and robust target detection. This invention can be applied to various tasks requiring UAV target detection, enabling efficient inspection of relevant areas under harsh conditions. Attached Figure Description

[0038] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly described below. Obviously, the drawings described below are only some embodiments recorded in this invention, and those skilled in the art can obtain other drawings based on these drawings.

[0039] Figure 1 This is an overall architecture diagram of an edge-assisted radar-visual fusion drone target detection system provided in an embodiment of the present invention.

[0040] Figure 2 This is a flowchart of multi-frame synthesis provided for an embodiment of the present invention.

[0041] Figure 3 This is a diagram of the detection head sub-network structure provided in an embodiment of the present invention.

[0042] Figure 4 This is a structural diagram of an integrated module provided in an embodiment of the present invention. Detailed Implementation

[0043] To better understand this technical solution, the method of the present invention will be described in detail below with reference to the accompanying drawings.

[0044] The edge-assisted radar-eye fusion unmanned aerial vehicle target detection system and method of the present invention, such as Figure 1 As shown, it includes the following steps:

[0045] (1) Multi-frame synthesis

[0046] When a drone is in motion, the point cloud generated by millimeter-wave radar suffers from severe sparsity and noise problems. To address this, this invention proposes a multi-frame synthesis method for millimeter-wave radar using a camera-assisted approach, specifically comprising the following two modules:

[0047] 1) Displacement Feature Extractor. First, a lightweight optical flow method is used to extract corresponding interest point pairs (pi) in the background of two adjacent frames i and i+1 in the camera branch. i ,p i+1 Then calculate each pair of interest points (p) separately. i,j ,p i+1,j The displacement vector d between ) i,j Then, clustering is performed on these displacement vectors, and the average value of the displacement vectors in the densest cluster is selected as the displacement vector d′ between the two frames. i Finally, a coordinate system transformation is performed to obtain the camera inter-frame displacement vector d in the world coordinate system between frame i and frame i+1. i .

[0048] 2) Frame generator. This is used to obtain the inter-frame displacement vector d of the camera. iThen, from the corresponding radar point cloud frame f i Subtract the displacement vector d from the middle i Generate the corresponding frame f′ after offset. i Subsequently, for 5 consecutive frames (f′) i -4,…,f′ i The radar point cloud frame F′ is generated by overlaying the points together. i Finally, the generated F′ i Clustering is performed to generate the final radar point cloud frame F. i The specific synthesis process is as follows: Figure 2 As shown:

[0049] Because a small number of synthesized frames cannot effectively enhance the quality of the target point cloud, while a large number of synthesized frames will further amplify the noise; at the same time, too many synthesized frames will lead to a greater computational load, which will reduce the execution efficiency of the module. Therefore, after experimental testing and analysis, this module chooses to synthesize 5 adjacent frames.

[0050] (2) Significant region extraction and encoding

[0051] Transmitting multi-sensor data from the edge to the terminal side results in high bandwidth consumption, leading to high transmission latency. Therefore, this invention designs a salient region extraction and encoding method that selectively segments image branches using target occupancy information provided by radar point clouds. Since salient regions clearly represent target information, we perform low-quality compression on the corresponding image branches. However, millimeter-wave radar can also experience missed detections; therefore, we do not compress other regions of the image branches to further improve target detection performance. It is important to note that during transmission, the radar point cloud is transmitted as a set of points rather than a point cloud image. This method specifically includes the following two modules:

[0052] 1) Significant region aid.

[0053] First, based on the synthesized point cloud frame F output by the multi-frame synthesis module i Based on its clustering information, determine the boundary range {(x)} of the k-th cluster. k,min ,y k,min ),(x k,max ,y k,max The specific calculation formula is as follows:

[0054] x k,min =min{x1,x2,…,x n},[x1,x2,…,x n ]∈F i,k

[0055] y k,min =min{y1,y2,…,yn},[y1,y2,…,y n ]∈F i,k

[0056] x k,max =max{x1,x2,…,x n},[x1,x2,…,x n ]∈F i,k

[0057] y k,max =max{y1,y2,…,y n},[y1,y2,…,y n ]∈F i,k ,

[0058] Among them, (x n ,y n ) is the synthesized point cloud frame F i Find the coordinates of the nth point in the k-th cluster; then, draw the salient region S corresponding to this cluster based on the boundary range. k The formula is as follows:

[0059] S k =[(x k,min ,y k,min ),(x k,max ,y k,max )],

[0060] This generates a synthetic point cloud frame F. i The corresponding set of salient regions {S k}

[0061] 2) Mapped region cutter.

[0062] The salient region generated based on radar point clouds cannot completely cover the corresponding target in the video, so it needs to be expanded. The expansion formula is as follows:

[0063] S′ k =[(x k,min -x padding ,y k,min -y padding ),(x k,max +x padding ,y k,max +y padding )].

[0064] However, the expansion operation may cause the salient regions of two adjacent objects to overlap. Therefore, this invention heuristically determines the fill value m in the x and y directions based on actual data, and then compares the distance d between two salient regions with the fill value m; when d > 2m, the two salient regions can be expanded and extracted separately; otherwise, the two salient regions are merged into one before expansion and extraction. The above expansion operation is performed on all salient regions in the current frame to generate an expanded set of salient regions {S′}. k Finally, by encoding the salient regions and leaving the other parts uncompressed, an encoded image frame is formed.

[0065] (3) Parallel transmission and inference

[0066] This invention shifts the intensive computation of deep neural networks to the edge for processing, requiring the transmission of collected camera and radar point cloud data from the edge, which significantly increases the overall end-to-end latency. To alleviate this latency issue, this invention proposes a parallel transmission and inference mechanism, effectively executing transmission and inference in parallel. Since transmission and inference have different functions and require different resources, they do not interfere with each other during parallel execution. Therefore, this mechanism can effectively utilize multiple resources to execute different tasks in parallel and can significantly reduce system latency. This mechanism mainly consists of three parts: a detection head sub-network, a fusion module, and an integration module. The workflow of these three modules is as follows:

[0067] 1) Detect the head network.

[0068] This module fuses confidence scores from the image branch with information from the millimeter-wave radar branch; it performs bounding box augmentation on the image-based target detector, thereby improving the accuracy of bounding box localization and classification. The specific structure is as follows: Figure 3 As shown, the feature map is first flattened, then a 512-dimensional fully connected layer is used, followed by two sibling fully connected layers for bounding box regression and classification. The subnetwork outputs a four-dimensional vector and a C+1-dimensional vector, where C is the number of classes.

[0069] 2) Integration module.

[0070] The radar point cloud generated in the multi-frame synthesis stage described above can provide a strong indication of whether there is a target in a specific area, which is a necessary supplement to the confidence score obtained by the detection head subnet. Therefore, this invention designs a fusion module that jointly considers the perception modes of two sensors and fuses the detection results output from the two detection branches with the confidence scores to better estimate occupancy information. Specifically, considering the characteristics of radar point clouds, this module uses intermediate convolutional layers to obtain confidence scores (the final convolutional layer size is consistent with the image branches, i.e., 7×7 and 1×1), then adds the confidence scores from the two branches and sends them to the activation function layer to obtain the fusion score.

[0071] 3) Integration module.

[0072] The outputs of the detection head subnet and the fusion module constitute a C+1 dimensional vector, and the outputs of the millimeter-wave radar and camera branches have the same dimension. To fuse the results of the two C+1 dimensional vectors and perform more reliable bounding box enhancement, this invention introduces an ensemble module based on a deep neural network. Figure 4 As shown, the first fully connected layer merges information from the two inputs of each class, the second fully connected layer captures global correlations between classes, and the final activation function layer outputs a 2D vector (P). foreground ,P background This allows the detection network to make decisions between foreground and background.

[0073] Since millimeter-wave radar cannot detect all targets without omission, this module utilizes an image-based detector as a supplement to achieve better detection results. Simultaneously, to obtain optimal detection results, this invention sets a threshold to determine whether to retain bounding boxes; intuitively, the more significant the difference between the categories of the two input vectors, the more consistent the classification results, and the higher the output confidence.

[0074] Through actual testing, the series of methods proposed in this invention can achieve robust and real-time target detection under various adverse conditions (rain / snow / fog and darkness).

[0075] This invention addresses the target detection problem of UAVs under harsh conditions by fusing camera and millimeter-wave radar. It proposes a series of methods to fully leverage the complementary advantages of cameras and millimeter-wave radar at three different levels.

[0076] A novel multi-frame synthesis method uses camera-assisted millimeter-wave radar to address the exacerbated point cloud sparsity and noise issues;

[0077] A salient region extraction and encoding method uses millimeter-wave radar to assist cameras to greatly reduce the amount of data transmitted and the corresponding offloading delay;

[0078] A parallel transport and inference method further reduces offloading latency by using parallelism, while using a lightweight bounding box augmentation scheme to ensure the accuracy-latency trade-off at the edge.

[0079] Furthermore, experimental results show that under fog, rain, snow, and darkness conditions, the system's detection accuracy can reach 91.03%, 92.61%, 92.62%, and 81.23%, respectively, and transmitting multi-sensor data only requires 6Mbps of bandwidth; target detection (edge-side) only occupies 15% of the UAV system's CPU resources.

[0080] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. However, these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for UAV target detection using edge-assisted radar-visual fusion, characterized in that, Includes the following steps: S1. Multi-frame synthesis: The camera video stream is acquired using a camera, and the camera video stream is transformed to obtain camera frames. The radar data stream is obtained using millimeter-wave radar, and the camera video stream is used to obtain the inter-frame displacement vector of the camera through a displacement feature extractor. The inter-frame displacement vector of the camera and the radar data stream are used to generate radar point cloud frames through a frame generator. S2. Saliency Region Extraction and Encoding: Design a saliency region helper to extract the region of interest from the image data using the target occupancy information of the radar point cloud; design a mapping region cutter to perform different proportions of padding on the initial regions in the x and y directions to form the final mapping region. The radar point cloud frame is processed into a radar frame by a salient region helper, and the camera frame and the processed radar frame are processed into an encoded camera frame by a mapped region cutter. The salient region helper extracts salient regions using the following method: First, based on the synthesized point cloud frames output by the multi-frame synthesis module... Based on its clustering information, determine the boundary range of the k-th cluster. The specific calculation formula is as follows: , , , , in, It is a composite point cloud frame The coordinates of the nth point in the kth cluster; Then, the salient region corresponding to the k-th cluster is drawn based on this boundary range. The formula is as follows: Finally, a composite point cloud frame is generated. The corresponding set of salient regions ; The mapping region shearer encodes salient regions using the following method: First, the salient region is expanded using the following formula: , Then, the padding values ​​*m* in the x and y directions are heuristically determined based on actual data. The distance *d* between two salient regions is then compared to the padding value *m*. When *d* > 2 *m*, the two salient regions are expanded and extracted separately; otherwise, they are merged into one and then expanded and extracted. This expansion operation is performed on all salient regions in the current frame to generate an expanded set of salient regions. ; Finally, by encoding the salient regions and leaving the other parts uncompressed, the encoded camera frame is formed. S3, Parallel Transmission and Inference: The encoding, transmission, decoding and inference of the processed radar frames and encoded camera frames are performed in parallel, while the bounding box enhancement method is executed during the inference process.

2. The edge-assisted radar-visual fusion method for UAV target detection according to claim 1, characterized in that, The displacement feature extraction method in step S1 is as follows: First, a lightweight optical flow method is used to extract two adjacent frames in the camera branch. and Corresponding points of interest in the background ; Then calculate each pair of interest points separately. Displacement vectors between Then, clustering is performed on the displacement vectors, and the average value of the displacement vectors in the densest cluster is selected as the displacement vector between the two frames. Finally, a coordinate system transformation is performed to obtain the frame. With frames The camera displacement vector between them in the world coordinate system .

3. The edge-assisted radar-visual fusion method for UAV target detection according to claim 1, characterized in that, Step S1, the radar point cloud frame generation method, is as follows: after obtaining the inter-frame displacement vector... Then, from the corresponding radar point cloud frame Subtract the displacement vector Generate the corresponding frame after offset. Subsequently, for 5 consecutive frames The radar point cloud frames are overlaid to generate the overlaid radar point cloud frames. Finally, the generated Clustering is performed to generate the final radar point cloud frame. .

4. The edge-assisted radar-visual fusion method for UAV target detection according to claim 1, characterized in that, Step S2 performs low-quality compression on the corresponding image branches of the salient regions, while leaving other regions of the image branches uncompressed. Meanwhile, during transmission, the radar point cloud is transmitted in the form of a set of points.

5. The edge-assisted radar-visual fusion method for UAV target detection according to claim 1, characterized in that, The bounding box enhancement method in step S3 is as follows: a detection head sub-network is used to obtain the detection results and confidence scores from the image branch and the millimeter-wave radar branch. The detection results and confidence scores output by the two detection branches are fused using the perception modes of the two sensors, and the outputs are vectors with the same dimension. A deep neural network is used to integrate the two vectors and perform bounding box enhancement.

6. The edge-assisted radar-visual fusion method for UAV target detection according to claim 5, characterized in that, The detection head sub-network first flattens the feature map, then uses a 512-dimensional fully connected layer, followed by two sibling fully connected layers for bounding box regression and classification, respectively. The output is a four-dimensional vector and a C+1-dimensional vector, where C is the number of categories.

7. The edge-assisted radar-visual fusion method for UAV target detection according to claim 5, characterized in that, The first fully connected layer of a deep neural network merges information from the two inputs of each class, the second fully connected layer captures global correlations between classes, and the final activation function layer outputs a 2D vector while setting a threshold to determine whether to retain the bounding box.

8. An edge-assisted radar-visual fusion unmanned aerial vehicle (UAV) target detection system, characterized in that, A method for implementing any one of claims 1-7 includes the following modules: The multi-frame synthesis module is used to convert the camera video stream obtained by the camera and the radar data stream obtained by the millimeter-wave radar into camera frames and radar point cloud frames. The multi-frame synthesis module includes a displacement feature extractor and a frame generator. The displacement feature extractor is used to obtain the inter-frame displacement vector of the camera, and the frame generator is used to generate radar point cloud frames by combining the inter-frame displacement vector of the camera and the radar data stream. The salient region extraction and encoding module includes a salient region helper and a mapped region cutter. The salient region helper uses target occupancy information from the radar point cloud to extract the region of interest from the image data. The mapped region cutter performs different proportions of padding on the initial regions in the x and y directions to form the final mapped regions. The radar point cloud frames are processed into radar frames by the salient region helper, and the camera frames and the processed radar frames are encoded into camera frames by the mapped region cutter. The salient region helper extracts salient regions using the following method: First, based on the synthesized point cloud frames output by the multi-frame synthesis module... Based on its clustering information, determine the boundary range of the k-th cluster. The specific calculation formula is as follows: , , , , in, It is a composite point cloud frame The coordinates of the nth point in the kth cluster; Then, the salient region corresponding to the k-th cluster is drawn based on this boundary range. The formula is as follows: Finally, a composite point cloud frame is generated. The corresponding set of salient regions ; The mapping region shearer encodes salient regions using the following method: First, the salient region is expanded using the following formula: , Then, the padding values ​​*m* in the x and y directions are heuristically determined based on actual data. The distance *d* between two salient regions is then compared to the padding value *m*. When *d* > 2 *m*, the two salient regions are expanded and extracted separately; otherwise, they are merged into one and then expanded and extracted. This expansion operation is performed on all salient regions in the current frame to generate an expanded set of salient regions. ; Finally, by encoding the salient regions and leaving the other parts uncompressed, the encoded camera frame is formed. The parallel transmission and inference module executes the encoding, transmission, decoding, and inference of processed radar frames and encoded camera frames in parallel, while simultaneously performing bounding box enhancement methods during inference. The parallel transmission and inference module includes a detection head sub-network, a fusion module, and an integration module. The detection head sub-network first flattens the feature maps, then uses a 512-dimensional fully connected layer, followed by two sibling fully connected layers for bounding box regression and classification. The fusion module uses intermediate convolutional layers to obtain confidence scores, then adds the confidence scores from the two branches and sends them to the activation function layer to obtain the fusion score. The integration module is based on a deep neural network; the first fully connected layer merges information from the two inputs of each class, the second fully connected layer captures global correlations between classes, and the final activation function layer outputs a 2D vector.

Citation Information

Patent Citations

  • Video encoding / decoding method and device using synthesized reference frame

    CN101272494A

  • Millimeter wave radar and vision fused three-dimensional target detection method based on attention mechanism

    CN114708585A