A multi-modal fusion-based unmanned aerial vehicle dynamic environment 3D target detection method and device
By using multimodal fusion of 4D millimeter-wave radar and camera, the problems of large point cloud changes and poor algorithm stability caused by the unstable pose of UAVs in dynamic environments are solved, achieving high-accuracy 3D target detection, which is suitable for UAV target detection in complex environments.
Patent Information
- Application Number
- CN202411890164.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2044-12-20
AI Technical Summary
Drones in dynamic environments suffer from unstable poses, leading to large changes in point cloud positions, poor algorithm stability, and low accuracy, especially under extreme weather conditions such as rain and fog.
A 4D millimeter-wave radar and camera fusion scheme is adopted. By calibrating the 4D millimeter-wave radar and camera, point cloud denoising and feature extraction are performed. Latent space normal estimation is used for multi-frame point cloud pose correction and fusion. Feature fusion is performed by combining dual-head channel attention mechanism and cross attention mechanism. An anchor box-based detection head module is designed to achieve 3D target detection.
It improves the accuracy and stability of target detection for UAVs in dynamic environments, enhances detection capabilities under adverse weather conditions, and reduces model complexity and computational costs.
Smart Images

Figure CN119942521B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of target detection, and particularly relates to a 3D target detection method for dynamic environment of unmanned aerial vehicle based on multi-modal fusion. BACKGROUND
[0002] As a low-altitude aircraft, the unmanned aerial vehicle has the advantages of small size, strong maneuverability, low cost and the ability to carry various devices, and is increasingly widely used in complex environments, such as urban traffic supervision, agricultural detection and forest rescue. In these application scenarios, target detection technology is particularly critical, especially in dynamic or occluded environments.
[0003] Although the traditional visual sensor can provide detailed image information, it performs poorly in low light, light pollution, rain, fog and other weather conditions. Laser radar has extremely high precision and resolution in aerial target detection, and can provide accurate three-dimensional point cloud data, unaffected by lighting conditions. However, the disadvantages of laser radar are high cost, low detection speed and performance degradation in bad weather (such as rain, snow, fog, etc.). 3D millimeter wave radar has high robustness in aerial target detection, and can work stably in bad weather, low light, fog and other environmental conditions, providing high-precision distance, speed and three-dimensional spatial data. It can also penetrate obstacles to detect hidden targets, making it suitable for target detection in complex environments. Its disadvantage is that the resolution is low and it cannot provide appearance or detail information of the target, making it difficult to classify the target.
[0004] Compared with traditional radar, 4D millimeter wave radar adds the measurement information of the pitch angle, providing higher point cloud density and higher angular resolution, and has more advantages in dynamic target detection and tracking. However, there are still problems of sparse and disordered point cloud data, and missing and false detection when performing multi-modal fusion. At the same time, the position and attitude of the unmanned aerial vehicle change in real time during movement, and it is difficult to align the obtained image and point cloud information, which greatly affects the accuracy of target detection. SUMMARY
[0005] The present application provides a 3D target detection method for dynamic environment of unmanned aerial vehicle based on multi-modal fusion, which effectively solves the problem of large change in point cloud position caused by unstable pose of unmanned aerial vehicle during movement, poor algorithm stability and low accuracy in extreme weather such as rain and fog.
[0006] To achieve the above purpose, the present application provides the following technical scheme:
[0007] A 3D target detection method for dynamic environment of unmanned aerial vehicle based on multi-modal fusion, comprising the following steps:
[0008] S1, calibrate the 4D millimeter wave radar and the camera to obtain the pose relationship between the millimeter wave radar and the camera;
[0009] S2, pre-process the 4D millimeter wave radar point cloud data to obtain a denoised point cloud image, and perform feature extraction on the denoised point cloud image;
[0010] S3, use hidden space normal estimation to obtain the pose matrix between adjacent frames, and then perform global discretization analysis to realize multi-frame point cloud attitude correction and fusion in a motion scene, and enhance the point cloud data;
[0011] S4, process the radar point cloud and the camera image, and obtain a BEV feature map through a feature extraction network;
[0012] S5, fuse the BEV features using a double-head channel attention mechanism and a cross attention mechanism;
[0013] S6, design an anchor box-based detection head module, and set different detection targets;
[0014] S7, verify using different data sets and different test scenes to obtain 3D target detection results.
[0015] Further, the step S1 is implemented by the following steps:
[0016] S11, calibrate the camera and the radar separately to obtain the internal and external parameters of the sensor;
[0017] S12, pre-process and normalize the homogeneous coordinates of the radar points and the image points to obtain the normalized coordinates, and the normalization formula is as follows:
[0018]
[0019] In the formula, and are the homogeneous coordinates of the radar points and the image points, respectively, H=[h ij ] 3×4 is a transformation matrix, T p and T q are the normalization matrices of and , m1, m2 and m3 are three elements of the average vector , n1 and n2 are two elements of the average vector , and d1 and d2 are the norm average values of and , respectively.
[0020] S13, substitute the radar camera data pair to solve the three-dimensional direct linear change (DLT) to obtain the radar camera inter-transformation matrix Because the millimeter wave radar used has three-dimensional spatial characteristics, the DLT change is extended from two dimensions to three dimensions, and the coordinate conversion formula is as follows:
[0021]
[0022] Further, the step S2 is specifically implemented by the following steps:
[0023] A statistical filtering algorithm is used to remove noise. Considering the sparsity of the point cloud, the statistical filtering algorithm defines that the point cloud in a region is invalid when the point cloud in the region is less than a certain density. The method calculates the average distance of each point to the adjacent k points, and the points exceeding the average distance are regarded as outliers. The outliers are removed from the data to obtain the denoised point cloud image. The formula is as follows:
[0024] D max =μ+std×σ (6)
[0025] In the formula, μ represents the average value of the distance of a certain point to other points, that is,
[0026]
[0027] In the formula, σ represents the standard deviation of the distance, and std represents a self-defined standard deviation coefficient used to control the influence of the distance standard deviation on the distance threshold.
[0028] The STEM module is used to extract and process the point cloud input data. The STEM structure first expands the dimension and then reduces the dimension, so as to guarantee the information extraction capability while reducing the calculation complexity of the model.
[0029] Further, the step S3 is specifically implemented by the following steps:
[0030] S31, each frame of point cloud image information is converted into a point cloud feature matrix, and a spatial feature node is established;
[0031] S32, the feature similarity of adjacent nodes is calculated, and a connection edge E={ε ij |i,j∈V,ε ij ≠0} is constructed according to the feature similarity, so as to obtain a three-dimensional space topology G=(V,E), and the calculation formula is as follows:
[0032]
[0033] In the formula, v i and v j are adjacent nodes, ψ are two linear transformation functions, the multiplication of which obtains an adjacency matrix, s represents the connection relationship between two nodes, and the obtained connection edge represents the time domain feature thereof.
[0034] S33, establish a radar point cloud front and back image pair, use hidden space normal estimation method to calculate the pose correlation of adjacent frames, and the connection edge established represents the time domain feature used to connect the space feature of the node;
[0035] S34, input the point cloud feature map into a space-time graph convolution model to obtain time dimension correlation, and the specific design details are as follows:
[0036]
[0037] In the formula, Z k represents the node attribute of the kth layer, W k is a parameter matrix, and A is a normalized adjacency matrix, represents L2 normalization, and sigma is a LeakyReLU activation function. After completing the graph convolution, a fully connected layer is used to obtain the final single-frame latent space normal estimation to obtain the gap between adjacent frames.
[0038] S35, calibrate and fuse multiple frames of point clouds in a motion scene to obtain enhanced point cloud data.
[0039] Further, the step S4 is specifically implemented by the following steps:
[0040] S41, parallel processing of the preprocessed radar point cloud and the camera image. The camera image is input into a ResNet network for multi-scale feature extraction, and an FPN network is used as a neck to fuse high-order and low-order feature information and realize feature enhancement.
[0041] S42, use orthogonal feature transformation to map two-dimensional features to three-dimensional space, and project the radar coordinate points to the camera coordinate system, and the formula is as follows:
[0042]
[0043] In the formula, the radar coordinate point represents (x r ,y y ,z r ), the camera coordinate point represents (x c ,y c ,z c ), and represent the radar-to-camera coordinate system rotation matrix and conversion vector.
[0044] S43, convert the camera coordinate point to the pixel coordinate system, and the formula is as follows:
[0045]
[0046] In the formula, f x = f / dx, f y= f / dy, f represents the focal length of the camera, dx and dy are the width and length of the image respectively, (c x ,c y ) is the pixel coordinate system principal point coordinate, (u c ,v c ) is the pixel coordinate system angle coordinate.
[0047] S44, feature sampling is performed to generate a voxel space, and an image BEV feature map is obtained through voxel coding. First, feature aggregation is performed by using cumulative summation, feature coding is performed for each voxel, if the coding of adjacent voxels is different, difference calculation is performed, redundant features are eliminated, and finally the z-axis is eliminated to obtain the image BEV feature map F c ∈R X×Y . The voxel space expression is as follows:
[0048] V(x,y,z)=D(u c ,v v ,d)×F(u c ,v v ) (12)
[0049] In the formula, V∈R C×Z×Y×X is the voxel space feature, (x,y,z) is the voxel coordinate, F represents the image information, D is the probability feature occupation of F in the voxel space, and d is the reference depth.
[0050] S45, input the 4D millimeter wave radar point cloud into the cylinder feature extraction network based on kernel density clustering, introduce adaptive kernel density estimation to enhance semantic features, and generate a cylinder space. Each point calculates the difference between itself and the surrounding points in the neighborhood in the multi-scale feature through the kernel density block to obtain the density estimation of the specific point. The Gaussian kernel function is as follows:
[0051]
[0052] In the formula, f(p) represents the density of each point, M p represents the number of other points in the neighborhood, R is the radius, Dop is the Doppler information, and K is the Gaussian kernel function.
[0053] S46, introduce the voxelization method, use three-dimensional sparse convolution to accelerate the inference process, project the cylinder space information to the BEV plane, and obtain the 4D millimeter wave radar BEV pseudo image.
[0054] Further, the step S5 is specifically implemented by the following steps:
[0055] S51, global compression and channel excitation are performed through a double-head channel attention mechanism, and the expression is as follows:
[0056]
[0057] In the formula, F' represents image information after feature mapping, F represents image information, M represents a mapping function c The channel attention module is composed of two parallel maximum pooling layers and average pooling layers, then an MLP network with two hidden layers is connected, and a sigmoid activation function is used for output, forming a channel attention mapping.
[0058] S52, fuse radar camera BEV features using cross attention mechanism, extract common features and difference features, and assign weights to query (Q)-key value (K, V);
[0059] S53, lightening the module, introducing deep separable convolution to downsample the generation of K and V, to obtain smaller K and V.
[0060] Further, the step S6 is specifically realized by the following steps: using an anchor box based detection head to set the direction of different categories of targets, such as pedestrians and vehicles, differently. In regression, multiple parameters are predicted for each anchor box, including the midpoint coordinates of the anchor box, the length, width, height and rotation angle of the anchor box, etc.
[0061] Further, the step S7 is specifically realized by the following steps: selecting images and radar data of different urban environments, weather and time periods as a data set, inputting the model for training, retaining the best model according to the evaluation result, inputting the test set into the best model to obtain the test result, and verifying the effectiveness of the application.
[0062] The second aspect of the application relates to a kind of unmanned aerial vehicle dynamic environment 3D target detection device based on multi-modal fusion, including memory and one or more processors, the executable code is stored in the memory, when the one or more processors execute the executable code, for realizing the application of a kind of unmanned aerial vehicle dynamic environment 3D target detection method based on multi-modal fusion.
[0063] The third aspect of the application relates to a computer readable storage medium, which stores a program, and the program is executed by a processor to realize the application of a kind of unmanned aerial vehicle dynamic environment 3D target detection method based on multi-modal fusion.
[0064] Compared with the prior art, the application has the following advantages:
[0065] The application provides a multi-modal fusion-based unmanned aerial vehicle dynamic environment 3D target detection method, which adopts a 4D millimeter wave radar and camera fusion scheme. BRIEF DESCRIPTION OF DRAWINGS
[0066] Figure 1 is a flowchart of the method of the application.
[0067] Figure 2 is a network structure diagram of the application.
[0068] Figure 3 is a STEM module structure diagram of the application.
[0069] Figure 4 is a radar point cloud multi-frame synthesis module flowchart of the application.
[0070] Figure 5 is a device schematic diagram of the application. DETAILED DESCRIPTION
[0071] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description only some embodiments of the application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor, and all belong to the protection scope of the application.
[0072] Embodiment 1
[0073] As Figure 1As shown, the unmanned aerial vehicle dynamic environment 3D target detection method based on multi-modal fusion of the application comprises the following steps: collecting multi-modal data with a calibrated camera and radar; preprocessing 4D radar point cloud, using statistical probability method for point cloud denoising, and using STEM network for feature extraction of the denoised image; converting the preprocessed radar point cloud information into a feature matrix, constructing a feature node, establishing a connection edge through similarity calculation, constructing a three-dimensional space graph, inputting adjacent multi-frame point clouds into a space-time graph convolution model to obtain a pose relationship, realizing multi-frame point cloud pose calibration and fusion in a motion scene; parallel processing of camera images and radar point clouds, using orthogonal feature changes to the voxel space of the camera image, using a kernel density clustering feature network to obtain a radar point cloud column space, eliminating the z-axis, and obtaining the BEV feature map of the image and the point cloud; using a double-head channel attention mechanism and a cross-attention mechanism to fuse global and local features; and finally obtaining the final 3D target detection result through an anchor box detection head.
[0074] As an example, the specific steps of the unmanned aerial vehicle dynamic environment 3D target detection method based on multi-modal fusion of the application are as follows:
[0075] S1, calibrate the 4D millimeter wave radar and the camera to obtain the pose relationship between the millimeter wave radar and the camera.
[0076] S11, individually calibrate the camera and the radar to obtain the internal and external parameters of the sensor;
[0077] S12, pre-process and normalize the radar points and image points in homogeneous coordinates to obtain the normalized coordinate pairs, and the normalization formula is as follows:
[0078]
[0079] In the formula, and are the homogeneous coordinates of the radar points and the image points, respectively, H = [h ij ] 3×4 is a transformation matrix, T p and T q are the normalization matrices of and , m1, m2 and m3 are three elements of the average vector , n1 and n2 are two elements of the average vector , and d1 and d2 are the norm average values of and , respectively.
[0080] S13, substitute the radar camera data pairs to solve the three-dimensional direct linear change (DLT) to obtain the radar camera inter-transformation matrix Because the millimeter wave radar used has three-dimensional spatial characteristics, the DLT change is extended from two dimensions to three dimensions, and the coordinate conversion formula is as follows:
[0081]
[0082] S2, the 4D millimeter wave radar point cloud data is preprocessed to obtain a denoised point cloud image, and the network structure of the application is as shown in Figure 2 The statistical filtering algorithm is used to remove noise. Considering the sparsity of the point cloud, the statistical filtering algorithm defines that when the point cloud in a region is less than a certain density, the point cloud of the region is invalid. The method calculates the average distance of each point to the adjacent k points, and the points exceeding the average distance are regarded as outliers. The outliers are removed from the data to obtain a denoised point cloud image. The formula is as follows:
[0083] D max = mu + std x sigma (6)
[0084] In the formula, mu represents the average value of the distance of a certain point to other points, that is,
[0085]
[0086] In the formula, sigma represents the standard deviation of the distance, and std represents a self-defined standard deviation coefficient used to control the influence of the distance standard deviation on the distance threshold.
[0087] As shown in Figure 3 , the STEM module is used for feature extraction and processing of the point cloud input data. The STEM structure first expands the dimension and then reduces the dimension, so as to ensure the information extraction capability while reducing the calculation complexity of the model.
[0088] S3, using hidden space normal estimation, obtaining the pose matrix between adjacent frames, and then through global discretization analysis, realizing multi-frame point cloud attitude correction and fusion in a motion scene, and enhancing the point cloud data, as shown in Figure 4 At present, the mainstream 4D millimeter wave radar imaging algorithm is relatively effective in a static environment, but in a dynamic environment, the coordinates and attitudes change at the same time, and the point cloud image cannot be accurately corrected. Therefore, the application proposes a space normal vector estimation method to realize point cloud enhancement of the unmanned aerial vehicle in a motion scene.
[0089] S31, converting each frame of point cloud image information into a point cloud feature matrix, and establishing a spatial feature node;
[0090] S32, a three-dimensional space graph G=(V, E) is constructed. V is a node set, which can be connected with other nodes in the neighborhood, E is a connection edge between adjacent nodes, which represents the connection relationship between nodes. Given a set of edges E={epsilon ij |i,j∈V,epsilonij ≠0}. After obtaining the node features, the connection edges are constructed according to the similarity of adjacent nodes. For v i ∈V, the adjacent k neighbor node set is selected, and the connection is made through ε ij ∈E. The calculation formula is as follows:
[0091]
[0092] In the formula, v i , v j are adjacent nodes, ψ are two linear transformation functions, the multiplication obtains the adjacency matrix, s represents the connection relationship between two nodes, and the obtained connection edge represents the time domain feature.
[0093] S33, the radar point cloud front and rear image pairs are established, the pose correlation of adjacent frames is calculated using the hidden space normal estimation method, and the connection edge established represents the time domain feature used to connect the spatial features of the nodes;
[0094] S34, the point cloud feature map is input into the space-time graph convolution model, and the time dimension correlation is obtained, and the specific design details are as follows:
[0095]
[0096] In the formula, Z k represents the node attribute of the kth layer, W k is a parameter matrix, A is a normalized adjacency matrix, represents L2 normalization, and σ is a LeakyReLU activation function. After completing the graph convolution, a fully connected layer is used to obtain the final single-frame hidden space normal estimation, and the gap between adjacent frame poses is obtained.
[0097] S35, the multi-frame point clouds in the motion scene are pose calibrated and fused to obtain enhanced point cloud data.
[0098] The final goal of multi-frame 4D point cloud space fusion is to realize the splicing and fusion of sparse point clouds in the motion scene while ensuring the confidence. In this example, by splicing the 4D point clouds of consecutive frames in the motion scene, the discontinuity caused by the spatial displacement and pose transformation of the unmanned aerial vehicle is reduced. Around an experimental scene, circular and square trajectories are used to collect consecutive frame point clouds, and 8 frames of images are obtained at an interval of 45° for fusion. The final test results of the example show that the scene point cloud image does not change in shape due to the motion of the unmanned aerial vehicle, and the multi-frame point cloud pose correction and fusion in the motion scene are realized.
[0099] S4, the radar point cloud and the camera image are processed in parallel, and the BEV feature map is obtained through the feature extraction network.
[0100] S41, the preprocessed radar point cloud and camera image are processed in parallel. The camera image is input into a ResNet network for multi-scale feature extraction, and a FPN network is used as a neck to fuse high-order and low-order feature information and realize feature enhancement;
[0101] S42, the two-dimensional features are mapped to a three-dimensional space using orthogonal feature transformation, and the radar coordinate points are projected to the camera coordinate system, and the formula is as follows:
[0102]
[0103] In the formula, the radar coordinate point is represented as (x r ,y y ,z r ), the camera coordinate point is represented as (x c ,y c ,z c ), and represent the rotation matrix and conversion vector of the radar to the camera coordinate system.
[0104] S43, the camera coordinate point is converted to the pixel coordinate system, and the formula is as follows:
[0105]
[0106] In the formula, f x =f / dx, f y =f / dy, f represents the focal length of the camera, dx and dy are the width and length of the image respectively, (c x ,c y ) is the principal point coordinate of the pixel coordinate system, and (u c ,v c ) is the angular coordinate of the pixel coordinate system.
[0107] S44, feature sampling is performed to generate a voxel space, and the image BEV feature map is obtained by voxel encoding. First, cumulative summation is used for feature aggregation, and feature encoding is performed for each voxel. If the adjacent voxels are different, difference calculation is performed to eliminate redundant features, and finally the z-axis is eliminated to obtain the image BEV feature map F c ∈R X×Y . The voxel space expression is as follows:
[0108] V(x,y,z)=D(u c ,v v ,d)×F(u c ,v v ) (12)
[0109] In the formula, V∈R C×Z×Y×Xis the voxel space feature, (x, y, z) is the voxel coordinate, F represents the image information, D is the probability feature of F in the voxel space, and d is the reference depth.
[0110] S45, input the 4D millimeter wave radar point cloud into the cylinder feature extraction network based on kernel density clustering, introduce adaptive kernel density estimation to enhance semantic features, and generate a cylinder space. The radar point cloud features include velocity, position and reflection intensity information. The feature extraction network creates a (D, P, N) tensor, where D is the point cloud feature, P is the number of cylinders, and N is the number of points in each cylinder.
[0111] Each point calculates the difference between itself and surrounding points in the neighborhood in multi-scale features through the kernel density block to obtain the density estimation of the specific point. The Gaussian kernel function is as follows:
[0112]
[0113] In the formula, f(p) represents the density of each point, M p represents the number of other points in the neighborhood, R is the radius, Dop is the Doppler information, and K is the Gaussian kernel function.
[0114] Feature extraction is performed on the tensorized point cloud data, which is linearized along the D axis, batch normalized and converted to a (C, P, N) tensor through a Relu layer, where C represents the mapped features. Then, maximum pooling and average pooling are performed along the N-dimensional channel to obtain a (C, P) tensor.
[0115] S46, introduce a voxelization method to obtain a shape tensor, extract deep features through three-dimensional sparse convolution, normalization layer and ReLU layer, and finally obtain a pseudo-image (C k , H, W) based on kernel density.
[0116] S5, use a double-headed channel attention mechanism and a cross-attention mechanism to fuse BEV features.
[0117] S51, perform global compression and channel excitation through the double-headed channel attention mechanism, and the expression is as follows:
[0118]
[0119] In the formula, F' represents the image information after feature mapping, F represents the image information, M c is a channel attention module, which is composed of two parallel maximum pooling layers and average pooling layers, then connected with an MLP network with two hidden layers, and the output adopts a sigmoid activation function to form a channel attention mapping M c (F)∈R C×1×1On the one hand, the attention weight of camera BEV creation can reduce the influence of 4D radar noise points on the detection result. On the other hand, the radar attention weight can enhance the accuracy of potential objects at a long distance in the radar camera and improve the detection performance.
[0120] S52, fuse radar camera BEV features using a cross-attention mechanism, extract common features and difference features, and assign weights to query (Q)-key value (K, V);
[0121] S53, module lightening, introducing deep separable convolution to downsample the generation of K and V, to obtain smaller K and V.
[0122] S6, in order to improve the accuracy of model detection, the application adopts an anchor box-based detection head, and differentiates the direction of different category targets such as pedestrians and vehicles, and sets a single size for each category of anchor box in the example: car [3.9, 1.6, 1.56], person [0.8, 0.6, 1.73], and person riding a bicycle [1.76, 0.6, 1.73]. In regression, multiple parameters are predicted for each anchor box, including the midpoint coordinates of the anchor box (x, y, z), the anchor box length, width and height (h, w, l), and the rotation angle θ.
[0123] S7, verify the multi-modal fusion unmanned aerial vehicle dynamic environment 3D target detection method. The experimental environment uses Ubuntu 22.04, the GPU is Nvidia RTX4090, and the CPU is Intel Core i9-13700K. In order to verify the feasibility of the application under different data sets, TJ4DRadSet data set and K-Radar data set are used for training and verification, both of which contain camera and 4D radar sensor data. TJ4DRadSet data set is collected in different traffic scenes, including trucks, cars and pedestrians of different categories. The K-Radar data set collects scenes including various roads (city, suburb, highway, etc.), multiple time periods (day, night), and various weather conditions (sunny, cloudy, rainy, foggy, snowy, etc.), which can verify the robustness of the application in bad weather such as rain and fog. After training, the overall average precision mAP of the model is significantly improved, which is better than other existing baseline models.
[0124] Through quantitative analysis, it can be known that the method proposed in the application has robustness and implementability, and compared with other real-time 3D target detection methods, the application meets the requirements of 3D target detection in the motion process of unmanned aerial vehicles.
[0125] Embodiment 2
[0126] Reference Figure 5The embodiment relates to a kind of unmanned aerial vehicle dynamic environment 3D target detection device based on multi-modal fusion, including memory and one or more processors, the executable code is stored in the memory, when the one or more processors execute the executable code, for implementing the 3D target detection method of unmanned aerial vehicle dynamic environment based on multi-modal fusion of embodiment 1.
[0127] Embodiment 3
[0128] The embodiment relates to a kind of computer readable storage medium, which stores program, the program is executed by processor, and a kind of 3D target detection method of unmanned aerial vehicle dynamic environment based on multi-modal fusion of the application is realized.
[0129] The above is only preferred specific embodiment of the application, but the protection scope of the application is not limited to this.For those skilled in the art, the technical solutions recorded in the foregoing examples can still be modified, or some technical features can be replaced equivalently.Any modification, equivalent replacement, etc.inside the spirit and principles of the application should be included in the protection scope of the application.
Claims
1. A method for detecting 3D targets in dynamic environment of a UAV based on multi-modal fusion, characterized in that, The method comprises the following detection steps: S1, calibrating the 4D millimeter wave radar and the camera to obtain the pose relationship between the millimeter wave radar and the camera; S2, preprocessing the 4D millimeter wave radar point cloud data to obtain a denoised point cloud image, and extracting features from the denoised point cloud image; S3, using hidden space normal estimation to obtain a pose matrix between adjacent frames, and then performing global discretization analysis to realize multi-frame point cloud pose correction and fusion in a motion scene and enhance the point cloud data; S4, processing the radar point cloud and the camera image to obtain a BEV feature map through a feature extraction network; S5, fusing the BEV features using a double-head channel attention mechanism and a cross attention mechanism; specifically comprising: S51, performing global compression and channel excitation through the double-head channel attention mechanism, and the expression is as follows: In the formula, F' represents the BEV feature map after feature mapping, M c is a channel attention module, and F represents a BEV feature map; S52, fusing the radar camera BEV feature map using the cross attention mechanism to extract common features and difference features, and assigning weights to the query (Q)-key value (K, V); S53, lightening the module, and downsampling K and V using a depth separable convolution to obtain smaller K and V; S6, designing an anchor box-based detection head module, and differentiating settings for different detection targets; S7, verifying using different data sets and different test scenes to obtain 3D target detection results.
2. The method of claim 1, wherein, The specific steps of S1 include: S11, separately calibrating the camera and the radar to obtain the internal and external parameters of the sensors; S12, preprocessing the radar points and image points in homogeneous coordinates to obtain normalized coordinates, and the normalization formula is as follows: In the formula, and These are the homogeneous coordinates of the radar point and the image point, respectively, H = [h ij ] 3×4 Let T be the transformation matrix. p and T q They are and The normalized matrices, m1, m2 and m3 are respectively The three elements of the average vector, n1 and n2, are respectively The two elements of the average vector, d1 and d2, are respectively and The norm average; S13, substituting the radar camera data pair to solve a three-dimensional direct linear transformation (DLT) to obtain a radar camera inter-transformation matrix, and the coordinate conversion formula is as follows: 。 3. The method of claim 1, wherein, The specific steps of S2 include: calculating the sparsity of the 4D millimeter wave radar point cloud using a statistical filtering algorithm to remove outliers from the data to obtain a denoised point cloud image, and inputting the denoised point cloud image into a STEM module to extract initial features of the data; the outlier calculation formula is as follows: D max = μ + std x σ (6) In the formula, μ represents the average distance of a certain point to other points, σ represents the standard deviation of the distance, and std represents the standard deviation coefficient.
4. The unmanned aerial vehicle dynamic environment 3D target detection method based on multi-modal fusion of claim 1, wherein, The specific steps of S3 include: S31, converting each frame of point cloud image information into a point cloud feature matrix to establish a spatial feature node; S32, calculating the similarity of adjacent node features, constructing a connection edge according to the feature similarity, obtaining a three-dimensional space topology, and the calculation formula is as follows: wherein, v i , v j is the adjacent node, ψ is two linear transformation functions, multiplied to get the adjacency matrix, s represents the connection between two nodes; S33, establishing a radar point cloud image pair before and after, using hidden space normal estimation method to calculate the pose correlation of adjacent frames; S34, inputting the point cloud feature map into a space-time graph convolution model to obtain time dimension correlation, and the specific design details are as follows: In the formula, Z k represent the node attributes of the kth layer, W k is a parameter matrix, A is a normalized adjacency matrix, represent L2 normalization, and σ is a LeakyReLU activation function. S35, performing pose correction and fusion on multiple frames of point cloud to obtain enhanced point cloud data.
5. The multi-modal fusion-based dynamic environment 3D target detection method for unmanned aerial vehicles according to claim 1, characterized in that, The specific steps of S4 include: S41, inputting the camera image into a ResNet network to extract multi-scale features; S42, projecting the radar coordinate points to the camera coordinate system using orthogonal feature transformation to realize feature enhancement, and the formula is as follows: wherein the radar coordinate point is represented by (x r ,y y ,z r ), the camera coordinate point is represented by (x c ,y c ,z c ), and represent the radar-to-camera coordinate system rotation matrix and translation vector, respectively. S43, converting the camera coordinate points to the pixel coordinate system, and the formula is as follows: where f x = f / dx, f y = f / dy, f represents the focal length of the camera, dx and dy are the width and length of the image respectively, (c x , c y ) are the principal point coordinates of the pixel coordinate system, and (u c , v c ) are the angle coordinates of the pixel coordinate system. S44, feature sampling is performed, a voxel space is generated, and an image BEV feature map is obtained through voxel coding; the voxel space expression is as follows: V(x, y, z) = D(u c ,v v ,d) x F(u c ,v v ) (12) In the formula, V∈R C×Z×Y×X is a voxel space feature, (x, y, z) is a voxel coordinate, F represents image information, D is a probability feature of F in the voxel space, and d is a reference depth. S45, the 4D millimeter wave radar point cloud is input into a column feature extraction network based on kernel density clustering, an adaptive kernel density estimation method is used, and a column space is generated; the Gaussian kernel function is as follows: where f(p) represents the density of each point, M p represents the number of other points in the neighborhood, R is the radius, Dop is the Doppler information, and K is a Gaussian kernel function. S46, a voxelization method is introduced, three-dimensional sparse convolution is used to accelerate the inference process, column space information is projected to a BEV plane, and a 4D millimeter wave radar BEV pseudo-image is obtained.
6. The unmanned aerial vehicle dynamic environment 3D target detection method based on multi-modal fusion of claim 1, wherein, The specific steps of step S6 include: different targets are set differently based on an anchor box detection head, and in regression, multiple parameters are predicted for each anchor box to improve the detection accuracy of the model.
7. The unmanned aerial vehicle dynamic environment 3D target detection method based on multi-modal fusion of claim 1, wherein, The specific steps of step S7 include: image and radar data of different urban environments, weather, and time periods are selected as a data set, input into the model for training, the best model is retained according to the evaluation result, and the test set is input into the best model to obtain the test result.
8. A multi-modal fusion-based unmanned aerial vehicle dynamic environment 3D target detection device, characterized in that, The memory stores executable code, and the one or more processors execute the executable code to implement the method of any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, The program is stored thereon and is executed by the processor to implement the method of any one of claims 1-7.