A Monocular 3D Object Detection Method, System, Medium and Terminal for Unmanned Aerial Vehicles
Through deep convolutional neural network and geometric prior transformation, the problem of three-dimensional object detection in the drone's perspective is solved, accurate object detection from two-dimensional images to three-dimensional space is achieved, and the drone's monocular 3D object detection effect is improved.
Patent Information
- Application Number
- CN202210657179.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-10
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-06-10
AI Technical Summary
The existing drone target detection is limited to two-dimensional image space and cannot effectively perceive the real three-dimensional physical space. The existing monocular 3D object detection method cannot be applied to diverse perspectives and serious deformation problems under the perspective of the drone.
A deep convolutional neural network is used to extract feature images, combine the height prediction module and geometric prior transformation to convert the feature image perspective into a bird's-eye view angle, generate three-dimensional bird's-eye view features through geometric prior deformable transformation, and supervise the detection results using a loss function.
It realizes accurate object detection from two-dimensional images to three-dimensional space, alleviates the diversity and deformation problems of the drone's perspective, and improves the monocular 3D object detection effect of the drone.
Smart Images

Figure CN117274835B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to a monocular 3D object detection method, system, medium and terminal for an unmanned aerial vehicle. Background Art
[0002] Unmanned aerial vehicles (UAVs) have greatly enhanced the ability of humans to perceive the world and have achieved remarkable success in a wide range of applications, including agriculture, aerial photography, air transportation, security, and disaster relief and rescue. A unique advantage of UAVs is their ability to move freely in three-dimensional space, which makes them have great potential in the understanding of three-dimensional scenes. However, the current object detection of UAVs is only limited to the two-dimensional bounding boxes obtained in the two-dimensional image space and does not have the effective perception of the real three-dimensional physical space.
[0003] There are three key challenges in achieving effective perception of the three-dimensional physical world based on a single image input from a drone's perspective: a well-organized dataset, a suitable 3D object representation based on the drone's perspective, and an effective 3D object detection method based on the bird's-eye view. First, currently, drone perception datasets only have 2D annotations within two-dimensional images and cannot be used for 3D object detection. Second, the 3D object bounding box representation commonly used in self-driving cars is not applicable to the drone's perspective because, from the bird's-eye view, the height of the object is negligible compared to the flight height of the drone, and the height of the object itself is almost impossible to estimate. On the other hand, from the self-driving car's perspective, the self-driving car and the object are on the same plane, and the elevation representation of the object is ignored. However, in the drone's perspective, it is crucial to locate the object's position in the real 3D physical world. Therefore, the elevation of the object and its position in the bird's-eye view together constitute the 3D object representation in the drone's perspective. Third, existing monocular 3D object detection methods based on self-driving cars are not applicable to drones because currently, more of them are based on the three-dimensional understanding of the vehicle's frontal view, while for drones, object detection from the bird's-eye view has more variable perspectives and more severe deformation problems. According to the imaging principle, the deformation of distant objects is more severe. Due to the large variations in the scale and direction of objects in the bird's-eye view image, many existing detection datasets and algorithms cannot be directly applied to bird's-eye view images. To alleviate the dataset problem, a large number of object detection datasets for bird's-eye view images have been proposed. Unfortunately, current bird's-eye view image object detection datasets and algorithms only focus on the two-dimensional view space and cannot directly achieve the understanding of three-dimensional scenes. Recently, in the field of autonomous driving scenarios, 3D object detection is also on the rise. To utilize the flexible maneuverability of drones to fill the huge gap in 3D object detection from the bird's-eye view, the present invention proposes a new task of monocular 3D object detection based on drones and provides a well-organized dataset. The goal of monocular 3D object detection is to detect objects in 3D space in a given two-dimensional image. These methods can be divided into three types: methods that directly regress from 2D representations, methods that perform 3D object detection based on point clouds after reconstructing pseudo-3D point clouds from estimated depths, and methods that perform 3D object detection based on grid-based 3D features. The direct method first detects the bounding box and then uses geometric constraints to regress the 3D box. Due to the lack of explicit depth information, it usually performs poorly. The depth-based method first estimates the depth map, then combines the depth map with image features to generate a pseudo-three-dimensional point cloud, and then uses the 3D object detection method based on point clouds to obtain the 3D object representation. Since the system is not end-to-end, errors in the first-stage depth estimation cannot be alleviated, and the overall inference process is relatively cumbersome, requiring two models and having low inference efficiency. The method based on grid-based 3D features infers the 3D representation of the scene based on 2D image features and the perspective of the image capture, and performs 3D object detection based on the recovered 3D scene representation.Current monocular 3D object detection methods, where the field-of-view image is parallel to the imaging plane and the scene, are not applicable to 3D object detection from the perspective of an unmanned aerial vehicle (UAV) with diverse perspective changes and severe deformation problems. Summary of the Invention
[0004] Aiming at the defects in the prior art, an object of the present invention is to provide a monocular 3D object detection method, system, medium and terminal for an unmanned aerial vehicle.
[0005] According to one aspect of the present invention, a monocular 3D object detection method for an unmanned aerial vehicle is provided, including:
[0006] Using a deep convolutional neural network to extract a feature image;
[0007] Using a height prediction module to predict the height of each feature point in the feature image, and transforming the perspective of the feature image into a bird's-eye view through geometric prior transformation to obtain the height of each feature point in the bird's-eye view;
[0008] Based on the feature image and the height of each feature point in the bird's-eye view, obtaining a three-dimensional bird's-eye view feature;
[0009] Decoding the three-dimensional bird's-eye view feature to obtain a detection result.
[0010] Preferably, the using a deep convolutional neural network to extract a feature image includes:
[0011] Using a backbone network to extract an image feature F from a data image (rv) , where
[0012]
[0013] where f backbone is a convolutional neural network based on DLA, and H R , W R , C are the length, width and channel dimensions of the feature map.
[0014] Preferably, the using a height prediction module to predict the height of each feature point in the feature image, and transforming the perspective of the feature image into a bird's-eye view through geometric prior transformation to obtain the height of each feature point in the bird's-eye view includes:
[0015] Using a height estimation module f altitude to predict the elevation height category of each image feature point, and performing geometric prior transformation to obtain the elevation height category A (bev) of each coordinate in the bird's-eye view:
[0016] A (bev) = f a1titude (F(rv) ) ∈ R {X×Y×Z} ,
[0017] where f altitude is the height estimation module, X and Y represent the sensing ranges along the X-axis and Y-axis respectively, Z is the number of height intervals, and each element A (bev) (x, y, z) reflects the confidence that the position of coordinates (x, y) in the bird's-eye view is in the z-th height interval.
[0018] Preferably, obtaining the three-dimensional bird's-eye view features based on the feature image and the height of each feature point in the bird's-eye view includes:
[0019] Obtaining the three-dimensional bird's-eye view feature F based on the deformable transformation with geometric prior, based on the 2D image features and the bird's-eye view height estimation: (bev) :
[0020] F (bev) = f deform (F (rv) , A (bev) ) ∈ R X×Y×C
[0021] where f deform is the deformable transformation with geometric prior.
[0022] Preferably, the deformable transformation f with geometric prior deform is to optimize the three-dimensional bird's-eye view features by using the geometric prior information transformation provided by the internal and external camera parameters and the learning transformation of deformable convolution;
[0023] Among them, the geometric prior information transformation provided by the internal and external camera parameters includes:
[0024] Defining the mapping between the global coordinates to the local image pixel coordinates through the camera projection matrix P,
[0025]
[0026] The relationship between the bird's-eye view and the image features is:
[0027]
[0028] where G (bev) (x, y, z) is the three-dimensional bird's-eye view feature in the coordinate system (x, y, z);
[0029] The learning transformation of the deformable convolution includes:
[0030] The deformable convolution DCN layer uses a trainable offset to alleviate the deformation problem in the perspective transformation,
[0031]
[0032] Among them, [;] represents concatenation, is a feature point under the x-axis; is a deformable feature obtained by learning;
[0033] The features of the geometric prior transformation and the deformable features obtained by learning Use a residual structure to obtain the final bird's-eye view feature F (bev) :
[0034]
[0035] Preferably, decoding the three-dimensional bird's-eye view feature to obtain a detection result includes:
[0036] Based on the feature decoding f decoder to obtain the object bounding box:
[0037] o i = f decoder (F (bev) , A (bev) )
[0038] where o i = (x i , y i , w i , l i , θ i , a i , c i ) represents an object with coordinates at (x i , y i ), width w i , length l i , deviation angle θ i , height a i , and object category c i . F (bev) is the bird's-eye view feature, and A (bev) is the elevation category of each coordinate in the bird's-eye view.
[0039] Preferably, two loss functions are used to separately supervise the elevation classification and object detection based on the bird's-eye view;
[0040] For the elevation classification, A (bev) is the estimated elevation category, and only foreground objects are supervised. The classification loss is
[0041] where is a mask of the foreground object, where 1 indicates the presence of the foreground object at that position and 0 indicates the absence of the foreground object. is the true elevation category;
[0042] For bird's-eye view object detection, (x, y, w, l, θ) is the detected object bounding box. is the true object bounding box, and the regression loss is:
[0043] where (x, y) represents the coordinates of the detected object, (w, l) represent the width and length of the detected object respectively, and θ represents the rotation angle of the detected object. represents the coordinates of the true object bounding box. respectively represent the width and length of the true object bounding box. represents the rotation angle of the true object bounding box.
[0044] According to the second aspect of the present invention, there is provided a drone-based monocular 3D object detection system for implementing the method of any one of the above, including:
[0045] An image feature extraction module for extracting features of the input drone bird's-eye view image using a deep convolutional neural network.
[0046] A height prediction module for predicting the height of each feature point in the image, obtaining the height of each feature point using a height classification module, and transforming the picture perspective to a bird's-eye view perspective through geometric prior transformation, thereby obtaining the height of each feature point in the bird's-eye view.
[0047] A deformable transformation module based on geometric prior for jointly converting the height of each feature point obtained by the height prediction module and the features of the two-dimensional image extracted by the image feature extraction module into three-dimensional bird's-eye view features.
[0048] An object bounding box classification and regression module for encoding the three-dimensional bird's-eye view features into object bounding boxes with object categories, where the two-dimensional picture features are decoded into two-dimensional object bounding boxes with categories, and the three-dimensional bird's-eye view features and the predicted height of each feature point in the bird's-eye view are jointly decoded into defined three-dimensional object bounding boxes.
[0049] According to the third aspect of the present invention, there is provided a terminal including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it can be used to execute the method of any one of the above, or run the above system.
[0050] According to a fourth aspect of the present invention, there is provided a computer-readable storage medium having stored thereon a computer program which, when executed by a processor, can be used for any of the described methods or for executing the described system.
[0051] Compared with the prior art, the present invention has the following beneficial effects:
[0052] A monocular 3D object detection method and system for an unmanned aerial vehicle in an embodiment of the present invention perform object detection from two perspectives simultaneously, combining the BEV and RV perspectives, and have the advantages of providing compensation information and promoting each other, thereby improving the monocular 3D object detection effect of the unmanned aerial vehicle.
[0053] A monocular 3D object detection method and system for an unmanned aerial vehicle in an embodiment of the present invention infer a 3D object representation from a 2D bird's-eye view image through altitude classification estimation and deformable transformation based on geometric priors; the deformable transformation based on geometric priors combines the stability of the geometric prior transformation and the learnable ability of the deformable module to effectively handle the diversity of bird's-eye view angle changes and alleviate the serious deformation problem, obtaining more accurate 3D features, thereby improving the 3D object detection effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] By reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings, other features, objects, and advantages of the present invention will become more apparent:
[0055] Figure 1 It is a flowchart of a monocular 3D object detection system based on an unmanned aerial vehicle according to an embodiment of the present invention;
[0056] Figure 2 It is a block diagram of a monocular 3D object detection system based on an unmanned aerial vehicle according to an embodiment of the present invention.
[0057] Reference numerals: 1 - image feature extraction module, 2 - altitude prediction module, 3 - deformable transformation module based on geometric priors, 4 - object box decoding module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several modifications and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0059] The present invention provides an embodiment, a monocular 3D object detection method for an unmanned aerial vehicle, including:
[0060] S100, extracting a feature image using a deep convolutional neural network;
[0061] S200. Use a height prediction module to predict the height of each feature point in the feature image obtained in S100, and change the perspective of the feature image in S100 to a bird's-eye view through a geometric prior transformation, so as to obtain the height of each feature point in the bird's-eye view;
[0062] S300. Based on the feature image obtained in S100 and the height of each feature point in the bird's-eye view obtained in S200, obtain three-dimensional bird's-eye view features;
[0063] S400. Decode the three-dimensional bird's-eye view features obtained in S300 to obtain the detection result.
[0064] Based on the above embodiments for further optimization, a preferred embodiment of the present invention provides a monocular 3D object detection method for an unmanned aerial vehicle, and its flowchart is as Figure 1 shown, including:
[0065] S11. For the input image, use a deep convolutional neural network to extract an image feature map;
[0066] S12. Use a height prediction module to predict the height of each feature point in the image feature map extracted in S11, and change the perspective of the image feature map in S11 after height prediction to a bird's-eye view through a geometric prior transformation, thereby obtaining the height of each feature point under the bird's-eye view; use a deformable transformation based on geometric prior to alleviate the serious deformation problem in the bird's-eye view transformation, and at the same time use the predicted height of each feature point under the bird's-eye view to convert the two-dimensional picture features into more accurate three-dimensional bird's-eye view features;
[0067] S13. Decode the picture features into two-dimensional object boxes with categories; decode the bird's-eye view features and the predicted height of each point under the bird's-eye view into three-dimensional object boxes.
[0068] In the above embodiments of the present invention, a deformable transformation based on geometric prior is proposed by combining the prior knowledge of camera pose parameters and the learning ability of deformable convolution, which specifically alleviates the problems of variable perspectives and serious deformations under the unmanned aerial vehicle's perspective, realizes an accurate transformation from the unmanned aerial vehicle's picture view to the 3D bird's-eye view, thereby obtaining a more accurate 3D object representation, and thus improving the effect of 3D object detection; at the same time, perform object detection from two perspectives, combining the BEV and RV perspectives, which has the advantages of providing compensation information and promoting each other, thereby improving the monocular 3D object detection effect of the unmanned aerial vehicle.
[0069] In a preferred embodiment of the present invention, S11 uses a backbone network to extract image (Range View, RV) features, and the specific relationship is
[0070]
[0071] where f backbone is the backbone network based on DLA, and H R , W R , and C are the height, width, and channel dimensions of the feature map.
[0072] In a preferred embodiment of the present invention, based on the above S11 to obtain the RV features of the image, S12 is implemented. The altitude category of each RV feature point is predicted based on the height estimation module, and geometric prior transformation is performed to obtain the altitude category of each coordinate in the Birds'Eye View (BEV). The specific relationship is:
[0073] A (bev) = f altitude (F (rv) ) ∈ R {X×Y×Z}
[0074] where f altimde is the height estimation module, X and Y represent the lengths of the perception fields along the X-axis and Y-axis, and Z is the number of altitude categories. Each element A (bev) (x, y, z) reflects the confidence of the z-th height in the (x, y)-th position bin in the BEV.
[0075] The altitude category estimation module proposed in the above embodiment of the present invention. Its specific objective is to locate objects on the Z-axis. Because the aerial view encounters serious long-distance problems, compared with the distance between the object and the drone, the height difference between different objects is relatively small. It is almost impossible to accurately locate objects in a continuous manner. To alleviate the difficulty of altitude, the height is divided into multiple levels, and the regression task is changed to a classification task, using a classification task to replace the regression task. The specific implementation process is as follows:
[0076] First, estimate the height level of each RV pixel. The image background can provide rich height cues. The specific relationship:
[0077]
[0078] where conv is a 1×1 convolutional layer;
[0079] Next, convert the picture view to a bird's-eye view, and t is the conversion function;
[0080] A (bev) = t(A rv ) ∈ R X×Y×Z
[0081] Then, the output A (bev) can reflect the height category information of each BEV position.
[0082] Furthermore, in S12, it is proposed to input the RV feature and the estimated height category A (bev) into a deformable transformation based on geometric priors to generate bird's-eye view features. The specific relationship is as follows:
[0083] F (bev) = f deform (F (rv) , A (bev) ) ∈ R X×Y×C
[0084] where f deform is the proposed deformable transformation based on geometric priors. At the same time, the geometric prior information provided by the internal and external camera parameters and the learning ability of deformable convolution (Deformable Convolution Network, DCN) are used to obtain more accurate bird's-eye view features. Referring to Figure 2 shown in the figure, it consists of a transformation module based on geometric prior information and a deformable convolution network module. These two modules are used to convert the given 2D RV features into 3D BEV features, thereby realizing 3D object detection. Since the input is 2D image features, the height dimension is missing without an additional depth sensor. To make up for the missing height information, this embodiment considers solutions from two aspects: 1) weighting along the Z-axis; 2) deforming features along the X and Y axes.
[0085] First, use the transformation of the geometric prior information obtained from the internal and external camera parameters to generate BEV representations at all possible heights; then, weight them with the height confidence estimated by the height classifier. The weighted BEV features are averaged along the height axis and folded into BEV features. Then, a trainable deformable convolution network (DCN) is used to adaptively correct the distortion of the BEV features caused by inaccurate heights. Using DCN in this view transformation stage can flexibly perform spatial sampling with additional offsets, which can help fine-tune the features of the geometric transformation. A residual structure is used to combine the stability of the geometric prior transformation and the adaptability of the deformable convolution.
[0086] The specific implementation process is as follows:
[0087] The transformation based on geometric prior information is a non-parametric view transformation method. The camera projection matrix P defines the mapping between the global coordinates to the local image pixel coordinates .
[0088]
[0089] At the same time, the relationship between BEV and RV can be expressed as
[0090]
[0091] Through geometric prior transformations at all possible heights, a flat but "stereoscopic" BEV feature is obtained. The geometric prior transformation is non-parametric and lacks learnable flexibility. Ideally, the BEV feature can fully represent the real world if the height is accurately predicted. However, the severe long-distance problem and the aerial perspective make the height estimation particularly difficult. Therefore, it is expected that the transformed BEV feature of the geometric prior information encounters spatial sampling noise.
[0092] To facilitate better transformation, a DCN layer is cascaded to enhance geometric space sampling with trainable offsets. This embodiment further connects the coordinates with the BEV feature to guide offset learning with the BEV feature. Since the coordinates can imply the geometric prior of the network, that is, the perception field increases with the increase of the distance between the object and the camera, which means that the perturbed area is large when far and small when near. Specifically, it is manifested as Finally, this embodiment uses a residual structure to combine the features of the geometric prior transformation and the adaptive deformable features to obtain the final BEV feature:
[0093]
[0094] In a preferred embodiment of the present invention, based on the BEV feature obtained in S12 above, and the image feature extracted by the backbone network, a dual-view object detection system is further proposed to implement S13. This system can simultaneously perceive objects in the two-dimensional image space and the three-dimensional physical space. Since these two views can promote each other, the two-dimensional image space can provide details of the object, such as color and shape, and the smooth image background can help understand the object. The three-dimensional space can provide more accurate spatial information. The implicit consistency of their information comes from the supervision of the two views, including RV and BEV, which can help reduce the errors of each other, so as to perceive the object more accurately.
[0095] The two views share the same backbone. The RV decoder locates the object in the two-dimensional image space, while the BEV decoder locates the object in the three-dimensional space by using the proposed altitude category estimation and geometric deformation transformation method. The process of obtaining the object box is as follows:
[0096] o i =f decoder (F (bev) ,A (bev) )
[0097] where o i =(x i ,y i ,w i, l i , θ i , a i , c i ) represents an object whose coordinates are located at (x i , y i ), with a width of w i , a length of l i , a deviation angle of θ i , a height of a i , and the object type is c i , where i represents the i-th detected object.
[0098] In other embodiments of the present invention, in order to train the entire system, two loss functions are used to supervise two tasks: altitude classification and BEV-based object detection; for altitude classification, let A (bev) be the estimated altitude category, and only foreground objects are supervised. The classification loss is The imbalance problem is alleviated by using the Focal loss function. For BEV object detection. (x, y, w, l, θ) is the detected object box, is the ground truth object box, then the regression loss is specifically as follows: where (x, y) represents the coordinates of the detected object, (w, l) represent the width and length of the detected object respectively, and θ represents the rotation angle of the detected object. represents the ground truth object box coordinates, represent the width and length of the ground truth object box respectively, represents the rotation angle of the ground truth object box.
[0099] Based on the same inventive concept, the present invention also provides a drone-based monocular 3D object detection system, which is used for the above-mentioned drone-based monocular 3D object detection method, and includes: an image feature extraction module, a height prediction module, a deformable transformation module based on geometric prior, and an object box classification and regression module; wherein,
[0100] The image feature extraction module is used to extract the features of the input drone aerial view image using a deep convolutional neural network;
[0101] The height prediction module is used to predict the height of each feature point in the image, obtain the height of each feature point using the height classification module, and transform the picture perspective to the bird's-eye view through geometric prior transformation, thereby obtaining the height of each feature point in the bird's-eye view;
[0102] The deformable transformation based on geometric prior is used to alleviate the serious deformation problem in the perspective transformation from the picture perspective to the bird's-eye view perspective. The height of each feature point obtained by the height prediction module and the two-dimensional picture feature points are jointly transformed into accurate three-dimensional bird's-eye view features;
[0103] The object box classification and regression module is used to decode the feature points into object boxes with object categories, where the two-dimensional picture features are decoded into two-dimensional object boxes with categories, and the three-dimensional bird's-eye view features and the height of each point under the predicted bird's-eye view are jointly decoded into defined three-dimensional object boxes.
[0104] Based on the same inventive concept, in other embodiments of the present invention, a terminal is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the above processor executes the above program, it can be used to execute the method of any one of the above, or, run the above system.
[0105] Based on the same inventive concept, in other embodiments of the present invention, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, it can be used for any one of the above methods, or, execute the above system.
[0106] It should be noted that the steps in the method provided by the present invention can be implemented by corresponding modules, devices, units, etc. in the system. Those skilled in the art can refer to the technical solutions of the system to implement the step flow of the method. That is, the embodiments in the system can be understood as preferred examples for implementing the method, and will not be elaborated here.
[0107] Those skilled in the art know that in addition to implementing the system and its various devices provided by the present invention in the form of pure computer-readable program codes, the method steps can be logically programmed to enable the system and its various devices provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc. to achieve the same functions. Therefore, the system and its various devices provided by the present invention can be regarded as a hardware component, and the devices included therein for implementing various functions can also be regarded as the structure within the hardware component; the devices for implementing various functions can also be regarded as both software modules for implementing the method and the structure within the hardware component.
[0108] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various deformations or modifications within the scope of the claims, which does not affect the essence of the present invention. The above preferred features can be combined arbitrarily without conflict.
Claims
1. A monocular 3D object detection method for an unmanned aerial vehicle, characterized in that, Including: Extracting a feature image using a deep convolutional neural network; Predicting the height of each feature point in the feature image using a height prediction module, and transforming the perspective of the feature image into a bird's-eye view through geometric prior transformation to obtain the height of each feature point in the bird's-eye view; Obtaining three-dimensional bird's-eye view features based on the height of each feature point in the feature image and the bird's-eye view; Decoding the three-dimensional bird's-eye view features to obtain a detection result; The step of using the height prediction module to predict the height of each feature point in the feature image, transforming the perspective of the feature image into a bird's-eye view through geometric prior transformation, and obtaining the height of each feature point in the bird's-eye view includes: Use the altitude estimation module f altitude Predict the altitude class of each image feature point and perform a geometric prior transformation to obtain the altitude class A of each coordinate in the bird's-eye view (bev) : A (bev) = f altitude (F (rv) ) ∈ R {X×Y×Z} , where f altitude is the height estimation module, X and Y represent the sensing ranges along the X-axis and Y-axis, Z is the number of height intervals, and each element A (bev) (x, y, z) reflects the confidence that the position of the coordinate (x, y) in the bird's-eye view is in the z-th height interval; The step of obtaining three-dimensional bird's-eye view features based on the height of each feature point in the feature image and the bird's-eye view includes: Deformable transformation based on geometric priors, obtaining the 3D bird's-eye view feature F from 2D image features and bird's-eye view height estimation (bev) : F (bev) = f deform (F (rv) , A (bev) ) ∈ R X×Y×C where f deform is a deformable transformation based on geometric priors; The deformable transformation f based on geometric prior deform optimizes the three-dimensional bird's-eye view features by using the geometric prior information transformation provided by the internal and external camera parameters and the learning transformation of deformable convolution; Among them, the geometric prior information transformation provided by the internal and external camera parameters includes: Define the global coordinates through the camera projection matrix P to the local image pixel coordinates between the mappings, Relationship between bird's-eye view and image features is as follows: Among which G (bev) (x, y, z) is a three-dimensional bird's-eye view feature in the coordinate system (x, y, z); The learnable transformation of the deformable convolution includes: The deformable convolution DCN layer uses a trainable offset to alleviate the deformation problem in perspective transformation; where [;] represents series connection, is a feature point under the x-axis; is a deformable feature obtained through learning; The features of the geometric prior transformation and the learned deformable features Use a residual structure to obtain the final bird's-eye view feature F (bev) :
2. The monocular 3D object detection method for a drone according to claim 1, characterized in that The step of extracting a feature image using a deep convolutional neural network includes: Extract image feature F from the data image using a backbone network from the data image (rv) for where f backbone is a convolutional neural network based on DLA, and H R , W R , and C are the length, width, and channel dimensions of the feature map.
3. The monocular 3D object detection method for an unmanned aerial vehicle according to claim 1, characterized in that, The step of decoding the three-dimensional bird's-eye view features to obtain a detection result includes: Feature-based decoding f decoder Obtain the object bounding box: o i = f decoder (F (bev) , A (bev) ) where o i =(x i , y i , w i , l i , θ i , a i , c i ) represents an object with coordinates located at (x i , y i ), width w i , length l i , rotation angle θ i , height a i , and object category c i . F (bev) is the bird's-eye view feature, A (bev) is the elevation category of each coordinate in the bird's-eye view, and i represents the i-th object detected.
4. The monocular 3D object detection method for an unmanned aerial vehicle according to claim 1, wherein Using two loss functions to respectively supervise the elevation category and object detection based on the bird's-eye view; For altitude class, A (bev) For the estimated altitude class, only the foreground objects are supervised, and the classification loss is wherein is a mask of the foreground object, where 1 indicates the presence of the foreground object at that position and 0 indicates the absence of the foreground object, is the true elevation class; For bird's-eye view object detection, (x, y, w, l, θ) is the detected object bounding box, is the ground truth object bounding box, and the regression loss is: Among them, (x, y) represents the coordinates of the detected object, (w, l) represent the width and length of the detected object respectively, and θ represents the rotation angle of the detected object. represents the coordinates of the true object bounding box. represent the width and length of the true object bounding box respectively. represents the rotation angle of the true object bounding box.
5. A drone-based monocular 3D object detection system for implementing the method according to any one of claims 1-4, characterized in that, Including: An image feature extraction module, which is used to extract the features of the input drone aerial image using a deep convolutional neural network; A height prediction module, which is used to predict the height of each feature point in the image, obtain the height of each feature point using a height classification module, and transform the picture perspective into a bird's-eye view through geometric prior transformation, thereby obtaining the height of each feature point under the bird's-eye view; A deformable transformation module based on geometric prior, which is used to jointly transform the height of each feature point obtained by the height prediction module and the features of the two-dimensional image extracted by the image feature extraction module into three-dimensional bird's-eye view features; An object box classification and regression module, which is used to decode the three-dimensional bird's-eye view features into an object box with an object category, where the two-dimensional picture features are decoded into a two-dimensional object box with a category, and the three-dimensional bird's-eye view features and the height of each feature point obtained by prediction under the bird's-eye view are jointly decoded into a defined three-dimensional object box.
6. A terminal, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it can be used to execute the method described in any one of claims 1-4, or, run the system described in claim 5.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it can be used to execute the method described in any one of claims 1-4, or, execute the system described in claim 5.
Citation Information
Patent Citations
Method and system for video-based positioning and mapping
CN110062871A
High-definition AR live video display method for unmanned aerial vehicle
CN110830815A