Three-dimensional object detection method, device, equipment and storage medium
Depth estimation and three-dimensional point cloud data processing are carried out through image data collected by binocular cameras. Combined with image fusion technology, high-precision detection of three-dimensional objects is achieved, solving the problem of low detection accuracy that does not rely on lidar data in the prior art.
Patent Information
- Application Number
- CN202510153613.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-12
AI Technical Summary
The prior art is difficult to achieve high-precision detection of three-dimensional objects without relying on lidar data.
By obtaining the left and right eye images of a three-dimensional object collected by a binocular camera, the depth estimation network is used to perform depth estimation to obtain a depth map and convert it into three-dimensional point cloud data. Then, the three-dimensional point cloud data is projected onto the bird's eye view and fused with the image data to detect it through the target three-dimensional object detection network to achieve high-precision detection of three-dimensional objects.
It realizes high-precision detection of three-dimensional objects without relying on lidar data, improving the accuracy and cost-effectiveness of the intelligent driving perception system.
Smart Images

Figure CN119625715B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of target detection technology, and in particular to three-dimensional object detection methods, devices, equipment and storage media. Background Art
[0002] At present, the rapid development of intelligent driving technology has accelerated the pace of technological upgrading and has put forward higher requirements for the 3D target detection capabilities of the perception system. At present, the mainstream vehicle 3D target detection methods are mainly based on two solutions: lidar and monocular vision sensors. The lidar solution is popular for its high-precision environmental perception information, but its high price increases the cost burden during mass production, making it impossible to be widely used in some low-cost models, thus limiting the popularity of intelligent driving functions. Although the monocular vision sensor solution has cost advantages, the perception information it provides is mostly one-dimensional and lacks depth information. It must rely on complex visual algorithms for data recovery, which not only increases the R&D cost, but is also easily affected by the environment. It is difficult to accurately restore the information of certain special and complex scenes, resulting in relatively low perception accuracy.
[0003] How to achieve high-precision detection of three-dimensional objects without relying on lidar data has become a problem to be solved.
[0004] The above contents are only used to assist in understanding the technical solution of the present application and do not constitute an admission that the above contents are prior art. Summary of the invention
[0005] The main purpose of this application is to provide a three-dimensional object detection method, device, equipment and storage medium, aiming to solve the technical problem of how to achieve high-precision detection of three-dimensional objects without relying on lidar data.
[0006] To achieve the above objectives, the present application proposes a three-dimensional object detection method, the method comprising:
[0007] Obtain a left-eye image and a right-eye image of a three-dimensional object captured by a binocular camera, and perform depth estimation on the left-eye image and the right-eye image through a binocular depth estimation network to obtain a depth map;
[0008] Acquire two-dimensional pixel data of the depth map and parameter information of the binocular camera, and convert the two-dimensional pixel data into three-dimensional point cloud data according to the parameter information;
[0009] Projecting the three-dimensional point cloud data onto a bird's-eye view, and fusing the left-eye image, the right-eye image, and the bird's-eye view to obtain three-dimensional object fusion data;
[0010] The three-dimensional object fusion data is detected through a target three-dimensional object detection network to obtain a three-dimensional object target detection result. The target three-dimensional object detection network includes a feature extraction layer, a spatial transformation layer, a reverse connection layer and a target detection layer. The spatial transformation layer includes a local network, a grid generator and a sampler.
[0011] In one embodiment, the step of detecting the three-dimensional object fusion data through a target three-dimensional object detection network to obtain a three-dimensional object target detection result includes:
[0012] Performing feature extraction on the three-dimensional object fusion data through the feature extraction layer to obtain an initial feature map;
[0013] Performing an affine transformation on the initial feature map through the spatial transformation layer to obtain a transformed feature map;
[0014] Combining the transformed feature map with the initial feature map through the reverse connection layer to obtain a feature enhancement map;
[0015] The target detection layer performs three-dimensional object detection on the feature enhancement image to obtain a three-dimensional object detection result.
[0016] In one embodiment, the step of performing an affine transformation on the initial feature map through the spatial transformation layer to obtain a transformed feature map includes:
[0017] Learning the features of the initial feature map through the local network to predict spatial transformation parameters;
[0018] The grid generator performs sampling transformation on the pixel coordinates of the initial feature map according to the spatial transformation parameters to obtain a sampling grid;
[0019] A transformed feature map is determined by the sampler according to the sampling grid and the initial feature map.
[0020] In one embodiment, after the step of determining the transformed feature map according to the sampling grid and the initial feature map by the sampler, the following steps are included:
[0021] Calculate the partial derivatives of the feature point abscissa and the feature point ordinate respectively according to the transformed feature map;
[0022] Deriving the sampler according to the partial derivative to obtain a back-propagation gradient;
[0023] The network parameters of the spatial transformation layer are updated according to the back-propagation gradient.
[0024] In one embodiment, the binocular depth estimation network includes a feature extraction layer and a feature matching layer;
[0025] The step of performing depth estimation on the left-eye image and the right-eye image through a binocular depth estimation network to obtain a depth map includes:
[0026] Performing feature extraction on the left image and the right image respectively through the feature extraction layer to obtain a first feature point descriptor of the left image and a second feature point descriptor of the right image;
[0027] Performing feature point matching on the first feature point descriptor and the second feature point descriptor through the feature matching layer to obtain a disparity prediction result;
[0028] A disparity map is generated according to the disparity prediction result, and the disparity map is converted into a depth map.
[0029] In one embodiment, the two-dimensional pixel data includes pixel point coordinate information and pixel point depth information, and the parameter information includes camera focal length information and camera principal point coordinate information;
[0030] The step of acquiring the two-dimensional pixel data of the depth map and the parameter information of the binocular camera, and converting the two-dimensional pixel data into three-dimensional point cloud data according to the parameter information comprises:
[0031] Obtaining the pixel coordinate information and the pixel depth information of the depth map;
[0032] Obtain the camera focal length information and the camera principal point coordinate information of the binocular camera;
[0033] The pixel point coordinate information and the pixel point depth information are converted into three-dimensional point cloud data according to the camera focal length information and the camera principal point coordinate information.
[0034] In one embodiment, before the step of acquiring a left-eye image and a right-eye image of a three-dimensional object captured by a binocular camera and performing depth estimation on the left-eye image and the right-eye image through a binocular depth estimation network to obtain a depth map, the step further includes:
[0035] Training a depth estimation network based on preset three-dimensional image data to obtain a reference depth estimation network, and obtaining a disparity matching error and a true error of the preset three-dimensional image data;
[0036] determining an estimated disparity of the preset three-dimensional image data according to the disparity matching error;
[0037] Calculating a disparity loss function of the estimated disparity and the true error;
[0038] The reference depth estimation network is optimized based on the disparity loss function to obtain a binocular depth estimation network.
[0039] In addition, to achieve the above objectives, the present application also proposes a three-dimensional object detection device, the three-dimensional object detection device comprising:
[0040] A depth estimation module is used to obtain a left-eye image and a right-eye image of a three-dimensional object captured by a binocular camera, and perform depth estimation on the left-eye image and the right-eye image through a binocular depth estimation network to obtain a depth map;
[0041] A coordinate conversion module, used to obtain the two-dimensional pixel data of the depth map and the parameter information of the binocular camera, and convert the two-dimensional pixel data into three-dimensional point cloud data according to the parameter information;
[0042] A data fusion module, used for projecting the three-dimensional point cloud data onto a bird's-eye view, and fusing the left-eye image, the right-eye image and the bird's-eye view to obtain three-dimensional object fusion data;
[0043] The target detection module is used to detect the three-dimensional object fusion data through a target three-dimensional object detection network to obtain a three-dimensional object target detection result. The target three-dimensional object detection network includes a feature extraction layer, a spatial transformation layer, a reverse connection layer and a target detection layer. The spatial transformation layer includes a local network, a grid generator and a sampler.
[0044] In addition, to achieve the above-mentioned purpose, the present application also proposes a three-dimensional object detection device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the three-dimensional object detection method as described above.
[0045] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the three-dimensional object detection method described above are implemented.
[0046] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of the three-dimensional object detection method described above are implemented.
[0047] One or more technical solutions proposed in this application have at least the following technical effects:
[0048] A left-eye image and a right-eye image of a three-dimensional object captured by a binocular camera are obtained, and depth estimation is performed on the left-eye image and the right-eye image through a binocular depth estimation network to obtain a depth map; two-dimensional pixel data of the depth map and parameter information of the binocular camera are obtained, and the two-dimensional pixel data is converted into three-dimensional point cloud data according to the parameter information; the three-dimensional point cloud data is projected onto a bird's-eye view, and the left-eye image, the right-eye image and the bird's-eye view are fused to obtain three-dimensional object fusion data; the three-dimensional object fusion data is detected through a target three-dimensional object detection network to obtain a three-dimensional object target detection result, the target three-dimensional object detection network includes a feature extraction layer, a spatial transformation layer, a reverse connection layer and a target detection layer, the spatial transformation layer includes a local network, a grid generator and a sampler, and depth estimation is performed using the left-eye and right-eye images captured by the binocular camera, and the two-dimensional pixel data is converted into three-dimensional point cloud data in combination with the parameter information of the camera, thereby achieving high-precision detection of three-dimensional objects. Specifically, the local network, mesh generator, and sampler in the spatial transformation layer help enhance the model's adaptability to objects from different perspectives, so that even without lidar data, accurate recognition and detection of three-dimensional objects can still be achieved by relying on image features collected by the binocular camera. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0050] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0051] Figure 1 A schematic diagram of a process flow provided for Embodiment 1 of the three-dimensional object detection method of the present application;
[0052] Figure 2 A schematic diagram of a network structure of a binocular depth estimation algorithm provided in Example 1 of the three-dimensional object detection method of the present application;
[0053] Figure 3 A schematic diagram of the structure of a pyramid pooling model provided in Example 1 of the three-dimensional object detection method of the present application;
[0054] Figure 4 A schematic diagram of the structure of a feature pyramid network model provided in Example 1 of the three-dimensional object detection method of the present application;
[0055] Figure 5A schematic diagram of the STN-RON model structure provided in Example 1 of the three-dimensional object detection method of the present application;
[0056] Figure 6 A schematic diagram of the DSAN module structure provided in Example 1 of the three-dimensional object detection method of the present application;
[0057] Figure 7 A schematic diagram of a flow chart provided for Embodiment 2 of the three-dimensional object detection method of the present application;
[0058] Figure 8 A schematic diagram of a simplified process of a three-dimensional object detection method provided in Embodiment 2 of the present application;
[0059] Fig. 9 This is a schematic diagram of the module structure of a three-dimensional object detection device according to an embodiment of the present application;
[0060] Fig.10 Schematic diagram of the device structure of the hardware operating environment involved in the three-dimensional object detection method in the embodiment of the present application.
[0061] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0062] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0063] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0064] The main solution of the embodiment of the present application is: obtaining the left-eye image and the right-eye image of the three-dimensional object collected by the binocular camera, and performing depth estimation on the left-eye image and the right-eye image through a binocular depth estimation network to obtain a depth map; obtaining two-dimensional pixel data of the depth map and parameter information of the binocular camera, and converting the two-dimensional pixel data into three-dimensional point cloud data according to the parameter information; projecting the three-dimensional point cloud data to a bird's-eye view, and fusing the left-eye image, the right-eye image and the bird's-eye view to obtain three-dimensional object fusion data; detecting the three-dimensional object fusion data through a target three-dimensional object detection network to obtain a three-dimensional object target detection result, wherein the target three-dimensional object detection network includes a feature extraction layer, a spatial transformation layer, a reverse connection layer and a target detection layer, and the spatial transformation layer includes a local network, a grid generator and a sampler.
[0065] In this embodiment, for ease of description, the following description is made by taking the three-dimensional object detection device as the execution subject.
[0066] The current mainstream vehicle 3D target detection methods are mainly based on two solutions: LiDAR and monocular vision sensors. The LiDAR solution is popular for its high-precision environmental perception information, but its high price increases the cost burden during mass production, making it impossible to be widely used in some low-cost models, thus limiting the popularity of intelligent driving functions. Although the monocular vision sensor solution has cost advantages, the perception information it provides is mostly one-dimensional and lacks depth information. It must rely on complex visual algorithms for data recovery, which not only increases R&D costs, but is also easily affected by the environment. It is difficult to accurately restore the information of certain special and complex scenes, resulting in relatively low perception accuracy.
[0067] This application provides a solution that uses the left and right images collected by the binocular camera to perform depth estimation, and combines the camera's parameter information to convert the two-dimensional pixel data into three-dimensional point cloud data, thereby achieving high-precision detection of three-dimensional objects. Specifically, the local network, grid generator, and sampler in the spatial transformation layer help enhance the model's adaptability to objects at different viewing angles, so that even without lidar data, accurate recognition and detection of three-dimensional objects can still be achieved by relying on the image features collected by the binocular camera.
[0068] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of realizing the above functions, a three-dimensional object detection device, etc. The following takes a three-dimensional object detection device as an example to illustrate this embodiment and the following embodiments.
[0069] Based on this, the present application embodiment provides a three-dimensional object detection method, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the three-dimensional object detection method of the present application.
[0070] In this embodiment, the three-dimensional object detection method includes steps S10 to S40:
[0071] Step S10, obtaining a left-eye image and a right-eye image of a three-dimensional object captured by a binocular camera, and performing depth estimation on the left-eye image and the right-eye image through a binocular depth estimation network to obtain a depth map;
[0072] It should be noted that the binocular camera consists of two parallel cameras with a certain distance between them, which are used to simulate the visual system of the human eye. Each camera can capture an image of a scene. The left and right images will have a certain parallax due to the different camera positions. The binocular depth estimation network is a preset deep learning network architecture. By inputting the left and right images, the parallax between the two can be learned, thereby generating the depth value of each pixel (i.e., the distance from the camera).
[0073] Specifically, refer to Figure 2 , Figure 2 A schematic diagram of the network structure of a binocular depth estimation algorithm provided in the first embodiment of the three-dimensional object detection method of this application. Figure 2 As shown in the figure, the feature extraction layer uses weight sharing to process the left image (i.e., left image) and the right image (i.e., left image) in turn through the convolution model, pyramid pooling model, and convolution layer to obtain the extracted left feature map and right feature map. After regularization of the extracted left feature map and right feature map, the regularized left feature map and right feature map are input into the feature matching layer. In the feature matching layer, the feature points in the left feature map and the right feature map are matched through the feature pyramid (Feature Pyramid Network, FPN) to obtain the feature point matching results, and the depth map is generated based on the feature point matching results. Compared with the traditional depth estimation algorithm, the binocular depth estimation network algorithm does not require manual design of feature extraction rules or adjustment of algorithm parameters. It directly uses the original image data as input and improves the accuracy and speed by establishing a feature learning network.
[0074] In a feasible implementation, the binocular depth estimation network includes a feature extraction layer and a feature matching layer; the step of performing depth estimation on the left eye image and the right eye image through the binocular depth estimation network in step S10 to obtain a depth map may include steps S11 to S13:
[0075] Step S11, performing feature extraction on the left image and the right image respectively through the feature extraction layer to obtain a first feature point descriptor of the left image and a second feature point descriptor of the right image;
[0076] It should be noted that in the feature extraction layer, the key local features (such as corners, edges, textures, etc.) in the left and right images are learned through the feature extraction network (including the convolution model, pyramid pooling model and convolution layer connected in sequence) to obtain the first feature point descriptor and the second feature point descriptor. The feature point descriptors of the left and right images specifically include the local feature information of the key points in the image. The descriptors provide a unique identifier for each key point for matching in the two images.
[0077] It should be noted that for situations where constraints with consistent application strength cannot be accurately estimated in special areas such as occluded areas, strongly reflective surfaces, and repeated areas, the pyramid pooling model in the feature extraction network of this application adopts a spatial pyramid pooling method to combine features of different scales to enhance the model's ability to detect and segment targets of different sizes in the image, and improves the performance and generalization ability of the model by retaining spatial information and providing multi-scale contextual features.
[0078] Specifically, refer to Figure 3 , Figure 3 Schematic diagram of the pyramid pooling model structure provided for the first embodiment of the three-dimensional object detection method of the present application. After the feature map obtained by the left or right image after preliminary feature extraction by the convolution model is input into the pyramid pooling model, the feature map will be downsampled to half of the original size (0.5x), and this process will be repeated at different levels to create feature representations of different scales. The downsampled feature map is subjected to multi-scale pooling operations through the spatial pyramid pooling layer (Special Pyramid Pooling, SPP) of the pyramid pooling model to obtain multiple feature maps of different resolutions, which are then fused and the fused features are predicted to obtain feature point descriptors of the feature map.
[0079] Step S12, performing feature point matching on the first feature point descriptor and the second feature point descriptor through the feature matching layer to obtain a disparity prediction result;
[0080] It should be understood that in the feature matching layer, the corresponding matching points in the two images will be found based on the feature point descriptors extracted from the left image and the right image (i.e., the first feature point descriptor and the second feature point descriptor). Specifically, matching can be performed by calculating the similarity between the descriptors (such as the Euclidean distance) to obtain the position difference of the corresponding feature points in the two images, i.e., the disparity, and the disparity prediction result including the disparity values of each pair of matching feature points for subsequent calculation of the depth map.
[0081] It should be noted that the present application introduces a feature pyramid network model in the feature matching layer, and adopts a top-down and lateral connection approach to fuse the low-layer high-resolution and low-semantic features with the high-layer low-resolution and high-semantic features, and at the same time fuses the original features so that the final feature maps of different scales have rich semantic information, that is, the top-level features are fused with the bottom-level features through upsampling, and each layer is independently predicted to obtain the feature point matching results and obtain the disparity prediction results of the disparity values of the matching feature points in the feature point matching results.
[0082] Specifically, refer to Figure 4 , Figure 4 A schematic diagram of the structure of a feature pyramid network model provided for the first embodiment of the three-dimensional object detection method of the present application. After the first feature point descriptor and the second feature point descriptor are input into the feature pyramid network model, the features of the first feature point descriptor and the second feature point descriptor are gradually extracted through continuous convolutional layers and pooling layers in a bottom-up path and the feature map size is reduced to obtain a high-level feature map. Starting from the high-level feature map in a top-down path, the feature map size is gradually restored and the low-level features are fused by upsampling (such as 2x up, i.e., doubling the feature map). The upper-level feature map is fused with the lower-level feature map in a lateral connection (laterial connection) usually through a 1x1 convolution (1x1conv) method to adjust the number of channels and fuse the features to obtain feature point matching results and disparity prediction results.
[0083] Step S13: generating a disparity map according to the disparity prediction result, and converting the disparity map into a depth map.
[0084] It should be understood that the disparity map contains the disparity value between each set of feature points extracted from the left and right images, that is, the horizontal offset of the same object in the left and right images. The disparity map can reflect the offset of each pixel under different viewing angles. By using the camera's intrinsic parameters (including focal length, principal point position, etc.) and the disparity map, the disparity value in the disparity map can be converted into depth information based on geometric relationships (such as triangulation) to obtain a depth map.
[0085] In a feasible implementation, before step S10, it also includes: training a depth estimation network based on preset three-dimensional image data to obtain a reference depth estimation network, and obtaining a disparity matching error and a true error of the preset three-dimensional image data; determining an estimated disparity of the preset three-dimensional image data according to the disparity matching error; calculating a disparity loss function of the estimated disparity and the true error; and optimizing the reference depth estimation network based on the disparity loss function to obtain a binocular depth estimation network.
[0086] It should be noted that the preset three-dimensional image data is the three-dimensional scene data used when training the depth estimation network, which is collected by a binocular camera. By dividing the preset three-dimensional image data, training data and test data can be obtained. The training data is used to train the depth estimation network to obtain a reference depth estimation network. The test data is used to verify the prediction results of the reference depth estimation network to obtain the disparity matching error and the true error. In the process of training the depth estimation network, after the feature matching is completed, a result with a certain disparity level can be obtained. Let its disparity level be E. The estimated disparity of the preset three-dimensional image data can be calculated through the soft argmax / argmin operation , the formula is as follows:
[0087]
[0088] Among them, E is the maximum disparity level, e is the current disparity level, represents the matching error at the disparity level of e. The matching error The probability obtained after the soft argmax / argmin function conversion. As can be seen from the formula, the matching error The larger it is, the smaller the probability that the current pixel disparity is the current disparity level d.
[0089] It should be understood that the disparity loss function for the estimated disparity and the true error is After the calculation, the reference depth estimation network can be optimized based on the disparity loss function to obtain a binocular depth estimation network.
[0090] For example, the SmoothL1 Loss function may be used as the loss function for calculation, and the formula is as follows:
[0091]
[0092] Where N is the total number of pixels, is the true disparity of the i-th pixel, is the estimated disparity of the i-th pixel, It is the SmoothL1 Loss function, which is used to calculate the difference between two values.
[0093] Step S20, acquiring two-dimensional pixel data of the depth map and parameter information of the binocular camera, and converting the two-dimensional pixel data into three-dimensional point cloud data according to the parameter information;
[0094] It should be noted that the two-dimensional pixel data of the depth map represents the plane information of the scene, including the two-dimensional coordinates of each pixel in the depth map in the camera coordinate system. and depth value The 3D point cloud data is pseudo-LiDAR point cloud data, which contains a set of 3D coordinates in the LiDAR coordinate system. The set of points is composed of a certain position in the scene, each of which represents the surface of an object. It can accurately represent the three-dimensional shape of the object and improve the accuracy of subsequent target detection.
[0095] In addition, it should be noted that the parameter information of the binocular camera can be obtained through camera calibration. Specifically, the intrinsic parameters of the camera can be estimated through images taken from multiple perspectives of a calibration object of known size (such as a chessboard) to obtain an intrinsic parameter matrix. The intrinsic parameter matrix contains the camera focal length information and the principal point coordinates (i.e., the pixel position coordinates of the camera center).
[0096] In a feasible implementation, the two-dimensional pixel data includes pixel coordinate information and pixel depth information, and the parameter information includes camera focal length information and camera principal point coordinate information; step S20 may include: obtaining the pixel coordinate information and the pixel depth information of the depth map; obtaining the camera focal length information and the camera principal point coordinate information of the binocular camera; converting the pixel coordinate information and the pixel depth information into three-dimensional point cloud data according to the camera focal length information and the camera principal point coordinate information.
[0097] It should be noted that the pixel coordinate information is the position of the pixel , the pixel depth information is the depth value of the pixel , the camera focal length information includes the horizontal focal length of the binocular camera and the vertical focal length of the camera After obtaining the above information, the coordinate relationship between the camera and the lidar sensor in space can be used to convert the pixel coordinate information and pixel depth information in the camera coordinate system into three-dimensional point cloud data in the lidar coordinate system according to the coordinate conversion formula.
[0098] Exemplarily, the coordinate transformation formula is as follows:
[0099]
[0100] Where r represents the reflectivity of the lidar.
[0101] Step S30, projecting the three-dimensional point cloud data onto a bird's-eye view, and fusing the left-eye image, the right-eye image, and the bird's-eye view to obtain three-dimensional object fusion data;
[0102] It should be noted that the bird's-eye view is a plane view formed by projecting the three-dimensional point cloud data from above, which is similar to looking down on the ground from a bird's eye view. By projecting the three-dimensional point cloud data onto the bird's-eye view, the three-dimensional data can be simplified, which is convenient for subsequent analysis and processing. By performing coordinate transformation and projection operations on the three-dimensional point cloud data and projecting it onto a two-dimensional plane, a bird's-eye view of the three-dimensional point cloud data can be obtained. The bird's-eye view provides a plan view of the three-dimensional scene, which is helpful for object recognition and detection. The left-eye image, the right-eye image and the bird's-eye view can be fused through the image fusion algorithm to obtain three-dimensional object fusion data. The three-dimensional object fusion data is image data that integrates the shape, depth and spatial position of the object, which is convenient for target detection.
[0103] Step S40, detecting the three-dimensional object fusion data through a target three-dimensional object detection network to obtain a three-dimensional object target detection result, wherein the target three-dimensional object detection network includes a feature extraction layer, a spatial transformation layer, a reverse connection layer and a target detection layer, and the spatial transformation layer includes a local network, a grid generator and a sampler.
[0104] It should be noted that the target three-dimensional object detection network includes a feature extraction layer, a spatial transformation layer, a reverse connection layer and a target detection layer, and the spatial transformation layer includes a local network, a grid generator and a sampler. Among them, the feature extraction layer extracts important features from the three-dimensional object fusion data, the spatial transformation layer performs spatial transformation and adjustment on the features through the local network, the grid generator and the sampler to make it more in line with the needs of object detection, the reverse connection layer is used to optimize the feature information in the network, adjust and correct the object detection results, and improve the detection accuracy. The target detection layer is used to perform the final object detection task on the optimized feature information, identify the three-dimensional object and output the detection results to obtain the three-dimensional object target detection results.
[0105] In addition, it should be noted that in order to effectively combine the visual information in the three-dimensional object fusion data with the pseudo-lidar point cloud data, the present application proposes an STN-RON model in the target three-dimensional object detection network, by embedding the spatial transformer network (STN) module into the backbone network of the RON algorithm, and constructing it after the feature extraction layer and before the reverse connection layer of the RON network. The spatial transformer network (STN) module is the spatial transformer layer, which can learn the spatial transformation parameters through a network including a local network, a grid generator and a sampler, and perform an affine transformation on the feature map based on the spatial transformation parameters, thereby optimizing the processing accuracy.
[0106] Specifically, refer to Figure 5 , Figure 5A schematic diagram of the STN-RON model structure provided for the first embodiment of the three-dimensional object detection method of the present application. Among them, the optical flow perception positioning network (DSAN) module is a DSAN model obtained by combining the spatial transformer STN and the reverse connection module of the RON algorithm, that is, an optical flow perception positioning network model. After receiving the input of the three-dimensional object fusion data, the convolution layer of the feature extraction layer (convolution layer 4Conv4, convolution layer 5Conv5, convolution layer 6Conv6, convolution layer 7Conv7) extracts the three-dimensional object features to obtain a feature map, and the local network, grid generator and sampler of the DSAN module can be used to learn the features of the local area information, and the learning results are input into the target detection layer (including the detection layer and the target layer) to obtain the three-dimensional object target detection result, wherein each detection layer (such as detection layer 4det4, detection layer 5det5, detection layer 6det6, detection layer 7det7) is responsible for target detection at different scales, and the target layer is used to obtain different targets output by the detection layer (such as target 1obj1, target 2obj12, target 3obj3).
[0107] Reference Figure 6 , Figure 6 Schematic diagram of the DSAN module structure provided for the first embodiment of the three-dimensional object detection method of the present application. The DSAN module includes a spatial transformer STN (i.e., a spatial transformer layer) and a reverse connection module (i.e., a reverse connection layer). The spatial transformer STN includes a local network, a grid generator, and a sampler, where θ is a spatial transformation parameter, T θ (G) represents the process of applying the spatial transformation parameter θ to the input feature map G for transformation. In the reverse connection module, the feature map input to the spatial transformer passes through a 3x3 convolutional layer with 512 output channels to extract deeper features. The output feature map of this convolutional layer is upsampled through a 2x2 deconvolutional layer with 512 output channels to match the spatial dimension of the input feature map. The upsampled feature map is added element by element to the original input feature map to form a residual connection, and the output of the residual block is passed to the next layer of the network. The embedding of DSAN enables more effective alignment and standardization of features when processing object deformation in images, thereby improving the accuracy and robustness of optical flow estimation, and significantly improving the detection ability of small-sized, low-contrast and partially occluded objects.
[0108] The present embodiment provides a three-dimensional object detection method, which includes obtaining a left-eye image and a right-eye image of a three-dimensional object collected by a binocular camera, and performing depth estimation on the left-eye image and the right-eye image through a binocular depth estimation network to obtain a depth map; obtaining two-dimensional pixel data of the depth map and parameter information of the binocular camera, and converting the two-dimensional pixel data into three-dimensional point cloud data according to the parameter information; projecting the three-dimensional point cloud data onto a bird's-eye view, and fusing the left-eye image, the right-eye image and the bird's-eye view to obtain three-dimensional object fusion data; detecting the three-dimensional object fusion data through a target three-dimensional object detection network to obtain a three-dimensional object target detection result, wherein the target three-dimensional object detection network includes a feature extraction layer, a spatial transformation layer, a reverse connection layer and a target detection layer, and the spatial transformation layer includes a local network, a grid generator and a sampler, and performing depth estimation on the left-eye and right-eye images collected by the binocular camera, and converting the two-dimensional pixel data into three-dimensional point cloud data in combination with the parameter information of the camera, thereby achieving high-precision detection of the three-dimensional object. Specifically, the local network, mesh generator, and sampler in the spatial transformation layer help enhance the model's adaptability to objects from different perspectives, so that even without lidar data, accurate recognition and detection of three-dimensional objects can still be achieved by relying on image features collected by the binocular camera.
[0109] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction, and will not be repeated in the following. Figure 7 , step S40 also includes steps S41 to S44:
[0110] Step S41, extracting features from the three-dimensional object fusion data through the feature extraction layer to obtain an initial feature map;
[0111] It should be noted that after receiving the input of the three-dimensional object fusion data, through a series of convolution and pooling operations in the feature extraction layer, the object surface, structure, edge and other features of the three-dimensional object fusion data can be extracted to form a preliminary feature representation and obtain the initial feature map as the input for subsequent processing.
[0112] Step S42, performing affine transformation on the initial feature map through the spatial transformation layer to obtain a transformed feature map;
[0113] It should be noted that the spatial transformation layer can perform affine transformations such as rotation, translation, and scaling on the initial feature map. During the affine transformation process, the angle and length of the initial feature map will be changed to obtain a transformed feature map.
[0114] In a feasible implementation, step S42 may include steps A11 to A13:
[0115] Step A11, learning the features of the initial feature map through the local network to predict spatial transformation parameters;
[0116] It should be noted that the local network in the spatial transformation layer is a network used to regress the spatial transformation parameter θ, which is used to perform affine transformation (such as translation, rotation, scaling, etc.) on the input feature image. After the initial feature map is input into the local network, a series of hidden network layers can learn the features of the initial feature map and predict the spatial transformation parameter θ, where the size of θ depends on the type of transformation. The generation formula of θ is as follows:
[0117]
[0118] Among them, U is the input feature image, that is, the initial feature map, Represents a local network function, which is used to regress the spatial transformation parameter θ. Specifically, θ may contain six parameters (such as two translation parameters, one rotation parameter, two scaling parameters, and one shear parameter).
[0119] Step A12, performing sampling transformation on the pixel coordinates of the initial feature map according to the spatial transformation parameters by the grid generator to obtain a sampling grid;
[0120] It should be noted that the grid generator in the spatial transformation layer can construct a sampling grid based on the predicted spatial transformation parameters θ. The sampling grid is an output image obtained after a set of points in the input image (i.e., the initial feature map) are sampled and transformed. The output image has undergone the required spatial transformation. Through the spatial transformation technology, it can more accurately locate and identify targets of different postures, improve the detection accuracy, and still maintain the operation speed of the original algorithm.
[0121] Exemplarily, the relationship mapping formula between the pixels in the input image and the output image is as follows:
[0122]
[0123] in, is the coordinate of each pixel of the input image, is the coordinate of each pixel of the corresponding output image, T θ is the spatial transformation function, G i is the i-th feature point in the input image, A θ is the transformation matrix of the transformation parameter θ, is the transformation parameter matrix.
[0124] Step A13: determining a transformed feature map according to the sampling grid and the initial feature map by the sampler.
[0125] It should be noted that the sampler in the spatial transformation layer can take the sampling grid and the initial feature map as input at the same time, and perform sampling according to the sampling grid and the input feature map (i.e., the initial feature map) to obtain the output feature map (i.e., the transformed feature map). Specifically, the formula for determining the transformed feature map is as follows:
[0126]
[0127] in, Represents the value of the cth channel at the ith position in the output feature map. m and n are the coordinates on the input feature map, which are used to determine the sampling position. m is the x-coordinate index of the input feature map, and n is the y-coordinate index of the input feature map.
[0128] In a feasible implementation, after step A13, it also includes: calculating the partial derivatives of the horizontal coordinate and the vertical coordinate of the feature point respectively according to the transformed feature map; deriving the sampler according to the partial derivative to obtain the back propagation gradient; and updating the network parameters of the spatial transformation layer according to the back propagation gradient.
[0129] It should be noted that in the process of training the target 3D object detection network, the back propagation of the loss needs to be considered, and the reverse derivative formula of the output feature map (i.e., the transformed feature map) to the sampler can be derived:
[0130]
[0131] in, Represents the partial derivative of the value of the cth channel at the i-th position in the output feature map V with respect to the value of the input feature map U at position (m, n) and channel c; Represents the partial derivative of the value of the cth channel at the i-th position in the output feature map V with respect to the output coordinate of the output feature map; Represents the partial derivative of the value of the cth channel at the i-th position in the output feature map V with respect to the output coordinates of the output feature map.
[0132] It should be understood that the sampler can be derived according to the partial derivatives of the feature point horizontal coordinate and the feature point vertical coordinate to obtain the back-propagation gradient, and the network parameters (such as weights and biases) of the spatial transformation layer can be updated according to the back-propagation gradient. The derivative formula of the sampler is as follows:
[0133]
[0134] in, Represents the partial derivative of the value of the cth channel at the ith position in the output feature map with respect to the spatial transformation parameter θ.
[0135] Step S43, combining the transformed feature map with the initial feature map through the reverse connection layer to obtain a feature enhancement map;
[0136] It should be noted that after obtaining the transformed feature map, the feature information can be transmitted in the neural network through the reverse connection layer, and the transformed feature map can be concatenated or weighted fused with the initial feature map to obtain an enhanced feature map containing more contextual information and spatial information, providing a more comprehensive input for subsequent target detection.
[0137] Step S44: Perform three-dimensional object detection on the feature enhancement image through the target detection layer to obtain a three-dimensional object target detection result.
[0138] It should be noted that the target detection layer includes the detection layer and the target layer, such as Figure 5 As shown, the detection layers (det 4, det5, det 6, det 7) are used to perform target detection on the feature enhancement map at different scales. Each detection layer outputs a set of feature map boxes that may contain three-dimensional object targets.
[0139] Specifically, the distribution of candidate object boxes can be designed so that specific feature map locations can be learned to respond to specific object scales, and each feature map box is recorded as ,but The expression formula is:
[0140]
[0141] in, Indicates the minimum scale.
[0142] It should be understood that after convolution processing of the candidate target box, the final detection result can be determined at the target layer (Obj 1, Obj 2, Obj 3), including the category and position of the three-dimensional object target, to obtain the three-dimensional object target detection result.
[0143] This embodiment provides a three-dimensional object detection method, which extracts features from the three-dimensional object fusion data through the feature extraction layer to obtain an initial feature map; performs affine transformation on the initial feature map through the spatial transformation layer to obtain a transformed feature map; combines the transformed feature map with the initial feature map through the reverse connection layer to obtain a feature enhancement map; performs three-dimensional object detection on the feature enhancement map through the target detection layer to obtain a three-dimensional object target detection result. The design of the spatial transformation layer in this solution can achieve higher-precision object detection and scene understanding, thereby narrowing the gap between the lidar solution and the binocular vision solution in environmental perception.
[0144] For example, to help understand the implementation process of the three-dimensional object detection method obtained by combining this embodiment with the above-mentioned embodiment 1, please refer to Figure 8 , Figure 8 A brief flowchart of a three-dimensional object detection method is provided, specifically:
[0145] The left and right images containing three-dimensional objects are collected by binocular cameras, and features are extracted from the left and right images through convolutional neural networks and pyramid pooling models (SPP). After feature matching through feature pyramid networks (FPN), depth maps are obtained. The depth maps are converted into pseudo-lidar point cloud data through the coordinate conversion module, and these pseudo-point cloud data are projected into a bird's-eye view (BEV). The bird's-eye view and image data (i.e., left and right images) are successively feature extracted, cropped, scaled, and fused. The fused data is then subjected to target detection and recognition through the STN-RON model to obtain the final three-dimensional object detection results.
[0146] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the three-dimensional object detection method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.
[0147] This application also provides a three-dimensional object detection device, please refer to Fig. 9 , the three-dimensional object detection device comprises:
[0148] The depth estimation module 10 is used to obtain a left-eye image and a right-eye image of a three-dimensional object captured by a binocular camera, and perform depth estimation on the left-eye image and the right-eye image through a binocular depth estimation network to obtain a depth map;
[0149] A coordinate conversion module 20, used to obtain the two-dimensional pixel data of the depth map and the parameter information of the binocular camera, and convert the two-dimensional pixel data into three-dimensional point cloud data according to the parameter information;
[0150] A data fusion module 30 is used to project the three-dimensional point cloud data onto a bird's-eye view, and fuse the left-eye image, the right-eye image and the bird's-eye view to obtain three-dimensional object fusion data;
[0151] The target detection module 40 is used to detect the three-dimensional object fusion data through a target three-dimensional object detection network to obtain a three-dimensional object target detection result. The target three-dimensional object detection network includes a feature extraction layer, a spatial transformation layer, a reverse connection layer and a target detection layer. The spatial transformation layer includes a local network, a grid generator and a sampler.
[0152] In one embodiment, the target detection module 40 is further used to perform feature extraction on the three-dimensional object fusion data through the feature extraction layer to obtain an initial feature map; perform affine transformation on the initial feature map through the spatial transformation layer to obtain a transformed feature map; combine the transformed feature map with the initial feature map through the reverse connection layer to obtain a feature enhancement map; perform three-dimensional object detection on the feature enhancement map through the target detection layer to obtain a three-dimensional object target detection result.
[0153] In one embodiment, the target detection module 40 is further used to learn the features of the initial feature map through the local network and predict the spatial transformation parameters; to sample and transform the pixel coordinates of the initial feature map according to the spatial transformation parameters through the grid generator to obtain a sampling grid; and to determine the transformation feature map according to the sampling grid and the initial feature map through the sampler.
[0154] In one embodiment, the target detection module 40 is also used to calculate the partial derivatives of the horizontal coordinate and the vertical coordinate of the feature point according to the transformed feature map; derive the sampler according to the partial derivative to obtain the back propagation gradient; and update the network parameters of the spatial transformation layer according to the back propagation gradient.
[0155] In one embodiment, the binocular depth estimation network includes a feature extraction layer and a feature matching layer; the depth estimation module 10 is further used to perform feature extraction on the left eye image and the right eye image respectively through the feature extraction layer to obtain a first feature point descriptor of the left eye image and a second feature point descriptor of the right eye image; perform feature point matching on the first feature point descriptor and the second feature point descriptor through the feature matching layer to obtain a disparity prediction result; generate a disparity map according to the disparity prediction result, and convert the disparity map into a depth map.
[0156] In one embodiment, the two-dimensional pixel data includes pixel coordinate information and pixel depth information, and the parameter information includes camera focal length information and camera principal point coordinate information; the coordinate conversion module 20 is also used to obtain the pixel coordinate information and the pixel depth information of the depth map; obtain the camera focal length information and the camera principal point coordinate information of the binocular camera; and convert the pixel coordinate information and the pixel depth information into three-dimensional point cloud data according to the camera focal length information and the camera principal point coordinate information.
[0157] In one embodiment, the depth estimation module 10 is also used to train a depth estimation network based on preset three-dimensional image data to obtain a reference depth estimation network, and obtain a disparity matching error and a true error of the preset three-dimensional image data; determine an estimated disparity of the preset three-dimensional image data based on the disparity matching error; calculate a disparity loss function of the estimated disparity and the true error; and optimize the reference depth estimation network based on the disparity loss function to obtain a binocular depth estimation network.
[0158] The three-dimensional object detection device provided by the present application adopts the three-dimensional object detection method in the above embodiment, which can solve the technical problem of how to achieve high-precision detection of three-dimensional objects without relying on laser radar data. Compared with the prior art, the beneficial effects of the three-dimensional object detection device provided by the present application are the same as the beneficial effects of the three-dimensional object detection method provided by the above embodiment, and other technical features in the three-dimensional object detection device are the same as the features disclosed in the above embodiment method, which will not be repeated here.
[0159] The present application provides a three-dimensional object detection device, which includes: at least one processor; and a memory that is communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the three-dimensional object detection method in the above-mentioned embodiment one.
[0160] Reference below Fig.10 , which shows a schematic diagram of the structure of a three-dimensional object detection device suitable for implementing the embodiment of the present application. The three-dimensional object detection device in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Fig.10 The three-dimensional object detection device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0161] like Fig.10As shown, the three-dimensional object detection device may include a processing device 1001 (such as a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM: Random Access Memory) 1004. In RAM1004, various programs and data required for the operation of the three-dimensional object detection device are also stored. The processing device 1001, ROM1002 and RAM1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the three-dimensional object detection device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows a three-dimensional object detection device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have alternatively.
[0162] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0163] The three-dimensional object detection device provided by the present application adopts the three-dimensional object detection method in the above embodiment, which can solve the technical problem of how to achieve high-precision detection of three-dimensional objects without relying on laser radar data. Compared with the prior art, the beneficial effects of the three-dimensional object detection device provided by the present application are the same as the beneficial effects of the three-dimensional object detection method provided by the above embodiment, and the other technical features in the three-dimensional object detection device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.
[0164] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0165] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0166] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, the computer-readable program instructions being used to execute the three-dimensional object detection method in the above-mentioned embodiment.
[0167] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM: Random Access Memory), a read-only memory (ROM: Read Only Memory), an erasable programmable read-only memory (EPROM: Erasable Programmable Read Only Memory or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM: CD-Read Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency: Radio Frequency), etc., or any suitable combination of the above.
[0168] The computer-readable storage medium may be included in the three-dimensional object detection device; or may exist independently without being assembled into the three-dimensional object detection device.
[0169] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by a three-dimensional object detection device, the three-dimensional object detection device: obtains a left-eye image and a right-eye image of a three-dimensional object captured by a binocular camera, and performs depth estimation on the left-eye image and the right-eye image through a binocular depth estimation network to obtain a depth map; obtains two-dimensional pixel data of the depth map and parameter information of the binocular camera, and converts the two-dimensional pixel data into three-dimensional point cloud data according to the parameter information; projects the three-dimensional point cloud data onto a bird's-eye view, and fuses the left-eye image, the right-eye image and the bird's-eye view to obtain three-dimensional object fusion data; detects the three-dimensional object fusion data through a target three-dimensional object detection network to obtain a three-dimensional object target detection result, wherein the target three-dimensional object detection network includes a feature extraction layer, a spatial transformation layer, a reverse connection layer and a target detection layer, and the spatial transformation layer includes a local network, a grid generator and a sampler.
[0170] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0171] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0172] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.
[0173] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned three-dimensional object detection method, and can solve the technical problem of how to achieve high-precision detection of three-dimensional objects without relying on laser radar data. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the three-dimensional object detection method provided in the above-mentioned embodiment, and will not be repeated here.
[0174] The present application also provides a computer program product, including a computer program, which implements the steps of the above-mentioned three-dimensional object detection method when executed by a processor.
[0175] The computer program product provided by the present application can solve the technical problem of how to achieve high-precision detection of three-dimensional objects without relying on laser radar data. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as the beneficial effects of the three-dimensional object detection method provided by the above embodiment, which will not be repeated here.
[0176] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A three-dimensional object detection method, characterized in that: The three-dimensional object detection method comprises: Acquire a left-eye image and a right-eye image of a three-dimensional object captured by a binocular camera, and perform depth estimation on the left-eye image and the right-eye image through a binocular depth estimation network to obtain a depth map; Acquire two-dimensional pixel data of the depth map and parameter information of the binocular camera, and convert the two-dimensional pixel data into three-dimensional point cloud data according to the parameter information; Projecting the three-dimensional point cloud data onto a bird's-eye view, and fusing the left-eye image, the right-eye image, and the bird's-eye view to obtain three-dimensional object fusion data; Detecting the three-dimensional object fusion data through a target three-dimensional object detection network to obtain a three-dimensional object target detection result, wherein the target three-dimensional object detection network includes a feature extraction layer, a spatial transformation layer, a reverse connection layer and a target detection layer, and the spatial transformation layer includes a local network, a grid generator and a sampler; The step of detecting the three-dimensional object fusion data through the target three-dimensional object detection network to obtain the three-dimensional object target detection result includes: Performing feature extraction on the three-dimensional object fusion data through the feature extraction layer to obtain an initial feature map; Performing an affine transformation on the initial feature map through the spatial transformation layer to obtain a transformed feature map; Combining the transformed feature map with the initial feature map through the reverse connection layer to obtain a feature enhancement map; Performing three-dimensional object detection on the feature enhancement map through the target detection layer to obtain a three-dimensional object target detection result; In the reverse connection layer, a 3x3 convolution layer with 512 output channels is used to perform deep feature extraction on the transformed feature map, and a 2x2 deconvolution layer with 512 output channels is used to upsample the output feature map after the deep feature extraction. The upsampled feature map is added element by element to the initial feature map to form a residual connection, so as to obtain the output of the residual block.
2. The method according to claim 1, characterized in that The step of performing affine transformation on the initial feature map through the spatial transformation layer to obtain a transformed feature map comprises: Learning the features of the initial feature map through the local network to predict spatial transformation parameters; The grid generator performs sampling transformation on the pixel coordinates of the initial feature map according to the spatial transformation parameters to obtain a sampling grid; A transformed feature map is determined by the sampler according to the sampling grid and the initial feature map.
3. The method according to claim 2, characterized in that After the step of determining the transformed feature map according to the sampling grid and the initial feature map by the sampler, the method further comprises: Calculate the partial derivatives of the feature point abscissa and the feature point ordinate respectively according to the transformed feature map; Deriving the sampler according to the partial derivative to obtain a back-propagation gradient; The network parameters of the spatial transformation layer are updated according to the back-propagation gradient.
4. The method according to claim 1, characterized in that The binocular depth estimation network includes a feature extraction layer and a feature matching layer; The step of performing depth estimation on the left-eye image and the right-eye image through a binocular depth estimation network to obtain a depth map includes: Performing feature extraction on the left image and the right image respectively through the feature extraction layer to obtain a first feature point descriptor of the left image and a second feature point descriptor of the right image; Performing feature point matching on the first feature point descriptor and the second feature point descriptor through the feature matching layer to obtain a disparity prediction result; A disparity map is generated according to the disparity prediction result, and the disparity map is converted into a depth map.
5. The method according to claim 1, characterized in that The two-dimensional pixel data includes pixel point coordinate information and pixel point depth information, and the parameter information includes camera focal length information and camera principal point coordinate information; The step of acquiring the two-dimensional pixel data of the depth map and the parameter information of the binocular camera, and converting the two-dimensional pixel data into three-dimensional point cloud data according to the parameter information comprises: Obtaining the pixel coordinate information and the pixel depth information of the depth map; Obtain the camera focal length information and the camera principal point coordinate information of the binocular camera; The pixel point coordinate information and the pixel point depth information are converted into three-dimensional point cloud data according to the camera focal length information and the camera principal point coordinate information.
6. The method according to any one of claims 1 to 5, characterized in that Before the step of acquiring a left-eye image and a right-eye image of a three-dimensional object captured by a binocular camera, and performing depth estimation on the left-eye image and the right-eye image through a binocular depth estimation network to obtain a depth map, the step further includes: Training a depth estimation network based on preset three-dimensional image data to obtain a reference depth estimation network, and obtaining a disparity matching error and a true error of the preset three-dimensional image data; determining an estimated disparity of the preset three-dimensional image data according to the disparity matching error; Calculating a disparity loss function of the estimated disparity and the true error; The reference depth estimation network is optimized based on the disparity loss function to obtain a binocular depth estimation network.
7. A three-dimensional object detection device, characterized in that: The device comprises: A depth estimation module is used to obtain a left-eye image and a right-eye image of a three-dimensional object captured by a binocular camera, and perform depth estimation on the left-eye image and the right-eye image through a binocular depth estimation network to obtain a depth map; A coordinate conversion module, used to obtain the two-dimensional pixel data of the depth map and the parameter information of the binocular camera, and convert the two-dimensional pixel data into three-dimensional point cloud data according to the parameter information; A data fusion module, used for projecting the three-dimensional point cloud data onto a bird's-eye view, and fusing the left-eye image, the right-eye image and the bird's-eye view to obtain three-dimensional object fusion data; A target detection module, used to detect the three-dimensional object fusion data through a target three-dimensional object detection network to obtain a three-dimensional object target detection result, wherein the target three-dimensional object detection network includes a feature extraction layer, a spatial transformation layer, a reverse connection layer and a target detection layer, and the spatial transformation layer includes a local network, a grid generator and a sampler; The target detection module is further used to perform feature extraction on the three-dimensional object fusion data through the feature extraction layer to obtain an initial feature map; perform affine transformation on the initial feature map through the spatial transformation layer to obtain a transformed feature map; combine the transformed feature map with the initial feature map through the reverse connection layer to obtain a feature enhancement map; perform three-dimensional object detection on the feature enhancement map through the target detection layer to obtain a three-dimensional object target detection result; The target detection module is also used to perform deep feature extraction on the transformed feature map in the reverse connection layer through a 3x3 convolution layer with 512 output channels, upsample the output feature map after the deep feature extraction through a 2x2 deconvolution layer with 512 output channels, add the upsampled feature map to the initial feature map element by element to form a residual connection, and obtain the output of the residual block.
8. A three-dimensional object detection device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the three-dimensional object detection method according to any one of claims 1 to 6.
9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the three-dimensional object detection method according to any one of claims 1 to 6 are implemented.