Three-dimensional (3D) object detection method and device, controller, vehicle and medium
By using a predetermined neural network model in three-dimensional (3D) object detection, combining 2D images and depth data, the problems of low detection accuracy and high computing resource consumption in the prior art are solved, and efficient and accurate 3D object detection is achieved, suitable for autonomous driving vehicles.
Patent Information
- Application Number
- CN202311848053.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-01
AI Technical Summary
The prior art has problems with low detection accuracy and high computing resource consumption in three-dimensional (3D) object detection, especially in autonomous vehicles, which are difficult to find a balance between efficiency and accuracy.
By obtaining two-dimensional (2D) images and depth data of the target 3D scene, 3D object detection is performed using a predetermined neural network model, and a single-level 3D object detection based on 2D images and depth data is realized.
It improves the efficiency and accuracy of 3D object detection, reduces the demand for computing resources, and enhances the application safety in autonomous driving vehicles.
Smart Images

Figure CN120236272A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of data processing, and more particularly, to methods, apparatuses, controllers, vehicles, and media for three-dimensional (3D) object detection. Background Art
[0002] With the development of intelligent devices, the demand for object detection in a scene is increasing, and in more and more applications, it is necessary to detect objects in a three-dimensional (3D) scene (e.g., roads, office areas, or playgrounds, etc.). For example, during the autonomous driving or assisted driving of a vehicle, it is necessary to detect objects (e.g., other vehicles or pedestrians, etc.) in a 3D driving scene for determining driving operations to be performed. For such applications, the efficiency and accuracy of object detection may directly affect the safety of users. Therefore, the efficiency and accuracy of object detection are very important. Summary of the Invention
[0003] Embodiments of the present disclosure provide a method, an apparatus, a controller, a vehicle, and a medium for three-dimensional (3D) object detection.
[0004] According to a first aspect of the present disclosure, there is provided a method for three-dimensional (3D) object detection. The method includes obtaining a two-dimensional (2D) image of a target 3D scene. The method further includes obtaining depth data corresponding to the 2D image. The method further includes using a predetermined neural network model to detect 3D objects in the target 3D scene based on the 2D image and the depth data.
[0005] According to a second aspect of the present disclosure, there is provided an apparatus for 3D object detection. The apparatus includes an obtaining unit configured to obtain a two-dimensional (2D) image of a target 3D scene. The apparatus further includes a generating unit configured to obtain depth data corresponding to the 2D image. The apparatus further includes a detecting unit configured to use a predetermined neural network model to detect 3D objects in the target 3D scene based on the 2D image and the depth data.
[0006] According to a third aspect of the present disclosure, there is provided a controller. The controller includes at least one processor; and a memory coupled to the at least one processor and having instructions stored thereon that, when executed by the at least one processor, cause the controller to implement the method according to the first aspect of the present disclosure.
[0007] According to a fourth aspect of the present disclosure, there is provided a vehicle. The vehicle includes the controller according to the third aspect of the present disclosure and a predetermined neural network model according to the first aspect of the present disclosure.
[0008] In a fifth aspect of the present disclosure, a computer-readable storage medium is provided. Computer-executable instructions are stored on the computer-readable storage medium, and the computer-executable instructions are executed by a processor to implement the method according to the first aspect of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] By describing the exemplary embodiments of the present disclosure in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become more apparent. In the exemplary embodiments of the present disclosure, the same reference numerals generally represent the same components.
[0010] Figure 1 A schematic diagram illustrating an example environment in which the device and / or method according to the embodiments of the present disclosure may be implemented.
[0011] Figure 2 A flowchart illustrating a method for three-dimensional (3D) object detection according to an embodiment of the present disclosure.
[0012] Figure 3 A schematic block diagram illustrating an example of a predetermined neural network model according to an embodiment of the present disclosure.
[0013] Figure 4A A schematic diagram illustrating an example of a process for obtaining a first neural network model according to an embodiment of the present disclosure.
[0014] Figure 4B A schematic diagram illustrating an example of a process for obtaining model anchors of a first neural network model according to an embodiment of the present disclosure.
[0015] Figure 5 A schematic diagram illustrating an example of a process for detecting 3D objects in a target 3D scene according to an embodiment of the present disclosure.
[0016] Figure 6 A schematic diagram illustrating an example of a process for obtaining a first model detection result by a first neural network model according to an embodiment of the present disclosure.
[0017] Figure 7 A schematic diagram illustrating an example of a process for determining 3D objects by a second neural network model according to an embodiment of the present disclosure.
[0018] Figure 8 A schematic block diagram illustrating a device for 3D object detection according to an embodiment of the present disclosure.
[0019] Figure 9 A schematic block diagram illustrating an example of an example device suitable for implementing the embodiments of the present disclosure.
[0020] In the respective drawings, the same or corresponding reference numerals denote the same or corresponding parts. Detailed implementation manners
[0021] Embodiments of the present disclosure will be described in more detail below with reference to the drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0022] In the description of the embodiments of the present disclosure, the term "including" and its like should be understood as an open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.
[0023] With the development of intelligent devices, the demand for object detection of objects in a scene is increasing. For example, for autonomous driving or assisted driving of a vehicle, the perception module of the vehicle is equivalent to the "eyes" of the vehicle to perceive the environment around the vehicle (for example, perform 3D object detection on objects in a three-dimensional (3D) road scene) to determine the driving operation to be performed, which is very important for the safety of the vehicle and the user. For example, the perception module can detect objects on the road (such as other vehicles, traffic lights, pedestrians, sidewalks, lane lines, and buildings or obstacles, etc.) based on the radar point cloud detected by the lidar or the image captured by the camera.
[0024] The lidar helps to perform accurate object detection due to its physical characteristics. However, the lidar is susceptible to weather conditions such as rain and snow, resulting in its inability to be well applicable to various driving environments. In addition, the cost of the lidar is relatively high, resulting in its inability to be installed in some vehicles that need to save costs. Therefore, how to perform 3D object detection based on the two-dimensional (2D) images captured by a monocular camera for a target 3D scene (such as a road scene) has attracted increasing attention.
[0025] Generally, a neural network model can be used for 3D object detection based on 2D images. Neural network models for such 3D object detection can be classified into an anchor-free type and an anchor-based type. For example, SMOKE (Single-Stage Monocular 3D Object Detection via Keypoint Estimation) is such an anchor-free type of neural network model. For example, M3D-RPN (Monocular 3D Region Proposal Network) is such an anchor-based type of neural network model. The anchors of this anchor-based type of neural network model can generally be classified into 2D anchors and 3D anchors. For this neural network model with 2D anchors, its input only includes 2D information, such as a 2D image; while for this neural network model with 3D anchors, its input, in addition to including 2D information, also needs to include 3D information about the object, such as position, size, and orientation information in the camera coordinate system, etc.
[0026] However, although the above-mentioned anchor-free neural network model is easy to develop and train without too many preset parameters, due to the limitations of the neural network model itself, its detection accuracy is low, and it is difficult to comprehensively detect objects in a 3D scene (i.e., the recall rate is low). Therefore, for example, in the process of autonomous driving or assisted driving of a vehicle, there may be potential safety hazards. For example, it may cause a collision between the vehicle and an obstacle because the obstacle on the road is not detected. In addition, this 3D object detection method reconstructs the 3D information of the object from a 2D image, and this reconstruction method itself will cause a large error.
[0027] In addition, the above-mentioned neural network model with 2D anchors also reconstructs the 3D information of the object only based on a 2D image, and in particular, it cannot accurately predict the depth of adjacent objects. In addition, the above-mentioned neural network model with 3D anchors requires a large amount of 3D information (such as position, size, and orientation information in the camera coordinate system), resulting in a large amount of computation, so that this neural network model needs to occupy a large amount of computing resources and cannot be applied to vehicles with limited computing resources, and cannot efficiently perform 3D object detection.
[0028] To address at least the above and other potential issues, embodiments of the present disclosure provide a method for 3D object detection. The method includes obtaining a 2D image of a target 3D scene. The method also includes obtaining depth data corresponding to the 2D image. The method further includes using a predetermined neural network model to detect 3D objects in the target 3D scene based on the 2D image and the depth data. According to the method of embodiments of the present disclosure, 3D object detection can be performed based only on the 2D image and the depth data by using the predetermined neural network model according to embodiments of the present disclosure, improving the efficiency and accuracy of 3D object detection.
[0029] Embodiments of the present disclosure will be described in further detail below with reference to the accompanying drawings. Figure 1 FIG. is a schematic diagram of an example environment 100 in which an apparatus and / or method according to an embodiment of the present disclosure may be implemented. As Figure 1 shown, vehicle 120 and vehicle 130 traveling on road 110 are shown in environment 100. In some embodiments, vehicle 120 and vehicle 130 may be vehicles capable of autonomous driving or assisted driving. Hereinafter, vehicle 120 will be used as an example for illustration. In some embodiments, vehicle 120 may have a camera 121 (e.g., a monocular camera), and the camera 121 may, for example, have a relatively large shooting angle F1 (e.g., 120°) to capture a 2D image of the target 3D scene (i.e., the scene of road 110) more comprehensively.
[0030] In some embodiments, vehicle 120 may further have a controller 122, and the controller 122 may be used to communicate with the camera 121 to receive the 2D image captured by the camera 121. In addition, in some embodiments, vehicle 120 may further have a unit (e.g., a radar) for obtaining depth data corresponding to the 2D image captured by the camera 121. The controller 122 may also communicate with this unit to receive depth data from this unit. In addition, vehicle 120 may further have a predetermined neural network model included in or coupled to the controller 122 (which will be further described in the examples below) so that the controller 122 can use this predetermined neural network model to perform 3D object detection based on the 2D image and the depth data. For example, the controller 122 can detect Figure 1 other vehicles 130, traffic lights 140, pedestrians 150, sidewalks 160, lane lines 170, etc. in the scene of the shown road 110. It should be understood that these detection objects are only examples, and depending on the different target 3D scenes, any other objects may also be detected, such as other obstacles or buildings, etc. The following will be combined with Figures 2 to 7 the shown examples to describe examples regarding 3D object detection.
[0031] Figure 2The flowchart of method 200 for 3D object detection according to an embodiment of the present disclosure is illustrated. Method 200 may be executed by any controller (e.g., Figure 1 the controller 122 shown), or an electronic device or a server, in combination with a predetermined neural network model according to an embodiment of the present disclosure. As Figure 2 shown, at block 202, a 2D image of a target 3D scene is obtained. For example, the target 3D scene may be Figure 1 the scene of road 110 shown, and the obtained 2D image may be a 2D image of the scene of road 110 captured by camera 121 on vehicle 120 in Figure 1 . It should be understood that the target 3D scene may also be any other scene, such as, for example, an office area, or a playground, etc. At block 204, depth data corresponding to the 2D image is obtained. In some embodiments, the depth data corresponding to the 2D object may be obtained via a radar on the vehicle. For example, the radar may be any type of radar.
[0032] At block 206, a predetermined neural network model is used to detect 3D objects in the target 3D scene based on the 2D image and the depth data. For example, other vehicles 130, traffic lights 140, pedestrians 150, sidewalks 160, lane lines 170, etc. shown in Figure 1 may be detected, or other obstacles or buildings, etc. may also be detected. According to the method of the embodiment of the present disclosure, 3D object detection can be performed only based on the 2D image and the depth data by using a predetermined neural network model according to the embodiment of the present disclosure, improving the efficiency and accuracy of 3D object detection. An example of a predetermined neural network model according to an embodiment of the present disclosure will be described below in conjunction with Figure 3 and FIG. 4.
[0033] Figure 3 The schematic block diagram of an example of a predetermined neural network model 300 according to an embodiment of the present disclosure is illustrated. As Figure 3 shown, the predetermined neural network model 300 according to an embodiment of the present disclosure may include a first neural network model 302 and a second neural network model 304. In some embodiments, the first neural network model 302 may be obtained based on a 2D object detection neural network model. In some embodiments, the second neural network model 304 may be a 3D object detection task head neural network model.
[0034] In some embodiments, the 2D object detection neural network model for obtaining the first neural network model 302 may be a neural network model with 2D anchors for single-stage 2D object detection (e.g., YOLOv5 or any other neural network model). Figure 4AA schematic diagram illustrating an example of a process 400 for obtaining a first neural network model 302 according to an embodiment of the present disclosure is shown. As Figure 4A shown, at block 402, predetermined depth data can be obtained. For example, the predetermined depth data can be training depth data in the training data for training the first neural network model 302 to be obtained, or can be other depth data different from the training depth data. At block 404, the predetermined depth data can be clustered to cluster the predetermined depth data into a plurality of clusters. For example, the K-means algorithm can be used to cluster the predetermined depth data.
[0035] At block 406, the first neural network model 302 can be obtained by setting 2D anchors into each of the plurality of clusters to obtain model anchors for each cluster. In some embodiments, the above setting of the 2D anchors can be such that for each cluster, the 2D anchor is associated with the central value of the predetermined depth data in the cluster, thereby obtaining a model anchor for the cluster. In some embodiments, the number of 2D anchors of the 2D object detection neural network model for obtaining the first neural network model 302 can be a first number. In some embodiments, the number of clusters into which the predetermined depth data is clustered can be a second number. In some embodiments, the number of all model anchors in the obtained first neural network model 302 can be the product of the first number and the second number. This can be further understood with reference to the following Figure 4B example of the process for obtaining model anchors described below.
[0036] Figure 4B A schematic diagram illustrating an example of a process 406A for obtaining model anchors of the first neural network model 302 according to an embodiment of the present disclosure is shown. As Figure 4B shown, the process 406A corresponds to Figure 4A block 406. Figure 4B The 2D anchors 410 shown can be the 2D anchors of the 2D object detection neural network model for obtaining the first neural network model 302. The 2D anchors 410 can include M 2D anchors 410-1 to 410-M, where M is an integer greater than or equal to 1. At Figure 4A block 406, the number of clusters into which the predetermined depth data is clustered can be N, where N is an integer greater than or equal to 1. As Figure 4B shown, the central values of the predetermined depth data in the N clusters obtained by clustering can be D1, D2,..., DN respectively.
[0037] As Figure 4BAs shown, 2D anchors 410 are respectively set in each cluster and associated with the center values D1, D2, …, DN of each cluster, so as to obtain corresponding model anchors 420-1, 420-2, …, 420-N. By generating these model anchors 420-1, 420-2, …, 420-N as above, a first neural network model 302 according to an embodiment of the present disclosure can be obtained. Thus, the first neural network model 302 can detect 2D objects in a 2D image and can also predict the predicted depth of the detected 2D objects. In some embodiments, the first neural network model 302 can be trained to be capable of performing desired 2D object detection and depth prediction.
[0038] In some embodiments, training depth data (for example, the above-mentioned predetermined depth data), corresponding training 2D images, and ground truth can be used to train the first neural network model 302. In some embodiments, during the training process, the first neural network model 302 can receive a training 2D image, perform feature extraction on the training 2D image to obtain a feature map, and then divide the feature map into multiple subgraphs (for example, grids). Then, each model anchor 420-1, 420-2, …, 420-N of the first neural network model 302 can output a 2D detection object (for example, output a 2D bounding box of the 2D detection object) and a corresponding predicted depth for the subgraph based on each subgraph and the corresponding training depth data in the training depth data. Then, based on the intersection over union of the 2D bounding box output by each model anchor and the corresponding bounding box of the ground truth, and the difference between the predicted depth output by each model anchor and the corresponding depth of the ground truth, the output of the model anchor with the output closest to the ground truth is selected as the final output of the first neural network model for the subgraph. In some embodiments, particularly in terms of the training of depth prediction, since there is input depth data, a linear prediction method shown in the following equation (1) can be used for depth prediction.
[0039] Z = pred[0] * depth_center_x_std + depth_center_x_mean (1)
[0040] In Equation (1), Z represents the predicted value of depth, and pred[0] represents the normalized depth value output by the first neural network model. depth_center_x_std represents the standard deviation of the training depth data for the x-th cluster, where x is an integer greater than or equal to 1 and less than or equal to N. depth_center_x_mean represents the mean value of the training depth data for the x-th cluster. In some embodiments, for example, during the training of depth prediction, when the change in the depth predicted by the first neural network with respect to the ground truth depth is less than a predetermined change threshold, or when the difference between the predicted depth and the ground truth depth is less than a predetermined difference threshold, the training can be ended. It should be understood that, according to actual needs, the training can also be ended according to other conditions, such as when the number of training rounds reaches a round threshold, etc. It should be understood that the training end conditions for 2D object detection training can be similar, or the depth prediction training and 2D object detection training can have associated training end conditions.
[0041] In addition, in some embodiments, Figure 3 the second neural network model 304 shown can also be pre-trained. The training data of the second neural network model 304 can include, for example, the training data corresponding to the output of the first neural network model 302, the anchor point parameters (such as the center value) of the model anchors of the first neural network model 302, and the camera parameters (which serve as the input of the second neural network model 304), and the training data can also include the corresponding 3D object ground truth. The training process of the second neural network model 304 is similar to that of the first neural network model 302 and will not be elaborated here. Below, with reference to Figures 5 to 7 the example of Figure 2 the method 200 according to an embodiment of the present disclosure shown, the process of 3D object detection performed by the trained first neural network model 302 and the trained second neural network model 304 according to an embodiment of the present disclosure will be described.
[0042] Figure 5 A schematic diagram illustrating an example of a process 500 for detecting 3D objects in a 3D scene according to an embodiment of the present disclosure. Figure 5 The process 500 shown corresponds to Figure 2 the block 206 shown. As Figure 5 shown, at block 502, the first neural network model 302 can be used, based on the 2D image obtained at Figure 2 block 202 and at Figure 2Depth data obtained at the frame 204 is used to obtain the first model detection result. In some embodiments, the first model detection result may include a 2D detection object (e.g., a bounding box of the 2D detection object) for the 2D image, and a predicted depth for the 2D detection object. At frame 504, a second neural network model 304 may be used to determine a 3D object corresponding to the 2D detection object based on the first model detection result and corresponding parameters. In some embodiments, the corresponding parameters may include anchor parameters of model anchors in the first neural network model. In some embodiments, the corresponding parameters may include camera parameters of a camera (e.g., Figure 1 the camera 121 shown) that captures the 2D image. The following describes frames 502 and 504 in further detail with reference to the examples shown in Figure 6 and Figure 7 .
[0043] Figure 6 FIG. illustrates an example of a process 600 for obtaining a first model detection result by a first neural network model 302 according to an embodiment of the present disclosure. Figure 6 The process 600 shown corresponds to Figure 5 the frame 502 shown. As shown in Figure 6 , at frame 602, a first neural network model 302 may be used to perform feature extraction on the 2D image to obtain a feature map 610 of the 2D image. At frame 604, the feature map 610 may be divided into multiple submaps. For example, as shown in Figure 6 , the feature map 610 is divided into multiple submaps of k rows and h columns, where both k and h are integers greater than or equal to 1. At frame 606, for each of the multiple submaps (for ease of description, the submap 610-i shown in Figure 6 is used as an example hereinafter, where i is an integer greater than or equal to 1 and less than or equal to k*h), all the model anchors 420-1, 420-2,..., 420-N may be used respectively to obtain corresponding anchor detection results.
[0044] In some embodiments, the corresponding anchor detection result may include a corresponding 2D detection object for the sub-graph 610-i. In some embodiments, the corresponding anchor detection result may further include a corresponding predicted depth for the corresponding 2D detection object. In some embodiments, for each of the model anchors 420-1, 420-2, …, 420-N in the first neural network model 302, the corresponding predicted depth is obtained based on the following: the mean and standard deviation of the corresponding depth data in the depth data that corresponds to the sub-graph 610-i. For example, the predicted depth may be obtained in a manner similar to the above equation (1), except that depth_center_x_std in equation (1) may represent the standard deviation of the above depth data for the x-th cluster here, and depth_center_x_mean may represent the mean of the above depth data for the x-th cluster here.
[0045] At block 608, based on the non-maximum suppression (NMS) algorithm, a target anchor detection result may be selected from all the anchor detection results as the sub-graph detection result for the sub-graph 610-i. Thus, the output of the first neural network model 302, i.e., the above first model detection result, can be obtained. In some embodiments, the above first model detection result may include the sub-graph detection result for each sub-graph. The first model detection result may be used by the second neural network model 304 to determine a 3D object.
[0046] Figure 7 A schematic diagram illustrates an example of a process 700 for determining a 3D object by the second neural network model 304 according to an embodiment of the present disclosure. Figure 7 The illustrated process 700 corresponds to Figure 5 the illustrated block 504. In some embodiments, the second neural network model 304 may be a 3D object detection task head neural network model with eight channels, which performs Figure 7 the illustrated process 700. As Figure 7 illustrated, at block 702, for each sub-graph 610-i among multiple sub-graphs (k*h sub-graphs), the second neural network model 304 may use the corresponding predicted depth included in the sub-graph detection result of the sub-graph 610-i as the predicted depth of the corresponding 3D object corresponding to the corresponding 2D detection object included in the sub-graph detection result. In some embodiments, the first channel of the second neural network model 304 is used to obtain the predicted depth of the corresponding 3D object.
[0047] At block 704, based on the predicted depth and the center point of the model anchor used to obtain the sub - graph detection result (for ease of description, hereinafter, the model anchor 410 - 2 shown in FIG. 4 is taken as an example for illustration, and it should be understood that such a model anchor can be any other model anchor), the orientation of the corresponding 3D object in the top view can be obtained. In some embodiments, the second and third channels of the second neural network model 304 can be used to obtain the orientation of the corresponding 3D object in the top view. In some embodiments, the second neural network model 304 can obtain the orientation of the corresponding 3D object in the top view through the following equations (2), (3), and (4).
[0048] θ = α z + arctan(XZ′) (2)
[0049] sin(α z ) = pred[1] (3)
[0050] cos(α z ) = pred[2] (4)
[0051] In the above equations (2) to (4), X represents the abscissa of the center point of the model anchor 420 - 2 defined above in the top view in the camera coordinate system, Z′ represents the ordinate of the center point of the model anchor 420 - 2 in the top view in the camera coordinate system. pred[1] represents the output of the second channel of the second neural network model 304, pred[2] represents the output of the third channel of the second neural network model 304, and θ represents the supervision signal of pred[1] and pred[2]. α z represents the predicted orientation of the corresponding 3D object in the top view.
[0052] At block 706, based on the preset average length, preset average width, and preset average height corresponding to the type of the above - mentioned 2D detection object (e.g., vehicle type), the predicted length, predicted width, and predicted height of the corresponding 3D object can be obtained. In some embodiments, the fourth to sixth channels of the second neural network model 304 can be used to obtain the predicted length, predicted width, and predicted height of the corresponding 3D object. In some embodiments, the second neural network model 304 can obtain the predicted length, predicted width, and predicted height of the corresponding 3D object through the following equations (5), (6), and (7).
[0053]
[0054]
[0055]
[0056] In the above equations (5) to (7), and respectively represent a preset average length, a preset average width, and a preset average height corresponding to the type of the above 2D detection object. For example, and can be set according to actual needs. pred[3] represents the output of the fourth channel of the second neural network model 304, pred[4] represents the output of the fifth channel of the second neural network model 304, and pred[5] represents the output of the sixth channel of the second neural network model 304. dim_l, dim_w, and dim_h respectively represent the predicted length, predicted width, and predicted height of the corresponding 3D object.
[0057] At block 708, based on the above predicted depth, predicted length, predicted width, predicted height, and camera parameters (e.g., the camera parameters of camera 121), the position of the corresponding 3D object in 3D space can be obtained. In some embodiments, the seventh and eighth channels of the second neural network model 304 can be used to obtain the position of the corresponding 3D object in 3D space. In some embodiments, the second neural network model 304 can obtain the position of the corresponding 3D object in 3D space through the following equations (8), (9), and (10).
[0058] proj_center_u = stride * (grid_u + pred[6]) (8)
[0059] proj_center_v = stride * (grid_v + pred[7]) (9)
[0060]
[0061] In the above equations (8) to (10), grid_u and grid_v represent the center points of the sub-graph 610-i of the feature map 610, and proj_center_u and proj_center_v represent the center points of the predicted 2D detection object for the sub-graph 610-i. pred[6] represents the output of the seventh channel of the second neural network model 304, and pred[7] represents the output of the eighth channel of the second neural network model 304. K 3×3 represents the camera parameter matrix of the camera. (X, Y, Z′) represents the coordinates of the corresponding 3D object in the above camera coordinates, that is, represents the position of the corresponding 3D object in 3D space.
[0062] In the above manner, by using the above-mentioned predetermined neural network model 300 according to the present disclosure, single-stage 3D object detection can be performed based only on 2D images and corresponding depth data, without the need for a large amount of 3D data to perform a large number of calculations, thereby improving the efficiency of 3D object detection. In addition, the above-mentioned predetermined neural network model according to an embodiment of the present disclosure uses a linear manner according to the present disclosure to predict depth without complex calculations, thus further improving the efficiency of 3D object detection. In addition, model anchors suitable for the above-mentioned predetermined neural network model according to the present disclosure are created, so that the recall rate of the above-mentioned predetermined neural network model for 3D object detection in the target 3D space is relatively high, that is, the probability that an actually existing 3D object is not detected can be greatly reduced. In addition, since the above-mentioned predetermined neural network model uses real depth data instead of inferring 3D information from 2D information, the accuracy of 3D object detection can be improved.
[0063] Figure 8 FIG. illustrates a schematic block diagram of a device 800 for 3D object detection according to an embodiment of the present disclosure. As Figure 8 shown, the device 800 includes an acquisition unit 802 configured to acquire a two-dimensional 2D image of a target 3D scene. The device 800 further includes a generation unit 804 configured to acquire depth data corresponding to the 2D image. The device 800 further includes a detection unit 806 configured to use a predetermined neural network model to detect 3D objects in the target 3D scene based on the 2D image and the depth data.
[0064] In some embodiments, the predetermined neural network model may include a first neural network model and a second neural network model. In some embodiments, the first neural network model may be obtained based on a 2D object detection neural network model. In some embodiments, the 2D object detection neural network model may be a neural network model with 2D anchors for single-stage 2D object detection. In some embodiments, the second neural network model may be a 3D object detection task head neural network model.
[0065] In some embodiments, the apparatus 800 may further include a model obtaining unit, which may be configured to obtain a first neural network model based on a 2D object detection neural network model. In some embodiments, the model obtaining unit may be configured to obtain predetermined depth data. In some embodiments, the model obtaining unit may be configured to cluster the predetermined depth data to cluster the predetermined depth data into a plurality of clusters. In some embodiments, the model obtaining unit may be configured to obtain the first neural network model by setting 2D anchors into each of the plurality of clusters to obtain model anchors for each cluster. In some embodiments, this setting may be such that for each cluster, the 2D anchor is associated with the central value of the predetermined depth data in the cluster to obtain the model anchor for the cluster. In some embodiments, the number of 2D anchors may be a first number. In some embodiments, the number of clusters may be a second number. In some embodiments, the number of all model anchors in the first neural network model may be the product of the first number and the second number.
[0066] In some embodiments, the detection unit 806 may be configured to use the first neural network model to obtain a first model detection result based on the 2D image and the depth data. In some embodiments, the first model detection result may include 2D detection objects for the 2D image. In some embodiments, the first model detection result may include predicted depths for the 2D detection objects.
[0067] In some embodiments, the detection unit 806 may be configured to use the first neural network model to extract features from the 2D image to obtain a feature map of the 2D image. In some embodiments, the detection unit 806 may be configured to divide the feature map into a plurality of submaps. In some embodiments, the detection unit 806 may be configured to, for each of the plurality of submaps, respectively use all the model anchors to obtain corresponding anchor detection results. In some embodiments, the corresponding anchor detection results may include corresponding 2D detection objects for the submaps. In some embodiments, the corresponding anchor detection results may include corresponding predicted depths for the corresponding 2D detection objects. In some embodiments, for each model anchor among all the model anchors in the first neural network model, the corresponding predicted depth may be obtained based on the following items: the average value and standard deviation of the corresponding depth data in the depth data corresponding to the submap. In some embodiments, the detection unit 806 may be configured to select a target anchor detection result from all the anchor detection results based on the non-maximum suppression algorithm as the submap detection result for the submap. In some embodiments, the first model detection result may include the submap detection result for each submap.
[0068] In some embodiments, the detection unit 806 may be configured to use a second neural network model to determine a 3D object corresponding to a 2D detection object based on the first model detection result and corresponding parameters. In some embodiments, the corresponding parameters may include the anchor parameters of the model anchors in the first neural network model. In some embodiments, the corresponding parameters may include the camera parameters of the camera for capturing the 2D image.
[0069] In some embodiments, for each sub-graph among a plurality of sub-graphs, the detection unit 806 may be configured to use the corresponding predicted depth included in the sub-graph detection result as the predicted depth of the corresponding 3D object corresponding to the corresponding 2D detection object included in the sub-graph detection result. In some embodiments, the detection unit 806 may be configured to obtain the orientation of the corresponding 3D object in the top view based on the predicted depth and the center point of the model anchor used to obtain the sub-graph detection result. In some embodiments, the detection unit 806 may be configured to obtain the predicted length, predicted width, and predicted height of the corresponding 3D object based on a preset average length, preset average width, and preset average height corresponding to the type of the 2D detection object. In some embodiments, the detection unit 806 may be configured to obtain the position of the corresponding 3D object in the 3D space based on the predicted depth, predicted length, predicted width, predicted height, and camera parameters.
[0070] The apparatus 800 according to an embodiment of the present disclosure can perform 3D object detection based only on a 2D image and depth data by using a predetermined neural network model according to an embodiment of the present disclosure, improving the efficiency and accuracy of 3D object detection.
[0071] Figure 9 A schematic block diagram of an exemplary device 900 suitable for implementing the embodiments of the present disclosure is shown. The controller described above may be implemented using the device 900. As shown, the device 900 includes a processor 901, which may execute various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 902 and loaded into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 may also be stored. The processor 901, ROM 902, and RAM 903 are connected to each other through a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0072] The various processes and treatments described above, such as method 200 and processes 400, 406A, 500, 600, 700, may be executed by processor 901. For example, in some embodiments, 200 and processes 400, 406A, 500, 600, 700 may be implemented as computer software programs tangibly embodied in a machine-readable medium. In some embodiments, part or all of the computer program may be loaded and / or installed onto device 900 via ROM 902. When the computer program is loaded into RAM 903 and executed by processor 901, one or more actions of method 200 and processes 400, 406A, 500, 600, 700 described above may be performed.
[0073] The present disclosure may be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for performing various aspects of the present disclosure.
[0074] The computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium may be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not to be construed as a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0075] The computer-readable program instructions described herein may be downloaded to each computing / processing device from a computer-readable storage medium or may be downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0076] The computer program instructions for performing the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partly on the user's computer, as a stand - alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of this disclosure.
Claims
1. A method for three-dimensional (3D) object detection, comprising: Obtaining a two-dimensional (2D) image of a target 3D scene; Obtaining depth data corresponding to the 2D image; And Using a predetermined neural network model, based on the 2D image and the depth data, to detect 3D objects in the target 3D scene.
2. The method according to claim 1, wherein the predetermined neural network model comprises a first neural network model and a second neural network model, wherein the first neural network model is obtained based on a 2D object detection neural network model, the second neural network model is a 3D object detection task head neural network model, and wherein the 2D object detection neural network model is a neural network model with 2D anchors for single-stage 2D object detection.
3. The method according to claim 2, further comprising obtaining the first neural network model based on the 2D object detection neural network model as follows: Obtaining predetermined depth data; Clustering the predetermined depth data to cluster the predetermined depth data into multiple clusters; and Obtaining the first neural network model by setting the 2D anchors into each of the multiple clusters to obtain model anchors for each cluster, wherein the setting associates the 2D anchors with the central values of the predetermined depth data in the cluster for each cluster to obtain the model anchors for the cluster.
4. The method according to claim 3, wherein using a predetermined neural network model, based on the 2D image and the depth data, to detect 3D objects in the target 3D scene comprises: Using the first neural network model, based on the 2D image and the depth data, to obtain a first model detection result; And Using the second neural network model, based on the first model detection result and corresponding parameters, to determine 3D objects corresponding to the 2D detection objects, wherein the first model detection result comprises 2D detection objects for the 2D image and predicted depths for the 2D detection objects, and wherein the corresponding parameters comprise: anchor parameters of the model anchors in the first neural network model and camera parameters of the camera for capturing the 2D image.
5. The method according to claim 4, wherein using the first neural network model, based on the 2D image and the depth data, to obtain a first model detection result comprises: Using the first neural network model to perform feature extraction on the 2D image to obtain a feature map of the 2D image; Dividing the feature map into multiple sub-maps; For each of the multiple sub-maps, respectively using all the model anchors to obtain corresponding anchor detection results; And Based on the non-maximum suppression (NMS) algorithm, selecting one target anchor detection result from all the anchor detection results as the sub-map detection result for the sub-map, wherein the corresponding anchor detection result comprises: corresponding 2D detection objects for the sub-map and corresponding predicted depths for the corresponding 2D detection objects, and The first model detection result includes sub - graph detection results for each sub - graph.
6. The method according to claim 5, wherein using the second neural network model, based on the first model detection result and corresponding parameters, to determine the 3D object corresponding to the 2D detection object includes: For each sub - graph among the multiple sub - graphs, taking the corresponding predicted depth included in the sub - graph detection result as the predicted depth of the corresponding 3D object corresponding to the corresponding 2D detection object included in the sub - graph detection result; Based on the predicted depth and the center point of the model anchor used to obtain the sub - graph detection result, obtaining the orientation of the corresponding 3D object in the top - view; Based on a preset average length, a preset average width, and a preset average height corresponding to the type of the 2D detection object, obtaining the predicted length, predicted width, and predicted height of the corresponding 3D object; And Based on the predicted depth, the predicted length, the predicted width, the predicted height, and the camera parameters, obtaining the position of the corresponding 3D object in 3D space.
7. The method according to claim 5, wherein for each model anchor among all model anchors in the first neural network model, obtaining the corresponding predicted depth based on: the mean and standard deviation of the corresponding depth data in the depth data corresponding to the sub - graph.
8. The method according to claim 3, wherein the number of 2D anchors is a first number, the number of clusters is a second number, and the number of all model anchors in the first neural network model is the product of the first number and the second number.
9. An apparatus for three - dimensional (3D) object detection, comprising: An acquisition unit configured to acquire a two - dimensional (2D) image of a target 3D scene; A generation unit configured to acquire depth data corresponding to the 2D image; And A detection unit configured to use a predetermined neural network model to detect 3D objects in the target 3D scene based on the 2D image and the depth data.
10. A controller, comprising: At least one processor; And A memory coupled to the at least one processor and having instructions stored thereon, the instructions, when executed by the at least one processor, cause the controller to execute the method according to any one of claims 1 - 8.
11. A vehicle, comprising the controller according to claim 10 and the predetermined neural network model.
12. A computer - readable storage medium having computer - executable instructions stored thereon, wherein the computer - executable instructions are executed by a processor to implement the method according to any one of claims 1 to 8.