Monocular three-dimensional target detection method, model training method and corresponding device

CN116721394BActive Publication Date: 2026-10-09ALIBABA (CHINA) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310623360.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-29
Publication Date
2026-10-09
Estimated Expiration
2043-05-29

AI Technical Summary

Technical Problem

因此现有的单目三维检测方法无法应用于路侧相机采集的图像

Benefits of technology

[0065] 1) This application adds the prediction of the vector representation of the road surface normal in the image to be detected, and uses the vector representation of the road surface normal to rotate the 3D target box to obtain the 3D target box in camera space. This technical solution enhances the perception capability of road surface information on the basis of 3D target detection, thereby reducing the impact of varying camera mounting angles on 3D target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116721394B_ABST
    Figure CN116721394B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a monocular three-dimensional target detection method, a model training method and corresponding devices. The main technical solution comprises: obtaining a to-be-detected image; performing feature extraction on the to-be-detected image to obtain a feature representation of the to-be-detected image; predicting a vector representation of a road surface normal in the to-be-detected image using the feature representation of the to-be-detected image; predicting a three-dimensional target frame in the to-be-detected image using the feature representation of the to-be-detected image; and performing rotation processing on the three-dimensional target frame using the vector representation of the road surface normal to obtain a three-dimensional target frame in a camera space. The technical solution enhances the perception ability of road surface information on the basis of three-dimensional target detection, thereby reducing the influence of the variable camera installation angle on three-dimensional target detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of artificial intelligence and autonomous driving technology, and in particular to a monocular 3D target detection method, model training method and corresponding device. Background Technology

[0002] Autonomous vehicles rely on the collaborative efforts of perception sensors, artificial intelligence, and global positioning systems to enable safe and autonomous driving. The system of an autonomous vehicle mainly comprises three modules: perception, decision-making, and execution. Among these, the perception module is a crucial component of autonomous driving technology. Currently, commonly used 3D perception devices include LiDAR and cameras. However, LiDAR has limited sensing range, is susceptible to interference from weather and environmental factors, and is expensive. Therefore, camera-based 3D perception has become an important research direction in autonomous driving technology, with monocular 3D target detection technology being a key area of ​​expertise.

[0003] Monocular 3D object detection technology is a technique that detects objects based on a single image captured by a camera, obtaining information such as the object's position, size, and pose in 3D space. Currently, monocular 3D object detection is mainly used for single-vehicle perception, that is, using images captured by the onboard cameras of autonomous vehicles for object detection. However, due to the limited perception range of single vehicles and severe occlusion problems, vehicle-to-infrastructure (V2I) cooperative perception has gradually developed. V2I cooperative perception refers to using images captured by the onboard cameras of autonomous vehicles and images captured by roadside cameras for object detection, thereby expanding the perception range.

[0004] However, the installation methods of onboard cameras and roadside cameras in autonomous vehicles differ significantly. Onboard cameras are mounted horizontally, so existing monocular 3D detection methods all assume that the road surface in the camera space is horizontal, and the target only has angular differences in its flight direction. Roadside cameras, on the other hand, are installed with pitch and roll angles to adjust to a suitable observation range. For example, roadside cameras mounted on poles are tilted downwards, resulting in a non-horizontal road surface in the camera space. Furthermore, different roadside cameras have varying installation angles. Therefore, existing monocular 3D detection methods cannot be applied to images acquired by roadside cameras. Summary of the Invention

[0005] In view of this, this application provides a monocular 3D target detection method, model training method and apparatus to reduce the impact of varying roadside camera installation angles on 3D target detection.

[0006] This application provides the following solution:

[0007] Firstly, a monocular three-dimensional target detection method is provided, the method comprising:

[0008] Acquire the image to be detected;

[0009] Feature extraction is performed on the image to be detected to obtain the feature representation of the image to be detected;

[0010] Using the feature representation of the image to be detected, predict the vector representation of the road surface normal in the image to be detected;

[0011] Using the feature representation of the image to be detected, predict the 3D bounding box in the image to be detected;

[0012] Using the vector representation of the road surface normal, the 3D target box is rotated to obtain the 3D target box in camera space.

[0013] According to one feasible method in an embodiment of this application, predicting a 3D target bounding box in the image to be detected using the feature representation of the image to be detected includes:

[0014] Using the feature representation of the image to be detected, predict the center point of the two-dimensional target bounding box, the center point offset information between the target attachment point and the two-dimensional target bounding box, the size information of the three-dimensional target bounding box and the heading angle information in the image to be detected;

[0015] Using the center point of the two-dimensional target bounding box and the center point offset information, the position information of the target attachment point in the image to be detected is determined;

[0016] The three-dimensional target bounding box in the image to be detected is determined by using the position information of the target point in the image to be detected, the size information of the three-dimensional target bounding box, and the heading angle information.

[0017] According to one achievable method in an embodiment of this application, the method further includes: predicting road surface depth information in the image to be detected using the feature representation of the image to be detected;

[0018] Using the position information of the target point in the image to be detected, the size information of the 3D target bounding box, and the heading angle information, the 3D target bounding box in the image to be detected is determined as follows:

[0019] The target depth information is determined using the position information of the target attachment point in the image to be detected and the road surface depth information;

[0020] Using the target depth information, the position information of the target attachment point in the image to be detected, and camera intrinsic parameters, the position of the target attachment point in camera space is determined;

[0021] Using the position of the target attachment point in camera space, the size information of the 3D target bounding box, and the heading angle information, the 3D target bounding box at the position of the target attachment point in camera space is determined.

[0022] According to one achievable method in an embodiment of this application, determining the 3D target bounding box at the location of the target attachment point in camera space using the position of the target attachment point in camera space, the size information of the 3D target bounding box, and the heading angle information includes:

[0023] Using the size information and heading angle information of the three-dimensional target bounding box, a three-dimensional target bounding box is established at the coordinate origin position in the camera space;

[0024] The established 3D target bounding box is translated to the position of the target attachment point in camera space.

[0025] According to one feasible method in an embodiment of this application, the three-dimensional target bounding box is rotated using the vector representation of the road surface normal to obtain the three-dimensional target bounding box in camera space, including:

[0026] The rotation matrix corresponding to the three-dimensional target box is determined using the vector representation of the three-dimensional target box and the road surface normal.

[0027] The three-dimensional target bounding box is rotated using the rotation matrix so that the bottom surface of the three-dimensional target bounding box is parallel to the road surface, thus obtaining the three-dimensional target bounding box in the camera space.

[0028] Secondly, a monocular 3D target detection method is provided, executed by a server, the method comprising:

[0029] Acquire the image to be detected captured by the roadside camera;

[0030] Feature extraction is performed on the image to be detected to obtain the feature representation of the image to be detected;

[0031] Using the feature representation of the image to be detected, predict the vector representation of the road surface normal in the image to be detected;

[0032] Using the feature representation of the image to be detected, predict the 3D bounding box in the image to be detected;

[0033] Using the vector representation of the road surface normal, the 3D target box is rotated to obtain the 3D target box in camera space;

[0034] Driving decision information is generated using the three-dimensional target bounding box in the camera space;

[0035] The driving decision information is sent to the autonomous vehicle.

[0036] Thirdly, a method for training a three-dimensional object detection model is provided, the method comprising:

[0037] Acquire training data including multiple training samples, wherein the training samples include image samples and labels annotating the image samples, and the labels include 3D target box labels and vector labels of road surface normals;

[0038] A 3D object detection model is trained using the training data; wherein the image sample is input into the 3D object detection model, and the 3D object detection model extracts features from the image sample to obtain the feature representation of the image sample; using the feature representation of the image sample, the vector representation of the road surface normal in the image sample is predicted; using the feature representation of the image sample, the 3D object bounding box in the image sample is predicted; using the vector representation of the road surface normal, the 3D object bounding box is rotated to obtain the 3D object bounding box in camera space;

[0039] The training objectives include: minimizing the difference between the 3D target bounding boxes in the camera space output by the 3D target detection model and the corresponding 3D target bounding box labels, and minimizing the difference between the vector representation of the road surface normal obtained by the 3D target detection model and the vector label of the corresponding road surface normal.

[0040] According to one feasible method in an embodiment of this application, predicting a 3D bounding box in the image sample using the feature representation of the image sample includes:

[0041] Using the feature representation of the image samples, predict the center point of the two-dimensional target bounding box, the center point offset information between the target attachment point and the two-dimensional target bounding box, the size information of the three-dimensional target bounding box, and the heading angle information in the image samples;

[0042] Using the center point of the two-dimensional target bounding box and the center point offset information, the position information of the target attachment point in the image sample is determined;

[0043] The three-dimensional target bounding box in the image sample is determined using the position information of the target point in the image sample, the size information of the three-dimensional target bounding box, and the heading angle information.

[0044] According to one achievable method in an embodiment of this application, the label further includes a road surface depth label;

[0045] Determining the 3D target bounding box in the image sample using the position information of the target attachment point in the image sample, the size information of the 3D target bounding box, and the heading angle information includes: predicting road surface depth information in the image sample using the feature representation of the image sample; determining target depth information using the position information of the target attachment point in the image sample and the road surface depth information; determining the position of the target attachment point in camera space using the target depth information, the position information of the target attachment point in the image sample, and camera intrinsic parameters; and determining the 3D target bounding box at the position of the target attachment point in camera space using the position of the target attachment point in camera space, the size information of the 3D target bounding box, and the heading angle information.

[0046] The training objective also includes minimizing the difference between the road surface depth information obtained by the three-dimensional object detection model and the corresponding road surface depth label.

[0047] According to one achievable method in an embodiment of this application, the label further includes a two-dimensional target box label;

[0048] The three-dimensional target detection model further utilizes the feature representation of the image sample to predict the two-dimensional target bounding box of the image sample;

[0049] The training objective also includes minimizing the difference between the two-dimensional bounding boxes obtained by the three-dimensional object detection model and the corresponding two-dimensional bounding box labels.

[0050] Fourthly, a monocular three-dimensional target detection device is provided, the device comprising:

[0051] The image acquisition module is configured to acquire the image to be detected;

[0052] The feature extraction module is configured to extract features from the image to be detected to obtain a feature representation of the image to be detected;

[0053] The normal prediction module is configured to predict the vector representation of the road surface normal in the image to be detected using the feature representation of the image to be detected;

[0054] The 3D bounding box prediction module is configured to predict 3D target bounding boxes in the image to be detected using feature representations of the image to be detected.

[0055] The rotation processing module is configured to rotate the 3D target box using the vector representation of the road surface normal to obtain the 3D target box in camera space.

[0056] Fifthly, an apparatus for training a three-dimensional object detection model is provided, the apparatus comprising:

[0057] The sample acquisition module is configured to acquire training data including multiple training samples, wherein the training samples include image samples and labels annotating the image samples, and the labels include 3D target box labels and vector labels of road surface normals;

[0058] The model training module is configured to train a 3D object detection model using the training data; wherein the image sample is input into the 3D object detection model, the 3D object detection model extracts features from the image sample to obtain a feature representation of the image sample; using the feature representation of the image sample, the vector representation of the road surface normal in the image sample is predicted; using the feature representation of the image sample, the 3D object bounding box in the image sample is predicted; using the vector representation of the road surface normal, the 3D object bounding box is rotated to obtain a 3D object bounding box in camera space;

[0059] The training objectives include: minimizing the difference between the 3D target bounding boxes in the camera space output by the 3D target detection model and the corresponding 3D target bounding box labels, and minimizing the difference between the vector representation of the road surface normal obtained by the 3D target detection model and the vector label of the corresponding road surface normal.

[0060] According to a sixth aspect, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps of the method described in any one of the first to third aspects.

[0061] According to the seventh aspect, an electronic device is provided, comprising:

[0062] One or more processors; and

[0063] A memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the method described in any one of the first to third aspects.

[0064] According to the specific embodiments provided in this application, the following technical effects are disclosed:

[0065] 1) This application adds the prediction of the vector representation of the road surface normal in the image to be detected, and uses the vector representation of the road surface normal to rotate the 3D target box to obtain the 3D target box in camera space. This technical solution enhances the perception capability of road surface information on the basis of 3D target detection, thereby reducing the impact of varying camera mounting angles on 3D target detection.

[0066] 2) This application predicts the three-dimensional target box in the image to be detected by combining the center point of the two-dimensional target box, the center point offset information, the size information of the three-dimensional target box, and the heading angle information. It also uses the vector representation of the road surface normal to rotate the three-dimensional target box in the image to be detected, thereby improving the fit between the three-dimensional target box and the target in the camera space.

[0067] 3) This application adds the prediction of road surface depth information in the image to be detected, and combines road surface depth information, size information of the three-dimensional target box, heading angle information, etc. to comprehensively determine the three-dimensional target box in the image to be detected, thereby enhancing the accuracy of target detection.

[0068] 4) In this application, the position of the target attachment point in the camera space, the size information of the three-dimensional target box, and the heading angle information are used to first establish a three-dimensional target box at the origin of the coordinate system in the camera space, and then translate the established three-dimensional target box to the position of the target attachment point in the camera space. This method can effectively reduce the computational load of three-dimensional target box prediction.

[0069] 5) This application can be executed by the server side, which enhances the perception of surface information by collecting images of the target from roadside cameras, reduces the impact of varying installation angles of roadside cameras on 3D target detection, and generates driving decision information using the detected 3D target boxes and provides it to autonomous vehicles, thereby realizing vehicle-road cooperative perception and enhancing the perception of the surrounding environment of autonomous vehicles by utilizing the wider field of view of roadside cameras.

[0070] 6) In the training process of the 3D object detection model, this application not only uses the vector labels of 3D object bounding boxes and road surface normals for supervised learning, but also combines road surface depth labels for supervised learning, and can also use 2D object bounding box labels for auxiliary supervised learning, thereby improving the detection effect of the 3D object detection model.

[0071] Of course, any product implementing this application does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description

[0072] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0073] Figure 1 This is a system architecture diagram applicable to the embodiments of this application;

[0074] Figure 2A flowchart of a monocular three-dimensional target detection method provided in the embodiments of this application;

[0075] Figure 3 A schematic diagram of the principle structure of the three-dimensional target detection model provided in the embodiments of this application;

[0076] Figure 4 A flowchart illustrating the method for training a 3D target detection model provided in this application embodiment;

[0077] Figure 5 A schematic diagram illustrating the principle of training a 3D target detection model provided in an embodiment of this application;

[0078] Figure 6 A schematic block diagram of a monocular three-dimensional target detection device provided in the embodiments of this application;

[0079] Figure 7 A schematic block diagram of an apparatus for training a 3D target detection model provided in an embodiment of this application;

[0080] Figure 8 A schematic block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0081] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0082] The terminology used in the embodiments of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The singular forms “a,” “the,” and “the” as used in the embodiments of this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0083] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0084] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."

[0085] To facilitate understanding of the embodiments of this application, the system architecture on which the embodiments of this application are based will be briefly described first. Figure 1 An exemplary system architecture that can be applied to embodiments of this application is shown, such as Figure 1 As shown, the system mainly includes a target detection device located on the server side, a roadside camera set on the roadside, and an autonomous vehicle.

[0086] Roadside cameras installed on the side of the road can capture images and transmit them to a target detection device on the server via a network. In this application embodiment, the camera refers to a vision sensor, a broad term encompassing instruments that use optical elements and imaging devices to acquire images of the external environment. This can include traditional cameras, digital cameras, webcams, camcorders, etc.

[0087] The term "autonomous vehicle" in this application is used in a broad sense and can refer to either driverless or driver-assisted vehicles. Autonomous vehicles are also equipped with cameras, referred to as onboard cameras.

[0088] As one application scenario, the target detection device can be like Figure 1 The settings shown can be configured on the server side, either on a single server, a server cluster consisting of multiple servers, or a cloud server. A cloud server, also known as a cloud computing server or cloud host, is a hosting product within the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Servers (VPS) services, such as high management difficulty and weak service scalability.

[0089] In this application scenario, as one possible approach, the target detection device receives an image captured by a roadside camera, uses this image as the image to be detected, and performs target detection on the image using the method provided in this embodiment (which uses a three-dimensional target detection model). The three-dimensional target bounding box is then determined as the target detection result. Driving decision information is generated using the target detection result and provided to the autonomous vehicle so that the autonomous vehicle can drive according to the driving decision information.

[0090] As another possible approach, the target detection device receives an image from a roadside camera, uses this image as the image to be detected, and performs target detection on the image using the method provided in this application embodiment, determining the three-dimensional target bounding box as the target detection result. The image and the target detection result are then sent to the autonomous vehicle, which generates driving decision information based on the image and the target detection result, and drives according to the driving decision information.

[0091] The two implementation methods described above can achieve vehicle-road cooperative perception, eliminate the impact of varying roadside camera installation angles on 3D target detection, and utilize the roadside camera's field of view for environmental perception of autonomous vehicles, effectively enhancing the driving safety of autonomous vehicles.

[0092] As another possible approach, the target detection device receives an image captured by an onboard camera, uses this image as the image to be detected, and performs target detection on the image using the method provided in this application embodiment, determining the three-dimensional target bounding box as the target detection result. Driving decision information is generated using the target detection result and provided to the autonomous vehicle so that the autonomous vehicle can drive according to the driving decision information.

[0093] As another possible approach, the target detection device receives an image captured by an onboard camera, uses this image as the image to be detected, and performs target detection on the image using the method provided in this application embodiment, determining the three-dimensional target bounding box as the target detection result. The image and the target detection result are then sent to the autonomous vehicle, which generates driving decision information based on the image and the target detection result, and drives according to the driving decision information.

[0094] The two implementation methods described above enable the perception of images captured by vehicle-mounted cameras on the server side, eliminating the impact of the vehicle-mounted camera's installation angle shifting due to factors such as bumps and collisions.

[0095] Apart from Figure 1 As shown, the target detection device is located outside the server, and the map generation device can also be located in the autonomous vehicle. The autonomous vehicle uses images captured by the onboard camera as the images to be detected, or images captured by the roadside camera as the images to be detected, and performs target detection on the images to be detected using the method provided in this application embodiment, determining the three-dimensional target bounding boxes as the target detection results. Driving decision information is generated using the target detection results, and driving is performed based on the driving decision information.

[0096] It should be understood that Figure 1The number of target detection devices, autonomous vehicles, vehicle-mounted cameras, and roadside cameras shown is merely illustrative. Any number of target detection devices, autonomous vehicles, vehicle-mounted cameras, and roadside cameras can be used depending on implementation needs.

[0097] Figure 2 This is a flowchart of a monocular 3D target detection method provided in an embodiment of this application. This process can be... Figure 1 The target detection device in the system shown is executing. For example... Figure 2 As shown, the method mainly includes the following steps:

[0098] Step 202: Obtain the image to be detected.

[0099] Step 204: Extract features from the image to be detected to obtain the feature representation of the image to be detected.

[0100] Step 206: Utilize the feature representation of the image to be detected to predict the vector representation of the road surface normal in the image to be detected.

[0101] Step 208: Predict the 3D target bounding box in the image to be detected using the feature representation of the image to be detected.

[0102] It should be noted that steps 206 and 208 can be executed in any order or in parallel.

[0103] Step 210: Using the vector representation of the road surface normal, rotate the 3D target box to obtain the 3D target box in camera space.

[0104] As can be seen from the above process, this application adds the prediction of the vector representation of the road surface normal in the image to be detected, and uses the vector representation of the road surface normal to perform rotation processing on the 3D target box to obtain the 3D target box in camera space. This technical solution enhances the perception capability of road surface information on the basis of 3D target detection, thereby reducing the impact of varying camera mounting angles on 3D target detection.

[0105] The steps in the above process are described in detail below. In this application embodiment, the image to be detected is an image captured by a camera. In a vehicle-road cooperative scenario, the image to be detected can be an image captured by a roadside camera. In a single-vehicle intelligent scenario, the image to be detected can be an image captured by an onboard camera. This application aims to detect 3D target boxes from a single image; therefore, the image to be detected obtained in step 202 is a single image, or although multiple images are obtained, each image is used as the image to be detected to perform the 3D target detection provided in this application embodiment. For example, each frame or keyframe image in a video captured by a roadside video is used as the image to be detected to perform the 3D target detection provided in this application embodiment.

[0106] Steps 204 to 210 in the above process can be implemented using a pre-established 3D target detection model, that is, inputting the image to be detected into the 3D target detection model and obtaining the 3D target bounding box in the camera space output by the 3D target detection model.

[0107] Step 204 above, namely "extracting features from the image to be detected to obtain the feature representation of the image to be detected," can be achieved by... Figure 3 The feature extraction module of the 3D object detection model shown is executed. This feature extraction module can employ convolutional neural networks such as ResNet (Residual Network) or DLA (Deep Layer Aggregation).

[0108] The feature extraction module can first perform embedding processing based on the tokens (elements) in the image to be detected to obtain the embedding representation of each token, and then encode the embedding representation of each token to obtain the feature representation of the image to be detected.

[0109] In an image, a token refers to an element that constitutes the image. For an image, dividing it into a sequence of non-overlapping blocks results in each block and the start symbol of the sequence being a token. A block can consist of one pixel or multiple pixels. The token-based embedding process described above includes at least two parts: image embedding and position embedding. Image embedding involves encoding each token into a graph vector representation. Position embedding involves encoding the position of each token within the image sequence to obtain a positional representation.

[0110] Step 206 above, namely "predicting the vector representation of the road surface normal in the image to be detected using the feature representation of the image to be detected", can be achieved by... Figure 3 The normal prediction module in the 3D target detection model shown is executed.

[0111] A road surface normal refers to a straight line perpendicular to the road surface. The vector representation of a road surface normal refers to using a vector to represent the road surface normal in the image to be detected. This can be represented using a unit vector (i.e., a magnitude of 1). For example, the vector representation n of the road surface normal can be [n...]. x ,n y ,n z Since the vector representation of the road surface normal is a unit vector, only n needs to be predicted. x n y and n z Two of them are sufficient.

[0112] Step 208 above, namely "predicting the 3D target bounding box in the image to be detected using the feature representation of the image to be detected", can be achieved by... Figure 3 The 3D bounding box prediction module in the 3D target detection model shown is executed. This 3D bounding box prediction module contains the following branches: a heatmap prediction submodule for predicting heatmaps, an offset prediction submodule for predicting offsets, a size prediction submodule for predicting dimensions, and a heading angle prediction submodule for predicting heading angles. Then, an integration submodule uses the prediction results of each branch (i.e., each submodule) to determine the 3D target bounding box.

[0113] In this context, heatmap prediction refers to predicting the probability that each token in the image to be detected belongs to a predefined target type. In autonomous driving scenarios, these predefined target types can include vehicles, pedestrians, cyclists, roadblocks, and so on. By predicting the heatmap, the target types included in the image to be detected, as well as the center point of the corresponding two-dimensional bounding box, can be determined.

[0114] The offset prediction involved in this application embodiment is the prediction of the center point offset information between the target contact point and the two-dimensional target bounding box. The target contact point refers to the intersection point between the target and the road surface.

[0115] Size prediction refers to predicting the size information of a 3D target bounding box, such as its length, width, and height.

[0116] Predicting the heading angle refers to predicting the angle by which the 3D target box rotates around the y-axis of the camera coordinate system.

[0117] One possible approach is to first use the predicted center point of the two-dimensional target bounding box and the aforementioned center point offset to determine the position information of the target attachment point in the image to be detected; then, using the position information of the target attachment point in the image to be detected, the size information of the three-dimensional target bounding box, and the heading angle information, the three-dimensional target bounding box in the image to be detected can be determined.

[0118] One possible approach to determining the 3D target bounding box in the image to be detected is to use the position information of the target point in the image to be detected, the size information of the 3D target bounding box, and the heading angle information. This approach involves establishing the 3D target bounding box at the position of the target point according to the size information and heading angle information of the 3D target bounding box.

[0119] To improve the accuracy of the established 3D target bounding box, this embodiment of the application can further enhance the prediction of road surface depth information, which can be achieved by... Figure 3The depth prediction module in the 3D object detection model shown is executed. The depth prediction module uses the feature representation of the image to predict the road surface depth information in the image. Road surface depth refers to the distance between the road surface and the camera; the prediction of road surface depth information is actually the depth information of each token on the road surface.

[0120] In one preferred embodiment, when determining the 3D target bounding box in the image to be detected using the position information of the target attachment point in the image to be detected, the size information of the 3D target bounding box, and the heading angle information, the target depth information can be determined first using the position information of the target attachment point in the image to be detected and the road surface depth information. For example, the road surface depth information corresponding to the token at the target attachment point can be used as the target depth information.

[0121] Then, using the target depth information, the target placement location in the image to be detected, and camera intrinsic parameters, the position of the target placement location in camera space is determined. Camera intrinsic parameters can include camera focal length, distortion parameters, etc. For example, the position of the target placement location in camera space can be determined using the following formula:

[0122]

[0123] Where x, y, z are the coordinates of the target placement point in camera space. K is the camera intrinsic parameter matrix. w =[x w ,y w [] represents the location information of the target patch in the image to be detected, which is a two-dimensional representation of the location. h represents the target depth information.

[0124] Finally, using the position of the target attachment point in camera space, the size information of the 3D target bounding box, and the heading angle information, the 3D target bounding box at the position of the target attachment point in camera space is determined.

[0125] By adding the prediction of road surface depth information in the image to be detected in the above process, and combining the road surface depth information, the size information of the three-dimensional target box, the heading angle information, etc., to comprehensively determine the three-dimensional target box in the image to be detected, the accuracy of target detection can be effectively enhanced.

[0126] When constructing a 3D target bounding box, it can be directly constructed at the target's attachment point in camera space, based on the box's size and heading angle information. However, since the target's attachment point is far from the origin of the camera space coordinate system, its coordinates may be large, resulting in high computational complexity. To effectively reduce computational complexity, this application provides a more preferred implementation method: first, a 3D target bounding box is constructed at the origin of the camera space coordinate system using the box's size and heading angle information; then, the constructed 3D target bounding box is translated to the target's attachment point's location in camera space.

[0127] Besides the methods provided in the embodiments of this application, other methods can also be used, such as predicting the center point offset between the two-dimensional target box and the three-dimensional target box during offset prediction. The center point of the three-dimensional target box is determined using the predicted center point of the two-dimensional target box and its offset; then, the three-dimensional target box in the image to be detected is determined using the center point of the three-dimensional target box, the size information of the three-dimensional target box, and the heading angle information.

[0128] Step 210 above, namely "using the vector representation of the road surface normal to rotate the 3D target box and obtain the 3D target box in camera space," can be achieved by... Figure 3 The rotation processing module of the 3D target detection model shown is executed.

[0129] The rotation processing module first uses the vector representations of the 3D target bounding box and the road surface normal to determine the rotation matrix corresponding to the 3D target bounding box. Then, it rotates the 3D target bounding box using this rotation matrix, making its bottom surface parallel to the road surface, thus obtaining the 3D target bounding box in camera space. While different cameras may have different mounting angles, this difference usually doesn't have a drastic impact on the 3D target bounding box; it generally only affects the angle between the 3D target bounding box and the ground. Therefore, the bottom surface of the 3D target bounding box can be considered the surface closest to the road surface, and this bottom surface has a certain angle with the ground surface. By rotating the 3D target bounding box, its bottom surface is made parallel to the road surface.

[0130] A rotation matrix is ​​a matrix that, when multiplied by a vector, changes the direction of the vector without changing its magnitude. In this embodiment, a rotation matrix needs to be determined such that, after multiplying the 3D target bounding box (whose pose can also be represented by a vector) by the rotation matrix, the bottom surface of the 3D target bounding box is parallel to the road surface, i.e., the bottom surface of the 3D target bounding box is perpendicular to the road surface normal. The derivation of the rotation matrix is ​​a well-known method and will not be described in detail here.

[0131] After obtaining the 3D target bounding box in the camera space, driving decisions for autonomous vehicles can be made based on this bounding box, allowing the autonomous vehicle to drive accordingly. Examples include obstacle avoidance and trajectory planning.

[0132] In some scenarios, it is necessary to use 3D bounding boxes in image space to generate driving decisions, such as lane positioning. In this case, camera intrinsics can be used to perform coordinate transformation on the 3D bounding boxes in camera space to obtain 3D bounding boxes in image space.

[0133] Figure 4 A flowchart of the method for training a 3D object detection model provided in the embodiments of this application is shown below. Figure 4 As shown, the method may include the following steps:

[0134] Step 402: Obtain training data including multiple training samples. The training samples include image samples and labels annotating the image samples. The labels include 3D target box labels and vector labels of road surface normals.

[0135] In this embodiment, image samples can be acquired to construct training samples, and these image samples can be labeled. This primarily includes labels for the vector representations of 3D target bounding boxes and road surface normals, i.e., 3D target bounding box labels and road surface normal vector labels. For example, image samples can be acquired using a camera on a data acquisition vehicle, while radar scans targets within the same area as the image samples to obtain target information. This target information can then be used to obtain the corresponding 3D target bounding box in the image, thereby labeling the 3D target bounding box.

[0136] Furthermore, the aforementioned labels can also include road surface depth labels, which can be used as one of the targets of supervised learning. For example, the radar of the data acquisition vehicle can scan the road surface to obtain the distance information between the road surface and the camera, and then label the depth information. Alternatively, the road surface may not be scanned. Since the target information has already been obtained through scanning, and targets such as vehicles, pedestrians, and cyclists will inevitably come into contact with the road surface, the distance between the target's contact point and the camera, or the distance between the target's bottom surface and the camera, can be used as ground depth information for labeling.

[0137] The aforementioned labels may further include two-dimensional target bounding box labels, which are used to assist learning and will be detailed in subsequent embodiments.

[0138] Step 404: Train a 3D object detection model using training data; wherein image samples are input into the 3D object detection model, and the 3D object detection model extracts features from the image samples to obtain feature representations of the image samples; using the feature representations of the image samples, the vector representations of the road surface normals in the image samples are predicted; using the feature representations of the image samples, the 3D object boxes in the image samples are predicted; using the vector representations of the road surface normals, the 3D object boxes are rotated to obtain the 3D object boxes in camera space; the training objectives include: minimizing the difference between the 3D object boxes in camera space output by the 3D object detection model and their corresponding 3D object box labels, and minimizing the difference between the vector representations of the road surface normals obtained by the 3D object detection model and their corresponding vector labels.

[0139] Specifically, a 3D target detection model may include a feature extraction module, a normal prediction module, and a 3D bounding box prediction module, and may also include a depth prediction module and a rotation processing module.

[0140] The feature extraction module is responsible for extracting features from image samples to obtain feature representations of the image samples.

[0141] The normal prediction module is responsible for using the feature representation of image samples to predict the vector representation of the road surface normal in the image samples.

[0142] The 3D bounding box prediction module is responsible for predicting 3D target boxes in image samples using the feature representations of the image samples.

[0143] As one feasible approach, the 3D bounding box prediction module in a 3D object detection model can predict the 3D bounding box in an image sample by first using the feature representation of the image sample to predict the center point of the 2D bounding box, the center point offset between the target attachment point and the 2D bounding box, the size information of the 3D bounding box, and the heading angle information. Then, using the center point and center point offset information of the 2D bounding box, the position information of the target attachment point in the image sample is determined. Finally, using the position information of the target attachment point in the image sample, the size information of the 3D bounding box, and the heading angle information, the 3D bounding box in the image sample is determined.

[0144] The depth prediction module is responsible for predicting road surface depth information in image samples using feature representations of the image samples. In one preferred implementation, when determining the 3D target bounding box in the image sample using the target's location information, the size of the 3D target bounding box, and the heading angle, the module can first predict the road surface depth information using the image sample's feature representation; then, using the target's location information and the road surface depth information, determine the target depth information; finally, using the target depth information, the target's location information, and camera intrinsic parameters, determine the target's location in camera space; and finally, using the target's location in camera space, the size of the 3D target bounding box, and the heading angle information, determine the 3D target bounding box at the target's location in camera space. In this implementation, the training objective can also include minimizing the difference between the road surface depth information obtained by the 3D target detection model and the corresponding road surface depth label. In other words, it uses the road surface depth label to enhance the supervised learning effect of the target detection model.

[0145] When constructing a 3D target bounding box, it can be directly constructed at the target's attachment point in camera space, based on the box's size and heading angle information. However, since the target's attachment point is far from the origin of the camera space coordinate system, its coordinates may be large, resulting in high computational complexity. To effectively reduce computational complexity, this application provides a more preferred implementation method: first, a 3D target bounding box is constructed at the origin of the camera space coordinate system using the box's size and heading angle information; then, the constructed 3D target bounding box is translated to the target's attachment point's location in camera space.

[0146] The rotation processing module is responsible for rotating the 3D target box using the vector representation of the road surface normal to obtain the 3D target box in camera space.

[0147] As one possible approach, the rotation processing module can first use the vector representation of the 3D target box and the road surface normal to determine the rotation matrix corresponding to the 3D target box; then use the rotation matrix to rotate the 3D target box so that the bottom surface of the 3D target box is parallel to the road surface, thus obtaining the 3D target box in camera space.

[0148] Other specific details regarding the model structure can be found in the embodiments of the monocular 3D target detection method. Figure 3 The relevant records will not be elaborated here.

[0149] In this embodiment, a two-dimensional bounding box prediction branch can be further added to the 3D object detection model. That is, the 2D bounding box prediction module uses the feature representation of the image sample to predict the 2D bounding box of the image sample. The training objective further includes minimizing the difference between the 2D bounding box obtained by the 3D object detection model and the corresponding 2D bounding box label. In other words, the prediction of the 2D bounding box is used as an auxiliary training task for the object detection model to enhance its learning effect. After training, the 2D bounding box prediction module is deleted; that is, there is no prediction branch for the 2D bounding box in the actual prediction process, and it is only used for auxiliary training during the model training phase.

[0150] In this embodiment, a loss function can be constructed based on the aforementioned training objective. In each iteration, the model parameters are updated using methods such as gradient descent, based on the value of the loss function, until a preset training termination condition is met. The training termination condition may include, for example, the value of the loss function being less than or equal to a preset loss function threshold, or the number of iterations reaching a preset threshold.

[0151] One possible approach is to construct a total loss function (L) composed of a first loss function (L1), a second loss function (L2), a third loss function (L3), and a fourth loss function (L4), for example, by weighted summing of the first, second, third, and fourth loss functions. Figure 5 As shown in the diagram. The first loss function (L1) reflects the difference between the 3D bounding box in camera space output by the 3D object detection model and its corresponding 3D bounding box label. The second loss function (L2) reflects the difference between the vector representation of the road surface normal obtained by the 3D object detection model and its corresponding vector label. The third loss function (L3) reflects the difference between the road surface depth information obtained by the 3D object detection model and its corresponding road surface depth label. The fourth loss function (L4) reflects the difference between the 2D bounding box obtained by the 3D object detection model and its corresponding 2D bounding box label.

[0152] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0153] According to another embodiment, a monocular three-dimensional target detection device is provided. Figure 6A schematic block diagram of a monocular three-dimensional target detection apparatus according to one embodiment is shown, the apparatus corresponding to Figure 1 The target detection device in the architecture shown. Figure 6 As shown, the device 600 includes: an image acquisition module 601, a feature extraction module 602, a normal prediction module 603, a 3D bounding box prediction module 604, and a rotation processing module 605, and may further include: a depth prediction module 606.

[0154] The image acquisition module 601 is configured to acquire the image to be detected.

[0155] The feature extraction module 602 is configured to extract features from the image to be detected, thereby obtaining a feature representation of the image to be detected.

[0156] The normal prediction module 603 is configured to predict the vector representation of the road surface normal in the image to be detected using the feature representation of the image to be detected.

[0157] The 3D bounding box prediction module 604 is configured to predict 3D bounding boxes in the image to be detected using feature representations of the image to be detected.

[0158] The rotation processing module 605 is configured to rotate the 3D target box using the vector representation of the road surface normal to obtain the 3D target box in camera space.

[0159] The above-mentioned device can be used Figure 3 The three-dimensional target detection model shown is implemented.

[0160] As one possible implementation, the 3D bounding box prediction module 604 can be specifically configured to: predict the center point of the 2D target bounding box, the center point offset information between the target attachment point and the 2D target bounding box, the size information and heading angle information of the 3D target bounding box in the image to be detected using the feature representation of the image to be detected; determine the position information of the target attachment point in the image to be detected using the center point of the 2D target bounding box and the center point offset information; and determine the 3D target bounding box in the image to be detected using the position information of the target attachment point in the image to be detected, the size information and heading angle information of the 3D target bounding box.

[0161] In one preferred implementation, the depth prediction module 606 is configured to predict road surface depth information in the image to be detected using the feature representation of the image to be detected. The 3D bounding box prediction module 604 can be specifically configured to: determine target depth information using the position information of the target attachment point in the image to be detected and the road surface depth information; determine the position of the target attachment point in camera space using the target depth information, the position information of the target attachment point in the image to be detected, and camera intrinsic parameters; and determine the 3D bounding box at the position of the target attachment point in camera space using the position of the target attachment point in camera space, the size information of the 3D bounding box, and the heading angle information.

[0162] As one possible implementation method, the 3D bounding box prediction module 604 can be specifically configured to: establish a 3D bounding box at the origin of the coordinate system in the camera space using the size information and heading angle information of the 3D bounding box; and translate the established 3D bounding box to the position of the target contact point in the camera space.

[0163] As one possible implementation method, the rotation processing module 605 can be specifically configured to: determine the rotation matrix corresponding to the 3D target box using the vector representation of the 3D target box and the road surface normal; rotate the 3D target box using the rotation matrix so that the bottom surface of the 3D target box is parallel to the road surface, thereby obtaining the 3D target box in camera space.

[0164] According to another embodiment, an apparatus for training a three-dimensional object detection model is provided. Figure 7 A schematic block diagram of an apparatus for training a 3D object detection model according to one embodiment is shown. Figure 7 As shown, the device 700 includes a sample acquisition module 701 and a model training module 702.

[0165] The sample acquisition module 701 is configured to acquire training data including multiple training samples, the training samples including image samples and labels annotating the image samples, the labels including 3D target box labels and vector labels of road surface normals;

[0166] The model training module 702 is configured to train a 3D object detection model using training data. Specifically, image samples are input into the 3D object detection model, which extracts features from the image samples to obtain feature representations. Using these feature representations, the model predicts the vector representations of road surface normals within the image samples. Finally, using the vector representations of road surface normals, the model predicts the 3D bounding boxes within the image samples. The model then rotates the 3D bounding boxes using the vector representations of the road surface normals to obtain the 3D bounding boxes in camera space.

[0167] The training objectives include minimizing the difference between the 3D bounding boxes in camera space output by the 3D object detection model and their corresponding 3D bounding box labels, and minimizing the difference between the vector representation of the road surface normal obtained by the 3D object detection model and its corresponding vector label.

[0168] As one feasible approach, when a 3D object detection model uses the feature representation of an image sample to predict the 3D bounding box in the image sample, it can use the feature representation of the image sample to predict the center point of the 2D bounding box, the center point offset information between the target attachment point and the 2D bounding box, the size information of the 3D bounding box, and the heading angle information; use the center point and center point offset information of the 2D bounding box to determine the position information of the target attachment point in the image sample; and use the position information of the target attachment point in the image sample, the size information of the 3D bounding box, and the heading angle information to determine the 3D bounding box in the image sample.

[0169] Furthermore, the aforementioned labels can also include road surface depth labels. When the 3D object detection model uses the position information of the target placement point in the image sample, the size information of the 3D target box, and the heading angle information to determine the 3D target box in the image sample, it can use the feature representation of the image sample to predict the road surface depth information in the image sample; use the position information of the target placement point in the image sample and the road surface depth information to determine the target depth information; use the target depth information, the position information of the target placement point in the image sample, and camera intrinsic parameters to determine the position of the target placement point in camera space; use the position of the target placement point in camera space, the size information of the 3D target box, and the heading angle information to determine the 3D target box at the position of the target placement point in camera space; the training objective at this time can also further include: minimizing the difference between the road surface depth information obtained by the 3D object detection model and the corresponding road surface depth label.

[0170] Furthermore, the aforementioned labels can also include two-dimensional bounding box labels. The three-dimensional object detection model further utilizes the feature representations of image samples to predict the two-dimensional bounding boxes of the image samples. The training objective at this point can further include minimizing the difference between the two-dimensional bounding boxes obtained by the three-dimensional object detection model and their corresponding two-dimensional bounding box labels.

[0171] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0172] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0173] In addition, embodiments of this application also provide a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method described in any of the foregoing method embodiments.

[0174] And an electronic device, comprising:

[0175] One or more processors; and

[0176] A memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the method described in any of the foregoing method embodiments.

[0177] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in any of the foregoing method embodiments.

[0178] in, Figure 8An exemplary architecture of an electronic device is shown, which may include a processor 810, a video display adapter 811, a disk drive 812, an input / output interface 813, a network interface 814, and a memory 820. The processor 810, video display adapter 811, disk drive 812, input / output interface 813, network interface 814, and memory 820 can communicate with each other via a communication bus 830.

[0179] The processor 810 can be implemented using a general-purpose CPU, microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits to execute relevant programs and implement the technical solution provided in this application.

[0180] The memory 820 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 820 can store the operating system 821 for controlling the operation of the electronic device 800, and the basic input / output system (BIOS) 822 for controlling the low-level operations of the electronic device 800. Additionally, it can store a web browser 823, a data storage management system 824, and a target detection device / model training device 825, etc. The aforementioned target detection device / model training device 825 can be the application program that specifically implements the aforementioned steps in this embodiment. In summary, when the technical solution provided in this application is implemented through software or firmware, the relevant program code is stored in the memory 820 and is called and executed by the processor 810.

[0181] The input / output interface 813 is used to connect input / output modules to enable information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.

[0182] Network interface 814 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0183] Bus 830 includes a pathway for transmitting information between various components of the device, such as processor 810, video display adapter 811, disk drive 812, input / output interface 813, network interface 814, and memory 820.

[0184] It should be noted that although the above-described device only shows the processor 810, video display adapter 811, disk drive 812, input / output interface 813, network interface 814, memory 820, bus 830, etc., in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the solution of this application, and does not necessarily include all the components shown in the figures.

[0185] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer program product. This computer program product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0186] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A monocular three-dimensional target detection method, characterized in that, The method includes: The image to be detected is acquired by a roadside camera, wherein the roadside camera is installed at a non-horizontal angle; Feature extraction is performed on the image to be detected to obtain the feature representation of the image to be detected; Using the feature representation of the image to be detected, predict the vector representation of the road surface normal in the image to be detected; Using the feature representation of the image to be detected, the position of the target sticker in the image to be detected is predicted, and based on the feature representation of the image to be detected and the position of the target sticker in the image to be detected, a three-dimensional target bounding box in the image to be detected is predicted; The rotation matrix corresponding to the three-dimensional target box is determined using the vector representation of the three-dimensional target box and the road surface normal. Using the rotation matrix, the three-dimensional target box is rotated so that the bottom surface of the three-dimensional target box is parallel to the road surface, thus obtaining the three-dimensional target box in camera space.

2. The method according to claim 1, characterized in that, The step of predicting the position of the target object in the image under test using the feature representation of the image under test, and predicting the 3D bounding box in the image under test based on the feature representation of the image under test and the position of the target object in the image under test, includes: Using the feature representation of the image to be detected, predict the center point of the two-dimensional target bounding box, the center point offset information between the target attachment point and the two-dimensional target bounding box, the size information of the three-dimensional target bounding box and the heading angle information in the image to be detected; Using the center point of the two-dimensional target bounding box and the center point offset information, the position information of the target attachment point in the image to be detected is determined; The three-dimensional target bounding box in the image to be detected is determined by using the position information of the target point in the image to be detected, the size information of the three-dimensional target bounding box, and the heading angle information.

3. The method according to claim 2, characterized in that, The method further includes: using the feature representation of the image to be detected to predict road surface depth information in the image to be detected; Using the position information of the target point in the image to be detected, the size information of the 3D target bounding box, and the heading angle information, the 3D target bounding box in the image to be detected is determined as follows: The target depth information is determined using the position information of the target attachment point in the image to be detected and the road surface depth information; Using the target depth information, the position information of the target attachment point in the image to be detected, and camera intrinsic parameters, the position of the target attachment point in camera space is determined; Using the position of the target attachment point in camera space, the size information of the 3D target bounding box, and the heading angle information, the 3D target bounding box at the position of the target attachment point in camera space is determined.

4. The method according to claim 3, characterized in that, Using the position of the target attachment point in camera space, the size information of the 3D target bounding box, and the heading angle information, determining the 3D target bounding box at the position of the target attachment point in camera space includes: Using the size information and heading angle information of the three-dimensional target bounding box, a three-dimensional target bounding box is established at the coordinate origin position in the camera space; The established 3D target bounding box is translated to the position of the target attachment point in camera space.

5. A monocular three-dimensional target detection method, executed by a server, characterized in that, The method includes: The image to be detected is acquired by a roadside camera, wherein the roadside camera is installed at a non-horizontal angle; Feature extraction is performed on the image to be detected to obtain the feature representation of the image to be detected; Using the feature representation of the image to be detected, predict the vector representation of the road surface normal in the image to be detected; Using the feature representation of the image to be detected, the position of the target sticker in the image to be detected is predicted, and based on the feature representation of the image to be detected and the position of the target sticker in the image to be detected, a three-dimensional target bounding box in the image to be detected is predicted; The rotation matrix corresponding to the three-dimensional target box is determined using the vector representation of the three-dimensional target box and the road surface normal. Using the rotation matrix, the three-dimensional target box is rotated so that the bottom surface of the three-dimensional target box is parallel to the road surface, thus obtaining the three-dimensional target box in camera space; Driving decision information is generated using the three-dimensional target bounding box in the camera space; The driving decision information is sent to the autonomous vehicle.

6. A method for training a three-dimensional object detection model, characterized in that, The method includes: Acquire training data including multiple training samples, wherein the training samples include image samples acquired by roadside cameras and labels annotating the image samples, wherein the labels include 3D target box labels and vector labels of road surface normals, and the roadside cameras are installed at non-horizontal angles; A 3D object detection model is trained using the training data. The image samples are input into the 3D object detection model, which extracts features from the image samples to obtain feature representations. Using these feature representations, the vector representations of the road surface normals in the image samples are predicted. The feature representations of the image samples are used to predict the position of the target object in the image to be detected. Based on the feature representations of the image to be detected and the position of the target object in the image to be detected, a 3D bounding box in the image samples is predicted. Using the 3D bounding box and the vector representations of the road surface normals, a rotation matrix corresponding to the 3D bounding box is determined. Using the rotation matrix, the 3D bounding box is rotated so that its bottom surface is parallel to the road surface, thus obtaining the 3D bounding box in camera space. The training objectives include: minimizing the difference between the 3D target bounding boxes in the camera space output by the 3D target detection model and the corresponding 3D target bounding box labels, and minimizing the difference between the vector representation of the road surface normal obtained by the 3D target detection model and the vector label of the corresponding road surface normal.

7. The method according to claim 6, characterized in that, The location of the target object in the image to be detected is predicted using the feature representation of the image samples, and the 3D bounding box in the image samples is predicted based on the feature representation of the image to be detected and the location of the target object in the image to be detected. Using the feature representation of the image samples, predict the center point of the two-dimensional target bounding box, the center point offset information between the target attachment point and the two-dimensional target bounding box, the size information of the three-dimensional target bounding box, and the heading angle information; Using the center point of the two-dimensional target bounding box and the center point offset information, the position information of the target attachment point in the image sample is determined; The three-dimensional target bounding box in the image sample is determined using the position information of the target point in the image sample, the size information of the three-dimensional target bounding box, and the heading angle information.

8. The method according to claim 7, characterized in that, The label also includes a road surface depth label; Determining the 3D target bounding box in the image sample using the position information of the target attachment point in the image sample, the size information of the 3D target bounding box, and the heading angle information includes: predicting road surface depth information in the image sample using the feature representation of the image sample; determining target depth information using the position information of the target attachment point in the image sample and the road surface depth information; determining the position of the target attachment point in camera space using the target depth information, the position information of the target attachment point in the image sample, and camera intrinsic parameters; and determining the 3D target bounding box at the position of the target attachment point in camera space using the position of the target attachment point in camera space, the size information of the 3D target bounding box, and the heading angle information. The training objective also includes minimizing the difference between the road surface depth information obtained by the three-dimensional object detection model and the corresponding road surface depth label.

9. The method according to claim 6, characterized in that, The label also includes a two-dimensional target bounding box label; The three-dimensional target detection model further utilizes the feature representation of the image sample to predict the two-dimensional target bounding box of the image sample; The training objective also includes minimizing the difference between the two-dimensional bounding boxes obtained by the three-dimensional object detection model and the corresponding two-dimensional bounding box labels.

10. A monocular three-dimensional target detection device, characterized in that, The device includes: The image acquisition module is configured to acquire the image to be detected captured by a roadside camera, wherein the roadside camera is installed at a non-horizontal angle; The feature extraction module is configured to extract features from the image to be detected to obtain a feature representation of the image to be detected; The normal prediction module is configured to predict the vector representation of the road surface normal in the image to be detected using the feature representation of the image to be detected; The 3D bounding box prediction module is configured to predict the position of the target object in the image to be detected using the feature representation of the image to be detected, and to predict the 3D bounding box in the image to be detected based on the feature representation of the image to be detected and the position of the target object in the image to be detected. The rotation processing module is configured to determine the rotation matrix corresponding to the 3D target box using the vector representation of the 3D target box and the road surface normal; and to rotate the 3D target box using the rotation matrix so that the bottom surface of the 3D target box is parallel to the road surface, thereby obtaining the 3D target box in camera space.

11. An apparatus for training a three-dimensional target detection model, characterized in that, The device includes: The sample acquisition module is configured to acquire training data including multiple training samples, wherein the training samples include image samples acquired by a roadside camera and labels annotating the image samples, wherein the labels include 3D target box labels and vector labels of road surface normals, and the roadside camera is installed at a non-horizontal angle. The model training module is configured to train a 3D object detection model using the training data. Specifically, the image samples are input into the 3D object detection model, which extracts features from the image samples to obtain feature representations. Using these feature representations, the vector representations of the road surface normals in the image samples are predicted. The feature representations of the image samples are used to predict the position of the target object's contact point in the image to be detected. Based on the feature representations of the image to be detected and the position of the target object's contact point in the image to be detected, a 3D bounding box in the image samples is predicted. Using the 3D bounding box and the vector representations of the road surface normals, a rotation matrix corresponding to the 3D bounding box is determined. Using the rotation matrix, the 3D bounding box is rotated so that its bottom surface is parallel to the road surface, thus obtaining the 3D bounding box in camera space. The training objectives include: minimizing the difference between the 3D target bounding boxes in the camera space output by the 3D target detection model and the corresponding 3D target bounding box labels, and minimizing the difference between the vector representation of the road surface normal obtained by the 3D target detection model and the vector label of the corresponding road surface normal.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 9.

13. An electronic device, characterized in that, include: One or more processors; as well as A memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Three-dimensional object detection method and device, electronic equipment and readable storage medium

    CN111612753A