Model training method, target detection method, device, equipment and storage medium
By combining a depth estimation module and transformer position embedding in monocular 3D object detection, and training the object detection model using a multi-task learning approach, the problem of inaccurate depth estimation is solved, thereby improving the performance of object detection and the accuracy of depth prediction.
Patent Information
- Application Number
- CN202311108541.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-30
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2043-08-30
AI Technical Summary
The inaccuracy of depth estimation in vision-based monocular 3D object detection algorithms makes object detection difficult.
By inputting sample images into the first neural network of the target detection model, a first target feature vector with three-dimensional position information is obtained. Combined with randomly generated three-dimensional position query information and voxel information of the three-dimensional sample region, the target detection model is trained using a multi-task learning approach, including a depth estimation module and transformer position embedding, to improve the accuracy of depth estimation.
It effectively overcomes the problem of insufficient depth supervision and improves the performance of target detection, especially the prediction accuracy of depth and 3D position.
Smart Images

Figure CN117011670B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of artificial intelligence, and more particularly, to a model training method, a target detection method, an apparatus, a device, and a storage medium. BACKGROUND
[0002] The task of target detection is to find all the targets (objects) of interest in an image, determine their categories and positions, and is one of the core problems in the field of computer vision. Due to different appearances, shapes and postures of various objects, and the interference of factors such as light and occlusion during imaging, target detection has always been the most challenging problem in the field of computer vision.
[0003] In the process of implementing the present disclosure, the inventors have found that at least the following problem exists in the related art: the difficulty of monocular 3D target detection algorithm based on vision mainly comes from the inaccuracy of depth estimation. SUMMARY
[0004] Therefore, the present disclosure provides a model training method, a target detection method, an apparatus, a device, and a storage medium.
[0005] One aspect of the present disclosure provides a target detection model training method, comprising: inputting a sample image into a first neural network of a target detection model to obtain a first target feature vector, wherein the first target feature vector has three-dimensional position information, and the sample image has a three-dimensional target bounding box label and a three-dimensional target detection label; inputting randomly generated first three-dimensional position query information and the first target feature vector into a second neural network of the target detection model to obtain first three-dimensional target bounding box information; inputting voxel information of a three-dimensional sample region and the first target feature vector into a third neural network of the target detection model to obtain a three-dimensional target detection result, wherein the three-dimensional sample region is determined according to the sample image; and training the target detection model according to the three-dimensional target bounding box label, the first three-dimensional target bounding box information, the three-dimensional target detection label, and the three-dimensional target detection result to obtain a trained target detection model.
[0006] Another aspect of the present disclosure provides a target detection method, comprising: obtaining an image to be processed; and inputting the image to be processed into a target detection model to obtain a target detection result, wherein the target detection model is trained by using the target detection model training method of the present disclosure.
[0007] One aspect of the present disclosure provides a device for training a target detection model, comprising: a first obtaining module configured to input a sample image into a first neural network of the target detection model to obtain a first target feature vector, wherein the first target feature vector has three-dimensional position information, and the sample image has a three-dimensional target bounding box label and a three-dimensional target label; a second obtaining module configured to input randomly generated first three-dimensional position query information and the first target feature vector into a second neural network of the target detection model to obtain first three-dimensional target bounding box information; a third obtaining module configured to input voxel information of a three-dimensional sample region and the first target feature vector into a third neural network of the target detection model to obtain a three-dimensional target detection result, wherein the three-dimensional sample region is determined according to the sample image; and a training module configured to train the target detection model according to the three-dimensional target bounding box label, the first three-dimensional target bounding box information, the three-dimensional target label and the three-dimensional target detection result to obtain a trained target detection model.
[0008] Another aspect of the present disclosure provides a device for target detection, comprising: a second obtaining module configured to obtain a to-be-processed image; and a fourth obtaining module configured to input the to-be-processed image into a target detection model to obtain a target detection result, wherein the target detection model is trained by the device for training a target detection model according to the present disclosure.
[0009] Another aspect of the present disclosure provides an electronic device, comprising: one or more processors; and a memory storing one or more programs, wherein the one or more programs, when executed by the one or more processors, cause the one or more processors to implement at least one of the method for training a target detection model and the method for target detection according to the present disclosure.
[0010] Another aspect of the present disclosure provides a computer-readable storage medium storing computer-executable instructions that, when executed, perform at least one of the method for training a target detection model and the method for target detection according to the present disclosure.
[0011] Another aspect of the present disclosure provides a computer program product comprising computer-executable instructions that, when executed, perform at least one of the method for training a target detection model and the method for target detection according to the present disclosure.
[0012] According to the embodiment of the present disclosure, because the first neural network of the target detection model is adopted to input the sample image to obtain the first target feature vector, the sample image has three-dimensional target bounding box labels and three-dimensional target detection labels, the first target feature vector has three-dimensional position information, the second neural network of the target detection model is adopted to input the randomly generated first three-dimensional position query information and the first target feature vector to obtain the first three-dimensional target bounding box information, the third neural network of the target detection model is adopted to input the voxel information of the three-dimensional sample region and the first target feature vector to obtain the three-dimensional target detection result, and the target detection model is trained according to the three-dimensional target bounding box labels, the first three-dimensional target bounding box information, the three-dimensional target detection labels and the three-dimensional target detection result, the technical means can learn the deep features in combination with the voxel information, at least partially overcomes the defect of insufficient depth supervision in target detection, and can be beneficial to improve the performance of target detection. BRIEF DESCRIPTION OF DRAWINGS
[0013] The above and other objects, features and advantages of the present disclosure will become more apparent from the following description of embodiments of the present disclosure taken in conjunction with the accompanying drawings, in which:
[0014] Figure 1 An exemplary system architecture to which at least one of the training method of the target detection model and the target detection method according to the embodiments of the present disclosure can be applied is schematically shown;
[0015] Figure 2 A flowchart of the training method of the target detection model according to the embodiments of the present disclosure is schematically shown;
[0016] Figure 3 An overall architecture diagram of the target detection model according to the embodiments of the present disclosure is schematically shown;
[0017] Figure 4 A flowchart of the target detection method according to the embodiments of the present disclosure is schematically shown;
[0018] Figure 5 A block diagram of the training device of the target detection model according to the embodiments of the present disclosure is schematically shown;
[0019] Figure 6 A block diagram of the target detection device according to the embodiments of the present disclosure is schematically shown; and
[0020] Figure 7 A block diagram of an electronic device suitable for implementing at least one of the training method of the target detection model and the target detection method according to the embodiments of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0021] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.
[0022] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0023] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0024] When using expressions such as "at least one of A, B, and C", they should generally be interpreted in accordance with the meaning that is commonly understood by a person skilled in the art (e.g., "a system having at least one of A, B, and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B, and C, etc.).
[0025] It should be noted that the collection, updating, analysis, processing, use, transmission, provision, disclosure, and storage of data (such as, but not limited to, user personal information) involved in the technical solution disclosed herein comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures have been taken to prevent unauthorized access to user personal information data and to safeguard user personal information security, network security, and national security.
[0026] In the embodiments disclosed herein, user authorization or consent is obtained before acquiring or collecting user personal information.
[0027] To address the inaccuracy of depth estimation, one approach is to add a depth estimation module to monocular 3D object detection. This can be implemented by pre-training the depth estimation module or by using it as an auxiliary supervision signal. Another approach is to build upon the pre-trained depth estimation by employing transformer position embedding to implicitly query the 2D feature vector at each 3D position, further improving the accuracy of depth estimation.
[0028] In realizing the concept disclosed herein, the inventors discovered that the two methods described above utilize depth estimation pre-training and transformer methods respectively for monocular 3D object detection. However, they suffer from the following drawbacks: In the first method, depth estimation performs depth constraints in 2D space, a low-dimensional space, and the supervision signal for depth estimation contains errors, primarily arising from calibration and projection errors. The second method uses a pre-trained model based on the first method and then uses transformer position embedding to query the 2D features of each 3D location. This leverages the advantages of both depth estimation pre-training and transformer querying to obtain features for each 3D location, further predicting obstacle positions. However, its final supervision signal only comes from the object detection bounding box, allowing only supervision of a few locations.
[0029] Embodiments of this disclosure provide a model training method, an object detection method, an apparatus, a device, and a storage medium. The method includes inputting a sample image into a first neural network of an object detection model to obtain a first object feature vector, wherein the first object feature vector has three-dimensional position information, and the sample image has a three-dimensional object detection bounding box label and a three-dimensional object detection label; inputting randomly generated first three-dimensional position query information and the first object feature vector into a second neural network of the object detection model to obtain first three-dimensional object detection bounding box information; inputting voxel information of a three-dimensional sample region and the first object feature vector into a third neural network of the object detection model to obtain a three-dimensional object detection result, wherein the three-dimensional sample region is determined based on the sample image; and training the object detection model based on the three-dimensional object detection bounding box label, the first three-dimensional object detection bounding box information, the three-dimensional object detection label, and the three-dimensional object detection result to obtain a trained object detection model.
[0030] Figure 1 An exemplary system architecture 100, illustrating at least one of the training methods for object detection models and object detection methods according to embodiments of this disclosure, is shown schematically. It should be noted that... Figure 1The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.
[0031] like Figure 1 As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0032] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, and / or social media platform software, etc. (for example only).
[0033] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0034] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0035] It should be noted that at least one of the training methods and target detection methods of the target detection model provided in this disclosure embodiment can generally be executed by server 105. Correspondingly, at least one of the training devices and target detection devices of the target detection model provided in this disclosure embodiment can generally be located in server 105. At least one of the training methods and target detection methods of the target detection model provided in this disclosure embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, at least one of the training devices and target detection devices of the target detection model provided in this disclosure embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Alternatively, at least one of the training methods and target detection methods of the target detection model provided in this embodiment of the disclosure may be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103, or by other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103. Accordingly, at least one of the training devices and target detection devices of the target detection model provided in this embodiment of the disclosure may be disposed in the first terminal device 101, the second terminal device 102, or the third terminal device 103, or in other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103.
[0036] For example, the sample image may originally be stored in any one of the first terminal device 101, the second terminal device 102, or the third terminal device 103 (e.g., the first terminal device 101, but not limited thereto), or it may be stored on an external storage device and imported into the first terminal device 101. Then, the first terminal device 101 may locally execute the training method of the object detection model provided in the embodiments of this disclosure, or send the sample image to other terminal devices, servers, or server clusters, and have the other terminal devices, servers, or server clusters that receive the sample image execute the training method of the object detection model provided in the embodiments of this disclosure.
[0037] For example, the image to be processed may originally be stored in any one of the first terminal device 101, the second terminal device 102, or the third terminal device 103 (e.g., the first terminal device 101, but not limited thereto), or it may be stored on an external storage device and imported into the first terminal device 101. Then, the first terminal device 101 may execute the target detection method provided in the embodiments of this disclosure locally, or send the image to be processed to other terminal devices, servers, or server clusters, and have the other terminal devices, servers, or server clusters that receive the image to be processed execute the target detection method provided in the embodiments of this disclosure.
[0038] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0039] Figure 2 A flowchart illustrating a method for training an object detection model according to an embodiment of the present disclosure is shown.
[0040] like Figure 2 As shown, the method includes operations S201 to S204.
[0041] In operation S201, the sample image is input into the first neural network of the target detection model to obtain the first target feature vector, wherein the first target feature vector has three-dimensional position information, and the sample image has a three-dimensional target detection box label and a three-dimensional target detection label.
[0042] According to embodiments of this disclosure, the sample images can be one or more images acquired by the same or different image acquisition devices, without limitation. For example, the sample images may include images from multiple perspectives obtained from a monocular camera at a fixed location.
[0043] According to embodiments of this disclosure, the object detection model can be a model built upon any one of the following object detection algorithms: R-CNN (Region-CNN, a deep learning-based object detection algorithm), SSD (Single Shot MultiBox Detector, a multi-object detection algorithm), YOLO (You Only Look Once, a deep learning-based regression method), PETRv2 (a unified framework for 3D perception from multi-camera images), and by adding branches to the network. It is not limited to these.
[0044] According to embodiments of this disclosure, for example, the sample image can be processed based on the PETRv2 algorithm to obtain a first target feature vector with three-dimensional position information.
[0045] It should be noted that the method for obtaining a first target feature vector with three-dimensional position information from a sample image is not limited to the method described above, but may also include other methods in the field, as long as they can achieve the goal of obtaining a first target feature vector with three-dimensional position information from a sample image.
[0046] According to embodiments of this disclosure, the target acquisition area can first be determined based on the sample image. A three-dimensional sample area is then determined based on the acquisition area. Next, three-dimensional target detection box information is determined based on the three-dimensional position information of the target object within the three-dimensional sample area. The three-dimensional target detection box information may include, for example, the position information of the center point of the three-dimensional target detection box (or a predefined point, such as a corner point of the three-dimensional target detection box, but not limited to this), the length, width, and height information of the three-dimensional target detection box, and the rotation angle relative to a preset coordinate system. This preset coordinate system can be a coordinate system customized according to business needs, or it can be a world coordinate system, and is not limited to these.
[0047] According to embodiments of this disclosure, after determining the three-dimensional sample region, a three-dimensional target detection label can be determined based on the area occupied by the target object within the three-dimensional sample region. The three-dimensional target detection label may include, but is not limited to, the positional information of each point on the target object.
[0048] In operation S202, the randomly generated first three-dimensional position query information and the first target feature vector are input into the second neural network of the target detection model to obtain the first three-dimensional target detection box information.
[0049] According to embodiments of this disclosure, for example, during image processing based on the PETRv2 algorithm, first three-dimensional location query information can be randomly generated based on the three-dimensional location information encoded for the first target feature vector. Then, the first target feature vector can be processed using the PETRv2 algorithm based on the first three-dimensional location query information to obtain first three-dimensional target detection box information. The first three-dimensional target detection box information can characterize the information of the three-dimensional target detection box predicted based on the first three-dimensional location query information.
[0050] It should be noted that the method for obtaining the first three-dimensional target detection box information based on the first three-dimensional location query information and the first target feature vector is not limited to the above-mentioned method, but may also include other methods in the art, as long as they can achieve the goal of obtaining the first three-dimensional target detection box information based on the first three-dimensional location query information and the first target feature vector.
[0051] In operation S203, the voxel information of the three-dimensional sample region and the first target feature vector are input into the third neural network of the target detection model to obtain the three-dimensional target detection result, wherein the three-dimensional sample region is determined based on the sample image.
[0052] According to embodiments of this disclosure, voxel information of a three-dimensional sample region can be obtained, for example, based on branches of the occupancy network. Then, the first target feature vector can be processed in units of voxels determined by the voxel information to obtain the predicted three-dimensional target detection result.
[0053] It should be noted that the method for obtaining the 3D target detection result based on voxel information and the first target feature vector is not limited to the method described above, but may also include other methods in this field, as long as they can achieve the 3D target detection result based on voxel information and the first target feature vector. The branches of the occupancy network in the target detection model described in this embodiment can also be replaced with any other form of convolutional network or deep learning network, as long as they can achieve the 3D target detection result based on voxel information and the first target feature vector.
[0054] In operation S204, the target detection model is trained based on the 3D target detection box label, the first 3D target detection box information, the 3D target detection label, and the 3D target detection result, resulting in a trained target detection model.
[0055] According to embodiments of this disclosure, a loss function can be constructed based on the 3D object detection box labels, first 3D object detection box information, 3D object detection labels, and 3D object detection results. Then, the parameters of the object detection model can be adjusted based on the value of the loss function until the network converges, thus obtaining the trained object detection model.
[0056] Through the above embodiments of this disclosure, since the target detection model is trained by combining the three-dimensional target detection results obtained by combining voxel information and the first target feature vector prediction, the target detection model can learn deep features better, which can effectively improve the performance of the trained target detection model.
[0057] The following describes specific embodiments. Figure 2 The method shown will be further explained.
[0058] According to embodiments of this disclosure, the sample images described above include images obtained from multiple perspectives using a monocular camera at a fixed location.
[0059] For example, the above sample images can be obtained by using a monocular camera installed on a vehicle (including autonomous vehicles) to capture images in various directions such as front, back, left, and right of the vehicle's driving direction.
[0060] It should be noted that the sampling scenarios for the sample images are not limited to those mentioned above, and may include any other scenarios, which are not limited here.
[0061] Through the above embodiments of this disclosure, combined with the training method of the improved target detection model, the monocular 3D target detection algorithm can be improved, thereby enhancing various indicators of monocular 3D target detection.
[0062] According to embodiments of this disclosure, the above operation S201 may include: extracting features from the sample image to obtain a two-dimensional feature vector of the sample image; and performing three-dimensional position encoding on the two-dimensional feature vector to obtain a first target feature vector.
[0063] According to embodiments of this disclosure, the target detection model may include a feature extraction module and a three-dimensional position encoding module. A sample image can first be input into the feature extraction module for processing to obtain a two-dimensional feature vector of the sample image. Then, the two-dimensional feature vector can be input into the three-dimensional position encoding module for processing to obtain a first target feature vector.
[0064] According to embodiments of this disclosure, the above operation S202 may include: querying a first feature vector related to the first three-dimensional location query information from a first target feature vector based on the first three-dimensional location query information; determining a target first feature vector characterizing the features of the target to be detected based on the decoding result of the first feature vector; determining target three-dimensional location query information corresponding to the target first feature vector from the first three-dimensional location query information; and determining first three-dimensional target detection box information based on the target three-dimensional location query information.
[0065] According to embodiments of this disclosure, the target detection model may further include a decoding module. A first feature vector can be input into the decoding module to obtain a decoding result of the first feature vector. After obtaining the decoding result of the first feature vector, the decoding result can be input into a first binary classification network to determine whether the first feature vector is a feature vector of the target to be detected. For first feature vectors whose determination result belongs to the target to be detected, they can be identified as the target's first feature vector. Then, target 3D position query information corresponding to the target's first feature vector can be obtained. Based on the 3D positions represented by the target's 3D position query information, the smallest 3D detection box that can cover these 3D positions can be determined to obtain the first 3D target detection box information.
[0066] According to embodiments of this disclosure, the first three-dimensional target detection box information can also be determined by combining the aforementioned target three-dimensional position query information and the remaining three-dimensional position query information whose judgment result is not to belong to the first feature vector of the target to be detected. For example, the first three-dimensional target detection box information can be obtained by determining the smallest three-dimensional detection box that can cover the three-dimensional position represented by the target three-dimensional position query information and does not cover the three-dimensional position represented by the remaining three-dimensional position query information.
[0067] According to embodiments of this disclosure, before performing the above-described operation S203, three-dimensional target detection labels and voxel information can be obtained first.
[0068] According to embodiments of this disclosure, obtaining a 3D target detection label may include: acquiring a preset number of frames of images corresponding to a sample image; determining the point cloud information of a 3D sample region based on the point cloud information of the sample images and the point cloud information of the acquired images; and dividing the point cloud information of the 3D sample region according to a preset voxel partitioning rule to obtain a 3D target detection label.
[0069] According to embodiments of this disclosure, the sample image can be a video frame from a video captured by a moving monocular camera. A predetermined number of preceding frames corresponding to the sample image can have at least the same acquisition angle as the sample image.
[0070] For example, one can first obtain the first 10 frames of sample images from the same acquisition perspective as the sample images, based on the video. Then, the point clouds of the first 10 frames can be overlaid on the point clouds of the corresponding frames of the sample images to obtain relatively dense 3D sample region point cloud information.
[0071] According to embodiments of this disclosure, the preset voxel partitioning rule may include: pre-setting the size of the voxels, such as including the length, width, and height information of the voxels. Then, based on the size of the voxels, the obtained 3D sample region point cloud information can be divided into meshes in units of voxels to obtain a first voxelized representation of the 3D sample region point cloud information. Subsequently, this first voxelized representation can be determined as a 3D target detection label.
[0072] For example, X, Y, and Z can be set as the lengths of the three-dimensional sample region in the length, width, and height directions, respectively. voxel_x, voxel_y, and voxel_z are the sizes of each voxel during voxelization, and the preset voxel division rules can be determined from these.
[0073] According to embodiments of this disclosure, obtaining voxel information may include: dividing a three-dimensional sample region according to a preset voxel division rule to obtain voxel information.
[0074] For example, a target region that can be captured by both the sample image and the captured image can be determined as a three-dimensional sample region based on the acquisition area targeted by the sample image and the acquisition area targeted by the captured image. This target region can be, for example, a cubic spatial region or a spherical spatial region, and is not limited to these.
[0075] It should be noted that, in cases where the three-dimensional sample region is determined solely based on the sample image, the target region described above can also be located based on the determined acquisition area of the sample image, combined with business requirements or predefined requirements such as area size and shape, in order to determine the three-dimensional sample region. This is not limited here.
[0076] According to embodiments of this disclosure, after determining the three-dimensional sample region, the three-dimensional sample region can be divided using the aforementioned preset voxel division rules to obtain a second voxelized representation of the three-dimensional sample region. Then, the voxel information of the three-dimensional sample region can be determined based on this second voxelized representation.
[0077] For example, in conjunction with the aforementioned embodiments, after voxelizing the three-dimensional sample region, a Cell can be set to contain (X / voxel_x*Y / voxel_y*Z / voxel_z) voxels, and the information corresponding to this Cell can be determined as the voxel information of the corresponding three-dimensional sample region.
[0078] Through the above embodiments of this disclosure, a relatively dense three-dimensional sample region point cloud information can be obtained by stacking multiple frames of image information. Based on the relatively dense three-dimensional sample region point cloud information, voxel information is determined, and combined with voxel-level three-dimensional target detection labels, a target detection model is trained, which can enable the trained target detection model to have high accuracy and performance.
[0079] According to embodiments of this disclosure, the above operation S203 may include: querying a second feature vector related to the voxel information from the target feature vector based on the voxel information; determining a target second feature vector characterizing the features of the target to be detected based on the decoding result of the second feature vector; determining target voxel information corresponding to the target second feature vector from the voxel information; and determining a three-dimensional target detection result based on the target voxel information.
[0080] According to embodiments of this disclosure, voxel information may include the three-dimensional voxel position information of the corresponding voxel. First, for each voxel in the three-dimensional cut-down region, the three-dimensional voxel position information represented by each voxel can be determined. Then, the three-dimensional voxel position information can be used as query information to retrieve a second feature vector related to each three-dimensional voxel position information from the target feature vector. By inputting the first feature vector into the decoding module of the aforementioned target detection model, the decoding result of the second feature vector can be obtained.
[0081] According to embodiments of this disclosure, after obtaining the decoding result of the second feature vector, the decoding result can be input into a second binary classification network to determine whether the corresponding second feature vector is a feature vector of the target to be detected. For second feature vectors whose determination result is that they belong to the target to be detected, they can be identified as the target's second feature vector. Then, based on the target voxel information corresponding to the target's second feature vector, the predicted three-dimensional target detection result can be determined.
[0082] It should be noted that the first and second classification networks mentioned above can be fully connected layer networks, and are not limited to this.
[0083] Through the above embodiments of this disclosure, each position in the entire three-dimensional sample region can be supervised on a voxel-by-voxel basis, achieving finer-grained supervision and enabling the model to learn depth and even 3D positional information more effectively.
[0084] Figure 3 An overall architecture diagram of the target detection model according to an embodiment of the present disclosure is shown.
[0085] like Figure 3 As shown, the object detection model 300 may include a feature extraction module 310, a 3D Position Encoder 320, a Decoder 330, a detection branch 340, and an occupancy branch 350. Corresponding to the detection branch 340, the decoding module 330 may receive Det queries 331 (detection query information). Corresponding to the occupancy branch 350, the decoding module 330 may receive Occupancy queries 332 (occupancy query information).
[0086] According to embodiments of this disclosure, occupancy branch 350 can be derived, for example, from an improved Det Head in PETRv2. For instance, the box prediction head of the Det Head can be replaced with a fully connected layer, and the output channels of the fully connected layer can be made equal to the number of voxels in the 3D sample region. This yields the occupancy head, which serves as occupancy branch 350.
[0087] According to embodiments of this disclosure, the target detection model 300, after adding an occupancy branch 350, can be transformed from a single-task model into a multi-task model of "detection" + "occupancy". The input to the target detection model 300 can be a sample image, such as a surround-view camera image 301 of an autonomous vehicle. The supervision signal can consist of two parts, such as: a 3D target detection bounding box label and a 3D target detection label represented by an occupancy mesh in point cloud voxelization.
[0088] According to embodiments of this disclosure, random sampling is performed based on the three-dimensional position information of the first target feature vector, for example, to obtain Det queries 331. Det queries 331 can serve as a supervision signal for the detection branch 340, and may include, for example, the aforementioned first three-dimensional position query information. Det queries 331 can be processed by the Decoder 330 and the detection branch 340 to predict whether the first target feature vector corresponding to each Det query does not represent the feature vector of the object to be detected, and to obtain three-dimensional target detection box information 341. The specific implementation method has been described in the foregoing embodiments and will not be repeated here.
[0089] According to embodiments of this disclosure, occupancy queries 332 can be generated, for example, by uniformly sampling within a three-dimensional sample region. Occupancy queries 332 can serve as a supervision signal for the occupancy branch 350, and may include, for example, voxel information of the aforementioned three-dimensional sample region. Based on this, the size of the supervision signal can, for example, be equal to the resolution of the voxel grid. Based on the foregoing embodiments, the occupancy queries 332 can be processed by the Decoder 330 and the occupancy branch 350 to predict whether each voxel corresponding to an occupancy query is occupied, and to obtain a voxel-level three-dimensional object detection result 342. Specific implementation methods have been described in the foregoing embodiments and will not be repeated here.
[0090] Through the embodiments described above, an occupancy network prediction branch is added to PetrV2, providing more fine-grained depth information supervision compared to PetrV2 detection boxes. Since the occupancy branch is predicted in 3D space, it provides stronger supervision of the model and places higher demands on it. Furthermore, the supervision of the occupancy branch is based on actual point clouds, resulting in a low ground truth error rate. Moreover, the supervision of the occupancy network is denser than that of the detection boxes, allowing the model to better learn depth and even 3D positional information.
[0091] According to embodiments of this disclosure, the above operation S204 may include: determining a first loss based on the 3D object detection box labels and first 3D object detection box information; determining a second loss based on the 3D object detection labels and 3D object detection results; and training the object detection model based on the first loss and the second loss.
[0092] For example, by inputting a sample image (such as a surround-view camera image) into the object detection model, the first neural network and the second neural network of the object detection model can be trained, and a first loss, loss_det1, can be obtained based on the 3D object detection box labels and the first 3D object detection box information. By inputting the sample image into the object detection model, the first neural network and the third neural network of the object detection model can be trained, and a second loss, loss_Occupancy, can be obtained based on the 3D object detection labels and the 3D object detection results. Based on this, a loss function as shown in Equation (1) can be constructed, and the object detection model can be trained according to the loss value determined by Equation (1).
[0093] Loss1=loss_det1+0.1*loss_Occupancy formula (1)
[0094] It should be noted that the loss function constructed based on the first and second losses is not limited to the form shown in formula (1). In practical applications, formula (1) can be adaptively adjusted according to business needs to determine the training loss that meets the actual needs. No limitation is made here.
[0095] Through the above embodiments of this disclosure, multi-task joint training can be achieved based on two constraints established according to the first loss and the second loss. Based on this method, the trained target detection model can have a more accurate depth estimation capability.
[0096] According to embodiments of this disclosure, operation S204 may further include: training the target detection model based on the 3D target detection box labels, first 3D target detection box information, 3D target detection labels, and 3D target detection results to obtain a pre-trained model; inputting sample images into the first neural network of the pre-trained model to obtain a second target feature vector; inputting randomly generated second 3D location query information and the second target feature vector into the second neural network of the pre-trained model to obtain second 3D target detection box information; training the first neural network and the second neural network in the pre-trained model based on the 3D target detection box labels and the second 3D target detection box information to obtain a trained first neural network and a trained second neural network; and determining the trained target detection model based on the trained first neural network and the trained second neural network.
[0097] According to embodiments of this disclosure, when the trained object detection model converges based on the 3D object detection box labels, first 3D object detection box information, 3D object detection labels, and 3D object detection results, the currently trained object detection model can be determined as a pre-trained model. See also... Figure 3As shown, the pre-trained model can have the same network structure as the object detection model 300. After obtaining the pre-trained model, for example, the occupancy branch 350 in the pre-trained model can be deleted and occupancy queries 432 can no longer be received. Only the feature extraction module 310, 3D Position Encoder 320, Decoder 330, and detection branch 340 are retained. Corresponding to the detection branch 340, the decoding module 330 continues to receive Detqueries 431, resulting in a single-task detection model that includes only the first and second neural networks. The single-task detection model is then further trained in the same way as the detection branch in the previous embodiment until the training of the single-task detection model converges, resulting in the trained object detection model.
[0098] According to embodiments of this disclosure, the second target feature vector has the same or similar features as the aforementioned second target feature vector. The second three-dimensional location query information has the same or similar features as the aforementioned first three-dimensional location query information. The second three-dimensional target detection box information has the same or similar features as the aforementioned first three-dimensional target detection box information. For embodiments describing how to obtain the second target feature vector by inputting the second target feature vector, the second three-dimensional location query information, the second three-dimensional target detection box information, and the second target detection box information by inputting the sample image into the first neural network of the pre-trained model, and how to obtain the second three-dimensional target detection box information by inputting the randomly generated second three-dimensional location query information and the second target feature vector into the second neural network of the pre-trained model, please refer to the foregoing embodiments, and will not be repeated here.
[0099] According to an embodiment of this disclosure, when training the first neural network and the second neural network in the pre-trained model based on the three-dimensional object detection box labels and the second three-dimensional object detection box information, the third loss loss_det2 can be determined based on the three-dimensional object detection box labels and the second three-dimensional object detection box information, and the loss function as shown in formula (2) can be constructed.
[0100] Loss2=loss_det2 formula (2)
[0101] It should be noted that the loss function constructed based on the third loss mentioned above is not limited to the form shown in formula (2). In practical applications, formula (2) can be adaptively adjusted according to business needs to determine the training loss that meets the actual needs. No limitation is made here.
[0102] Then, based on the third loss, the first and second neural networks in the pre-trained model can be trained until the networks converge, resulting in the trained first and second neural networks. Finally, based on the trained first and second neural networks, the trained object detection network is determined.
[0103] Through the embodiments described above, targeted enhancement training for the single task used to determine the object detection box in the object detection task can compensate for the negative impact of competition during multi-task training and help alleviate the problem of limited performance of the corresponding single task during joint training. Ultimately, through joint training and single-task training, the accuracy of the object detection model in predicting 3D object detection boxes, especially the depth accuracy, can be effectively improved.
[0104] According to embodiments of this disclosure, an object detection model is constructed and trained based on the aforementioned method. A voxel grid, accumulated from multiple frames and cloudified into dense points, is used as a dense depth supervision signal, which can compensate for the insufficient depth supervision in monocular 3D object detection bounding boxes. Employing a joint training approach involving occupancy networks and detection improves the model's ability to predict 3D spatial depth, significantly enhancing the performance of monocular 3D object detection. The improvement effect is shown in Table 1.
[0105] Table 1:
[0106]
[0107] Table 1 shows the improvement effect of the +OccupanCy Network on the object detection model after adding the occupancy branch for supervision. According to Table 1, adding the occupancy branch effectively improves the NDS and mAP metrics of the object detection model. Specifically, mAP improved by 1.6 percentage points, NDS improved by 0.8 percentage points, and mATE (the evaluation metric for position estimation error) decreased by 3 percentage points. The experimental results strongly demonstrate that the supervision of the occupancy network improves the performance of 3D position estimation in object detection. It also shows that the supervision of the occupancy network is more effective than the supervision of the depth map in improving the model's position prediction ability.
[0108] Figure 4 A flowchart illustrating a target detection method according to an embodiment of the present disclosure is shown schematically.
[0109] like Figure 4 As shown, the method includes operations S401 to S402.
[0110] In operation S401, the image to be processed is acquired.
[0111] In operation S402, the image to be processed is input into the target detection model to obtain the target detection result.
[0112] According to embodiments of this disclosure, the object detection model is trained using the object detection model training method described above.
[0113] According to embodiments of this disclosure, the image to be processed may have the same or similar features as the aforementioned sample image, which will not be repeated here.
[0114] The target detection results obtained through the above embodiments of this disclosure can have good depth features. They can demonstrate good applicability in fields such as monocular 3D target detection.
[0115] Figure 5 A block diagram of a training apparatus for an object detection model according to an embodiment of the present disclosure is shown schematically.
[0116] like Figure 5 As shown, the training device 500 for the target detection model includes a first acquisition module 510, a second acquisition module 520, a third acquisition module 530, and a training module 540.
[0117] The first acquisition module 510 is used to input the sample image into the first neural network of the target detection model to obtain the first target feature vector, wherein the first target feature vector has three-dimensional position information, and the sample image has a three-dimensional target detection box label and a three-dimensional target detection label.
[0118] The second acquisition module 520 is used to input the randomly generated first three-dimensional position query information and the first target feature vector into the second neural network of the target detection model to obtain the first three-dimensional target detection box information.
[0119] The third acquisition module 530 is used to input the voxel information of the three-dimensional sample region and the first target feature vector into the third neural network of the target detection model to obtain the three-dimensional target detection result, wherein the three-dimensional sample region is determined based on the sample image.
[0120] The training module 540 is used to train the target detection model based on the 3D target detection box labels, the first 3D target detection box information, the 3D target detection labels, and the 3D target detection results, so as to obtain the trained target detection model.
[0121] According to embodiments of this disclosure, the third obtaining module includes a first query unit, a first determining unit, a second determining unit, and a third determining unit.
[0122] The first query unit is used to query the second feature vector related to the voxel information from the target feature vector based on the voxel information.
[0123] The first determining unit is used to determine the target second feature vector representing the features of the target to be detected based on the decoding result of the second feature vector.
[0124] The second determining unit is used to determine the target voxel information corresponding to the second feature vector of the target from the voxel information.
[0125] The third determining unit is used to determine the three-dimensional target detection result based on the target voxel information.
[0126] According to embodiments of this disclosure, the training apparatus for the target detection model further includes a first acquisition module, a determination module, and a first partitioning module.
[0127] The first acquisition module is used to acquire a preset number of frames of images corresponding to the sample image.
[0128] The determination module is used to determine the point cloud information of the three-dimensional sample region based on the point cloud information of the sample image and the point cloud information of the acquired image.
[0129] The first segmentation module is used to segment the point cloud information of the three-dimensional sample region according to the preset voxel segmentation rules to obtain three-dimensional target detection labels.
[0130] According to embodiments of this disclosure, the training apparatus for the object detection model further includes a second partitioning module.
[0131] The second partitioning module is used to partition the three-dimensional sample region according to the preset voxel partitioning rules to obtain voxel information.
[0132] According to embodiments of this disclosure, the training module includes a fourth determining unit, a fifth determining unit, and a first training unit.
[0133] The fourth determining unit is used to determine the first loss based on the three-dimensional target detection box label and the first three-dimensional target detection box information.
[0134] The fifth determining unit is used to determine the second loss based on the 3D target detection labels and the 3D target detection results.
[0135] The first training unit is used to train the object detection model based on the first loss and the second loss.
[0136] According to embodiments of this disclosure, the training module includes a second training unit, a first acquisition unit, a second acquisition unit, a third training unit, and a sixth determination unit.
[0137] The second training unit is used to train the object detection model based on the 3D object detection box labels, the first 3D object detection box information, the 3D object detection labels, and the 3D object detection results, to obtain a pre-trained model.
[0138] The first acquisition unit is used to input the sample image into the first neural network of the pre-trained model to obtain the second target feature vector.
[0139] The second acquisition unit is used to input the randomly generated second three-dimensional location query information and the second target feature vector into the second neural network of the pre-trained model to obtain the second three-dimensional target detection box information.
[0140] The third training unit is used to train the first neural network and the second neural network in the pre-trained model based on the labels of the three-dimensional object detection boxes and the information of the second three-dimensional object detection boxes, so as to obtain the trained first neural network and the trained second neural network.
[0141] The sixth determining unit is used to determine the trained target detection model based on the trained first neural network and the trained second neural network.
[0142] According to embodiments of this disclosure, the sample images include images obtained from multiple perspectives using a monocular camera at a fixed location.
[0143] According to embodiments of this disclosure, the first obtaining module includes a third obtaining unit and a fourth obtaining unit.
[0144] The third acquisition unit is used to extract features from the sample image to obtain a two-dimensional feature vector of the sample image.
[0145] The fourth acquisition unit is used to perform three-dimensional position encoding on the two-dimensional feature vector to obtain the first target feature vector.
[0146] According to embodiments of this disclosure, the second obtaining module includes a second query unit, a seventh determining unit, an eighth determining unit, and a ninth determining unit.
[0147] The second query unit is used to query the first feature vector related to the first three-dimensional location query information from the first target feature vector based on the first three-dimensional location query information.
[0148] The seventh determining unit is used to determine the target first feature vector representing the features of the target to be detected based on the decoding result of the first feature vector;
[0149] The eighth determining unit is used to determine, from the first three-dimensional position query information, the target three-dimensional position query information corresponding to the target's first feature vector; and
[0150] The ninth determining unit is used to determine the first three-dimensional target detection box information based on the target's three-dimensional position query information.
[0151] Figure 6 A block diagram of a target detection apparatus according to an embodiment of the present disclosure is shown schematically.
[0152] like Figure 6 As shown, the target detection device 600 includes a second acquisition module 610 and a fourth acquisition module 620.
[0153] The second acquisition module 610 is used to acquire the image to be processed.
[0154] The fourth obtaining module 620 is used to input the image to be processed into the target detection model to obtain the target detection result, wherein the target detection model is trained using the training device of the target detection model of this disclosure.
[0155] Any one or more of the modules or units according to embodiments of this disclosure, or at least a portion thereof, may be implemented in a single module. Any one or more of the modules or units according to embodiments of this disclosure may be implemented by dividing them into multiple modules. Any one or more of the modules or units according to embodiments of this disclosure may be at least partially implemented as hardware circuitry, such as a Field-Programmable Gate Array (FPGA), a Programmable Logic Array (PLA), a System-on-Chip, a System-on-Substrate, a System-on-Package, an Application-Specific Integrated Circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging the circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three methods. Alternatively, one or more of the modules or units according to embodiments of this disclosure may be at least partially implemented as computer program modules, which, when run, can perform corresponding functions.
[0156] For example, any plurality of the first acquisition module 510, the second acquisition module 520, the third acquisition module 530, and the training module 540, or the second acquisition module 610 and the fourth acquisition module 620, can be combined into one module / unit, or any one of these modules / units can be split into multiple modules / units. Alternatively, at least part of the functionality of one or more of these modules / units can be combined with at least part of the functionality of other modules / units and implemented in one module / unit. According to embodiments of this disclosure, at least one of the first acquisition module 510, the second acquisition module 520, the third acquisition module 530, and the training module 540, or the second acquisition module 610 and the fourth acquisition module 620, can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging the circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the first acquisition module 510, the second acquisition module 520, the third acquisition module 530, and the training module 540, or the second acquisition module 610 and the fourth acquisition module 620, can be at least partially implemented as a computer program module that can perform corresponding functions when the computer program module is run.
[0157] It should be noted that the training device part of the object detection model in the embodiments of this disclosure corresponds to the training method part of the object detection model in the embodiments of this disclosure. For a detailed description of the training device part of the object detection model, please refer to the training method part of the object detection model, and it will not be repeated here. Similarly, the object detection device part in the embodiments of this disclosure corresponds to the object detection method part in the embodiments of this disclosure. For a detailed description of the object detection device part, please refer to the object detection method part, and it will not be repeated here.
[0158] Figure 7 A block diagram of an electronic device suitable for implementing a training method for an object detection model and at least one of the object detection methods according to embodiments of the present disclosure is illustrated. Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0159] like Figure 7 As shown, an electronic device 700 according to an embodiment of the present disclosure includes a processor 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage portion 708 into a random access memory (RAM) 703. The processor 701 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 701 may also include onboard memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0160] RAM 703 stores various programs and data required for the operation of electronic device 700. Processor 701, ROM 702, and RAM 703 are interconnected via bus 704. Processor 701 performs various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 702 and / or RAM 703. It should be noted that the programs may also be stored in one or more memories other than ROM 702 and RAM 703. Processor 701 may also perform various operations of the method flow according to embodiments of the present disclosure by executing programs stored in said one or more memories.
[0161] According to embodiments of this disclosure, the electronic device 700 may further include an input / output (I / O) interface 705, which is also connected to a bus 704. The system 700 may also include one or more of the following components connected to the input / output (I / O) interface 705: an input section 706 including a keyboard, mouse, etc.; an output section 707 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the input / output (I / O) interface 705 as needed. A removable medium 711, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 710 as needed so that computer programs read from it can be installed into the storage section 708 as needed.
[0162] According to embodiments of this disclosure, the method flow according to embodiments of this disclosure can be implemented as a computer software program. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via communication section 709, and / or installed from removable medium 711. When the computer program is executed by processor 701, it performs the functions defined in the system of embodiments of this disclosure. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0163] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.
[0164] According to embodiments of this disclosure, the computer-readable storage medium can be a non-volatile computer-readable storage medium. Examples include, but are not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0165] For example, according to embodiments of this disclosure, a computer-readable storage medium may include the ROM 702 and / or RAM 703 described above and / or one or more memories other than ROM 702 and RAM 703.
[0166] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods provided in the embodiments of this disclosure. When the computer program product is run on an electronic device, the program code is used to enable the electronic device to implement the methods provided in the embodiments of this disclosure.
[0167] When the computer program is executed by the processor 701, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0168] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 709, and / or installed from a removable medium 711. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0169] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can execute entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0170] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions. Those skilled in the art will understand that the features recited in the various embodiments and / or claims of this disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not expressly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure may be combined and / or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.
[0171] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. The scope of this disclosure is defined by the appended claims and their equivalents. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.
Claims
1. A method for training an object detection model, comprising: The sample image is input into the first neural network of the target detection model to obtain the first target feature vector, wherein the first target feature vector has three-dimensional position information, and the sample image has a three-dimensional target detection box label and a three-dimensional target detection label; The randomly generated first three-dimensional location query information and the first target feature vector are input into the second neural network of the target detection model to obtain the first three-dimensional target detection box information; The voxel information of the three-dimensional sample region and the first target feature vector are input into the third neural network of the target detection model to obtain the three-dimensional target detection result, wherein the three-dimensional sample region is determined based on the sample image; and The target detection model is trained based on the 3D target detection box label, the first 3D target detection box information, the 3D target detection label, and the 3D target detection result to obtain a trained target detection model. The step of inputting randomly generated first three-dimensional location query information and the first target feature vector into the second neural network of the target detection model to obtain first three-dimensional target detection box information includes: querying a first feature vector related to the first three-dimensional location query information from the first target feature vector based on the first three-dimensional location query information; determining a target first feature vector representing the features of the target to be detected based on the decoding result of the first feature vector; determining target three-dimensional location query information corresponding to the target first feature vector from the first three-dimensional location query information; and determining the first three-dimensional target detection box information based on the target three-dimensional location query information. The step of inputting the voxel information of the three-dimensional sample region and the first target feature vector into the third neural network of the target detection model to obtain the three-dimensional target detection result includes: querying the second feature vector related to the voxel information from the target feature vector based on the voxel information; determining the target second feature vector representing the feature of the target to be detected based on the decoding result of the second feature vector; determining the target voxel information corresponding to the target second feature vector from the voxel information; and determining the three-dimensional target detection result based on the target voxel information. The step of training the target detection model based on the 3D target detection box labels, the first 3D target detection box information, the 3D target detection labels, and the 3D target detection results to obtain a trained target detection model includes: training the target detection model based on the 3D target detection box labels, the first 3D target detection box information, the 3D target detection labels, and the 3D target detection results to obtain a pre-trained model; inputting the sample image into the first neural network of the pre-trained model to obtain a second target feature vector; inputting randomly generated second 3D location query information and the second target feature vector into the second neural network of the pre-trained model to obtain second 3D target detection box information; training the first neural network and the second neural network in the pre-trained model based on the 3D target detection box labels and the second 3D target detection box information to obtain a trained first neural network and a trained second neural network; and determining the trained target detection model based on the trained first neural network and the trained second neural network.
2. The method according to claim 1, further comprising: Acquire a preset number of frames of images corresponding to the sample image; The point cloud information of the three-dimensional sample region is determined based on the point cloud information of the sample image and the point cloud information of the acquired image. as well as According to the preset voxel division rules, the point cloud information of the three-dimensional sample region is divided to obtain the three-dimensional target detection label.
3. The method according to claim 2, further comprising: Before inputting the voxel information of the three-dimensional sample region and the first target feature vector into the third neural network of the target detection model to obtain the three-dimensional target detection result. The three-dimensional sample region is divided according to the preset voxel division rules to obtain the voxel information.
4. The method according to claim 1, wherein, The step of training the target detection model based on the 3D target detection box label, the first 3D target detection box information, the 3D target detection label, and the 3D target detection result includes: The first loss is determined based on the three-dimensional target detection box label and the first three-dimensional target detection box information; Based on the 3D target detection labels and the 3D target detection results, a second loss is determined; and The target detection model is trained based on the first loss and the second loss.
5. The method according to claim 1, wherein, The sample images include images obtained from multiple perspectives using a monocular camera at a fixed location.
6. The method according to claim 1, wherein, The step of inputting the sample image into the first neural network of the target detection model to obtain the first target feature vector includes: Feature extraction is performed on the sample image to obtain a two-dimensional feature vector of the sample image; and The two-dimensional feature vector is encoded in three dimensions to obtain the first target feature vector.
7. A target detection method, comprising: Obtain the image to be processed; as well as The image to be processed is input into the target detection model to obtain the target detection result, wherein the target detection model is trained using the method described in any one of claims 1-6.
8. A training device for an object detection model, comprising: The first acquisition module is used to input the sample image into the first neural network of the target detection model to obtain the first target feature vector, wherein the first target feature vector has three-dimensional position information, and the sample image has a three-dimensional target detection box label and a three-dimensional target detection label; The second acquisition module is used to input the randomly generated first three-dimensional position query information and the first target feature vector into the second neural network of the target detection model to obtain the first three-dimensional target detection box information. The third acquisition module is used to input the voxel information of the three-dimensional sample region and the first target feature vector into the third neural network of the target detection model to obtain the three-dimensional target detection result, wherein the three-dimensional sample region is determined based on the sample image; and The training module is used to train the target detection model based on the 3D target detection box label, the first 3D target detection box information, the 3D target detection label, and the 3D target detection result, so as to obtain the trained target detection model. The second obtaining module includes: a second query unit, configured to query a first feature vector related to the first three-dimensional location query information from the first target feature vector based on the first three-dimensional location query information; a seventh determining unit, configured to determine a target first feature vector representing the features of the target to be detected based on the decoding result of the first feature vector; an eighth determining unit, configured to determine target three-dimensional location query information corresponding to the target first feature vector from the first three-dimensional location query information; and a ninth determining unit, configured to determine the first three-dimensional target detection box information based on the target three-dimensional location query information. The third obtaining module includes: a first query unit, configured to query a second feature vector related to the voxel information from the target feature vector based on the voxel information; a first determining unit, configured to determine a target second feature vector characterizing the features of the target to be detected based on the decoding result of the second feature vector; a second determining unit, configured to determine target voxel information corresponding to the target second feature vector from the voxel information; and a third determining unit, configured to determine the three-dimensional target detection result based on the target voxel information. The training module includes: a second training unit, used to train the target detection model based on the 3D target detection box labels, the first 3D target detection box information, the 3D target detection labels, and the 3D target detection results to obtain a pre-trained model; a first obtaining unit, used to input the sample image into the first neural network of the pre-trained model to obtain a second target feature vector; a second obtaining unit, used to input randomly generated second 3D location query information and the second target feature vector into the second neural network of the pre-trained model to obtain second 3D target detection box information; a third training unit, used to train the first neural network and the second neural network in the pre-trained model based on the 3D target detection box labels and the second 3D target detection box information to obtain a trained first neural network and a trained second neural network; and a sixth determining unit, used to determine the trained target detection model based on the trained first neural network and the trained second neural network.
9. A target detection device, comprising: The second acquisition module is used to acquire the image to be processed; as well as The fourth obtaining module is used to input the image to be processed into the target detection model to obtain the target detection result, wherein the target detection model is trained using the device as described in claim 8.
10. An electronic device, comprising: One or more processors; Memory, used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method of any one of claims 1-7.
11. A computer-readable storage medium having stored thereon executable instructions that, when executed by a processor, cause the processor to perform the method of any one of claims 1-7.
12. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-7.
Citation Information
Patent Citations
Model training method and device, target detection method and device, equipment and storage medium
CN115719436A
Three-dimensional target detection method, device and equipment and storage medium
CN116229451A