A three-dimensional object detection method, system, storage medium and electronic device
By performing feature extraction and parallax estimation on images from different perspectives, the problem of high cost of three-dimensional object detection is solved and efficient and low-cost three-dimensional object detection is achieved.
Patent Information
- Application Number
- CN202111563416.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-20
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-12-20
AI Technical Summary
The existing three-dimensional object detection methods rely on lidar to cause high costs, while lidar with fewer lines leads to degradation of detection accuracy and performance.
By acquiring two original images from different perspectives, image feature extraction and parallax estimation are generated, and pseudo-point cloud data is used to box the three-dimensional target to be detected, replacing traditional lidar detection.
Without reducing the recognition speed and accuracy, the cost of three-dimensional object detection is significantly reduced, while maintaining category richness and recognition speed.
Smart Images

Figure CN114387573B_ABST
Abstract
Description
Background Art
[0002] In recent years, with the development of computer and deep learning technologies, numerous impressive methods have been proposed. 3D object detection plays an important role in many applications. A key task in scene understanding is to detect 3D objects, which has become a hot research topic in various application fields such as autonomous driving. The purpose of 3D object detection technology is to detect and locate the 3D bounding box of the object to be detected from the input sensor data. It is applied to fields such as unmanned driving, robot environmental perception technology, and augmented reality.
[0003] Currently, most methods use lidar point clouds as input. Lidar has the characteristics of high measurement accuracy and strong anti-interference ability. The more lines and the denser the lidar, the more environmental information the data contains. However, with the increase in the number of lidar lines, while bringing high performance, the corresponding cost will also be higher. To reduce costs, people tend to use lidar with fewer lines, but in this case, the accuracy and performance of 3D object detection will also decline. Summary of the Invention
[0004] To solve the above technical problems, the present invention provides a 3D object detection method, system, storage medium, and electronic device.
[0005] The technical solution of a 3D object detection method of the present invention is as follows:
[0006] S1. Obtain and perform image feature extraction on two original images from different perspectives to obtain first image processing data, and perform disparity estimation processing on the two original images from different perspectives to obtain second image processing data, where both of the two original images contain the 3D object to be detected;
[0007] S2. Map and generate pseudo-point cloud data according to the first image processing data and the second image processing data;
[0008] S3. Select the 3D object to be detected in at least one original image according to the pseudo-point cloud data.
[0009] The beneficial effects of a 3D object detection method of the present invention are as follows:
[0010] The method of the present invention obtains and extracts image features from two original images from different perspectives to obtain first image processing data, and performs disparity estimation processing on the two original images from different perspectives to obtain second image processing data; then maps and generates pseudo-point cloud data according to the first image processing data and the second image processing data; finally, frames the three-dimensional target to be detected in at least one original image according to the pseudo-point cloud data. Therefore, the method of the present invention greatly reduces the dependence of three-dimensional target detection on traditional lidar detection methods, and while having performance no less than that of relying on lidar point cloud recognition in terms of rich category recognition, fast recognition speed, high accuracy, etc., it greatly reduces the cost.
[0011] On the basis of the above solution, a three-dimensional target detection method of the present invention can be further improved as follows.
[0012] Further, the two original images from different perspectives include: a left image and a right image, and the S1 specifically includes:
[0013] S11. Use a residual network model to extract image features from the left image to obtain local proposal data, and use a self-attention mechanism to extract image features from the right image to obtain global proposal data, and generate the first image processing data based on the local proposal data and the global proposal data;
[0014] S12. Use the disparity estimation network model to compare the left image and the right image to obtain disparity depth data, and use the self-attention mechanism to extract image features from the disparity depth data to obtain the second image processing data.
[0015] Further, the generating the first image processing data based on the local proposal data and the global proposal data specifically includes:
[0016] Fuse the local proposal data and the global proposal data to generate fused proposal data, and process the fused proposal data after pyramid pooling with an RCNN network model to generate the first image processing data.
[0017] Further, the S2 specifically includes:
[0018] S21. Fuse the features of the first image processing data and the second image processing data to obtain feature fusion data;
[0019] S22. Perform ground truth supervised learning on the feature fusion data to generate the pseudo-point cloud data.
[0020] Further, the S3 specifically includes:
[0021] S31. Use the PoinNet network model to extract features from the pseudo-point cloud data to obtain center point feature data;
[0022] S32. Process the center point data in the 3D object detector to obtain additional point feature data;
[0023] S33. Classify and regress the center point feature data and the additional point feature data to generate a 3D bounding box of the 3D object to be detected.
[0024] The technical solution of a 3D object detection system of the present invention is as follows:
[0025] It includes: a processing module, a generating module, and a detecting module;
[0026] The processing module is used to: obtain and perform image feature extraction on two original images from different perspectives to obtain first image processing data, and perform disparity estimation processing on the two original images from different perspectives to obtain second image processing data, where both of the two original images contain the 3D object to be detected;
[0027] The generating module is used to: map and generate pseudo-point cloud data according to the first image processing data and the second image processing data;
[0028] The detecting module is used to: frame out the 3D object to be detected in at least one original image according to the pseudo-point cloud data.
[0029] The beneficial effects of a 3D object detection system of the present invention are as follows:
[0030] The system of the present invention obtains and performs image feature extraction on two original images from different perspectives to obtain first image processing data, and performs disparity estimation processing on the two original images from different perspectives to obtain second image processing data; then maps and generates pseudo-point cloud data according to the first image processing data and the second image processing data; finally, frames out the 3D object to be detected in at least one original image according to the pseudo-point cloud data. Therefore, the system of the present invention greatly reduces the dependence on the traditional lidar detection method for 3D object detection, and while having performance no less than that of relying on lidar point cloud recognition in terms of rich category recognition, fast recognition speed, high accuracy, etc., it also greatly reduces the cost.
[0031] On the basis of the above solution, a 3D object detection system of the present invention can also be improved as follows.
[0032] Further, the two original images from different perspectives include: a left image and a right image, then the processing module is specifically used to:
[0033] Extract image features from the left image using a residual network model to obtain local proposal data, and extract image features from the right image using a self-attention mechanism to obtain global proposal data, and generate the first image processing data based on the local proposal data and the global proposal data;
[0034] Compare the left image and the right image using the disparity estimation network model to obtain disparity depth data, and extract image features from the disparity depth data using the self-attention mechanism to obtain the second image processing data.
[0035] Further, the generating the first image processing data based on the local proposal data and the global proposal data specifically includes:
[0036] Fuse the local proposal data and the global proposal data to generate fused proposal data, and process the fused proposal data after pyramid pooling using an RCNN network model to generate the first image processing data.
[0037] The technical solution of a storage medium of the present invention is as follows:
[0038] Instructions are stored in the storage medium, and when the computer reads the instructions, the computer is caused to execute the steps of a three-dimensional object detection method of the present invention.
[0039] The technical solution of an electronic device of the present invention is as follows:
[0040] It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. It is characterized in that when the processor executes the computer program, the computer is caused to execute the steps of a three-dimensional object detection method of the present invention. Description of the Drawings
[0041] Figure 1 It is a schematic flowchart of a three-dimensional object detection method according to an embodiment of the present invention;
[0042] Figure 2 It is a schematic structural diagram of a three-dimensional object detection system according to an embodiment of the present invention. Detailed Embodiments
[0043] As Figure 1 shown, a three-dimensional object detection method according to an embodiment of the present invention includes:
[0044] S1. Obtain and extract image features from two original images with different perspectives to obtain first image processing data, and perform disparity estimation processing on the two original images with different perspectives to obtain second image processing data, wherein both of the two original images contain a three-dimensional object to be detected;
[0045] Among them, the two original images from different perspectives are: two original images from different perspectives collected in the same scene. The image feature extraction process is as follows: after feature extraction by a residual network model and an attention mechanism, first image processing data is obtained.
[0046] Among them, the two original images for disparity estimation processing are the same as the original images for feature extraction described above.
[0047] Among them, the original images all contain the three-dimensional target to be detected. For example, the three-dimensional target to be detected can be a vehicle.
[0048] S2. Generate pseudo point cloud data according to the mapping of the first image processing data and the second image processing data;
[0049] Among them, the mapping generation process in image processing belongs to conventional technical means and will not be elaborated here.
[0050] S3. Select the three-dimensional target to be detected in at least one original image according to the pseudo point cloud data.
[0051] Preferably, the two original images from different perspectives include: a left image and a right image. Then S1 specifically includes:
[0052] S11. Use a residual network model to perform image feature extraction on the left image to obtain local proposal data, and use a self-attention mechanism to perform image feature extraction on the right image to obtain global proposal data. Generate the first image processing data based on the local proposal data and the global proposal data;
[0053] Among them, the left image and the right image in this step are the images after obtaining the RGB image features of the two original images from different perspectives collected in the same scene.
[0054] Among them, the residual network model is a type of convolutional neural network model. The characteristics of the residual network model are easy to optimize and can improve the accuracy by increasing the appropriate depth. The internal residual blocks use skip connections, alleviating the problem of gradient disappearance caused by increasing the depth in deep neural networks.
[0055] Among them, the self-attention mechanism is a variant of the attention mechanism, which reduces the dependence on external information and is better at capturing the internal correlation of data or features.
[0056] S12. Use the disparity estimation network model to compare the left image and the right image to obtain disparity depth data, and use the self-attention mechanism to perform image feature extraction on the disparity depth data to obtain the second image processing data.
[0057] Among them, the left image and the right image in this step are the images after obtaining depth image features from two original images with different perspectives collected in the same scene.
[0058] Among them, the disparity estimation network model is a type of convolutional neural network model.
[0059] Preferably, generating the first image processing data based on the local proposal data and the global proposal data specifically includes:
[0060] Fusing the local proposal data and the global proposal data to generate fused proposal data, and processing the fused proposal data after pyramid pooling with an RCNN network model to generate the first image processing data.
[0061] Preferably, S2 specifically includes:
[0062] S21. Feature-fusing the first image processing data and the second image processing data to obtain feature-fused data;
[0063] S22. Performing ground truth supervised learning on the feature-fused data to generate the pseudo point cloud data.
[0064] Preferably, S3 specifically includes:
[0065] S31. Using a PoinNet network model to extract features from the pseudo point cloud data to obtain center point feature data;
[0066] S32. Processing the center point data in a 3D object detector to obtain additional point feature data;
[0067] S33. Classifying and regressing the center point feature data and the additional point feature data to generate the 3D bounding box of the 3D object to be detected.
[0068] Specifically, first, a PoinNet network model is used to extract features from the pseudo point cloud data, and then the features are input into a center point 3D object detector. A key point detector and regression are used to object center to other attributes, including 3D size, 3D orientation, and velocity. Secondly, it uses additional point features on the object. At the center point, 3D object tracking is simplified to greedy nearest point matching. Finally, under the supervision of the laser point cloud ground truth, classification and regression are performed, and finally the predicted 3D bounding box of the 3D object to be detected is generated.
[0069] The technical solution of this embodiment extracts image features from two original images with different perspectives to obtain first image processing data, and performs disparity estimation processing on the two original images with different perspectives to obtain second image processing data; then maps and generates pseudo-point cloud data according to the first image processing data and the second image processing data; finally, frames the three-dimensional target to be detected in at least one original image according to the pseudo-point cloud data. Therefore, the technical solution of this embodiment greatly reduces the dependence of three-dimensional target detection on traditional lidar detection methods, and while having performance no less than that of relying on lidar point cloud recognition in terms of rich category recognition, fast recognition speed, high accuracy, etc., it also greatly reduces the cost.
[0070] As Figure 2 shown, a three-dimensional target detection system 200 according to an embodiment of the present invention includes: a processing module 210, a generation module 220, and a detection module 230;
[0071] The processing module 210 is configured to: extract image features from two original images with different perspectives to obtain first image processing data, and perform disparity estimation processing on the two original images with different perspectives to obtain second image processing data, where both of the two original images include the three-dimensional target to be detected;
[0072] The generation module 220 is configured to: map and generate pseudo-point cloud data according to the first image processing data and the second image processing data;
[0073] The detection module 230 is configured to: frame the three-dimensional target to be detected in at least one original image according to the pseudo-point cloud data.
[0074] Preferably, the two original images with different perspectives include: a left image and a right image, and the processing module 210 is specifically configured to:
[0075] Use a residual network model to extract image features from the left image to obtain local proposal data, and use a self-attention mechanism to extract image features from the right image to obtain global proposal data, and generate the first image processing data based on the local proposal data and the global proposal data;
[0076] Use the disparity estimation network model to compare the left image and the right image to obtain disparity depth data, and use the self-attention mechanism to extract image features from the disparity depth data to obtain the second image processing data.
[0077] Preferably, the generating the first image processing data based on the local proposal data and the global proposal data specifically includes:
[0078] Fuse the local proposal data and the global proposal data to generate fused proposal data, and process the fused proposal data after pyramid pooling using an RCNN network model to generate the first image processing data.
[0079] The technical solution of this embodiment obtains and extracts image features from two original images with different perspectives to obtain the first image processing data, and performs disparity estimation processing on the two original images with different perspectives to obtain the second image processing data; then maps and generates pseudo point cloud data according to the first image processing data and the second image processing data; finally, frames the three-dimensional target to be detected in at least one original image according to the pseudo point cloud data. Therefore, the technical solution of this embodiment greatly reduces the dependence of three-dimensional target detection on traditional lidar detection methods, and while having performance no less than that of relying on laser point cloud recognition in terms of rich category recognition, fast recognition speed, high accuracy, etc., it also greatly reduces the cost.
[0080] For the parameters and corresponding functions implemented by each module in the three-dimensional target detection system of the present invention described above, reference can be made to the parameters and steps in the embodiment of the three-dimensional target detection method in the above text, which will not be elaborated here.
[0081] A storage medium provided by an embodiment of the present invention includes: instructions are stored in the storage medium, and when a computer reads the instructions, the computer is made to execute the steps of a three-dimensional target detection method as described above. Specifically, reference can be made to the parameters and steps in the embodiment of the three-dimensional target detection method in the above text, which will not be elaborated here.
[0082] Computer storage media such as: USB flash drives, external hard drives, etc.
[0083] An electronic device provided by an embodiment of the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor, when executing the computer program, makes the computer execute the steps of a three-dimensional target detection method as described above. Specifically, reference can be made to the parameters and steps in the embodiment of the three-dimensional target detection method in the above text, which will not be elaborated here.
[0084] Those skilled in the art know that the present invention can be implemented as a method, a system, a storage medium, and an electronic device.
[0085] Accordingly, the present invention may be embodied in the following forms, namely: it may be entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, which is generally referred to herein as "circuit", "module" or "system". In addition, in some embodiments, the present invention may also be implemented in the form of a computer program product in one or more computer-readable media, which contain computer-readable program code. Any combination of one or more computer-readable media may be adopted. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium may be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A three-dimensional object detection method, characterized in that, Including: S1. Obtain and perform image feature extraction on two original images from different perspectives to obtain first image processing data, and perform disparity estimation processing on the two original images from different perspectives to obtain second image processing data, where both of the two original images contain a three-dimensional target to be detected; S2. Map and generate pseudo point cloud data according to the first image processing data and the second image processing data; S3. Select the three-dimensional target to be detected in at least one original image according to the pseudo point cloud data; The two original images from different perspectives include: a left image and a right image, then S1 specifically includes: S11. Use a residual network model to perform image feature extraction on the left image to obtain local proposal data, and use a self-attention mechanism to perform image feature extraction on the right image to obtain global proposal data, and generate the first image processing data based on the local proposal data and the global proposal data; S12. Use a disparity estimation network model to compare the left image and the right image to obtain disparity depth data, and use the self-attention mechanism to perform image feature extraction on the disparity depth data to obtain the second image processing data.
2. The three-dimensional object detection method according to claim 1, characterized in that The generating the first image processing data based on the local proposal data and the global proposal data specifically includes: Fuse the local proposal data and the global proposal data to generate fused proposal data, and use an RCNN network model to process the fused proposal data after pyramid pooling processing to generate the first image processing data.
3. A three-dimensional object detection method according to claim 1, characterized in that, S2 specifically includes: S21. Perform feature fusion on the first image processing data and the second image processing data to obtain feature fusion data; S22. Perform ground truth supervised learning on the feature fusion data to generate the pseudo point cloud data.
4. A three-dimensional object detection method according to claim 1, characterized in that S3 specifically includes: S31. Use a PointNet network model to perform feature extraction on the pseudo point cloud data to obtain center point feature data; S32. Process the center point feature data in a three-dimensional target detector to obtain additional point feature data; S33. Perform classification regression on the center point feature data and the additional point feature data to generate a three-dimensional bounding box of the three-dimensional target to be detected.
5. A three-dimensional object detection system, characterized in that, Including: A processing module, a generating module, and a detecting module; The processing module is used for: obtaining and performing image feature extraction on two original images from different perspectives to obtain first image processing data, and performing disparity estimation processing on the two original images from different perspectives to obtain second image processing data, where both of the two original images contain a three-dimensional target to be detected; The generating module is used for: mapping and generating pseudo point cloud data according to the first image processing data and the second image processing data; The detecting module is used for: selecting the three-dimensional target to be detected in at least one original image according to the pseudo point cloud data; The two original images from different perspectives include: a left image and a right image, then the processing module is specifically used for: Extract image features from the left image using a residual network model to obtain local proposal data, and extract image features from the right image using a self-attention mechanism to obtain global proposal data, and generate the first image processing data based on the local proposal data and the global proposal data; Compare the left image and the right image using a disparity estimation network model to obtain disparity depth data, and extract image features from the disparity depth data using the self-attention mechanism to obtain the second image processing data.
6. The three-dimensional target detection system according to claim 5, characterized in that, The generating the first image processing data based on the local proposal data and the global proposal data specifically includes: Fuse the local proposal data and the global proposal data to generate fused proposal data, and process the fused proposal data after pyramid pooling using an RCNN network model to generate the first image processing data.
7. A storage medium, characterized in that, Instructions are stored in the storage medium, and when the computer reads the instructions, the computer is caused to execute a three-dimensional object detection method according to any one of claims 1 to 4.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the electronic device is caused to execute a three-dimensional object detection method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Three-dimensional target detection method, device and system based on binocular vision
CN112287824A
Obstacle detection and marking method and device for automatic driving and storage medium
CN112419494A
Vehicle recognition method and device, computer equipment and storage medium
CN113763717A