Target detection method, device, equipment and storage medium

By performing 3D target detection and feature fusion on each frame of multiple image sets, and utilizing a monocular 3D target detection algorithm and a self-attention mechanism, the problem of low cross-camera target detection accuracy in autonomous driving scenarios is solved, achieving higher detection accuracy.

CN114708583BActive Publication Date: 2025-12-09GUANGZHOU WERIDE TECH LTD CO
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210171913.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-24
Publication Date
2025-12-09
Estimated Expiration
2042-02-24

AI Technical Summary

Technical Problem

Existing technologies have low accuracy in cross-camera target detection in autonomous driving scenarios, requiring multiple cameras and multiple frames of information to predict the motion information of the target.

Method used

By performing 3D object detection on each frame of images from multiple image sets, extracting 3D feature maps and performing feature fusion, and filtering out object detection boxes, the monocular 3D object detection algorithm and self-attention mechanism are used to filter object candidate boxes.

Benefits of technology

It improves the accuracy of cross-camera object detection, identifying complete and non-overlapping object detection boxes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114708583B_ABST
    Figure CN114708583B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of artificial intelligence, and discloses a target object detection method, device and equipment and a storage medium, which are used for improving the accuracy of cross-camera target object detection. The target object detection method comprises the following steps: performing 3D target detection on each frame of image in a plurality of image sets to obtain a plurality of target object candidate boxes of each frame of image, one image set corresponds to one camera, and each image set comprises a plurality of frames of images collected by the camera; performing 3D space feature extraction on each frame of image in the plurality of image sets to obtain a 3D feature map corresponding to each frame of image; performing feature fusion on the 3D feature map corresponding to each frame of image to obtain a target fusion feature map; extracting fusion feature information corresponding to each target object candidate box of each frame of image from the target fusion feature map, and screening all target object candidate boxes according to the fusion feature information corresponding to each target object candidate box of each frame of image to obtain at least one target object detection box.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a target object detection method and device, equipment and a storage medium. BACKGROUND

[0002] With the development of computer vision processing technology, cameras have become important sensing elements for autonomous driving perception, and can provide rich details and texture information.

[0003] The prior art is usually based on the image view itself to predict the actual position of each 2D target object in 3D, but in the autonomous driving scenario, multiple cameras are usually needed to observe the target object completely, and multiple frames of information are needed to predict the motion information (such as speed, acceleration, etc.) of the target object, so the prior art has the technical problem of low accuracy when processing cross-camera target object detection. SUMMARY

[0004] The present application provides a target object detection method, device, equipment and storage medium, which can improve the accuracy of cross-camera target object detection.

[0005] The first aspect of the present application provides a target object detection method, comprising:

[0006] performing 3D target detection on each frame of image in a plurality of image sets to obtain a plurality of target object candidate boxes of each frame of image, one image set corresponding to one camera, and each image set comprising a plurality of frames of image collected by the camera;

[0007] performing 3D spatial feature extraction on each frame of image in the plurality of image sets to obtain a 3D feature map corresponding to each frame of image;

[0008] performing feature fusion on the 3D feature map corresponding to each frame of image to obtain a target fusion feature map;

[0009] extracting fusion feature information corresponding to each target object candidate box of each frame of image from the target fusion feature map, and performing screening on all target object candidate boxes according to the fusion feature information corresponding to each target object candidate box of each frame of image to obtain at least one target object detection box.

[0010] Optionally, in the first implementation manner of the first aspect of the present application, the 3D spatial feature extraction on each frame of image in the plurality of image sets to obtain a 3D feature map corresponding to each frame of image comprises:

[0011] performing 3D spatial conversion on each frame of image in the plurality of image sets to obtain a 3D space map corresponding to each frame of image;

[0012] Obtain target feature information corresponding to each frame of image, and project the target feature information corresponding to each frame of image to a 3D space graph corresponding to each frame of image to obtain a 3D feature graph corresponding to each frame of image.

[0013] Optionally, in the second implementation manner of the first aspect of the present application, the 3D space conversion of each frame of image in the plurality of image sets to obtain a 3D space graph corresponding to each frame of image comprises:

[0014] Perform pixel-by-pixel depth estimation on each frame of image in the plurality of image sets to obtain a 3D space graph corresponding to each frame of image, wherein each 3D point in the 3D space graph corresponding to each frame of image corresponds to a 3D space coordinate information.

[0015] Optionally, in the third implementation manner of the first aspect of the present application, the obtaining of the target feature information corresponding to each frame of image and the projecting of the target feature information corresponding to each frame of image to the 3D space graph corresponding to each frame of image to obtain a 3D feature graph corresponding to each frame of image comprises:

[0016] Reading target feature information corresponding to each frame of image, wherein the target feature information corresponding to each frame of image comprises at least one of laser radar feature information, millimeter wave radar feature information, ultrasonic wave feature information and image feature information of each frame of image;

[0017] Obtaining feature coordinate information, wherein the feature coordinate information is used to indicate coordinate information of the target feature information corresponding to each frame of image in the corresponding frame of image;

[0018] According to the feature coordinate information, mapping the target feature information corresponding to each frame of image to the corresponding 3D space graph to obtain a 3D feature graph corresponding to each frame of image.

[0019] Optionally, in the fourth implementation manner of the first aspect of the present application, the feature fusion of the 3D feature graph corresponding to each frame of image to obtain a target fusion feature graph comprises:

[0020] Performing bird's eye view feature synthesis on the 3D feature graphs corresponding to the same frame of image in all image sets to obtain a bird's eye view feature graph corresponding to the same sequence frame of image;

[0021] Performing feature superposition on the bird's eye view feature graphs corresponding to the same sequence frame of image to obtain a target fusion feature graph.

[0022] Optionally, in the fifth implementation manner of the first aspect of the present application, the feature fusion of the 3D feature graph corresponding to each frame of image to obtain a target fusion feature graph further comprises:

[0023] Performing feature superposition on the 3D feature graphs corresponding to each frame of image in each image set to obtain an initial fusion feature graph corresponding to each image set;

[0024] The initial fusion feature maps corresponding to all the image sets are subjected to bird's-eye view feature synthesis to obtain a target fusion feature map.

[0025] Optionally, in the sixth implementation manner of the first aspect of the present application, the 3D feature map corresponding to each frame of image in each image set is subjected to feature superposition to obtain the initial fusion feature map corresponding to each image set, including:

[0026] The 3D feature map corresponding to each frame of image in each image set is subjected to 3D point alignment transformation according to the vehicle pose information at the time of image acquisition to obtain an aligned feature map corresponding to each frame of image in each image set;

[0027] The aligned feature map corresponding to each frame of image in each image set is subjected to 3D point-by-3D point feature superposition to obtain the initial fusion feature map corresponding to each image set.

[0028] Optionally, in the seventh implementation manner of the first aspect of the present application, the initial fusion feature maps corresponding to all the image sets are subjected to bird's-eye view feature synthesis to obtain a target fusion feature map, including:

[0029] The initial fusion feature maps corresponding to all the image sets are subjected to the same 3D point detection to obtain bird's-eye view splicing position information;

[0030] The initial fusion feature maps corresponding to all the image sets are subjected to feature superposition and splicing of the same 3D points according to the bird's-eye view splicing position information to obtain a target fusion feature map.

[0031] Optionally, in the eighth implementation manner of the first aspect of the present application, the 3D target detection is performed on each frame of image in the multiple image sets to obtain multiple target object candidate boxes of each frame of image, including:

[0032] The 2D detection box generation and 3D detection box regression are performed on each frame of image in the multiple image sets by a preset monocular 3D target detection algorithm to obtain multiple target object candidate boxes of each frame of image.

[0033] Optionally, in the ninth implementation manner of the first aspect of the present application, the fusion feature information corresponding to each target object candidate box of each frame of image is extracted from the target fusion feature map, and all the target object candidate boxes are subjected to screening according to the fusion feature information corresponding to each target object candidate box of each frame of image to obtain at least one target object detection box, including:

[0034] The fusion feature information corresponding to each target object candidate box of each frame of image is extracted from the target fusion feature map according to the 3D space coordinate information corresponding to each target object candidate box of each frame of image;

[0035] The target object information corresponding to each target object candidate box of each image frame is predicted through a preset self-attention mechanism, so as to obtain the target object information corresponding to each target object candidate box of each image frame.

[0036] The target object screening is performed on all target object candidate boxes according to the target object information corresponding to each target object candidate box of each image frame, so as to obtain at least one target object detection box.

[0037] The second aspect of the present application provides a target object detection device, comprising:

[0038] The detection module is configured to perform 3D target detection on each image frame in the plurality of image sets to obtain a plurality of target object candidate boxes of each image frame, one image set corresponding to one camera, and each image set comprising a plurality of image frames collected by the camera.

[0039] The extraction module is configured to perform 3D spatial feature extraction on each image frame in the plurality of image sets to obtain a 3D feature map corresponding to each image frame.

[0040] The fusion module is configured to perform feature fusion on the 3D feature map corresponding to each image frame to obtain a target fusion feature map.

[0041] The screening module is configured to extract fusion feature information corresponding to each target object candidate box of each image frame from the target fusion feature map, and perform screening on all target object candidate boxes according to the fusion feature information corresponding to each target object candidate box of each image frame to obtain at least one target object detection box.

[0042] Optionally, in the first implementation manner of the second aspect of the present application, the extraction module comprises:

[0043] The conversion unit is configured to perform 3D spatial conversion on each image frame in the plurality of image sets to obtain a 3D spatial map corresponding to each image frame.

[0044] The projection unit is configured to obtain target feature information corresponding to each image frame, and project the target feature information corresponding to each image frame to the 3D spatial map corresponding to each image frame to obtain a 3D feature map corresponding to each image frame.

[0045] Optionally, in the second implementation manner of the second aspect of the present application, the conversion unit is specifically configured to:

[0046] Perform pixel-by-pixel depth estimation on each image frame in the plurality of image sets to obtain a 3D spatial map corresponding to each image frame, each 3D point in the 3D spatial map corresponding to each image frame corresponding to a 3D spatial coordinate information.

[0047] Optionally, in the third implementation manner of the second aspect of the present application, the projection unit is specifically configured to:

[0048] read the target feature information corresponding to each frame of image, the target feature information corresponding to each frame of image includes at least one of laser radar feature information, millimeter wave radar feature information, ultrasonic wave feature information and image feature information of each frame of image;

[0049] obtain feature coordinate information, the feature coordinate information is used for indicating coordinate information of target feature information corresponding to each frame of image in the corresponding frame of image;

[0050] according to the feature coordinate information, the target feature information corresponding to each frame of image is mapped to the corresponding 3D space graph, and 3D feature map corresponding to each frame of image is obtained.

[0051] optionally, in the fourth implementation manner of the second aspect of the present application, the fusion module comprises:

[0052] the first synthesis unit is configured to perform bird's eye view feature synthesis on the 3D feature maps corresponding to the same frame of image in all image sets, and obtain the bird's eye view feature map corresponding to the same sequence frame of image;

[0053] the first superposition unit is configured to perform feature superposition on the bird's eye view feature map corresponding to the same sequence frame of image, and obtain the target fusion feature map.

[0054] optionally, in the fifth implementation manner of the second aspect of the present application, the fusion module further comprises:

[0055] the second superposition unit is configured to perform feature superposition on the 3D feature maps corresponding to each frame of image in each image set, and obtain the initial fusion feature map corresponding to each image set;

[0056] the second synthesis unit is configured to perform bird's eye view feature synthesis on the initial fusion feature maps corresponding to all image sets, and obtain the target fusion feature map.

[0057] optionally, in the sixth implementation manner of the second aspect of the present application, the second superposition unit is specifically configured to:

[0058] perform 3D point alignment transformation on the 3D feature maps corresponding to each frame of image in each image set according to the vehicle pose information when each frame of image is collected, and obtain the alignment feature map corresponding to each frame of image in each image set;

[0059] perform 3D point feature superposition on the alignment feature maps corresponding to each frame of image in each image set, and obtain the initial fusion feature map corresponding to each image set.

[0060] optionally, in the seventh implementation manner of the second aspect of the present application, the second synthesis unit is specifically configured to:

[0061] The initial fusion feature maps corresponding to each image set are subjected to the same 3D point detection to obtain bird's-eye view splicing position information.

[0062] According to the bird's-eye view splicing position information, the initial fusion feature maps corresponding to each image set are subjected to feature superposition and splicing of the same 3D points to obtain a target fusion feature map.

[0063] Optionally, in an eighth implementation manner of the second aspect of the present application, the detection module is specifically configured to:

[0064] The preset monocular 3D target detection algorithm is used to generate a 2D detection frame and regress a 3D detection frame for each image in the multiple image sets, so as to obtain multiple target object candidate frames of each image.

[0065] Optionally, in a ninth implementation manner of the second aspect of the present application, the screening module is specifically configured to:

[0066] According to the 3D space coordinate information corresponding to each target object candidate frame of each image, the fusion feature information corresponding to each target object candidate frame of each image is extracted from the target fusion feature map;

[0067] The preset self-attention mechanism is used to predict target object information of the fusion feature information corresponding to each target object candidate frame of each image, so as to obtain the target object information corresponding to each target object candidate frame of each image;

[0068] According to the target object information corresponding to each target object candidate frame of each image, target object screening is performed on all target object candidate frames, so as to obtain at least one target object detection frame.

[0069] The third aspect of the present application provides a target object detection device, which comprises a memory and at least one processor, the memory stores a computer program, and the at least one processor invokes the computer program in the memory to enable the target object detection device to execute the target object detection method described above.

[0070] The fourth aspect of the present application provides a computer readable storage medium, which stores a computer program, and when the computer program is run on a computer, the computer is enabled to execute the target object detection method described above.

[0071] In the technical solution provided by this invention, 3D target detection is performed on each frame of images in multiple image sets to obtain multiple target object candidate boxes for each frame of images. One image set corresponds to one camera, and each image set includes multiple frames of images captured by the camera. 3D spatial features are extracted from each frame of images in the multiple image sets to obtain a 3D feature map corresponding to each frame of images. Feature fusion is performed on the 3D feature maps corresponding to each frame of images to obtain a target fusion feature map. Fusion feature information corresponding to each target object candidate box of each frame of images is extracted from the target fusion feature map, and all target object candidate boxes are filtered according to the fusion feature information corresponding to each target object candidate box of each frame of images to obtain at least one target object detection box. In this embodiment of the invention, in order to improve the accuracy of target object detection, multiple target object candidate boxes are identified in each frame of an image set acquired by multiple cameras. Since there may be incomplete or overlapping target object detection boxes in the multiple target object candidate boxes of each frame, in order to accurately filter out complete and non-overlapping target object detection boxes from the target object candidate boxes, after extracting the 3D feature map corresponding to each frame image, all 3D feature maps are fused to obtain a target fusion feature map. Then, the fusion feature information corresponding to each target object candidate box is extracted from the target fusion feature map. Finally, the target object candidate boxes are filtered using the fusion feature information to obtain at least one target object candidate box. This invention uses the fusion features of multi-camera and multi-frame images to filter target objects, which can improve the accuracy of cross-camera target object detection. Attached Figure Description

[0072] Figure 1 This is a schematic diagram of one embodiment of the target object detection method in this invention;

[0073] Figure 2 This is a schematic diagram of one embodiment of the target object detection device in this invention;

[0074] Figure 3 This is a schematic diagram of another embodiment of the target object detection device in this invention;

[0075] Figure 4 This is a schematic diagram of one embodiment of the target object detection device in this invention. Detailed Implementation

[0076] This invention provides a method, apparatus, device, and storage medium for detecting target objects, which can improve the accuracy of target object detection across cameras.

[0077] The terms "first", "second", "third", "fourth" and the like in the description and in the claims of the present application, and above-mentioned drawings, if any, are used to distinguish between similar objects and not necessarily for describing a specific sequential or chronological order. It is to be understood that the use of these terms here is not meant to limit the scope of the application to the precise embodiments disclosed but rather the scope of the application encompasses all embodiments that are within the scope of the claims. Furthermore, the terms "comprising" or "including" and any of their derivatives, are intended to be construed as encompassing not only the listed items but also any additional steps or elements that could be added to the process, method, system, product or apparatus without departing from the scope of the application.

[0078] It can be understood that the execution subject of the present application can be a target object detection device, and can also be a terminal or a server, which is not limited here. The embodiments of the present application take the server as the execution subject for example.

[0079] For the convenience of understanding, the specific process of the embodiments of the present application is described below. Please refer to Figure 1 One embodiment of the target object detection method in the embodiments of the present application includes:

[0080] 101. 3D target detection is performed on each frame of image in a plurality of image sets to obtain a plurality of target object candidate boxes of each frame of image, one image set corresponds to one camera, and each image set includes a plurality of frames of images collected by the camera;

[0081] It can be understood that, in order to improve the completeness of target object observation, a plurality of cameras are arranged on the autonomous vehicle in advance, which are used to collect environment images in different perspectives, each camera collects a plurality of frames of images within 1 second to obtain an image set corresponding to each camera, for example, the front direction of the vehicle is the front direction, one camera is arranged at the left front, right front, left rear and right rear of the vehicle respectively, assuming that each camera collects 25 frames of images within 1 second, then the image set A corresponding to the left front camera includes 25 frames of images collected by the left front camera at the current time, the image set B corresponding to the right front camera includes 25 frames of images collected by the right front camera at the current time, and so on. Each camera corresponding image set contains a plurality of frames of images collected by the corresponding camera at the same time, which is used for feature fusion of multiple camera multiple frames of images, so as to improve the accuracy of target object detection.

[0082] In an embodiment, to improve the accuracy of target detection, step 101 comprises: performing 2D detection box generation and 3D detection box regression on each frame of image in the plurality of image sets by a preset monocular 3D target detection algorithm to obtain a plurality of target candidate boxes of each frame of image, one image set corresponding to one camera, and each image set comprising a plurality of frames of image captured by the camera. The monocular 3D target detection algorithm includes but is not limited to monocular 3D detection algorithms such as single-stage monocular 3D detection algorithm and two-stage monocular 3D detection algorithm. In another embodiment, before performing 3D target detection on each frame of image in the plurality of image sets by the preset monocular 3D target detection algorithm, it further comprises: performing multi-scale feature extraction on each frame of image in the plurality of image sets by a feature pyramid to obtain image feature information of each frame of image. Then performing 3D target detection on the image feature information of each frame of image by the preset monocular 3D target detection algorithm to obtain a plurality of target candidate boxes of each frame of image. It should be noted that the target candidate box is the smallest circumscribed 3D rectangular detection box of the target, and each target candidate box in each frame of image comprises 3D spatial coordinate information, size information, rotation information, category information, etc. of the target. This embodiment can improve the accuracy of target candidate box detection, and further improve the accuracy of target detection.

[0083] 102. performing 3D spatial feature extraction on each frame of image in the plurality of image sets to obtain a 3D feature map corresponding to each frame of image;

[0084] It should be noted that, since it is difficult for the monocular 3D target detection algorithm to fuse the image feature information of multiple cameras and multiple frames, the target candidate boxes obtained by performing 3D target detection on a single frame of image have a large amount of noise data, i.e. there may be overlapping or incomplete target detection boxes in all target candidate boxes. To accurately eliminate the noise data in the target candidate boxes and obtain non-overlapping and complete target detection boxes, 3D spatial feature extraction is performed on each frame of image in the plurality of image sets to obtain a 3D feature map corresponding to each frame of image, and the 3D feature maps corresponding to each frame of image are fused to obtain a target fusion feature map, which comprises feature information of a plurality of complete vehicle environment observation images and is used for screening the target candidate boxes to obtain accurate target detection boxes, so as to improve the accuracy of target detection.

[0085] In an embodiment, the step of performing 3D spatial feature extraction on each frame of image in the plurality of image sets to obtain a 3D feature map corresponding to each frame of image comprises: obtaining target feature information corresponding to each frame of image, and projecting the target feature information corresponding to each frame of image to a 3D space to obtain a 3D feature map corresponding to each frame of image. In another embodiment, the step of performing 3D spatial feature extraction on each frame of image in the plurality of image sets to obtain a 3D feature map corresponding to each frame of image further comprises: performing 3D spatial conversion on each frame of image in the plurality of image sets to obtain a 3D spatial map corresponding to each frame of image; obtaining target feature information corresponding to each frame of image, and projecting the target feature information corresponding to each frame of image to the 3D spatial map corresponding to each frame of image to obtain a 3D feature map corresponding to each frame of image. The target feature information can be 2D feature information or 3D feature information, and the order of 3D spatial conversion (projection) of the image or the feature is not limited. The target feature information is 3D feature information or not depends on the target feature information. The embodiment can flexibly obtain 2D or 3D feature information, so that the subsequent target fusion feature map contains multi-dimensional feature information, thereby improving the accuracy of target object candidate box screening and the accuracy of target object detection.

[0086] Based on the above, in order to convert each frame of image to a 3D space, the step of performing 3D spatial conversion on each frame of image in the plurality of image sets to obtain a 3D spatial map corresponding to each frame of image comprises: performing pixel-by-pixel depth estimation on each frame of image in the plurality of image sets to obtain a 3D spatial map corresponding to each frame of image, each 3D point in the 3D spatial map corresponding to each frame of image corresponding to a 3D spatial coordinate information. Specifically, the pixel-by-pixel depth estimation on each frame of image in the plurality of image sets is performed by a monocular depth estimation model to obtain a 3D spatial map corresponding to each frame of image. In addition to the 3D spatial conversion of the image by depth estimation, in another embodiment, the step of performing 3D spatial conversion on each frame of image in the plurality of image sets to obtain a 3D spatial map corresponding to each frame of image further comprises: obtaining pixel values of each pixel point in each frame of image in the plurality of image sets, and performing pixel point correlation prediction on each frame of image in the plurality of image sets according to the pixel values of each pixel point to obtain a prediction result, and converting each frame of image in the plurality of image sets to a 3D space according to the prediction result to obtain a 3D spatial map corresponding to each frame of image. The embodiment can quickly convert a 2D image to a 3D image, thereby improving the efficiency of target object detection.

[0087] Based on the above, in order to fuse more feature information to improve the accuracy of target detection, the execution steps of acquiring target feature information corresponding to each frame of image and projecting the target feature information corresponding to each frame of image onto the 3D space map corresponding to each frame of image to obtain the 3D feature map corresponding to each frame of image include: reading the target feature information corresponding to each frame of image, which includes, but is not limited to, at least one of the following: lidar feature information, millimeter-wave radar feature information, ultrasonic feature information, and image feature information of each frame of image; acquiring feature coordinate information, which is used to indicate the coordinate information of the target feature information corresponding to each frame of image in the corresponding frame image; and mapping the target feature information corresponding to each frame of image to the corresponding 3D space map according to the feature coordinate information to obtain the 3D feature map corresponding to each frame of image. It is understood that the target feature information corresponding to each frame of image includes feature information from multiple sensors, such as LiDAR, millimeter-wave radar, ultrasound, and cameras. Therefore, the target feature information corresponding to each frame of image includes at least one of the following: LiDAR feature information, millimeter-wave radar feature information, ultrasound feature information, and image feature information. The position information corresponding to the target feature information of each frame of image is then converted into the coordinate information of the corresponding frame image to obtain feature coordinate information. Finally, based on the feature coordinate information, all target feature information corresponding to each frame of image is projected onto the corresponding 3D spatial map to obtain the 3D feature map corresponding to each frame of image. This implementation can acquire feature information from multiple sensors for environmental detection, enabling the subsequent fused feature map to contain more comprehensive feature information, thereby improving the accuracy of target detection.

[0088] Based on the above, the image feature information in the target feature information includes semantic segmentation information for each pixel in the corresponding frame image, such as semantic segmentation information based on binary classification ("obstacle-non-obstacle"), or semantic segmentation information based on multi-class classification ("person-vehicle-bicycle-static object-animal-road surface-sky-plant-other"), etc., without specific limitations here. This embodiment can obtain image feature information by performing semantic segmentation on the image, which can be used to improve the accuracy of subsequent target candidate box selection, thereby improving the accuracy of target detection.

[0089] 103. Perform feature fusion on the 3D feature maps corresponding to each frame of the image to obtain the target fused feature map;

[0090] In an implementation, since each frame of image in the plurality of image sets is captured within 1 second, the similarity of each frame of image in each image set is high, that is, there are more same pixel points between each frame of image in each image set, then the same pixel points of the 3D feature maps corresponding to each frame of image in the same image set are fused to obtain an initial fusion feature map corresponding to each image set, and then the initial fusion feature maps corresponding to each image set are combined into a feature map of a panoramic view to obtain a target fusion feature map, the target fusion feature map contains feature information of multiple cameras, multiple frames and multiple sensors, so that the accuracy of subsequent target object candidate box screening through the target fusion feature map is improved, and the accuracy of target object detection is improved.

[0091] In an implementation, after obtaining the target fusion feature map, the method further includes: performing fusion feature extraction on the target fusion feature map through a preset convolutional neural network model to obtain fusion feature information in the target fusion feature map, which is used for subsequent target object candidate box screening, and can further improve the accuracy of target object detection.

[0092] By way of example but not limitation, in the feature fusion process of the 3D feature map, a same-camera image feature superposition step and a cross-camera image synthesis step are included, and the order of the two steps can be reversed, which is not specifically limited here. In an embodiment, the cross-camera image synthesis step is performed first, and then the same-camera image feature superposition step is performed, i.e., step 103 includes: performing bird's eye view feature synthesis on the 3D feature maps corresponding to the same frame image in all image sets to obtain bird's eye view feature maps corresponding to the same sequence frame image; and performing feature superposition on the bird's eye view feature maps corresponding to the same sequence frame image to obtain a target fusion feature map. For example, assuming that two monocular cameras 1 and 2 with different viewing angles are provided on an autonomous vehicle, the monocular camera 1 corresponds to an image set A, and the monocular camera 2 corresponds to an image set B. The image set A includes three frames of images collected by the monocular camera 1, and the 3D feature maps corresponding to the three frames of images are feature map a1, feature map a2, and feature map a3, respectively. The image set B includes three frames of images collected by the monocular camera 2, and the 3D feature maps corresponding to the three frames of images are feature map b1, feature map b2, and feature map b3, respectively. In this embodiment, first, the 3D feature maps corresponding to the same frame image in all image sets are subjected to bird's eye view feature synthesis to obtain bird's eye view feature maps corresponding to the same sequence frame image, i.e., the feature map a1 corresponding to the first frame image in the image set A and the feature map b1 corresponding to the first frame image in the image set B are subjected to bird's eye view feature synthesis to obtain a bird's eye view feature map X corresponding to the first frame image. Then, the feature map a2 corresponding to the second frame image in the image set A and the feature map b2 corresponding to the second frame image in the image set B are subjected to bird's eye view feature synthesis to obtain a bird's eye view feature map Y corresponding to the second frame image. Finally, the feature map a3 corresponding to the third frame image in the image set A and the feature map b3 corresponding to the third frame image in the image set B are subjected to bird's eye view feature synthesis to obtain a bird's eye view feature map Z corresponding to the third frame image. Next, the bird's eye view feature maps corresponding to the same sequence frame image are subjected to feature superposition to obtain a target fusion feature map, i.e., the bird's eye view feature map X corresponding to the first frame image, the bird's eye view feature map Y corresponding to the second frame image, and the bird's eye view feature map Z corresponding to the third frame image are subjected to feature superposition to obtain a target fusion feature map. This embodiment can fuse multi-camera and multi-frame feature information, so that the subsequent target object candidate box screening is more accurate, thereby improving the accuracy of target object detection.

[0093] Based on the above, specifically, the execution steps of synthesizing the 3D feature maps corresponding to the same frame image in all image sets into the bird's eye view feature maps corresponding to the same sequence frame image include: performing the same 3D point detection on the 3D feature maps corresponding to the same frame image in all image sets to obtain the bird's eye view splicing position information corresponding to each sequence frame image; and performing the feature superposition and splicing of the same 3D points on the 3D feature maps corresponding to the same frame image in all image sets according to the bird's eye view splicing position information corresponding to each sequence frame image to obtain the bird's eye view feature maps corresponding to the same sequence frame image. For example, based on the above example, the same 3D point detection is performed on the feature map a1 corresponding to the first frame image in the image set A and the feature map b1 corresponding to the first frame image in the image set B to obtain the bird's eye view splicing position information corresponding to the first frame image, the same 3D point detection is performed on the feature map a2 corresponding to the second frame image in the image set A and the feature map b2 corresponding to the second frame image in the image set B to obtain the bird's eye view splicing position information corresponding to the second frame image, and the same 3D point detection is performed on the feature map a3 corresponding to the third frame image in the image set A and the feature map b3 corresponding to the third frame image in the image set B to obtain the bird's eye view splicing position information corresponding to the third frame image. Then, the feature superposition and splicing of the same 3D points are performed on the feature map a1 corresponding to the first frame image in the image set A and the feature map b1 corresponding to the first frame image in the image set B according to the bird's eye view splicing position information corresponding to the first frame image to obtain the bird's eye view feature map X corresponding to the first frame image, the feature superposition and splicing of the same 3D points are performed on the feature map a2 corresponding to the second frame image in the image set A and the feature map b2 corresponding to the second frame image in the image set B according to the bird's eye view splicing position information corresponding to the second frame image to obtain the bird's eye view feature map Y corresponding to the second frame image, and the feature superposition and splicing of the same 3D points are performed on the feature map a3 corresponding to the third frame image in the image set A and the feature map b3 corresponding to the third frame image in the image set B according to the bird's eye view splicing position information corresponding to the third frame image to obtain the bird's eye view feature map Z corresponding to the third frame image.

[0094] Based on the above, the execution steps of obtaining the target fusion feature map by performing feature superposition on the bird's eye view feature maps corresponding to the same sequence frame images include: performing 3D point alignment transformation on the bird's eye view feature maps corresponding to the same sequence frame images according to the vehicle pose information at the time of collecting each frame image, to obtain the alignment feature maps corresponding to each sequence frame image; and performing feature superposition on the alignment feature maps corresponding to each sequence frame image by 3D point, to obtain the target fusion feature map. For example, based on the above example, the bird's eye view feature map X corresponding to the first frame image, the bird's eye view feature map Y corresponding to the second frame image, and the bird's eye view feature map Z corresponding to the third frame image are subjected to 3D point alignment transformation, to obtain the alignment feature map X' corresponding to the first frame image, the alignment feature map Y' corresponding to the second frame image, and the alignment feature map Z' corresponding to the third frame image. Finally, the alignment feature map X', the alignment feature map Y', and the alignment feature map Z' are subjected to feature superposition by 3D point, to obtain the target fusion feature map.

[0095] Based on the above, the same camera image feature superposition step can be performed first, and then the cross-camera image synthesis step is performed, that is, step 103 further includes: performing feature superposition on the 3D feature maps corresponding to each frame image in each image set, to obtain the initial fusion feature maps corresponding to each image set; and performing bird's eye view feature synthesis on the initial fusion feature maps corresponding to all image sets, to obtain the target fusion feature map. For example, based on the above example, first, the 3D feature maps corresponding to each frame image in each image set are subjected to feature superposition, to obtain the initial fusion feature maps corresponding to each image set, that is, the feature map a1, the feature map a2, and the feature map a3 are subjected to feature superposition, to obtain the initial fusion feature map M corresponding to the image set A, and then the feature map b1, the feature map b2, and the feature map b3 are subjected to feature superposition, to obtain the initial fusion feature map N corresponding to the image set B. Then, the initial fusion feature maps corresponding to all image sets are subjected to bird's eye view feature synthesis, to obtain the target fusion feature map, that is, the initial fusion feature map M corresponding to the image set A and the initial fusion feature map N corresponding to the image set B are subjected to bird's eye view feature synthesis, to obtain the target fusion feature map. The present embodiment can fuse multi-camera and multi-frame feature information, so that the subsequent target object candidate box screening is more accurate, thereby improving the accuracy of target object detection.

[0096] Based on the above, specifically, the execution steps of performing feature superposition on the 3D feature maps corresponding to each frame of image in each image set to obtain the initial fusion feature maps corresponding to each image set include: performing 3D point alignment transformation on the 3D feature maps corresponding to each frame of image in each image set according to the vehicle pose information at the time of collecting each frame of image, to obtain the aligned feature maps corresponding to each frame of image in each image set; and performing 3D point-by-point feature superposition on the aligned feature maps corresponding to each frame of image in each image set, to obtain the initial fusion feature maps corresponding to each image set. For example, based on the above example, first, perform 3D point alignment transformation on the feature map a1, the feature map a2 and the feature map a3 corresponding to each frame of image in the image set A according to the vehicle pose information at the time of collecting each frame of image, to obtain the aligned feature map a1', the aligned feature map a2' and the aligned feature map a3' corresponding to each frame of image in the image set A, and perform 3D point alignment transformation on the feature map b1, the feature map b2 and the feature map b3 corresponding to each frame of image in the image set B, to obtain the aligned feature map b1', the aligned feature map b2' and the aligned feature map b3' corresponding to each frame of image in the image set B; and then perform 3D point-by-point feature superposition on the aligned feature map a1', the aligned feature map a2' and the aligned feature map a3' corresponding to each frame of image in the image set A, to obtain the initial fusion feature map M corresponding to the image set A, and perform 3D point-by-point feature superposition on the aligned feature map b1', the aligned feature map b2' and the aligned feature map b3' corresponding to each frame of image in the image set B, to obtain the initial fusion feature map N corresponding to the image set B.

[0097] Based on the above, specifically, the execution steps of performing bird's-eye view feature synthesis on the initial fusion feature maps corresponding to all image sets to obtain the target fusion feature map include: performing same 3D point detection on the initial fusion feature maps corresponding to each image set to obtain bird's-eye view splicing position information; and performing feature superposition and splicing on the same 3D points of the initial fusion feature maps corresponding to each image set according to the bird's-eye view splicing position information, to obtain the target fusion feature map. For example, based on the above example, perform same 3D point detection on the initial fusion feature map M corresponding to the image set A and the initial fusion feature map N corresponding to the image set B, to obtain bird's-eye view splicing position information, and then perform feature superposition and splicing on the same 3D points of the initial fusion feature map M corresponding to the image set A and the initial fusion feature map N corresponding to the image set B according to the bird's-eye view splicing position information, to obtain the target fusion feature map.

[0098] 104、From the target fusion feature map, extract fusion feature information corresponding to each target object candidate box of each frame of image, and according to the fusion feature information corresponding to each target object candidate box of each frame of image, screen all target object candidate boxes to obtain at least one target object detection box.

[0099] It should be noted that, since the target fusion feature map contains feature information of multiple cameras and multiple frames of images, the target fusion feature map contains fusion feature information of all target object bounding boxes. The fusion feature information corresponding to each target object bounding box of each frame of image is extracted from the target fusion feature map, and all target object bounding boxes are screened according to the fusion feature information corresponding to each target object bounding box of each frame of image, to obtain at least one target object detection box. The target object detection box is a target object detection box meeting a preset condition. As an example but not limitation, the target object detection box can be a target object detection box of an obstacle type such as a pedestrian, a roadblock, a car, etc., can be a target object detection box with a distance from the current autonomous vehicle less than a preset threshold, can be a target object detection box of a non-crossable type, etc., which is not limited here. The embodiment can accurately screen the target object bounding box based on the feature information of multiple cameras, multiple frames and multiple sensors, thereby improving the accuracy of target object detection.

[0100] In an embodiment, step 104 includes: extracting, from the target fusion feature map, fusion feature information corresponding to each target object bounding box of each frame of image according to 3D spatial coordinate information corresponding to each target object bounding box of each frame of image; performing target object information prediction on the fusion feature information corresponding to each target object bounding box of each frame of image by a preset self-attention mechanism, to obtain target object information corresponding to each target object bounding box of each frame of image; and performing target object screening on all target object bounding boxes according to the target object information corresponding to each target object bounding box of each frame of image, to obtain at least one target object detection box. In the embodiment, after the fusion feature information corresponding to each target object bounding box of each frame of image is extracted from the target fusion feature map according to the 3D spatial coordinate information corresponding to each target object bounding box of each frame of image, the target object information corresponding to each target object bounding box of each frame of image is obtained by performing target object information prediction on the fusion feature information corresponding to each target object bounding box of each frame of image by a preset self-attention mechanism. Specifically, the correlation between each target object bounding box of each frame of image and other target object bounding boxes is calculated by an inner product algorithm in the preset self-attention mechanism, to obtain cross feature information corresponding to each target object bounding box of each frame of image, which contains the features of each other target object bounding box. Then, the target object information corresponding to each target object bounding box of each frame of image is predicted according to the cross feature information corresponding to each target object bounding box of each frame of image, so that the accuracy of target object information prediction is improved, and the accuracy of target object detection is further improved. The target object information includes but is not limited to existence information, category information, geometric information and position information of the target object.

[0101] In the embodiment of the present application, in order to improve the accuracy of target detection, a plurality of target candidate boxes of each frame of image in the image set collected by the plurality of cameras are identified. Since there may be incomplete or overlapping target detection boxes in the plurality of target candidate boxes of each frame of image, in order to accurately screen out complete and non-overlapping target detection boxes from the target candidate boxes, after extracting the 3D feature map corresponding to each frame of image, all 3D feature maps are fused to obtain a target fusion feature map, and then the fusion feature information corresponding to each target candidate box is extracted from the target fusion feature map, and finally the target candidate boxes are screened according to the fusion feature information to obtain at least one target detection box. The present application can improve the accuracy of cross-camera target detection.

[0102] The above describes the target detection method in the embodiment of the present application. The target detection device in the embodiment of the present application is described below. Please refer to Figure 2 An embodiment of the target detection device in the embodiment of the present application includes:

[0103] The detection module 201 is configured to perform 3D target detection on each frame of image in the plurality of image sets to obtain a plurality of target candidate boxes of each frame of image. One image set corresponds to one camera, and each image set includes a plurality of frames of image collected by the camera.

[0104] The extraction module 202 is configured to perform 3D spatial feature extraction on each frame of image in the plurality of image sets to obtain a 3D feature map corresponding to each frame of image.

[0105] The fusion module 203 is configured to fuse the 3D feature maps corresponding to each frame of image to obtain a target fusion feature map.

[0106] The screening module 204 is configured to extract fusion feature information corresponding to each target candidate box of each frame of image from the target fusion feature map, and screen all target candidate boxes according to the fusion feature information corresponding to each target candidate box of each frame of image to obtain at least one target detection box.

[0107] In the embodiment of the present application, in order to improve the accuracy of target detection, multiple target candidate boxes of each frame of image in the image set collected by the plurality of cameras are recognized. Since there may be incomplete or overlapping target detection boxes in the multiple target candidate boxes of each frame of image, in order to accurately screen out complete and non-overlapping target detection boxes from the target candidate boxes, after extracting the 3D feature map corresponding to each frame of image, the feature fusion of all 3D feature maps is performed to obtain a target fusion feature map, and then the fusion feature information corresponding to each target candidate box is extracted from the target fusion feature map, and finally the target candidate boxes are screened through the fusion feature information to obtain at least one target candidate box. The present application screens the target based on the fusion features of the multi-camera multi-frame images, which can improve the accuracy of cross-camera target detection.

[0108] Please refer to Figure 3 Another embodiment of the target detection device in the embodiment of the present application includes:

[0109] The detection module 201 is configured to perform 3D target detection on each frame of image in the plurality of image sets to obtain multiple target candidate boxes of each frame of image. One image set corresponds to one camera, and each image set includes multiple frames of image collected by the camera.

[0110] The extraction module 202 is configured to perform 3D spatial feature extraction on each frame of image in the plurality of image sets to obtain a 3D feature map corresponding to each frame of image.

[0111] The fusion module 203 is configured to perform feature fusion on the 3D feature map corresponding to each frame of image to obtain a target fusion feature map.

[0112] The screening module 204 is configured to extract fusion feature information corresponding to each target candidate box of each frame of image from the target fusion feature map, and screen all target candidate boxes according to the fusion feature information corresponding to each target candidate box of each frame of image to obtain at least one target detection box.

[0113] Optionally, the extraction module 202 includes:

[0114] The conversion unit 2021 is configured to perform 3D spatial conversion on each frame of image in the plurality of image sets to obtain a 3D space map corresponding to each frame of image.

[0115] The projection unit 2022 is configured to obtain target feature information corresponding to each frame of image, and project the target feature information corresponding to each frame of image to the 3D space map corresponding to each frame of image to obtain a 3D feature map corresponding to each frame of image.

[0116] Optionally, the conversion unit 2021 is specifically configured to:

[0117] The depth estimation of each frame of image in the plurality of image sets is performed pixel by pixel, and a 3D space map corresponding to each frame of image is obtained, wherein each 3D point in the 3D space map corresponding to each frame of image corresponds to 3D space coordinate information.

[0118] Optionally, the projection unit 2022 is specifically used for:

[0119] The target feature information corresponding to each frame of image is read, and the target feature information corresponding to each frame of image includes at least one of laser radar feature information, millimeter wave radar feature information, ultrasonic wave feature information and image feature information of each frame of image.

[0120] The feature coordinate information is obtained, and the feature coordinate information is used to indicate the coordinate information of the target feature information corresponding to each frame of image in the corresponding frame of image.

[0121] According to the feature coordinate information, the target feature information corresponding to each frame of image is mapped to the corresponding 3D space map, and a 3D feature map corresponding to each frame of image is obtained.

[0122] Optionally, the fusion module 203 includes:

[0123] The first synthesis unit 2031 is configured to perform bird's eye view feature synthesis on the 3D feature maps corresponding to the same frame of image in all image sets, and obtain a bird's eye view feature map corresponding to the same sequence frame of image.

[0124] The first superposition unit 2032 is configured to perform feature superposition on the bird's eye view feature maps corresponding to the same sequence frame of image, and obtain a target fusion feature map.

[0125] Optionally, the fusion module 203 further includes:

[0126] The second superposition unit 2033 is configured to perform feature superposition on the 3D feature maps corresponding to each frame of image in each image set, and obtain an initial fusion feature map corresponding to each image set.

[0127] The second synthesis unit 2034 is configured to perform bird's eye view feature synthesis on the initial fusion feature maps corresponding to all image sets, and obtain a target fusion feature map.

[0128] Optionally, the second superposition unit 2033 is specifically used for:

[0129] According to the vehicle pose information when each frame of image is collected, 3D point alignment transformation is performed on the 3D feature maps corresponding to each frame of image in each image set, and an aligned feature map corresponding to each frame of image in each image set is obtained.

[0130] The aligned feature maps corresponding to each frame of image in each image set are superimposed on each 3D point, and an initial fusion feature map corresponding to each image set is obtained.

[0131] Optionally, the second synthesis unit 2034 is specifically used for:

[0132] The same 3D point detection is performed on the initial fusion feature maps corresponding to each image set to obtain aerial view splicing position information.

[0133] According to the aerial view splicing position information, the feature superposition and splicing of the same 3D points are performed on the initial fusion feature maps corresponding to each image set to obtain a target fusion feature map.

[0134] Optionally, the detection module 201 is specifically used for:

[0135] The 2D detection frame generation and 3D detection frame regression are performed on each frame of image in the multiple image sets by using a preset monocular 3D target detection algorithm to obtain multiple target object candidate frames of each frame of image.

[0136] Optionally, the screening module 204 is specifically used for:

[0137] The fusion feature information corresponding to each target object candidate frame of each frame of image is extracted from the target fusion feature map according to the 3D space coordinate information corresponding to each target object candidate frame of each frame of image.

[0138] The target object information corresponding to each target object candidate frame of each frame of image is predicted by using a preset self-attention mechanism on the fusion feature information corresponding to each target object candidate frame of each frame of image to obtain the target object information corresponding to each target object candidate frame of each frame of image.

[0139] The target object screening is performed on all target object candidate frames according to the target object information corresponding to each target object candidate frame of each frame of image to obtain at least one target object detection frame.

[0140] In the embodiment of the application, in order to improve the accuracy of target object detection, multiple target object candidate frames of each frame of image in the image set collected by multiple cameras are identified. Since there may be incomplete or overlapping target object detection frames in the multiple target object candidate frames of each frame of image, in order to accurately screen out complete and non-overlapping target object detection frames from the target object candidate frames, after the 3D feature map corresponding to each frame of image is extracted, the feature fusion of all 3D feature maps is performed to obtain a target fusion feature map, the fusion feature information corresponding to each target object candidate frame is extracted from the target fusion feature map, and finally the target object candidate frames are screened by using the fusion feature information to obtain at least one target object candidate frame. The target object screening based on the fusion features of the multiple camera multiple frame images can improve the accuracy of cross-camera target object detection.

[0141] The above Figure 2 and Figure 3The target object detection device in the embodiments of the present application is described in detail from the perspective of a modular functional entity, and the target object detection device in the embodiments of the present application is described in detail from the perspective of hardware processing.

[0142] Figure 4 Fig. 4 is a structural schematic diagram of a target object detection device provided by the embodiments of the present application. The target object detection device 400 can be different in configuration or performance and can include one or more processors (central processing units, CPU) 410 (for example, one or more processors) and a memory 420, and one or more storage media 430 (for example, one or more mass storage devices) storing application programs 433 or data 432. The memory 420 and the storage media 430 can be temporary storage or persistent storage. The programs stored in the storage media 430 can include one or more modules (not shown in the figure), and each module can include a series of computer program operations in the target object detection device 400. Further, the processor 410 can be configured to communicate with the storage media 430 and execute a series of computer program operations in the storage media 430 on the target object detection device 400.

[0143] The target object detection device 400 can further include one or more power supplies 440, one or more wired or wireless network interfaces 450, one or more input / output interfaces 460, and / or one or more operating systems 431, such as Windows Serve, Mac OS X, Unix, Linux, FreeBSD, and the like. Those skilled in the art can understand that the target object detection device 400 can include more or fewer components than those shown in the figure, or some components can be combined, or different components can be arranged. Figure 4 The target object detection device structure shown does not constitute a limitation on the target object detection device, and the target object detection device can include more or fewer components than those shown in the figure, or some components can be combined, or different components can be arranged.

[0144] The present application also provides a computer device including a memory and a processor, the memory storing a computer readable computer program, and the computer readable computer program is executed by the processor to make the processor execute the steps of the target object detection method in each of the embodiments.

[0145] The present application also provides a computer readable storage medium, which can be a non-volatile computer readable storage medium or a volatile computer readable storage medium, and the computer readable storage medium stores a computer program, and the computer program makes a computer execute the steps of the target object detection method when the computer program is executed on the computer.

[0146] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be described here again.

[0147] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art, or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several computer programs for making a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0148] The above-described and above-embodied examples are only used to illustrate the technical solutions of the present application, rather than limit the same. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some technical features; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method of detecting a target object, characterized by, The detection method of the target object comprises: 3D target detection is performed on each image in a plurality of image sets to obtain a plurality of target object candidate boxes of each image, one image set corresponds to one camera, and each image set comprises a plurality of images collected by the camera; 3D spatial feature extraction is performed on each image in the plurality of image sets according to target feature information corresponding to each image to obtain a 3D feature map corresponding to each image, the target feature information corresponding to each image comprises feature information corresponding to a plurality of sensors, and the camera is one of the sensors; feature fusion is performed on the 3D feature map corresponding to each image to obtain a target fusion feature map; fusion feature information corresponding to each target object candidate box of each image is extracted from the target fusion feature map, and all target object candidate boxes are screened according to the fusion feature information corresponding to each target object candidate box of each image to obtain at least one target object detection box; the fusion feature information corresponding to each target object candidate box of each image is extracted from the target fusion feature map according to 3D spatial coordinate information corresponding to each target object candidate box of each image; target object information prediction is performed on the fusion feature information corresponding to each target object candidate box of each image through a preset self-attention mechanism to obtain target object information corresponding to each target object candidate box of each image; target object screening is performed on all target object candidate boxes according to the target object information corresponding to each target object candidate box of each image to obtain at least one target object detection box. the 3D spatial feature extraction performed on each image in the plurality of image sets to obtain a 3D feature map corresponding to each image comprises:

2. The method of claim 1, wherein 3D spatial conversion is performed on each image in the plurality of image sets to obtain a 3D space map corresponding to each image; target feature information corresponding to each image is obtained, and the target feature information corresponding to each image is projected onto the 3D space map corresponding to each image to obtain a 3D feature map corresponding to each image. the 3D spatial conversion performed on each image in the plurality of image sets to obtain a 3D space map corresponding to each image comprises:

3. The method of claim 2, wherein pixel-by-pixel depth estimation is performed on each image in the plurality of image sets to obtain a 3D space map corresponding to each image, and each 3D point in the 3D space map corresponding to each image corresponds to a 3D spatial coordinate information. the target feature information corresponding to each image is obtained, and the target feature information corresponding to each image is projected onto the 3D space map corresponding to each image to obtain a 3D feature map corresponding to each image, comprising:

4. The method of claim 2, wherein reading target feature information corresponding to each image, the target feature information corresponding to each image comprising at least one of laser radar feature information, millimeter wave radar feature information, ultrasonic feature information and image feature information of each image; ​ Obtain feature coordinate information, the feature coordinate information is used for indicating coordinate information of target feature information corresponding to each frame image in the corresponding frame image; Map the target feature information corresponding to each frame image to the corresponding 3D space graph according to the feature coordinate information, and obtain the 3D feature graph corresponding to each frame image.

5. The method of claim 1, wherein The feature fusion of the 3D feature graph corresponding to each frame image obtains the target fusion feature graph, including: The 3D feature graphs corresponding to the same frame image in all image sets are synthesized into bird's eye view feature, and the bird's eye view feature corresponding to the same sequence frame image is obtained. The bird's eye view features corresponding to the same sequence frame image are superimposed, and the target fusion feature graph is obtained.

6. The method of claim 1, wherein The feature fusion of the 3D feature graph corresponding to each frame image obtains the target fusion feature graph, and further includes: The 3D feature graphs corresponding to each frame image in each image set are superimposed, and the initial fusion feature graph corresponding to each image set is obtained. The initial fusion feature graphs corresponding to all image sets are synthesized into bird's eye view feature, and the target fusion feature graph is obtained.

7. The method for detecting a target object according to claim 6, characterized in that, The feature fusion of the 3D feature graph corresponding to each frame image obtains the target fusion feature graph, and further includes: According to the vehicle pose information when each frame image is collected, the 3D point alignment transformation is performed on the 3D feature graph corresponding to each frame image in each image set, and the aligned feature graph corresponding to each frame image in each image set is obtained. The aligned feature graphs corresponding to each frame image in each image set are superimposed by 3D points, and the initial fusion feature graph corresponding to each image set is obtained.

8. The method for detecting a target object according to claim 6, characterized in that, The feature fusion of the 3D feature graph corresponding to each frame image obtains the target fusion feature graph, and further includes: The same 3D point detection is performed on the initial fusion feature graphs corresponding to each image set, and the bird's eye view splicing position information is obtained. According to the bird's eye view splicing position information, the feature superposition and splicing of the same 3D point are performed on the initial fusion feature graphs corresponding to each image set, and the target fusion feature graph is obtained.

9. The method of claim 1, wherein The 3D target detection of each frame image in the plurality of image sets obtains a plurality of target object candidate boxes of each frame image, including: Through the preset monocular 3D target detection algorithm, the 2D detection frame generation and 3D detection frame regression are performed on each frame image in the plurality of image sets, and a plurality of target object candidate boxes of each frame image are obtained.

10. A device for detecting a target, characterized by The target object detection device includes: A detection module is configured to perform 3D target detection on each frame image in a plurality of image sets to obtain a plurality of target object candidate boxes of each frame image, one image set corresponding to one camera, and each image set including a plurality of frame images collected by the camera. An extraction module is configured to perform 3D space feature extraction on each frame image in the plurality of image sets according to target feature information corresponding to each frame image to obtain a 3D feature graph corresponding to each frame image, the target feature information corresponding to each frame image including feature information corresponding to a plurality of sensors, and the camera being one of the sensors. A fusion module is configured to perform feature fusion on the 3D feature graph corresponding to each frame image to obtain a target fusion feature graph. The screening module is configured to: extract the fusion feature information corresponding to each target object candidate box of each frame of image from the target fusion feature map according to the 3D space coordinate information corresponding to each target object candidate box of each frame of image; perform target object information prediction on the fusion feature information corresponding to each target object candidate box of each frame of image through a preset self-attention mechanism to obtain target object information corresponding to each target object candidate box of each frame of image; perform target object screening on all target object candidate boxes according to the target object information corresponding to each target object candidate box of each frame of image to obtain at least one target object detection box. The target object detection device comprises a memory and at least one processor, and the memory stores a computer program.

11. A device for detecting a target object, characterized by comprising: The at least one processor invokes the computer program in the memory, so that the target object detection device performs the target object detection method in any one of claims 1-9. The computer program is executed by the processor to implement the target object detection method in any one of claims 1-9.

12. A computer-readable storage medium having stored thereon a computer program, characterized in that ​

Citation Information

Patent Citations

  • Omnibearing obstacle detection method based on multi-sensor fusion

    CN111583337A

  • Intersection multi-view target detection method and system based on angular point pooling

    CN113673444A

  • Target detection model training method and device, target detection method and device, equipment and medium

    CN113902897A