Detection Model Training Method and Device, Object Detection Method and Device

Through multi-frame multi-modal data fusion processing, combined with image and point cloud data, the accuracy of target detection is improved, the problem of lack of comprehensiveness of data acquisition by a single sensor is solved, and the safety of the autonomous driving system is improved.

CN115019034BActive Publication Date: 2025-07-22HANGZHOU ZHIHUI MANTU TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210615907.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-31
Publication Date
2025-07-22
Estimated Expiration
2042-05-31

AI Technical Summary

Technical Problem

In the prior art, single-frame environmental data collected based on a single sensor leads to low accuracy of target detection, especially in complex urban environments, with limited detection accuracy, missed target detection, and safety hazards.

Method used

Multi-frame multi-modal data is used for target detection, combined with image information collected by image sensors and point cloud data collected by lidar, and a feature extraction network and detection network are used to process multi-frame sample point clouds and images, update the model parameters of the detection model, and improve detection accuracy.

Benefits of technology

Through the fusion processing of multi-frame and multi-modal data, the accuracy and comprehensiveness of object detection are effectively improved, the error of object detection is reduced, and the safety of the autonomous driving system is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115019034B_ABST
    Figure CN115019034B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a method and device for training a detection model. The method includes: obtaining at least one set of training data, where the training data includes multiple frames of sample point clouds, multiple frames of sample images, and sample object detection results corresponding to the sample point clouds and sample images. Processing the multiple frames of sample point clouds and multiple frames of sample images according to the feature extraction network in the detection model to obtain feature information corresponding to the multiple frames of sample point clouds and multiple frames of sample images. Processing the feature information according to the detection network in the detection model to obtain a first object detection result output by the object detection model. Updating the model parameters of the detection model according to the first object detection result and the sample object detection result. The method provided by the present application can effectively improve the accuracy of object detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to image processing technologies, and in particular, to a method and device for training a detection model, and a method and device for object detection. Background Art

[0002] With the continuous development of image processing technologies, current environmental perception has become an important application. For example, the target detection module in environmental perception can effectively detect the targets existing in the environment.

[0003] Currently, when performing target detection in the prior art, environmental data at the current moment is usually collected by a separate sensor, and then target detection is implemented based on the environmental data collected at the current moment. That is to say, in the prior art, target detection is usually implemented based on a single-frame image collected by a single sensor.

[0004] However, the single-frame environmental data collected by an independent single sensor usually lacks data comprehensiveness, which may lead to low accuracy of target detection. Summary of the Invention

[0005] The embodiments of the present application provide a method and device for training a detection model, and a method and device for object detection to overcome the problem of low accuracy of target detection.

[0006] In a first aspect, the embodiments of the present application provide a method for training a detection model, including:

[0007] Obtaining at least one set of training data, where the training data includes multiple frames of sample point clouds, multiple frames of sample images, and sample object detection results corresponding to the sample point clouds and the sample images;

[0008] Processing the multiple frames of sample point clouds and the multiple frames of sample images according to a feature extraction network in the detection model to obtain feature information corresponding to the multiple frames of sample point clouds and the multiple frames of sample images;

[0009] Processing the feature information according to a detection network in the detection model to obtain a first object detection result output by the object detection model;

[0010] Updating model parameters of the detection model according to the first object detection result and the sample object detection results.

[0011] In a second aspect, the embodiments of the present application provide an object detection method, including:

[0012] Obtaining a first point cloud and a first image collected at a first moment;

[0013] Obtaining multiple frames of second point clouds and multiple frames of second images collected before the first moment;

[0014] Process the first point cloud, the first image, the multiple frames of second point clouds, and the multiple frames of second images according to the detection model to obtain the object detection results corresponding to the first point cloud and the first image.

[0015] Wherein, the detection model is a model trained according to the method described in the first aspect above.

[0016] In a third aspect, an embodiment of the present application provides a detection model training device, including:

[0017] An acquisition module, configured to acquire at least one set of training data, wherein the training data includes multiple frames of sample point clouds, multiple frames of sample images, and sample object detection results corresponding to the sample point clouds and the sample images;

[0018] A first processing module, configured to process the multiple frames of sample point clouds and the multiple frames of sample images according to the feature extraction network in the detection model to obtain feature information corresponding to the multiple frames of sample point clouds and the multiple frames of sample images;

[0019] A second processing module, configured to process the feature information according to the detection network in the detection model to obtain a first object detection result output by the object detection model;

[0020] An update module, configured to update the model parameters of the detection model according to the first object detection result and the sample object detection result.

[0021] In a fourth aspect, an embodiment of the present application provides an object detection device, including:

[0022] A first acquisition module, configured to acquire a first point cloud and a first image collected at a first moment;

[0023] A second acquisition module, configured to acquire multiple frames of second point clouds and multiple frames of second images collected before the first moment;

[0024] A processing module, configured to process the first point cloud, the first image, the multiple frames of second point clouds, and the multiple frames of second images according to the detection model to obtain the object detection results corresponding to the first point cloud and the first image.

[0025] Wherein, the detection model is a model trained according to the method described in the first aspect above.

[0026] In a fifth aspect, an embodiment of the present application provides an electronic device, including:

[0027] A memory, configured to store programs;

[0028] A processor for executing the program stored in the memory, and when the program is executed, the processor is configured to execute the method described in the first aspect or the second aspect above.

[0029] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, including instructions that, when running on a computer, cause the computer to execute the method described in the first aspect or the second aspect above.

[0030] In a seventh aspect, an embodiment of the present application provides a computer program product, including a computer program that is executed by a processor to perform the method described in the first aspect or the second aspect above.

[0031] An embodiment of the present application provides a method and apparatus for training a detection model. The method includes: obtaining at least one set of training data, where the training data includes multiple frames of sample point clouds, multiple frames of sample images, and sample object detection results corresponding to the sample point clouds and the sample images. Processing the multiple frames of sample point clouds and the multiple frames of sample images according to a feature extraction network in the detection model to obtain feature information corresponding to the multiple frames of sample point clouds and the multiple frames of sample images. Processing the feature information according to a detection network in the detection model to obtain a first object detection result output by the object detection model. Updating the model parameters of the detection model according to the first object detection result and the sample object detection result. By obtaining at least one set of training data, each set of training data includes multiple frames of sample point clouds, multiple frames of sample images, and sample object detection results corresponding to the sample point clouds and the sample images, and then training the detection model according to these training data. Specifically, through sequential processing by the feature extraction network and the detection network in the detection model, a first object detection result output by the object detection model is obtained, and then, according to the first object detection result and the sample object detection result in the training data, the model parameters in the detection model are updated, so that it is effectively possible to train the detection model according to multiple frames of point cloud data and multiple frames of image data, and further ensure that the trained detection model can process multiple frames of point clouds and multiple frames of images to achieve object detection. Because there are multiple frames of multi-modal data as support, the accuracy of object detection can be effectively improved.

[0032] In addition, an object detection method and apparatus are provided in an embodiment of the present application. The method includes: obtaining a first point cloud and a first image collected at a first moment; obtaining multiple frames of second point clouds and multiple frames of second images collected before the first moment; processing the first point cloud, the first image, the multiple frames of second point clouds, and the multiple frames of second images according to a detection model to obtain an object detection result corresponding to the first point cloud and the first image; processing multiple frames of images and multiple frames of point clouds through the detection model trained as described above, where the multiple frames of images include the first image collected at the first moment and the multiple frames of second images collected before the first moment, and the multiple frames of point clouds include the first point cloud collected at the first moment and the multiple frames of second point clouds collected before the first moment, so as to output an object detection result. Since the object detection result is determined based on multiple frames of multi-modal environmental data, the comprehensiveness and richness of the data on which the output object detection result depends can be effectively ensured, and thus the accuracy and effectiveness of object detection can be effectively improved. Description of the Drawings

[0033] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following briefly introduces the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0034] Figure 1 It is a schematic diagram of the scenario of object detection provided in an embodiment of the present application;

[0035] Figure 2 It is a flowchart of the detection model training method provided in an embodiment of the present application;

[0036] Figure 3 It is the process of the detection model training method provided in an embodiment of the present application Figure 2 ;

[0037] Figure 4 It is a schematic structural diagram of the detection model provided in an embodiment of the present application;

[0038] Figure 5 It is a schematic diagram of the corresponding relationship between multiple frames of sample point clouds and multiple frames of sample images provided in an embodiment of the present application;

[0039] Figure 6 It is a schematic diagram of the feature points included in the grid provided in an embodiment of the present application;

[0040] Figure 7 It is a schematic diagram of the implementation of the regional division of the feature map provided in an embodiment of the present application;

[0041] Figure 8Schematic diagram for implementing the determination of the region set provided by the embodiments of the present application;

[0042] Figure 9 Schematic diagram for implementing the downsampling process provided by the embodiments of the present application;

[0043] Figure 10 Flowchart of the detection model training method provided by the embodiments of the present application Figure 3 ;

[0044] Figure 11 Flowchart of the object detection method provided by the embodiments of the present disclosure;

[0045] Figure 12 Schematic diagram of the flow of the detection model training method and the object detection method provided by the embodiments of the present application;

[0046] Figure 13 Schematic diagram of the structure of the detection model training device provided by the embodiments of the present application;

[0047] Figure 14 Schematic diagram of the structure of the object detection device provided by the embodiments of the present application;

[0048] Figure 15 Schematic diagram of the hardware structure of the electronic device provided by the embodiments of the present application. Detailed implementation manners

[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, rather than all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0050] To better understand the technical solutions of the present application, the related technologies involved in the present application will be further introduced in detail below.

[0051] In related fields such as autonomous driving and robotics, environmental perception is an important part of the algorithm system. Moreover, 3D object detection is the core algorithm module in the environmental perception system. Taking autonomous driving as an example, 3D object detection can detect the dynamic objects and static obstacles existing around the autonomous driving vehicle in real time. The dynamic objects include, for example, people, vehicles, etc., and the static obstacles include, for example, road signs, road piles, etc. By detecting the objects existing around the autonomous driving vehicle, safe, reasonable, and reliable route planning and prediction can be provided for the autonomous driving vehicle. The object detection introduced in the present application can also be understood as object detection, which has the same meaning.

[0052] Therefore, high-precision 3D object detection is the cornerstone of the automatic environment perception ability of autonomous driving. Among them, there are many sensors in autonomous vehicles to obtain environmental information about the surrounding environment. For example, the sensors can include lidar, image sensors, ultrasonic waves, millimeter-wave radars, and so on. Among them, lidar and image sensors have been widely used in autonomous driving systems.

[0053] For example, there is a 3D object detection algorithm based on single-frame lidar point cloud in the prior art. This method locates and classifies objects in three-dimensional space through the point cloud information captured by a multi-line lidar, and can generally be divided into point-based methods and grid-based methods. The point-based method takes the original point cloud as input, uses a typical point cloud feature extraction network to extract the representation information of the point cloud, and then regresses the category and position of the object at the point level; the grid-based method first projects the point cloud into the corresponding grid coordinate system such as a bird's-eye Figure 2 D plane, 3D voxel, etc., and then uses a 2D object detection algorithm such as Faster Region Convolutional Neural Networks (Faster-RCNN), Single Shot Multibox Detector (SSD), YOLO, etc. or a 3D sparse convolutional network to process the rasterized point cloud information to generate the final detection result. Compared with the point-based method, the grid-based method requires less computing resources and has higher overall detection accuracy, and has been widely used in current autonomous driving systems. However, the detection method based on single-frame lidar point cloud only relies on the point cloud information at the current moment for judgment, losing a large amount of useful historical information, so the detection accuracy in complex urban environments is limited.

[0054] In addition, there is also a 3D object detection algorithm based on single-frame image information in the prior art, which will not be elaborated here.

[0055] That is to say, in the prior art, object detection can be based on single-frame lidar point cloud, or it can also be based on single-frame image information. However, because lidar has characteristics such as all-weather, high ranging accuracy, and rich three-dimensional information, but lacks important semantic information; image sensors have characteristics such as rich color texture information and complete semantic information, but lack important depth information. Therefore, using only the single-frame environmental information collected by a single sensor independently for object detection has the problem of lack of comprehensiveness of data, which will in turn lead to low accuracy of object detection.

[0056] Furthermore, due to the low accuracy of object detection, problems such as missed detection, false detection, and inaccurate estimation of objects may occur. In the context of autonomous driving, this may further lead to risks such as unreasonable deceleration and sudden braking of autonomous vehicles, posing serious safety hazards to autonomous vehicles.

[0057] To address the problems in the prior art, the present application proposes the following technical concept: Since the information collected by a single sensor lacks comprehensiveness, it is possible to consider fusing the information collected by multiple sensors to enhance the comprehensiveness of the information. For example, the image information collected by an image sensor and the point cloud data collected by a lidar can be fused for object detection. Among them, the method of image and point cloud fusion can largely solve the problem of multi-sensor information complementarity. However, due to the inevitable calibration error between the image sensor and the lidar, it will cause confusion in the point cloud image information, thereby limiting the effect of object detection. At the same time, due to the complexity of the overall design of information fusion for object detection, it has high requirements for the computing power of the perception system.

[0058] At the same time, since the information provided by single-frame environmental data lacks richness, it is also possible to consider using multi-frame environmental data for object detection. However, similarly, if only multi-frame data collected by a single sensor is used for object detection, the accuracy of object detection will still be low due to the lack of comprehensiveness of the data collected by a single sensor.

[0059] Therefore, considering the above comprehensively, the present application proposes the idea of using multi-frame multi-modal data for object detection. Among them, multi-frame means using the data collected at the current moment and the data collected at historical moments for object detection. The data collected at historical moments can make up for the insufficient observation at the current moment to improve the detection accuracy. And multi-modal means using data collected by multiple sensors for object detection. For example, the image data collected by an image sensor and the point cloud data collected by a lidar can be fused for object detection, thereby effectively avoiding the problem of lack of comprehensiveness of the data collected by a single sensor and effectively improving the accuracy of object detection.

[0060] For example, it can be combined with Figure 1 for understanding. Figure 1 This is a schematic diagram of the object detection scenario provided by the embodiment of the present application.

[0061] As Figure 1 shown, when performing object detection, multiple frames of images and multiple frames of point clouds can be obtained. For example, the images 1, 2, 3,... and point clouds 1, 2, 3,... shown in Figure 1 etc. Then, object detection can be performed based on the multiple frames of images and multiple frames of point clouds to obtain the detection result.

[0062] Based on the above introduction, the method provided by this application will be introduced in detail below in combination with specific embodiments.

[0063] It can be understood that the detection model training method and object detection method provided in this application include two parts. One part is the detection model training method, which is for the training process of the detection model; the other part is the object detection method, which is for the application of the trained detection model. The following will introduce these two parts respectively.

[0064] Before introducing the specific method implementation, it should also be noted that the execution subject of each embodiment in this application can be a local server, a cloud server, a processor, a chip, etc., which are devices with data processing functions. The specific execution subject can be selected and set according to actual needs. This embodiment does not limit this, and any device with data processing functions can be used as the execution subject in this embodiment.

[0065] The following will first introduce the part of the detection model training method, that is, the training process of the detection model. Figure 2 It is a flowchart of the detection model training method provided by the embodiment of this application.

[0066] As Figure 2 shown, this method includes:

[0067] S201. Obtain at least one set of training data, where the training data includes multiple frames of sample point clouds, multiple frames of sample images, and sample object detection results corresponding to the sample point clouds and sample images.

[0068] In this embodiment, in order to train the detection model, for example, at least one set of training data can be obtained first. Among them, in any set of training data, there can be multiple frames of sample point clouds, multiple frames of sample images, and sample object detection results corresponding to these multiple frames of sample point clouds and multiple frames of sample images.

[0069] It can be understood that the sample object detection results in the training data can be, for example, manually labeled, or can also be machine-labeled. However, for the case of machine labeling, the premise is that its labeling correctness can be guaranteed, that is, the sample object detection results in the training data can be guaranteed to be accurate as a reference for training.

[0070] And in a possible implementation, for any set of training data, the multiple-frame sample point clouds therein, for example, can be multiple frames of point clouds collected by a lidar on the same device at different times. And, the multiple-frame sample images therein, for example, can be multiple frames of images collected by an image sensor on the same device at different times. And, there can be multiple image sensors.

[0071] Furthermore, for example, the multiple-frame sample point clouds can include the sample point cloud collected at the first moment, and multiple-frame sample point clouds collected before the first moment. And, the multiple-frame sample images can include the sample image collected at the first moment, and multiple-frame sample images collected before the first moment. And, the sample object detection result in this set of training data can be the detection result corresponding to the sample point cloud and sample image collected at the first moment.

[0072] It can be understood that the first moment can be understood as the moment when object detection needs to be performed currently. Then, it can also be understood that for the image sensor and lidar point cloud of the same device, the sample image and sample point cloud collected at the first moment can be obtained, and multiple-frame sample point clouds and multiple-frame sample images taken at historical moments (before the first moment) can be obtained, and the object detection result corresponding to the sample image and sample point cloud collected at the first moment can be used as the sample object detection result, so as to obtain a set of training data.

[0073] S202. Process the multiple-frame sample point clouds and multiple-frame sample images according to the feature extraction network in the detection model to obtain the feature information corresponding to the multiple-frame sample point clouds and multiple-frame sample images.

[0074] After obtaining the training data, the detection model can be trained according to the training data. The detection model in this embodiment can process multiple-frame point clouds and multiple-frame images, and thus output the object detection results corresponding to the point clouds and images.

[0075] In a possible implementation, the detection model in this embodiment can include a feature extraction network. Among them, the feature extraction network can process the multiple-frame sample point clouds and multiple-frame sample images to obtain the feature information corresponding to the multiple-frame sample point clouds and multiple-frame sample images.

[0076] It should be noted here that what is output after the feature extraction network processes is the feature information jointly corresponding to the multiple-frame sample point clouds and multiple-frame sample images, rather than the feature information corresponding to the multiple-frame sample point clouds and multiple-frame sample images respectively.

[0077] S203. Process the feature information according to the detection network in the detection model to obtain the first object detection result output by the object detection model.

[0078] Moreover, a detection network is also included in the detection model. After the feature extraction network processes to obtain the feature information corresponding to multiple frames of sample point clouds and multiple frames of sample images, the detection network in the detection model can process the feature information to obtain the first object detection result output by the object detection model.

[0079] In a possible implementation, the first object detection result may include the position and classification information of each object in the sample point cloud and sample image collected at the first moment among the multiple frames of sample images and multiple frames of sample point clouds.

[0080] Among them, the specific implementation in the detection network is to process the extracted feature information to output the position and classification information of each object in the sample point cloud and sample image collected at the first moment. In the actual implementation process, the specific implementation of the detection network can be selected and set according to actual needs, as long as it is a network structure that can implement the object detection function.

[0081] S204. Update the model parameters of the detection model according to the first object detection result and the sample object detection result.

[0082] After the detection model outputs the first object detection result, the model parameters of the detection model can be updated according to the first object detection result and the sample object detection result in the training data, so as to realize the training of the detection model.

[0083] It can be understood that when updating the model parameters of the detection model, the first object detection result is obtained by the detection model, while the sample object detection result is pre-annotated. Therefore, the sample object detection result can ensure its correctness.

[0084] Therefore, in a possible implementation, for example, a preset loss function can be used to process the first object detection result and the sample object detection result to determine the loss function value. The specific implementation of the preset calculation function can be selected and set according to actual needs, as long as the loss function value can reflect the gap between the first object detection result and the sample object detection result.

[0085] It can be understood that the model optimization goal in this embodiment is to make the first object detection result output by the detection model as close as possible to the sample object detection result in the training data. Therefore, after determining the loss function value, the model parameters of the detection model can be updated according to the loss function value, so as to optimize the detection model according to the loss function value, narrow the distance between the first object detection result output by the detection model and the sample object detection result, and further effectively ensure the correctness of the first sample object detection result output by the detection model.

[0086] Moreover, there are multiple sets of training data in this embodiment. When training and optimizing the detection model, the same training process will be executed for each set of training data, so as to realize multiple rounds of training of the detection model. In a possible implementation manner, when it is determined that the number of training rounds of the detection model reaches a preset number of rounds, or when it is determined that the detection accuracy of the detection model reaches a preset accuracy, it can be determined that the training of the detection model is completed, so as to obtain a trained detection model. The trained model after recognition can then perform object detection on the point cloud data and image data.

[0087] The detection model training method provided by the embodiment of the present application includes: obtaining at least one set of training data, where the training data includes multiple frames of sample point clouds, multiple frames of sample images, and sample object detection results corresponding to the sample point clouds and sample images. Process the multiple frames of sample point clouds and multiple frames of sample images according to the feature extraction network in the detection model to obtain feature information corresponding to the multiple frames of sample point clouds and multiple frames of sample images. Process the feature information according to the detection network in the detection model to obtain a first object detection result output by the object detection model. Update the model parameters of the detection model according to the first object detection result and the sample object detection result. By obtaining at least one set of training data, each set of training data includes multiple frames of sample point clouds, multiple frames of sample images, and sample object detection results corresponding to the sample point clouds and sample images, and then train the detection model according to these training data. Specifically, through the sequential processing of the feature extraction network and the detection network in the detection model, a first object detection result output by the object detection model is obtained, and then according to the first object detection result and the sample object detection result in the training data, the model parameters in the detection model are updated, so as to effectively realize the training of the detection model according to multiple frames of point cloud data and multiple frames of image data, and further ensure that the trained detection model can process multiple frames of point clouds and multiple frames of images to realize object detection. Because there are multiple frames of multi-modal data as support, the accuracy of object detection can be effectively improved.

[0088] Based on the above introduction, the specific model structure and processing process in the detection model of the present application will be further introduced in detail below in combination with specific embodiments, in combination with Figures 3 to 8 for illustration, Figure 3 is the flow chart of the detection model training method provided by the embodiment of the present application Figure 2 , Figure 4 is the structural schematic diagram of the detection model provided by the embodiment of the present application, Figure 5 is the schematic diagram of the corresponding relationship between multiple frames of sample point clouds and multiple frames of sample images provided by the embodiment of the present application, Figure 6 is the schematic diagram of the feature points included in the grid provided by the embodiment of the present application,Figure 7 Schematic diagram for implementing regional division of feature maps provided by embodiments of this application Figure 8 Schematic diagram for implementing determination of a region set provided by embodiments of this application

[0089] As Figure 3 shown, the method includes:

[0090] S301. Obtain at least one set of training data, where the training data includes multiple frames of sample point clouds, multiple frames of sample images, and sample object detection results corresponding to the sample point clouds and sample images

[0091] Among them, the implementation manner of S301 is similar to that of S201 introduced above, and will not be elaborated here

[0092] S302. For any frame of sample image, project the sample image onto the corresponding sample point cloud according to the calibration parameters between the image acquisition device and the point cloud acquisition device, to obtain the projected image information corresponding to the sample image

[0093] Based on the above introduction, it can be determined that after determining at least one set of training data, the multiple frames of sample point clouds and multiple frames of sample images can be processed according to the feature extraction network in the detection model, to obtain the feature information corresponding to the multiple frames of sample point clouds and multiple frames of sample images

[0094] In a possible implementation manner, for example, it can be combined with Figure 4 to understand the model structure of the detection model. As Figure 4 shown, the detection model includes a feature extraction network and a detection network, and the feature extraction network includes a feature encoding unit and a feature processing unit

[0095] For any set of training data, after inputting the multiple frames of sample images and multiple frames of sample point clouds into the detection model, for example, it can be first processed by the feature encoding unit in the feature extraction network

[0096] There are some differences in the processing of the multiple frames of sample images and multiple frames of sample point clouds by the feature encoding unit. First, the processing process of the multiple frames of sample images will be introduced below. Among them, the processing of each frame of sample image is similar. Therefore, below, any one frame of the multiple frames of sample images will be taken as an example to illustrate the processing process of the sample image, and the processing of the remaining sample images will not be elaborated

[0097] First, it is necessary to combine Figure 5Introduce the relationship between multiple-frame sample images and multiple-frame sample point clouds. Based on the above introduction, it can be understood that for any set of training data, among the multiple-frame sample images and multiple-frame sample point clouds, both include the sample point cloud and sample image collected at the first moment, as well as multiple-frame sample point clouds and multiple-frame sample images collected before the first moment.

[0098] For example, referring to Figure 5 , taking an autonomous driving vehicle as an example, assume that autonomous driving vehicle A collects Image 1 and Point Cloud 1 at time t1, and assume that object detection needs to be performed on Image 1 and Point Cloud 1 currently. Then time 1 can be the first moment.

[0099] In addition, multiple-frame historical point clouds and multiple-frame historical images before time t1 are also required as training data currently. For example, it can be obtained Figure 5 as shown, Image 2 and Point Cloud 2 at time t2, Image 3 and Point Cloud 3 at time t3, Image 4 and Point Cloud 4 at time t4, Image 5 and Point Cloud 5 at time t5,..., Image j and Point Cloud j at time tj, and so on, where j can be an integer greater than or equal to 2.

[0100] It can be understood that times t2, t3, t4, etc. are all historical times before time t1, and for each moment, there are collected image data and point cloud data. Therefore, for each frame of sample image in this embodiment, there is a corresponding sample point cloud, and the corresponding relationship here means that they are collected at the same moment.

[0101] After understanding the corresponding relationship between the sample image and the sample point cloud, when processing the sample image, in one possible implementation manner, for example, according to the calibration parameters between the image acquisition device and the point cloud acquisition device, the sample image can be projected onto the corresponding sample point cloud, so as to obtain the projected image information corresponding to the sample image.

[0102] Among them, the image acquisition device is used to collect image data. The image acquisition device can be, for example, the image sensor introduced above. And the point cloud acquisition device is used to collect point cloud data. The point cloud acquisition device can be, for example, the lidar introduced above.

[0103] It can be understood that due to the different installation positions of the image acquisition device and the point cloud acquisition device, the image data collected by the image acquisition device and the point cloud data collected by the point cloud acquisition device are in different coordinate systems. Therefore, for example, the calibration parameters between the image acquisition device and the point cloud acquisition device can be determined according to the installation position of the point cloud acquisition device, the installation position of the image acquisition device, the appearance design parameters of the autonomous vehicle, and the projection parameters of each camera. Among them, the calibration parameters can indicate the correspondence between image pixels and points in the laser point cloud. The specific implementation of determining the calibration parameters between different sensors can be selected and set according to actual needs, and this embodiment does not limit it.

[0104] Among them, when projecting the sample image onto the corresponding sample point cloud, for example, the sample image can be directly projected onto the corresponding sample point cloud to obtain the projected image information corresponding to the sample image. Or, the target detection network can be first used to extract effective feature information from the sample image, and then the feature information of the sample image can be projected onto the corresponding sample point cloud according to the calibration parameters, so as to obtain the projected image information corresponding to the sample image.

[0105] S303. For any frame of sample image, project the projected image information corresponding to the sample image onto the target image to obtain the second projection map corresponding to the sample image, and perform feature extraction on the second projection map to obtain the second feature map corresponding to the sample image, where the second feature map includes at least one second grid.

[0106] After obtaining the projected image information corresponding to the sample image, the feature map of the sample image can be determined according to the projected image information. The following also takes any frame of sample image as an example for illustration, and the remaining sample images will not be elaborated.

[0107] In a possible implementation manner, after determining the projected image information corresponding to the sample image, for example, the projected image information corresponding to the sample image can be projected onto the target image. The target image in this embodiment can be a 2D bird's-eye view.

[0108] After projection, the second projection map corresponding to the sample image can be obtained. Then, feature extraction can be performed on the second projection map to obtain the second feature map corresponding to the sample image. It can be understood that the 2D bird's-eye view includes multiple grids. Therefore, after performing feature extraction on the second projection map obtained after projection, the second feature map can include at least one second grid.

[0109] Based on the calibration parameters between the image acquisition device and the point cloud acquisition device, multiple frames of sample images are projected onto the corresponding sample point clouds respectively to obtain the projected image information corresponding to each sample image. Then, according to the projected image information, the second feature map corresponding to the sample image is determined. Since the processing is based on multiple frames of sample images, it can avoid the confusion of point cloud image information caused by inevitable calibration errors during the processing of a single frame of image, thereby effectively improving the overall model effect.

[0110] In addition, by projecting the projected image information corresponding to multiple frames of sample images onto a 2D bird's-eye view respectively, the projected maps corresponding to each sample image can be effectively obtained. Then, based on the projected maps, the second feature maps corresponding to each frame of sample images can be simply and effectively determined.

[0111] S304. For any frame of sample point cloud, project the sample point cloud onto a target image to obtain the first projected map corresponding to the sample point cloud, and perform feature extraction on the first projected map to obtain the first feature map corresponding to the sample point cloud. Among them, the first feature map includes at least one first grid.

[0112] In this embodiment, projection processing can also be performed on multiple frames of sample point clouds. The processing of multiple frames of sample point clouds is similar. Therefore, below, taking any one frame of the multiple frames of sample point clouds as an example, the processing process of the sample point cloud is introduced, and the processing of the remaining sample point clouds will not be elaborated.

[0113] In a possible implementation manner, for example, the sample point cloud can be projected onto a target image, and the target image here can also be a 2D bird's-eye view. Among them, the sample point cloud can be represented as {x, y, z, intensity}, where x, y, and z are the three-dimensional coordinate information of the point cloud, and intensity is the intensity information of the point cloud.

[0114] After projection, the first projected map corresponding to the sample point cloud can be obtained. Then, feature extraction can be performed on the first projected map to obtain the first feature map corresponding to the sample point cloud. Similarly, it can be understood that the 2D bird's-eye view includes multiple grids. Therefore, after performing feature extraction on the first projected map obtained after projection, the first feature map obtained can include at least one first grid.

[0115] By projecting multiple frames of sample point clouds onto a 2D bird's-eye view respectively, the projected maps corresponding to each sample point cloud can be effectively obtained. Then, based on the projected maps, the first feature maps corresponding to each frame of sample point clouds can be simply and effectively determined.

[0116] It should also be noted that in the network design, the original information of multi-frame sample point clouds and multi-frame sample images are extracted independently. This can retain the three-dimensional information of the point cloud and the semantic information of the image to the greatest extent, so as to effectively improve the effectiveness of model training, while ensuring the accuracy and comprehensiveness of the detection results output by the model.

[0117] S305. For any first grid in the first feature map, obtain multiple feature points in the first grid.

[0118] Similarly, taking any sample point cloud as an example, after obtaining the first feature map corresponding to the sample point cloud, the first feature map may include multiple first grids, and each first grid may include multiple feature points. Therefore, in this embodiment, for any first grid in the first feature map, multiple feature points in the first grid may be obtained.

[0119] For example, you can refer to Figure 6 , such as Figure 6 As shown, for example, assuming that 9 first grids are currently included in the first feature map, each first grid may include multiple feature points. For example, taking the first first grid as an example, it includes feature point a, feature point b, feature point c, and feature point d.

[0120] in, Figure 6 The illustration is merely exemplary and is intended to facilitate understanding of the relationship between feature maps, grids, and feature points. In actual implementation, the specific representation relationship between feature maps, grids, and feature points can be selected and set according to actual needs.

[0121] S306: Determine a correlation parameter corresponding to each feature point, wherein the correlation parameter is used to indicate a correlation degree between the feature point and the first grid.

[0122] Afterwards, for each feature point, its corresponding correlation parameter can be determined, wherein the correlation parameter is used to indicate the degree of correlation between the feature point and the first grid to which it belongs. Alternatively, it can be understood that the correlation parameter is used to indicate the degree of contribution of the feature point to the feature of the first grid to which it belongs.

[0123] In a possible implementation, for example, a lightweight learnable multilayer neural network (Multilayer Perceptron, MLP) can be used to process each feature point in each first grid separately, so as to obtain the correlation parameter corresponding to each feature point.

[0124] Among them, the use of MPL network in determining the correlation parameters can not only ensure the efficient dynamic interaction of temporal multimodal information, but also avoid excessive computational burden on the system.

[0125] Further, by determining the correlation parameters of each feature point in the first grid, and then using the correlation parameters as the weights of the feature points to fuse each feature point, the grid feature corresponding to the first grid can be obtained, so as to effectively determine the overall first grid feature of the first feature map.

[0126] S307. Obtain the grid feature corresponding to the first grid according to the correlation parameter corresponding to each feature point, where the first grid feature includes the grid features of multiple first grids in the first feature map.

[0127] Taking any one of the first grids in the first feature map as an example, the grid feature corresponding to the first grid can be obtained for the correlation parameter corresponding to each feature point in the first grid. In a possible implementation manner, for example, the correlation parameters of each feature point in the first grid can be used as weights, and then each feature point is fused to obtain the grid feature of the first grid.

[0128] The above processing is performed for each first grid in the first feature map, so that the grid features of each first grid can be obtained. Further, the first grid feature in this embodiment includes the grid features of multiple first grids in the first feature map.

[0129] It can also be understood that in this embodiment, for each frame of sample point cloud, its corresponding first feature map will be processed. Then, similarly, for the first feature map of each frame of sample point cloud, its corresponding first grid feature will be obtained according to the process introduced above.

[0130] S308. For any one of the second grids in the second feature map, obtain multiple feature points in the second grid.

[0131] The above introduces the implementation process of determining the first grid feature of the first feature map. The implementation process of determining the second grid feature for the second feature map is similar.

[0132] Taking any one of the sample images as an example, after obtaining the second feature map corresponding to the sample image, the second feature map can include multiple second grids, and each second grid can include multiple feature points. In this embodiment, for any one of the second grids in the second feature map, multiple feature points in the second grid can be obtained.

[0133] S309. Determine the correlation parameter corresponding to each feature point, where the correlation parameter is used to indicate the correlation degree between the feature point and the second grid.

[0134] Moreover, for each feature point, its corresponding relevance parameter can be determined. The implementation of determining the relevance parameter of the feature points in the second grid is similar to that of determining the relevance parameter of the feature points in the first grid introduced in S306 above, and will not be elaborated here.

[0135] S310. According to the relevance parameters corresponding to the respective feature points, the grid feature corresponding to the second grid is obtained, where the second grid feature includes the grid features of the multiple second grids in the second feature map.

[0136] After obtaining the relevance parameters of each feature point in the second grid, taking any one of the second grids as an example, the grid feature of the second grid can be determined according to the relevance parameters of each feature point in the second grid. By determining the grid feature of each second grid in the second feature map, the second grid feature of the second feature map can be obtained, and its implementation method is similar to that introduced in S307 above, and will not be elaborated here.

[0137] Similarly, in this embodiment, for each frame of sample image, its corresponding second feature map will be processed. Then, similarly, for the second feature map of each frame of sample image, its corresponding second grid feature will be obtained according to the process introduced above.

[0138] Moreover, by determining the respective relevance parameters of each feature point in the second grid, and then using the relevance parameters as the weights of the feature points to fuse each feature point, the grid feature corresponding to the second grid can be obtained, so as to effectively determine the overall second grid feature of the second feature map.

[0139] S311. For any one of the first feature maps, the first feature map is divided into N×M first regions, where N and M are integers greater than or equal to 1.

[0140] After obtaining the first feature maps corresponding to each sample point cloud, further, the first feature maps can be divided into regions. Among them, the processing processes of each first feature map are similar. Therefore, taking any one of the first feature maps as an example below, the processing of the first feature map will be introduced, and the processing of the remaining first feature maps will not be elaborated.

[0141] In a possible implementation manner, the first feature map can be divided into regions, so as to obtain N×M first regions, where both N and M are integers greater than or equal to 1. In the actual implementation process, the specific values of N and M can be selected and set according to actual needs. The values of N and M determine how many first regions the first feature map is specifically divided into. This embodiment does not limit the specific values of N and M.

[0142] For example, it can be combined withFigure 7 For understanding, as Figure 7 shown, assume that currently, for the first feature map 701, regional division is performed, and the first feature map 701 is divided into 4×4 first regions, thus obtaining Figure 7 the 16 first regions, namely Region 1 to Region 16 as shown

[0143] S312. For any second feature map, perform regional division on the second feature map to obtain N×M second regions

[0144] Moreover, in this embodiment, regional division can also be performed on the second feature map of the sample image. The regional division method of the second feature map is similar to the regional division of the first feature map introduced above, and will not be elaborated here

[0145] It should be emphasized that for the regional division of the second feature map, the number of the obtained second regions is also N×M, that is to say, the division method of the second feature map is the same as that of the first feature map

[0146] S313. According to the first regions of each first feature map and the second regions of each second feature map, determine each first region and each second region at the same position as a region set

[0147] After performing regional division on the first feature maps of each sample point cloud and the second feature maps of each sample image, since the regional division methods of each first feature map and each second feature map are the same. For example, the first feature map is divided into 4×4 first regions, and the second feature map is divided into 4×4 second regions, then the regions after the division of these feature maps can all be corresponding to each other

[0148] Therefore, in this embodiment, according to the first regions of each first feature map and the second regions of each second feature map, each first region and each second region at the same position can be determined as a region set

[0149] For example, in the example introduced above, assume that each first feature map and each second feature map are both divided into 4×4 regions, then there are a total of 16 regions, that is to say, there are regions at 16 positions. Then, for these 16 positions, each first region and each second region at the same position are determined as a region set

[0150] For example, it can be combined with Figure 8 for understanding, as Figure 8 shown, assume that currently there are the first feature map of sample point cloud 1, the first feature map of sample point cloud 2, the second feature map of sample image 1, and the second feature map of sample image 2. Assume that for these 4 feature maps, they are all divided intoFigure 7 The 4×4 area shown

[0151] After that, each first area and each second area at the same position are determined as a set of areas. For example, referring to Figure 8 , each area at position 4 among them (the area 4 of the first feature map of sample point cloud 1, the area 4 of the first feature map of sample point cloud 2, the area 4 of the second feature map of sample image 1, the area 4 of the second feature map of sample image 2) is determined as a set of areas. Then, for example, the set of areas 4 shown in Figure 7 can be obtained

[0152] The same operation is performed for each position, and then N×M sets of areas can be obtained. For example, in the illustration in Figure 8 , 16 sets of areas can be obtained

[0153] S314. For any set of areas, the first grid features corresponding to each first area in the set of areas and the second grid features corresponding to each second area in the set of areas are input into the self-attention network, so that the self-attention network outputs the sub-feature information corresponding to the set of areas

[0154] After obtaining multiple sets of areas, each set of areas is processed separately. Since the processing processes of each set of areas are similar, any set of areas is taken as an example for introduction, and the processing of the remaining sets of areas is similar and will not be elaborated

[0155] Based on the above introduction, it can be determined that for any set of areas, it includes multiple first areas and multiple second areas. And, there are corresponding first grid features for the first feature map. Then, after the area division of the first feature map, the first areas therein have corresponding first grid features, that is, the area part of the first areas corresponds to part of the first grid features in the overall first grid of the first feature map. And similarly, there are corresponding second grid features for the second feature map. Therefore, after the area division of the second feature map, the second areas therein also have corresponding second grid features

[0156] In a possible implementation, currently, in order to determine the sub-feature information of the set of areas, the first grid features corresponding to each first area in the set of areas and the second grid features corresponding to each second area in the set of areas can be input into the self-attention network. So that the self-attention network processes the above input data and outputs the sub-feature information corresponding to the set of areas

[0157] Among them, the self-attention network can be, for example, a Transformer self-attention network, or it can also be other self-attention networks. This embodiment does not limit this. It can be understood that the self-attention network performs unified modeling on the mutual relationships of features in different frames and different modalities through multiple layers of non-linear transformations, so as to effectively ensure the effectiveness and accuracy of the common feature information of the obtained multiple sample point clouds and multiple sample images.

[0158] The same processing is performed on each region set, so that the sub-feature information corresponding to each region set can be obtained.

[0159] S315. Concatenate the sub-feature information of each region set to obtain feature information.

[0160] The above is the processing of the region sets composed of the first region and the second region respectively after dividing each region, so as to obtain the sub-feature information of each region set. Further, in order to obtain the feature information corresponding to the multiple sample point clouds and the multiple sample images, the sub-feature information of each region set can be concatenated to obtain the feature information corresponding to the multiple sample point clouds and the multiple sample images.

[0161] By dividing the first feature map, multiple first regions after the division of the first feature map are obtained, and by dividing the second feature map, multiple second regions after the division of the second feature map are obtained. Then, multiple first regions and multiple second regions at the same position are determined as a region set, and each region set is processed separately, so as to effectively reduce the amount of data processed at one time, and improve the overall calculation accuracy and efficiency of the processing system. And when processing each region set, a self-attention network is used to fuse the grid features of each region in each region set, so as to effectively realize the unified modeling of the mutual relationships of features in different frames and different modalities, and further effectively obtain the unified feature information of multi-frame and multi-modal environmental data.

[0162] S316. Process the feature information according to the detection network in the detection model to obtain the first object detection result output by the object detection model.

[0163] S317. Update the model parameters of the detection model according to the first object detection result and the sample object detection result.

[0164] Among them, the implementation manners of S316 and S317 are similar to the implementation manners of S203 and S204 introduced above, and will not be elaborated here.

[0165] The detection model training method provided by the embodiments of this application trains the detection model based on multiple frames of sample images and multiple frames of sample point clouds. During the specific training process, through the calibration parameters between the image acquisition device and the point cloud acquisition device, multiple frames of sample images are projected onto the corresponding sample point clouds, and then the sample images are processed. This can avoid the confusion of point cloud image information caused by inevitable calibration errors when processing single-frame images, and thus effectively improve the overall model effect. Then, the projected image information corresponding to multiple frames of sample images and multiple frames of point clouds are projected onto a 2D bird's-eye view to obtain the projection maps corresponding to each sample image and each sample point cloud. Then, feature extraction is performed based on the projection maps, so that the feature maps corresponding to each frame of sample point cloud and each frame of sample image can be simply and effectively determined. Then, for each feature map, the grid features of each grid in the feature map are determined according to the relevance parameters of multiple feature points in the grid, and then the overall grid features corresponding to each feature map are obtained. Then, by dividing each feature map into regions, multiple regions after division of each feature map are obtained, and then regions at the same position are determined as a region set. For each region set, a self-attention network is used for processing respectively to determine the sub-feature information of each region set, and then the sub-feature information of each region set is spliced to obtain the unified feature of multi-frame multi-modal data, so that the overall calculation accuracy and efficiency of the processing system can be effectively improved. Then, the result of object detection is determined according to the processed feature information, and then the detection model is trained according to the object detection result output by the model and the sample object detection result in the training data, so that it can be effectively ensured that the trained detection model can effectively process multi-frame multi-modal environmental data to output the object detection result.

[0166] Based on the above introduction, in another possible implementation manner, after determining the feature information, further, the currently determined feature information can be further processed to obtain the final feature information. The following combines Figure 9 to understand this processing process. Figure 9 It is a schematic diagram of the implementation of downsampling processing provided by the embodiments of this application.

[0167] Based on the above introduction, it can be determined that after dividing each first feature map and second feature map into regions, N×M region sets can be obtained. For example, these N×M region sets can be determined as the original layer.

[0168] For example, it can be combined with Figure 9For understanding, assume that each first feature map and each second feature map are divided into 4×4 regions. Then, 4×4 region sets can be obtained, and these 4×4 region sets can form Figure 9 the original layer shown as 901 in

[0169] After determining the original layer, T downsampling processes can be performed on the N×M region sets in the original layer to obtain T downsampled layers. Among them, the i-th downsampled layer includes P i ×Q i region sets, where T is an integer greater than or equal to 1, and P i and Q i are integers greater than or equal to 1, and P i is less than N, Q i is less than M, and the value of i ranges from 1 to T.

[0170] Among them, after each downsampling, a downsampled layer will be obtained. The downsampling process can be understood as downsampling multiple region sets in the original layer into one region set. Therefore, the i-th downsampled layer includes P i ×Q i region sets, where P i and Q i are integers greater than or equal to 1, and P i is less than N, Q i is less than M.

[0171] In a possible implementation, for example, 4 region sets in the original layer can be successively downsampled into 1 region set in the downsampled layer. Therefore, P i for example, can be equal to and Q i for example, can be equal to

[0172] For example, with reference to Figure 9 for understanding, assume that the original layer includes 4×4 region sets. Then, the first downsampling is performed, and 4 region sets in the original layer are downsampled into 1 region set. After the first downsampling, the first downsampled layer shown as 902 in Figure 9 can be obtained, where the size of the original layer 902 is H 1 / 2 ×W 1 / 2 . As Figure 9 shown, the first downsampled layer 902 obtained includes 2×2 region sets.

[0173] In addition, a second downsampling can be further performed to downsample the four region sets in the first downsampling layer into one region set. After the second downsampling, the second downsampling layer shown as 903 in Figure 9 can be obtained, where the size of the original layer 903 is H 1 / 4 ×W 1 / 4 . As shown in Figure 9 , the obtained second downsampling layer 903 includes 1×1 region sets.

[0174] After that, for the i-th downsampling layer among the T downsampling layers, the sub-feature information of each of the P i ×Q i region sets in the downsampling layer can be determined. The implementation of determining the sub-feature information of the region set is similar to that introduced above and will not be elaborated here.

[0175] Then, the sub-feature information of each of the P i ×Q i region sets in the region sets can be concatenated to obtain the intermediate feature information of the i-th downsampling layer.

[0176] Furthermore, the intermediate feature information of the i-th downsampling layer can be mapped to the size of the feature information of the original layer to obtain the adjusted intermediate feature information, where the feature information of the original layer is the feature information corresponding to the multi-frame sample point cloud and the multi-frame sample images determined above.

[0177] Then, the adjusted intermediate feature information and the feature information of the original layer can be fused to obtain the fused feature information.

[0178] For example, it can be understood with reference to Figure 9 . As shown in Figure 9 , for example, the intermediate feature information can be determined for the first downsampling layer 902, and then the size of the intermediate feature information of the first downsampling layer 902 can be adjusted to obtain the adjusted intermediate feature information corresponding to the first downsampling layer 902. In addition, the intermediate feature information can be determined for the second downsampling layer 903, and then the size of the intermediate feature information of the second downsampling layer 903 can be adjusted to obtain the adjusted intermediate feature information corresponding to the second downsampling layer 903.

[0179] At this time, the adjusted intermediate feature information corresponding to one downsampling layer 902, the adjusted intermediate feature information corresponding to the second downsampling layer 903, and the feature information corresponding to the original layer 901 are of the same size, and then these three feature information are fused to obtain the fused feature information.

[0180] Afterwards, the fused feature information is determined as the feature information corresponding to the multi-frame sample point cloud and the multi-frame sample image, and then processing is performed based on this feature information to obtain the first object detection result output by the object detection model.

[0181] It can be understood that by designing the above multi-scale sparse self-attention mechanism (as Figure 9 shown), this mechanism first performs T downsamplings on the original layer to obtain T downsampled layers. Then, the self-attention module is independently used on each layer to determine the intermediate feature information corresponding to each downsampled layer. Then, the intermediate feature information and the feature information of the original layer are fused to obtain the final feature information, thereby enhancing the ability to perceive information at different scales. At the same time, due to the high sparsity of point cloud data in space, this mechanism performs downsampling processing, so sparse coding is also performed on the features, avoiding repeated calculations in the fully sparse region, and greatly improving the overall operation efficiency of the system. Therefore, through the above-introduced process, the accuracy and efficiency of model processing can be effectively improved.

[0182] Furthermore, on the basis of the above-introduced content, the implementation of obtaining at least one set of training data will be described below.

[0183] It can be understood that when obtaining training data, at least one set of original training data can be obtained first through network data or local data. However, there may be a situation of insufficient training data in network data or local data. Therefore, after obtaining the original training data, synthetic training data can be further obtained based on the original training data.

[0184] The following will be described in conjunction with Figure 10 which is Figure 10 the flow of the detection model training method provided by the embodiment of the present application Figure 3 .

[0185] As Figure 10 shown, the method includes:

[0186] S1001. Obtain at least one set of original training data.

[0187] In this embodiment, at least one set of original training data can be obtained first. Among them, the original training data is similar to the training data introduced above. The original training data may include multi-frame sample point clouds, multi-frame sample images, and sample object detection results corresponding to the sample point clouds and sample images.

[0188] S1002. In the original training data, determine the sample point clouds and sample images of at least one target object.

[0189] After determining the original training data, it can be understood that the original training data includes multiple frames of point clouds and multiple frames of images, and at least one object is included in these multiple frames of point clouds and multiple frames of images. For example, any one of the objects can be determined as the target object, and then the sample point cloud and sample image of this at least one target object can be intercepted from these multiple frames of sample point clouds and multiple frames of sample images.

[0190] S1003. Obtain at least one scene point cloud and scene image.

[0191] In this embodiment, for example, it can also be pre-determined that there are at least one scene point cloud and at least one scene image. The scene point cloud and scene image here are used to provide different scenes, such as outdoor scenes, indoor scenes, rainy-day scenes, sunny-day scenes, etc. The specific scene selection can be selected and set according to actual needs.

[0192] Correspondingly, the scene point cloud and scene image are the point cloud data and scene data collected for these scenes, which can be obtained through the network, or can also be obtained through local data, etc. This embodiment does not limit this.

[0193] S1004. For any one target object, perform synthesis processing on the sample point cloud of the target object and the scene point cloud to obtain a synthesized point cloud, and perform synthesis processing on the sample image of the target object and the scene image to obtain a synthesized image.

[0194] In the actual implementation process, any object existing in the sample point cloud and sample image can be understood as the target object, and the processing for each target object is similar. Therefore, only one target object will be introduced below, and the implementation methods for the remaining objects are similar.

[0195] Specifically, the sample point cloud of the target object obtained above and the scene point cloud introduced above can be synthesized to obtain a synthesized point cloud. And the sample image of the target object obtained above and the scene image can also be synthesized to obtain a synthesized image.

[0196] S1005. Determine the consistency parameter corresponding to each synthesized point cloud, and determine the consistency parameter corresponding to each synthesized image.

[0197] After obtaining the synthesized point cloud, in order to avoid conflicts between the synthesized sample point cloud and scene point cloud, and problems such as inconsistent multi-frame occlusion relationships of image data, it is also necessary to further determine the consistency parameter of each synthesized point cloud and determine the consistency parameter of each synthesized image.

[0198] Among them, the consistency parameter is a parameter used to indicate the consistency of the synthesized point cloud data and the synthesized image data. For example, it can be processed by a data consistency network, or it can be processed point by point and pixel by pixel to determine the consistency parameter. This embodiment does not limit this, and it can be selected according to actual needs.

[0199] S1006. Determine the synthesized point cloud whose consistency parameter meets the first preset condition as the target synthesized point cloud, and determine the synthesized image whose consistency parameter meets the second preset condition as the target synthesized image.

[0200] After determining the consistency parameter, the synthesized point cloud whose consistency parameter meets the first preset condition can be determined as the target synthesized point cloud, and the synthesized image whose consistency parameter meets the second preset condition can be determined as the target synthesized image.

[0201] Among them, if the consistency parameter can be, for example, a numerical parameter, then the first preset condition can be, for example, that the consistency parameter is greater than or equal to the first threshold; or, if the consistency parameter can also be, for example, a binary parameter indicating whether the synthesized image has consistency, then the first preset condition can also be, for example, that the consistency parameter indicates that the synthesized point cloud has consistency. This embodiment does not limit the specific implementation of the first preset condition, and it can be selected according to actual needs, as long as it is a condition used to limit that the target synthesized point cloud needs to have consistency.

[0202] And, the second preset condition is similar to the first preset condition introduced above. The difference is that the second preset condition is a condition set for the synthesized image, and the specific implementation of the second preset condition will not be elaborated here.

[0203] S1007. Determine the target synthesized point cloud and the target synthesized image as a set of synthesized training data.

[0204] After determining the target synthesized point cloud and the target synthesized image, the target synthesized point cloud and the target synthesized image can be determined as a set of synthesized training data. Since there can be multiple target synthesized point clouds and target synthesized images, at least one set of synthesized training data can be obtained.

[0205] S1008. Determine at least one set of original training data and at least one set of synthesized training data as at least one set of training data.

[0206] After that, determine at least one set of original training data and at least one set of synthesized training data as at least one set of training data, so as to obtain the training data used for training the detection model.

[0207] It is understandable that data-driven deep neural networks rely to a great extent on the quantity and quality of the data used. Although the amount of data collected for autonomous driving is huge, it still cannot cover all possible situations that may occur in road traffic. To improve data diversity with limited data collection, the present embodiment provides the above-described temporal multi-modal data augmentation scheme, which can ensure the coherence of the generated synthetic point clouds and synthetic image data in terms of time series and the consistency in cross-modal data.

[0208] The detection model training method provided by the embodiments of the present application first intercepts all the point cloud and image data of a certain target object in consecutive multiple frames of an autonomous driving scenario; then after subjecting the segment of point cloud image data to projection transformation and random perturbation processing, it is pasted into a new autonomous driving scenario; after pasting, it is judged frame by frame whether the point cloud data conflicts with the original scenario and whether the multi-frame occlusion relationship of the image data is consistent, and finally only the consistent target synthetic data is retained. Therefore, it can be understood that the above-described implementation process can effectively synthesize available training data to enhance the richness of the training data in scenarios, thereby effectively improving the accuracy and generalization of the detection model. At the same time, this scheme can also be directly applied to any autonomous driving collected data to provide richer autonomous driving scenario data.

[0209] The above embodiments describe the training process for the detection model. After the detection model is trained, the multi-frame point cloud data and multi-frame image data can be processed according to the detection model to obtain the object detection result.

[0210] Therefore, the present disclosure also provides an object detection method. The object detection method provided in the present disclosure will be introduced below in combination with specific embodiments. First, in combination with Figure 11 for illustration. Figure 11 is the flowchart of the object detection method provided by the embodiments of the present disclosure.

[0211] As Figure 11 shown, the method includes:

[0212] S1101. Obtain the first point cloud and the first image collected at the first moment.

[0213] In this embodiment, it is assumed that the first point cloud and the first image are obtained at the first moment. The acquisition moments of the first point cloud and the first image are the same, so there is a corresponding relationship between them.

[0214] The first moment can be understood as the moment when object detection needs to be performed, that is, currently object detection needs to be performed on the first point cloud and the first image collected at the first moment.

[0215] S1102. Obtain multiple frames of second point clouds and multiple frames of second images collected at the first moment.

[0216] Based on the above introduction, it can be determined that when the detection model in this embodiment performs object detection, in addition to processing the first point cloud and the first image collected at the first moment when object detection is required, it will also rely on historical point clouds and historical images.

[0217] Therefore, in this embodiment, multiple frames of second point clouds and multiple frames of second images collected before the first moment can also be obtained.

[0218] In a possible implementation, for example, all the point clouds and all the images collected by the current device (such as an autonomous vehicle) within a preset time period before the first moment can be collected to determine multiple frames of second point clouds and multiple frames of second images. Or, it can also be that within the preset time period before the first moment when the current device is collected, partial point clouds and partial images are collected. Here, the partial point clouds and partial images, for example, are data obtained intermittently at intervals of the first time period, or partial point clouds and partial images can also be randomly collected. This embodiment does not limit the specific implementation of obtaining the second point clouds and the second images. As long as the second point clouds and the second images are collected before the first moment and there are multiple frames.

[0219] It can be understood that these multiple frames of second point clouds and multiple frames of second images currently obtained also have a corresponding relationship in time sequence, similar to the above introduction.

[0220] S1103. Process the first point cloud, the first image, multiple frames of second point clouds, and multiple frames of second images according to the detection model to obtain the object detection results corresponding to the first point cloud and the first image.

[0221] Among them, the detection model is a model trained according to the method of any one of claims 1 to 13.

[0222] After determining the first point cloud, the first image, multiple frames of second point clouds, and multiple frames of second images introduced above, these data can be processed according to the detection model. The detection model is trained according to the embodiment introduced above, so it can effectively implement the processing of multiple frames of point clouds and multiple frames of images, and thus output the object detection results corresponding to the first point cloud and the first image. Among them, the output object detection results can include the position and classification information of each object in the first point cloud and the first image.

[0223] It should be noted that during the application process of the detection model, the internal processing process is similar to the processing process when training the detection model introduced above. The only difference is that during the application process of the detection model, there is no need to additionally synthesize training data.

[0224] The object detection method provided by the embodiment of the present application includes: obtaining a first point cloud and a first image collected at a first moment. Obtaining multiple frames of second point clouds and multiple frames of second images collected before the first moment. Processing the first point cloud, the first image, the multiple frames of second point clouds, and the multiple frames of second images according to a detection model to obtain an object detection result corresponding to the first point cloud and the first image. By processing multiple frames of images and multiple frames of point clouds through the detection model trained as described above, where the multiple frames of images include the first image collected at the first moment and multiple frames of second images collected before the first moment, and the multiple frames of point clouds include the first point cloud collected at the first moment and multiple frames of second point clouds collected before the first moment, so as to output an object detection result. Since the object detection result is determined based on multiple frames of multi-modal environmental data, the comprehensiveness and richness of the data on which the output object detection result depends can be effectively ensured, and thus the accuracy and effectiveness of object detection can be effectively improved.

[0225] Based on the above-described various embodiments, the method provided by the embodiment of the present application will be further described in a systematic and complete manner below in combination with Figure 12 a systematic and complete description of the method provided by the embodiment of the present application. Figure 12 It is a schematic flowchart of the detection model training method and the object detection method provided by the embodiment of the present application.

[0226] As Figure 12 shown, multiple frames of lidar point clouds can be obtained first. Among them, a single frame of lidar point cloud can be obtained by a multi-line rotational lidar or a solid-state lidar, and multiple frames of lidar point clouds are accumulated by the point clouds observed historically.

[0227] In addition, multiple frames of camera images can also be obtained. Among them, a single frame of camera image can be obtained by an in-vehicle camera, mainly including front and rear view cameras, surround view cameras, other supplementary cameras, etc., and multiple frames of camera images are accumulated by the images observed historically.

[0228] After that, the calibration parameters between the image acquisition device and the point cloud acquisition device can be obtained through a camera calibration projection unit.

[0229] During the model training process, the multiple frames of lidar point clouds and the multiple frames of camera images can be used as the original training data. In addition, at least one set of synthetic training data can be obtained through a data synthesis unit according to the original training data. The specific implementation can refer to the above introduction, so as to obtain at least one set of training data. Among them, by converting the target projection in the existing training data to different scenarios, synthetic training data is obtained, thereby greatly enriching the diversity of data scenarios and improving the speed and accuracy of the detection network training.

[0230] Then, the original training data, the synthetic training data, and the calibration parameters obtained above are input into the feature extraction network to determine the feature information corresponding to the multi-frame lidar point cloud and the multi-frame camera images. After that, the feature information is input into the detection network to obtain the detection result output by the detection model.

[0231] The process described above can be understood as the processing process of the detection model during the training of the model. After the detection model is trained, during the specific application process of the detection model, the processing process of the detection model is similar, except that there is no step of synthetic training data described above. For a more detailed implementation method, reference can be made to the above description, which will not be elaborated here.

[0232] In summary, the detection model training method and the object detection method provided by the embodiments of the present application provide an efficient deep learning architecture, realizing the full and efficient fusion of multi-frame point clouds and multi-frame image information, thereby improving the overall accuracy of perception object detection. In addition, an effective data augmentation mechanism is provided, which improves the training efficiency and test accuracy of the network according to the data characteristics of multi-frame multi-modal data.

[0233] Figure 13 It is a schematic structural diagram of a detection model training device provided by an embodiment of the present application. As Figure 13 shown, the device 130 includes: an acquisition module 1301, a first processing module 1302, a second processing module 1303, and an update module 1304.

[0234] The acquisition module 1301 is configured to acquire at least one set of training data, where the training data includes multi-frame sample point clouds, multi-frame sample images, and sample object detection results corresponding to the sample point clouds and the sample images;

[0235] The first processing module 1302 is configured to process the multi-frame sample point clouds and the multi-frame sample images according to the feature extraction network in the detection model to obtain the feature information corresponding to the multi-frame sample point clouds and the multi-frame sample images;

[0236] The second processing module 1303 is configured to process the feature information according to the detection network in the detection model to obtain a first object detection result output by the object detection model;

[0237] The update module 1304 is configured to update the model parameters of the detection model according to the first object detection result and the sample object detection result.

[0238] In a possible design, the feature extraction network includes a feature encoding unit and a feature processing unit;

[0239] The first processing module 1302 is specifically configured to:

[0240] Process the multi-frame sample point cloud and the multi-frame sample images according to the feature encoding unit to obtain first grid features corresponding to each of the sample point clouds and second grid features corresponding to each of the sample images;

[0241] Process each of the first grid features and each of the second grid features according to the feature processing unit to obtain the feature information.

[0242] In a possible design, the first processing module 1302 is specifically configured to:

[0243] For any one frame of the sample images, project the sample image onto the corresponding sample point cloud according to the calibration parameters between the image acquisition device and the point cloud acquisition device to obtain the projected image information corresponding to the sample image;

[0244] Obtain first feature maps corresponding to each of the sample point clouds and second feature maps corresponding to each of the sample images according to the multi-frame sample point clouds and the projected image information corresponding to each of the multi-frame sample images;

[0245] Obtain first grid features corresponding to each of the sample point clouds according to the first feature maps;

[0246] Obtain second grid features corresponding to each of the sample images according to the second feature maps.

[0247] In a possible design, the first processing module 1302 is specifically configured to:

[0248] For any one frame of the sample point clouds, project the sample point cloud onto a target image to obtain a first projection map corresponding to the sample point cloud, where at least one first grid is included in the first feature map;

[0249] Extract features from the first projection map to obtain a first feature map corresponding to the sample point cloud;

[0250] For any one frame of the sample images, project the projected image information corresponding to the sample image onto the target image to obtain a second projection map corresponding to the sample image, where at least one second grid is included in the second feature map;

[0251] Extract features from the second projection map to obtain a second feature map corresponding to the sample image.

[0252] In a possible design, the first processing module 1302 is specifically configured to:

[0253] For any one of the first grids in the first feature map, obtain a plurality of feature points in the first grid;

[0254] Determine the correlation parameter corresponding to each of the feature points, where the correlation parameter is used to indicate the degree of correlation between the feature point and the first grid;

[0255] According to the correlation parameter corresponding to each of the feature points, obtain the grid feature corresponding to the first grid, where the first grid feature includes the grid features of a plurality of first grids in the first feature map.

[0256] In a possible design, the first processing module 1302 is specifically configured to:

[0257] For any one of the second grids in the second feature map, obtain a plurality of feature points in the second grid;

[0258] Determine the correlation parameter corresponding to each of the feature points, where the correlation parameter is used to indicate the degree of correlation between the feature point and the second grid;

[0259] According to the correlation parameter corresponding to each of the feature points, obtain the grid feature corresponding to the second grid, where the second grid feature includes the grid features of a plurality of second grids in the second feature map.

[0260] In a possible design, the second processing module 1303 is specifically configured to:

[0261] For any one of the first feature maps, perform region division on the first feature map to obtain N×M first regions, where N and M are integers greater than or equal to 1;

[0262] For any one of the second feature maps, perform region division on the second feature map to obtain N×M second regions;

[0263] According to the first regions of the first feature maps, the second regions of the second feature maps, the first grid features, and the second grid features, obtain the feature information.

[0264] In a possible design, the second processing module 1303 is specifically configured to:

[0265] According to the first regions of the first feature maps and the second regions of the second feature maps, determine the first regions and the second regions at the same position as a region set;

[0266] For any one of the said region sets, input the first grid features corresponding to each of the first regions in the region set and the second grid features corresponding to each of the second regions into the self-attention network, so that the self-attention network outputs sub-feature information corresponding to the region set;

[0267] Concatenate the sub-feature information of each region set to obtain the feature information.

[0268] In a possible design, the second processing module is further configured to:

[0269] After determining, according to the first regions of the first feature maps and the second regions of the second feature maps, each of the first regions and each of the second regions at the same position as a region set, determine the N×M region sets as the original layer;

[0270] Perform T downsampling processes on the N×M region sets in the original layer to obtain T downsampled layers, where the i-th downsampled layer includes P i ×Q i region sets, where T is an integer greater than or equal to 1, P i and Q i are integers greater than or equal to 1, and P i is less than N, Q i is less than M, and i ranges from 1 to T.

[0271] In a possible design, the second processing module 1303 is further configured to:

[0272] After concatenating the sub-feature information of each region set to obtain the feature information, for the i-th downsampled layer among the T downsampled layers, determine the sub-feature information of each of the P i ×Q i region sets in the downsampled layer;

[0273] Concatenate the sub-feature information of each of the P i ×Q i region sets in the region sets to obtain the intermediate feature information of the i-th downsampled layer;

[0274] Map the intermediate feature information of the i-th downsampled layer to the size of the feature information of the original layer to obtain the adjusted intermediate feature information, where the feature information of the original layer is the feature information corresponding to the multi-frame sample point cloud and the multi-frame sample images;

[0275] Fuse the adjusted intermediate feature information and the feature information of the original layer to obtain the fused feature information.

[0276] In a possible design, the obtaining module 1301 is specifically configured to:

[0277] Obtain at least one set of original training data;

[0278] In the original training data, determine the sample point cloud and sample image of at least one target object;

[0279] Obtain at least one scene point cloud and scene image;

[0280] Determine the at least one set of training data according to the at least one set of original training data, the sample point cloud and sample image of the at least one target, the scene point cloud, and the scene image.

[0281] In a possible design, the obtaining module 1301 is specifically configured to:

[0282] For any one of the target objects, perform synthesis processing on the sample point cloud of the target object and the scene point cloud to obtain a synthesized point cloud, and perform synthesis processing on the sample image of the target object and the scene image to obtain a synthesized image;

[0283] Determine the consistency parameter corresponding to each synthesized point cloud and determine the consistency parameter corresponding to each synthesized image;

[0284] Determine the synthesized point cloud whose consistency parameter meets the first preset condition as the target synthesized point cloud, and determine the synthesized image whose consistency parameter meets the second preset condition as the target synthesized image;

[0285] Determine the target synthesized point cloud and the target synthesized image as a set of synthesized training data;

[0286] Determine the at least one set of original training data and the at least one set of synthesized training data as the at least one set of training data.

[0287] In a possible design, the multi-frame sample point clouds include the sample point cloud collected at the first moment and multiple frames of sample point clouds collected before the first moment, the multi-frame sample images include the sample image collected at the first moment and multiple frames of sample images collected before the first moment, and the sample object detection result is the detection result corresponding to the sample point cloud and sample image collected at the first moment.

[0288] The device provided in this embodiment can be used to execute the technical solutions of the above method embodiments. The implementation principles and technical effects are similar, and will not be elaborated here in this embodiment.

[0289] Figure 14 It is a schematic structural diagram of the object detection device provided in an embodiment of this application. As Figure 14 shown, the device 140 includes: a first acquisition module 1401, a second acquisition module 1402, and a processing module 1403.

[0290] The first acquisition module 1401 is configured to acquire a first point cloud and a first image collected at a first moment;

[0291] The second acquisition module 1402 is configured to acquire multiple frames of second point clouds and multiple frames of second images collected before the first moment;

[0292] The processing module 1403 is configured to process the first point cloud, the first image, the multiple frames of second point clouds, and the multiple frames of second images according to a detection model, so as to obtain an object detection result corresponding to the first point cloud and the first image,

[0293] wherein, the detection model is a model trained according to the detection model training method described in the above embodiments.

[0294] The device provided in this embodiment can be used to execute the technical solutions of the above method embodiments. The implementation principles and technical effects are similar, and will not be elaborated here in this embodiment.

[0295] Figure 15 It is a schematic hardware structure diagram of an electronic device provided in an embodiment of this application. As Figure 15 shown, the electronic device 150 in this embodiment includes: a processor 1501 and a memory 1502; wherein

[0296] The memory 1502 is used to store computer execution instructions;

[0297] The processor 1501 is configured to execute the computer execution instructions stored in the memory to implement each step executed by the electronic method in the above embodiments. For details, please refer to the relevant descriptions in the foregoing method embodiments.

[0298] Optionally, the memory 1502 can be either independent or integrated with the processor 1501.

[0299] When the memory 1502 is independently provided, the electronic device further includes a bus 1503 for connecting the memory 1502 and the processor 1501.

[0300] The embodiments of the present application also provide a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the detection model training method or the object detection method executed by the above electronic device is implemented.

[0301] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the device or module can be in an electrical, mechanical or other form.

[0302] The integrated modules implemented in the form of software function modules described above can be stored in a computer-readable storage medium. The above software function modules are stored in a storage medium, including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor (English: processor) to execute some steps of the methods described in the embodiments of the present application.

[0303] It should be understood that the above processor may be a central processing unit (English: Central Processing Unit, abbreviated as: CPU), and may also be other general-purpose processors, digital signal processors (English: Digital Signal Processor, abbreviated as: DSP), application specific integrated circuits (English: Application Specific Integrated Circuit, abbreviated as: ASIC), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0304] The memory may include a high-speed RAM memory, and may also include a non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk or an optical disc, etc.

[0305] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the buses in the drawings of this application are not limited to only one bus or one type of bus.

[0306] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disc. The storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0307] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0308] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for training a detection model, characterized in that, Including: Obtaining at least one set of training data, where the training data includes multiple frames of sample point clouds, multiple frames of sample images, and sample object detection results corresponding to the sample point clouds and the sample images; Processing the multiple frames of sample point clouds and the multiple frames of sample images according to a feature extraction network in a detection model to obtain feature information corresponding to the multiple frames of sample point clouds and the multiple frames of sample images; Processing the feature information according to a detection network in the detection model to obtain a first object detection result output by the detection model; Updating model parameters of the detection model according to the first object detection result and the sample object detection result; The feature extraction network includes a feature encoding unit and a feature processing unit; The processing the multiple frames of sample point clouds and the multiple frames of sample images according to a feature extraction network in a detection model to obtain feature information corresponding to the multiple frames of sample point clouds and the multiple frames of sample images includes: For any one frame of the sample images, projecting the sample image onto a corresponding sample point cloud according to calibration parameters between an image acquisition device and a point cloud acquisition device to obtain projected image information corresponding to the sample image; Obtaining a first feature map corresponding to each of the sample point clouds and a second feature map corresponding to each of the sample images according to the multiple frames of sample point clouds and the projected image information corresponding to each of the multiple frames of sample images; Obtaining a first grid feature corresponding to each of the sample point clouds according to the first feature map; Obtaining a second grid feature corresponding to each of the sample images according to the second feature map; Processing each of the first grid features and each of the second grid features according to the feature processing unit to obtain the feature information.

2. The method according to claim 1, wherein The obtaining a first feature map corresponding to each of the sample point clouds and a second feature map corresponding to each of the sample images according to the multiple frames of sample point clouds and the projected image information corresponding to each of the multiple frames of sample images includes: For any one frame of the sample point clouds, projecting the sample point cloud onto a target image to obtain a first projection map corresponding to the sample point cloud, where at least one first grid is included in the first feature map; Performing feature extraction on the first projection map to obtain a first feature map corresponding to the sample point cloud; For any one frame of the sample images, projecting the projected image information corresponding to the sample image onto the target image to obtain a second projection map corresponding to the sample image, where at least one second grid is included in the second feature map; Performing feature extraction on the second projection map to obtain a second feature map corresponding to the sample image.

3. The method according to claim 2, wherein The obtaining a first grid feature corresponding to each of the sample point clouds according to the first feature map includes: For any one of the first grids in the first feature map, obtaining multiple feature points in the first grid; Determining a relevance parameter corresponding to each of the feature points, where the relevance parameter is used to indicate the degree of relevance between the feature point and the first grid; Obtain the grid feature corresponding to the first grid according to the relevance parameters corresponding to each of the feature points, where the first grid feature includes the grid features of multiple first grids in the first feature map.

4. The method according to claim 2, characterized in that, The obtaining of the second grid feature corresponding to each of the sample images according to the second feature map includes: For any one of the second grids in the second feature map, obtain multiple feature points in the second grid; Determine the relevance parameter corresponding to each of the feature points, where the relevance parameter is used to indicate the degree of relevance between the feature point and the second grid; Obtain the grid feature corresponding to the second grid according to the relevance parameters corresponding to each of the feature points, where the second grid feature includes the grid features of multiple second grids in the second feature map.

5. The method according to any one of claims 1 to 4, characterized in that, The obtaining of the feature information by processing each of the first grid features and each of the second grid features by the feature processing unit includes: For any one of the first feature maps, perform region division on the first feature map to obtain N×M first regions, where N and M are integers greater than or equal to 1; For any one of the second feature maps, perform region division on the second feature map to obtain N×M second regions; Obtain the feature information according to the first regions of each of the first feature maps, the second regions of each of the second feature maps, each of the first grid features, and each of the second grid features.

6. The method according to claim 5, characterized in that The obtaining of the feature information according to the first regions of each of the first feature maps, the second regions of each of the second feature maps, each of the first grid features, and each of the second grid features includes: According to the first regions of each of the first feature maps and the second regions of each of the second feature maps, determine each of the first regions and each of the second regions at the same position as a region set; For any one of the region sets, input the first grid feature corresponding to each of the first regions in the region set and the second grid feature corresponding to each of the second regions into the self-attention network, so that the self-attention network outputs the sub-feature information corresponding to the region set; Concatenate the sub-feature information of each region set to obtain the feature information.

7. The method according to claim 6, wherein After determining each of the first regions and each of the second regions at the same position as a region set according to the first regions of each of the first feature maps and the second regions of each of the second feature maps, the method further includes: Determine the N×M region sets as the original layer; Perform T downsampling operations on the set of N×M regions in the original layer to obtain T downsampled layers, where the i-th downsampled layer includes P i ×Q i sets of regions, where T is an integer greater than or equal to 1, and P i and Q i are integers greater than or equal to 1, and P i is less than N, Q i is less than M, and i ranges from 1 to T.

8. The method according to claim 7, characterized in that After concatenating the sub-feature information of each region set to obtain the feature information, the method further includes: For the i-th downsampling layer among the T downsampling layers, determine the sub-feature information of each of the P i ×Q i region sets in the downsampling layer; Concatenate the sub-feature information of each of the i ×Q i region sets in the region sets to obtain the intermediate feature information of the i-th downsampling layer; Map the intermediate feature information of the i-th downsampling layer to the size of the feature information of the original layer to obtain the adjusted intermediate feature information, where the feature information of the original layer is the feature information corresponding to the multi-frame sample point cloud and the multi-frame sample image; Fuse according to the adjusted intermediate feature information and the feature information of the original layer to obtain the fused feature information.

9. The method according to any one of claims 1-4, 6-7, characterized in that The obtaining of at least one set of training data includes: Obtaining at least one set of original training data; In the original training data, determining sample point clouds and sample images of at least one target object; Obtaining at least one scene point cloud and scene image; Determining the at least one set of training data according to the at least one set of original training data, the sample point clouds and sample images of the at least one target, the scene point cloud, and the scene image.

10. The method according to claim 9, characterized in that, Determining the at least one set of training data according to the at least one set of original training data, the sample point clouds and sample images of the at least one target, the scene point cloud, and the scene image includes: For any one of the target objects, performing synthesis processing on the sample point cloud of the target object and the scene point cloud to obtain a synthesized point cloud, and performing synthesis processing on the sample image of the target object and the scene image to obtain a synthesized image; Determining the consistency parameter corresponding to each of the synthesized point clouds and determining the consistency parameter corresponding to each of the synthesized images; Determining the synthesized point cloud whose consistency parameter meets the first preset condition as the target synthesized point cloud, and determining the synthesized image whose consistency parameter meets the second preset condition as the target synthesized image; Determining the target synthesized point cloud and the target synthesized image as a set of synthesized training data; Determining the at least one set of original training data and the at least one set of synthesized training data as the at least one set of training data.

11. The method according to any one of claims 1-4, 6-8, 10, characterized in that, The multi-frame sample point clouds include the sample point cloud collected at the first moment and multiple frames of sample point clouds collected before the first moment, the multi-frame sample images include the sample image collected at the first moment and multiple frames of sample images collected before the first moment, and the sample object detection result is the detection result corresponding to the sample point cloud and sample image collected at the first moment.

12. An object detection method, characterized in that, It includes: Obtaining a first point cloud and a first image collected at the first moment; Obtaining multiple frames of second point clouds and multiple frames of second images collected before the first moment; Processing the first point cloud, the first image, the multiple frames of second point clouds, and the multiple frames of second images according to a detection model to obtain an object detection result corresponding to the first point cloud and the first image, wherein, the detection model is a model trained according to the method of any one of claims 1 to 11.

13. A detection model training device, characterized in that, It includes: An obtaining module, configured to obtain at least one set of training data, wherein the training data includes multiple frames of sample point clouds, multiple frames of sample images, and the sample object detection results corresponding to the sample point clouds and the sample images; A first processing module, configured to process the multiple frames of sample point clouds and the multiple frames of sample images according to a feature extraction network in the detection model to obtain feature information corresponding to the multiple frames of sample point clouds and the multiple frames of sample images; the feature extraction network includes a feature encoding unit and a feature processing unit; A second processing module, configured to process the feature information according to a detection network in the detection model to obtain a first object detection result output by the detection model; An update module, configured to update model parameters of the detection model according to the first object detection result and the sample object detection result; A first processing module, specifically configured to, for any frame of the sample images, project the sample image onto a corresponding sample point cloud according to calibration parameters between an image acquisition device and a point cloud acquisition device, to obtain projected image information corresponding to the sample image; obtain a first feature map corresponding to each of the sample point clouds and a second feature map corresponding to each of the sample images according to the multi-frame sample point clouds and the projected image information corresponding to the multi-frame sample images respectively; obtain a first grid feature corresponding to each of the sample point clouds according to the first feature map; obtain a second grid feature corresponding to each of the sample images according to the second feature map; and obtain the feature information according to the feature processing unit processing each of the first grid features and each of the second grid features.

14. An object detection device, characterized in that, Comprising: A first acquisition module, configured to acquire a first point cloud and a first image acquired at a first moment; A second acquisition module, configured to acquire multiple frames of second point clouds and multiple frames of second images acquired before the first moment; A processing module, configured to process the first point cloud, the first image, the multiple frames of second point clouds, and the multiple frames of second images according to a detection model, to obtain an object detection result corresponding to the first point cloud and the first image, wherein the detection model is a model trained according to the method of any one of claims 1 to 11.

15. An electronic device, characterized in that, Comprising: A memory, configured to store a program; A processor, configured to execute the program stored in the memory, and when the program is executed, the processor is configured to execute the method of any one of claims 1 to 11 or claim 12.

16. A computer-readable storage medium, characterized in that, Comprising instructions which, when running on a computer, cause the computer to execute the method of any one of claims 1 to 11 or claim 12.

17. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 11 or claim 12.

Citation Information

Patent Citations

  • Sample generation method and device, neural network training method and device, and data processing method and device

    CN112163643A

  • Neural network training method, object detection method, device and equipment

    CN113052295A