Target detection method and device based on driving scene, equipment, medium and program product
By adopting the fusion of multiple feature extraction network models and multi-scale feature fusion in three-dimensional object detection, the problems of loss of feature information and high calculation costs in the prior art are solved, and more accurate and efficient object detection is achieved.
Patent Information
- Application Number
- CN202510095080.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-21
AI Technical Summary
The existing three-dimensional object detection method based on voxelized convolution has lost feature information, which affects the feature extraction of objects with smaller sizes, and is highly computationally cost.
The fusion of multiple feature extraction network models is adopted, combined with multi-scale feature fusion, and object detection and feature extraction are carried out on three-dimensional point cloud data. The specific methods include obtaining three-dimensional point cloud data for preprocessing, establishing a training data set, training feature extraction network models through single-stage detection technology, and fusing point cloud feature learning network models through voxelized convolutional network models for feature extraction.
It improves the detailed feature extraction ability of the target, enhances the network's prediction accuracy of object positions, improves the overall detection effect, and reduces the calculation cost.
Smart Images

Figure CN119992500A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of three-dimensional target detection, and more specifically, to a target detection method, device, equipment, medium and program product based on driving scenarios. Background Art
[0002] The collection of three-dimensional point cloud data can usually be obtained directly from the LiDAR camera, and the overall data set construction is completed through manual data annotation. However, the problem with existing voxelized convolution-based methods is that feature information is lost during the point cloud voxelization process, which affects the network's feature extraction of smaller objects. At the same time, the model relies on the manually set volume size of the voxelized sampling. A larger volume setting is prone to loss of detail information and affects the estimation of the object's position. Therefore, many methods need to further refine the volume of the voxels to improve the representation ability of the network, but this will also increase the complexity of the calculation. For point learning-based methods, since the network extracts features for each point, the overall computational cost is high and the model takes a long time. Summary of the invention
[0003] The purpose of the embodiments of the present application is to provide a target detection method, device, equipment, medium and program product based on driving scenarios, so as to solve the problems of poor feature extraction and prediction accuracy and high computational cost of existing target detection methods.
[0004] In a first aspect, an embodiment of the present application provides a target detection method based on a driving scene, comprising:
[0005] Acquire three-dimensional point cloud data of the driving scene and pre-process the three-dimensional point cloud data;
[0006] Based on the preprocessed 3D point cloud data and the autonomous driving scene object detection database, a training data set is established;
[0007] Train and verify the training data set to obtain a feature extraction network model;
[0008] By fusing multiple feature extraction network models and combining multi-scale feature fusion, target detection and feature extraction are performed on 3D point cloud data.
[0009] In the above implementation process, the embodiment of the present application obtains three-dimensional point cloud data of the driving scene and preprocesses the three-dimensional point cloud data; establishes a training data set based on the preprocessed three-dimensional point cloud data in combination with an autonomous driving scene target detection database; trains and verifies the training data set to obtain a feature extraction network model; performs target detection and feature extraction on the three-dimensional point cloud data through the fusion of multiple feature extraction network models and in combination with multi-scale feature fusion; proposes a network feature fusion method to enhance the ability to extract detailed features of the target, which helps the network learn more discriminative features, and can refine the position prediction accuracy and improve the overall detection effect.
[0010] Furthermore, the obtaining of three-dimensional point cloud data of the driving scene and preprocessing the three-dimensional point cloud data includes:
[0011] The three-dimensional point cloud data of the driving scene is obtained, and the farthest point sampling technology is used to perform preliminary sampling on the three-dimensional point cloud data.
[0012] In the above implementation process, the farthest point sampling technology is used to screen data points while ensuring the correct characterization of the object shape. This can better maintain the shape characteristics of the object, reduce the loss of texture detail information, and will not destroy the spatial structure and direction information of the object.
[0013] Furthermore, the fusion of multiple feature extraction network models and multi-scale feature fusion to perform target detection and feature extraction on three-dimensional point cloud data includes:
[0014] The voxelized convolutional network model is fused with the point cloud feature learning network model to extract key point features of 3D point cloud data, and feature fusion is performed by combining multi-scale feature fusion.
[0015] In the above implementation process, the feature extraction of the voxelized convolutional network can provide certain prior information for the point cloud feature learning network, speeding up the calculation speed; effectively combining the deep and shallow layers of the neural network to reduce the impact of background interference and scale.
[0016] Furthermore, the voxelized convolutional network model is used to fuse the point cloud feature learning network model to extract key point features from the three-dimensional point cloud data, and feature fusion is performed by combining multi-scale feature fusion, including:
[0017] Processing the three-dimensional point cloud data to obtain voxelized point cloud data;
[0018] Train the voxelized point cloud data to obtain a voxelized convolutional network model;
[0019] The convolution features trained by the voxelized convolutional network model are processed by the point cloud feature learning network model and combined with multi-scale feature fusion to extract key point features;
[0020] The key point features are fused based on the voxelized convolutional network model, and the shallow feature information is fused with the deep features of the deep network.
[0021] In the above implementation process, the feature extraction of the voxelized convolutional network can provide certain prior information for the point cloud feature learning network, speeding up the calculation speed; effectively combining the deep and shallow layers of the neural network to reduce the impact of background interference and scale.
[0022] Furthermore, the training and verification of the training data set to obtain a feature extraction network model includes:
[0023] Using the single-stage detection technology as the basic framework, the feature extraction network model is obtained by improving the position frame generation method, using the prior 3D object position information as an auxiliary, and adopting the multi-scale target prediction method to train the training data set;
[0024] The training dataset is validated to generate candidate location boxes at multiple scales.
[0025] In the above implementation process, a single-stage detection method is used as the basic framework, that is, only one position frame regression and classification calculation is performed. However, by improving the position frame generation method, the prior three-dimensional object position information is used as an auxiliary, and an attempt is made to generate candidate position frames in the three-dimensional data space to reduce feature loss in the feature conversion process; a multi-scale target prediction method is used to make the overall detection model more sensitive to object scale changes and improve detection accuracy.
[0026] Furthermore, it also includes:
[0027] Conduct field scenario tests on the feature extraction network model, receive feedback results, and adjust the feature extraction network model based on the feedback results.
[0028] In the above implementation process, the feature extraction network model is tested in field scenarios, the model is verified, and the method is optimized and adjusted based on the feedback results.
[0029] In a second aspect, an embodiment of the present application provides a target detection device based on a driving scene, comprising:
[0030] A data acquisition module, used to acquire three-dimensional point cloud data of the driving scene and pre-process the three-dimensional point cloud data;
[0031] A data processing module is used to establish a training data set based on the pre-processed 3D point cloud data and the autonomous driving scene object detection database;
[0032] The model training module is used to train and verify the training data set to obtain a feature extraction network model;
[0033] The target detection module is used to perform target detection and feature extraction on 3D point cloud data by fusing multiple feature extraction network models and combining multi-scale feature fusion.
[0034] In a third aspect, an embodiment of the present application provides an electronic device, including:
[0035] A processor, a memory and a bus, wherein the processor is connected to the memory via the bus, and the memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, they are used to implement the target detection method based on the driving scene as described above.
[0036] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a server, the target detection method based on the driving scene as described above is implemented.
[0037] In a fifth aspect, an embodiment of the present invention provides a computer program product, which includes instructions, and when the instructions are executed by a computer, the computer implements the target detection method based on the driving scene as described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0039] Figure 1 A schematic diagram of a process flow of a target detection method based on a driving scenario provided in an embodiment of the present application;
[0040] Figure 2 A schematic diagram of a feature fusion process of a target detection method based on a driving scenario provided in an embodiment of the present application;
[0041] Figure 3 A schematic diagram of a cross-layer multi-scale feature fusion process of a target detection method based on a driving scenario provided in an embodiment of the present application;
[0042] Figure 4is a schematic structural diagram of a target detection device based on a driving scenario provided in an embodiment of the present application;
[0043] Figure 5 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0044] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0045] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0046] In view of the difficulty in processing and acquiring three-dimensional point cloud data, the embodiment of the present application proposes a more reasonable data sampling and processing method to ensure the efficiency of model training and improve the effect of deep learning detection model; considering that the target detection task in the autonomous driving scene is affected by different environments, the embodiment of the present application improves the robustness of the three-dimensional target detection model based on deep learning by establishing a three-dimensional target database with a wider range of environmental conditions, providing prerequisites for future practical applications. In addition, the three-dimensional target detection database in the current autonomous driving scene has many data types, messy annotation methods, and no unified standard method, which in turn affects the problem that each detection model cannot be expanded to practical applications; in order to ensure the robustness and scalability of the three-dimensional detection model, the embodiment of the present application needs to solve the problems of formulating standards for training and testing databases, unifying annotation formats, and multi-model verification capabilities. At the same time, the embodiment of the present application designs a reasonable and effective three-dimensional object point cloud data annotation method, which is also an effective guarantee for improving the discrimination ability of the deep learning detection model.
[0047] The problem with the voxelized convolution-based method is that feature information is lost during the voxelization process of the point cloud, which affects the network's feature extraction of smaller objects. At the same time, the model relies on the manually set volume size of the voxelized sampling. A larger volume setting is prone to loss of detail information, affecting the estimation of the object's position. Therefore, many methods require a more refined voxel volume to improve the network's representation capabilities, but this will also increase the complexity of the calculation. For the point learning-based method, since the network extracts features for each point, the overall computational cost is high and the model requires a high time consumption. In response to the above method, the embodiment of the present application believes that integrating the two types of feature learning methods will help the network learn more discriminative features, and can refine the position prediction accuracy, which is more suitable for the practical application of autonomous driving target detection methods.
[0048] For the two-stage target detection framework, this type of method needs to complete two position regression calculations and two classification calculations, and the overall accuracy is relatively high. Since the number of features of three-dimensional point cloud data or RGB-D data is much larger than that of two-dimensional images, the computational cost of the two-stage method is relatively high, and the overall speed will be slowed down by one step. For the single-stage target detection framework, the overall computational amount of this method is greatly reduced, and the overall detection speed is improved. However, this method also brings about a lot of accuracy loss. Although it has good practicality, the detection accuracy is often difficult to meet application requirements. For the current three-dimensional target detection method, the embodiment of the present application optimizes the object position estimation regression and classification calculation from the perspective of balancing detection accuracy and speed on the basis of the single-stage detection algorithm, so that the overall model can maintain a faster detection speed and improve position prediction and classification accuracy.
[0049] Please see Figure 1 , Figure 1 A flowchart of a target detection method based on a driving scene provided in an embodiment of the present application. Figure 1 , the target detection method based on driving scene includes:
[0050] 100. Acquire three-dimensional point cloud data of the driving scene and pre-process the three-dimensional point cloud data.
[0051] Specifically, three-dimensional point cloud data of the driving scene is obtained, and the farthest point sampling technology is used to perform preliminary sampling on the three-dimensional point cloud data.
[0052] Optionally, 3D point cloud data can be obtained through sensors such as lidar, radar, sonar or depth camera. These sensors measure the distance and position of objects by emitting light waves or sound waves and receiving reflected signals in driving scenarios, thereby constructing a 3D model of the environment.
[0053] For example, the embodiment of the present application uses a laser radar to obtain accurate three-dimensional spatial position information and presents environmental data in the form of a point cloud, wherein the camera is responsible for capturing rich texture and color information for identifying the appearance of objects, and the millimeter wave radar is used to supplement the target detection capability in bad weather. In addition, the vehicle is equipped with multiple types of cameras, such as a front-view camera, a rear-view camera, a surround-view camera, etc., in order to collect visual information in all directions.
[0054] Among them, the sensor works in real time while the vehicle is driving, controls the laser radar to scan periodically, and the camera shoots at a certain frame rate, such as the common 30 frames per second or 60 frames per second, to adapt to environmental changes. At the same time, the vehicle is controlled to collect data under different road conditions (highways, urban roads, etc.), weather (sunny, rainy, snowy, etc.) and time (daytime, nighttime, etc.) to improve the adaptability of the system.
[0055] Optionally, farthest point sampling is a greedy algorithm for selecting a set of representative points from a point cloud that are as dispersed in space as possible. Specifically, a starting point is randomly selected from the point cloud as the first sampling point; in each iteration, the algorithm calculates the distances of all points to the nearest point in the current sampling set, selects the point with the farthest distance as the next sampling point; adds the newly selected point to the sampling set; repeatedly selects the next sampling point and adds it to the sampling set until the predetermined number of sampling points is reached.
[0056] Therefore, through the farthest point sampling technology, data points can be screened under the premise of correctly describing the shape of the object, which can better maintain the shape characteristics of the object, reduce the loss of texture detail information, and will not destroy the spatial structure and direction information of the object.
[0057] Exemplarily, the embodiment of the present application first analyzes the existing autonomous driving database. In the deep learning algorithm, the quality of the data has a certain impact on the accuracy of the detection model. Therefore, the existing target detection database is analyzed, the characteristics and test conditions of each database are summarized, the categories and feature information of the objects in the data are analyzed, and the challenges and difficulties of the problem are summarized. Then, based on the existing database, an autonomous driving scene target detection database that can be used in the embodiment of the present application can be established according to the needs.
[0058] Exemplarily, for various databases, the embodiments of the present application include KITTI database, NuScenes database, and Waymo Open Dataset database. KITTI database is currently the most widely used autonomous driving scene target detection database in the world. It contains 14,999 2D images and point cloud data, and the annotations contain 8 categories, such as cars, pedestrians, bicycles, etc. This data set lays the foundation for the establishment of a larger autonomous driving scene target detection database. NuScenes database contains 1.4 million images and 1.1 million three-dimensional bounding boxes, which poses a more severe challenge to the autonomous driving scene target detection algorithm, and also promotes the improvement and development of deep learning-based target detection algorithms. The Waymo Open Dataset database improves the problem of a single driving environment in the current data. At the same time, the database not only faces weather and driving scene problems, but also introduces time changes, such as night, day, dusk, etc., so that the data is more in line with the actual environment. The database contains 3,000 driving clips, which are not only oriented to detection problems, but also involve tracking and segmentation tasks. The embodiments of the present application intend to analyze the characteristics and methods of constructing autonomous driving target detection data from these databases. Since the scenarios that these data are intended for are too broad, the embodiments of the present application refine the challenges and conditions of the target requirements and only target driving scenarios to reduce data collection costs and improve overall efficiency.
[0059] Exemplarily, after acquiring the data, it is necessary to process the data to a certain extent. The data points collected by the laser radar camera are relatively large, usually 40,000-50,000 points. If the data is directly input, the calculation cost is high and the operation efficiency is low. Therefore, the embodiment of the present application intends to perform preliminary sampling of the data according to the characteristics of the data, reduce the number of data points, and improve the overall calculation efficiency. For the sampling method, the embodiment of the present application intends to screen the data points by farthest point sampling (Farthest Point Sampling, FPS) under the premise of ensuring the correct characterization of the shape of the object. The farthest point sampling method calculates the distance from all data points in the space to the initial point, selects the maximum value as the next sampling point, and obtains the number of points that can cover the entire object shape through iterative cycles. On the one hand, this method can better maintain the shape characteristics of the object and reduce the loss of texture detail information. On the other hand, this method does not destroy the spatial structure and directional information of the object. In addition, the embodiment of the present application will also use more methods for comparison during experimental verification, and then perform relevant optimization.
[0060] 200. Establish a training data set based on the preprocessed three-dimensional point cloud data and the autonomous driving scene target detection database.
[0061] Optionally, after acquiring the three-dimensional point cloud data, the data is further preprocessed, including data cleaning, labeling and classification, and data enhancement.
[0062] Data cleaning includes removing outliers in the lidar point cloud data and invalid frames (blur, abnormal exposure, etc.) of the camera image, and processing duplicate information to reduce data redundancy. Labeling and classification include manually labeling object categories, locations, bounding boxes and other information, and classifying them into training sets, validation sets and test sets. Data enhancement includes rotating, scaling, translating, flipping, color transformation and other operations on image data, and performing random point sampling, rotation, translation and noise addition on point cloud data to increase data diversity and model generalization capabilities.
[0063] 300. Train and verify the training data set to obtain a feature extraction network model.
[0064] Specifically, a single-stage detection technology is used as the basic framework, and the location box generation method is improved. Prior three-dimensional object location information is used as an aid, and a multi-scale target prediction method is adopted to train the training data set to obtain a feature extraction network model; the training data set is verified to generate candidate location boxes at multiple scales.
[0065] Exemplarily, single-stage detection technology can use methods such as YOLO-6D and PointPillars to divide all data into multiple groups. Each group presets multiple rectangular candidate boxes, and directly predicts the position and category of the target through candidate box regression and classification within the group. The overall computational complexity is greatly reduced and the overall detection speed is improved.
[0066] Therefore, a single-stage detection method is used as the basic framework, that is, only one position box regression and classification calculation is performed, but by improving the position box generation method, using prior three-dimensional object position information as an aid, and trying to generate candidate position boxes in the three-dimensional data space, so as to reduce feature loss in the feature conversion process; using a multi-scale target prediction method, the overall detection model is made more sensitive to object scale changes and the detection accuracy is improved.
[0067] Illustratively, the embodiments of the present application provide a more detailed representation of 3D objects by fusing multiple feature extraction network models, and introduce the concept of multi-scale feature fusion to effectively combine deep and shallow neural network networks to reduce the impact of background interference and scale.
[0068] Exemplarily, in the current 3D target detection method based on deep learning, the commonly used method is a two-stage detection method based on region generation, that is, on the feature map output by the feature extraction network, a candidate position frame is generated, some frames with background areas are screened out, and then the position frame is regressed and the target category classification calculation is performed to complete the final detection result output. For 3D target detection, candidate position frames are usually given on the generated bird's-eye view or top view. Typical methods include Point R-CNN, Part-A2Net, etc. According to literature and actual test analysis, the use of 2D feature maps to generate candidate regions will cause a certain spatial position loss, affecting the regression calculation of the 3D position frame by the subsequent detection method. At the same time, the two-stage method has a large amount of calculation and is difficult to be applied in practice. Therefore, the embodiment of the present application intends to use a single-stage detection method as the basic framework, that is, only one regression and classification calculation of the position frame is performed, but by improving the position frame generation method, using the prior 3D object position information as an auxiliary, and trying to generate candidate position frames in the three-dimensional data space to reduce the feature loss in the feature conversion process. Furthermore, the embodiments of the present application intend to use a multi-scale target prediction method to address the problem of object scale changes and small-sized object detection, generate candidate position boxes at multiple scales, cope with the scale changes of 3D objects, make the overall detection model more sensitive to object scale changes, and improve detection accuracy. In addition, the single-stage detection method has a faster running speed and a smaller amount of calculation, which has certain advantages in the implementation of the method and the subsequent expansion and improvement of the method.
[0069] 400. Through the fusion of multiple feature extraction network models and combined with multi-scale feature fusion, target detection and feature extraction are performed on 3D point cloud data.
[0070] Specifically, a voxelized convolutional network model is fused with a point cloud feature learning network model to extract key point features of three-dimensional point cloud data, and feature fusion is performed by combining multi-scale feature fusion.
[0071] Therefore, the feature extraction of the voxelized convolutional network can provide certain prior information for the point cloud feature learning network, speeding up the calculation speed; effectively combining the deep and shallow layers of the neural network to reduce the impact of background interference and scale.
[0072] Among them, the voxelized convolutional network model converts point clouds into three-dimensional voxel matrices. To address the problem of point cloud disorder, rectangular voxel blocks are set to convert disordered point clouds into spatial matrices, and then convolutional neural networks (CNNs) are used to extract and characterize features, ultimately achieving position estimation and classification of three-dimensional objects. Common methods such as VoxelNet and Voxel-FPN all use voxelization methods to enable point clouds to use convolutional neural networks for feature extraction, and combine them with detection frameworks for position prediction and category estimation.
[0073] Among them, the point cloud feature learning network model directly processes the point cloud data, uses the Full Connection Neural Networks to construct a Multi0-Layer Perception (MLP), extracts the relationship between point clouds, and thus completes the target representation of three-dimensional objects. Common methods include PoinNet, PointNet++, and Frustum PointNet. Compared with the voxelized convolution method, this method can further refine the feature relationship between points, and also maintain the spatial position relationship of each point, and the overall position estimation accuracy is also improved.
[0074] In some embodiments, the voxelized convolutional network model is fused with the point cloud feature learning network model to extract key point features from the three-dimensional point cloud data, and feature fusion is performed by combining multi-scale feature fusion, including:
[0075] 410. Process the three-dimensional point cloud data to obtain voxelized point cloud data.
[0076] 420. Train the voxelized point cloud data to obtain a voxelized convolutional network model.
[0077] The convolution features trained by the voxelized convolutional network model are processed by the point cloud feature learning network model and combined with multi-scale feature fusion to extract key point features.
[0078] 440. Based on the voxelized convolutional network model, the key point features are fused and the shallow feature information is fused with the deep features of the deep network.
[0079] In the above implementation process, the feature extraction of the voxelized convolutional network can provide certain prior information for the point cloud feature learning network, speeding up the calculation speed; effectively combining the deep and shallow layers of the neural network to reduce the impact of background interference and scale.
[0080] For example, for multi-model fusion, since the current voxelized convolutional network model loses a lot of detail information in the process of extracting 3D point cloud information features, which will affect the judgment of the object's position frame and category, the embodiment of the present application combines the point cloud feature learning network (PointNet) to make up for the loss of the current feature extraction process, which can effectively improve the representation ability of 3D objects. At the same time, although the point cloud feature learning network has a high amount of computation and complexity, only using part of the network can keep the computation low and meet the corresponding speed requirements. As follows Figure 2 As shown in the figure, this is the process of feature fusion of the two models. The feature extraction of the voxelized convolutional network can provide certain prior information for the point cloud feature learning network and also speed up the calculation speed.
[0081] Exemplarily, for multi-scale feature fusion, according to the concept of neural networks, shallow neural networks can obtain geometric information such as texture, color and shape of objects. As the network layer deepens, the background information of the image is gradually screened out. In the deep network, the abstract information and semantic information of the object can be obtained. At this time, the classification of the object and the prediction of the position box can be completed. However, for 3D objects, if the network layer is too deep, it is easy to cause the loss of detail information. Especially for point cloud data, a network that is too deep will affect the subsequent prediction of the position box. At the same time, there are more scale changes in 3D data, and the network is not sensitive to the scale changes of objects. Therefore, the scale changes of objects in the image space are likely to cause a significant increase in the false detection rate. Therefore, the embodiment of the present application intends to propose a cross-layer feature fusion method suitable for 3D objects based on the multi-scale feature extraction method that is currently a hot topic in research, effectively fuse the shallow feature information with the deep features of the network, and explore the fusion method to achieve the optimal 3D object target feature extraction and expression, and lay the foundation for the early feature modeling for subsequent target detection tasks. As follows Figure 3 This is an example diagram of the cross-layer multi-scale feature fusion method to be adopted in the embodiments of the present application.
[0082] In order to solve the practical limitation of deep learning-based three-dimensional detection algorithms, a fast detection framework based on candidate boxes is proposed. The speed is improved while ensuring the detection accuracy, and the detection accuracy in harsh environmental conditions is improved. 3D object detection requires the use of object feature information of different scales. Since 3D point cloud data lacks information such as color and spatial structure, it is necessary to combine local object features with global features for analysis. Therefore, in the process of modeling the detection algorithm, in order to better judge the location area and semantic category information of candidate objects based on feature information, it will directly affect the performance of the overall method. For 3D object detection models, the essence lies in establishing a mapping relationship between 3D feature information and object semantic categories and positions. However, the complexity of the mapping relationship is related to the detection accuracy and speed of the model. The key technology of the 3D object detection model involved in the embodiment of the present application is the linear regression and semantic category discrimination of the position box, which can not only ensure the introduction of feature information of different scales in the modeling process, but also ensure that the entire model has a clear linear function calculation method (position box regression calculation). Therefore, in order to solve the problem of balancing accuracy and speed in detection tasks, the embodiment of the present application can not only apply the 3D object detection model to actual products, but also provide a solution for the accelerated computing strategy of the deep model.
[0083] As described above, the embodiment of the present application obtains three-dimensional point cloud data of the driving scene and preprocesses the three-dimensional point cloud data; establishes a training data set based on the preprocessed three-dimensional point cloud data in combination with an autonomous driving scene target detection database; trains and verifies the training data set to obtain a feature extraction network model; performs target detection and feature extraction on the three-dimensional point cloud data through the fusion of multiple feature extraction network models and in combination with multi-scale feature fusion; proposes a network feature fusion method to enhance the ability to extract detailed features of the target, which helps the network learn more discriminative features, and can refine the position prediction accuracy and improve the overall detection effect.
[0084] On the basis of the above embodiments, the embodiments of the present application may also be embodied as follows:
[0085] Conduct field scenario tests on the feature extraction network model, receive feedback results, and adjust the feature extraction network model based on the feedback results.
[0086] In the above implementation process, the feature extraction network model is tested in field scenarios, the model is verified, and the method is optimized and adjusted based on the feedback results.
[0087] For example, regarding the core technology of target detection, on the one hand, there are a variety of ways to utilize data, including methods based on point cloud data, multi-sensor fusion (such as lidar cameras, depth cameras, structured light scanners, ToF cameras, etc.), visual laser fusion, and visual depth cognition to obtain and process data, providing rich information for target detection. For example, the uniqueness of point cloud data is used for target detection, or the advantages of multiple sensor data are integrated to complement each other to improve detection accuracy. On the other hand, there are advanced network architectures and algorithms, involving deep learning architectures such as PointNet, PointNet++, convolutional neural networks (CNNs), Transformer networks, and graph convolutions, as well as methods combined with traditional machine learning methods (such as SVM, decision trees). These architectures and algorithms are used for feature extraction, pose estimation and other links, such as three-dimensional target detection based on graph convolution to achieve visual laser fusion, and using deep learning to learn pose estimation directly from data.
[0088] For example, regarding the integration with autonomous driving, in terms of target detection related to intelligent driving, the combination of three-dimensional target detection and intelligent driving methods is emphasized, and the target detection results are applied to the driving strategy of autonomous driving to ensure the safety and efficiency of autonomous driving. For example, after detecting the target, the vehicle can adjust the driving speed and direction according to the target information; in terms of the evaluation and guarantee of the capabilities of the autonomous driving system, it includes the autonomous driving capability detection method, the construction of the driving task test scenario library, and the autonomous driving simulation test method. By testing the autonomous driving system in different scenarios, quantitatively evaluating its capabilities, and simulating various situations in the simulation environment, the reliability and stability of the autonomous driving system in actual operation are ensured.
[0089] For example, regarding the support of related devices and equipment, in terms of target detection devices and equipment, it covers hardware such as devices, controllers, electronic devices, and media for storing and running related algorithms corresponding to various target detection methods. These hardware and media constitute the physical basis for the implementation of target detection technology on autonomous driving vehicles, ensuring the efficient operation of detection algorithms; in terms of software platforms and hardware acceleration technologies, it includes operating systems such as Ubuntu and ROS, robot development frameworks such as MoveIt and RoboticsMiddleware, and hardware acceleration technologies such as FPGA and dedicated AI chips. Appropriate software platforms and hardware acceleration can improve computing efficiency and optimize the overall performance of target detection and autonomous driving systems.
[0090] For example, regarding data processing and model training optimization, in addition to the farthest point sampling, random sampling consistency, Voxel Grid filter and other downsampling methods are also used in data processing and downsampling technology. At the same time, data enhancement techniques such as rotation, scaling, and shearing are used to increase the diversity of data sets, reduce overfitting, and improve data quality and model generalization capabilities. In terms of target detection model training, attention is paid to the training methods and devices of target detection models. By optimizing the training process, such as adjusting parameters and combining pseudo-labeling information, the model can better adapt to the complex and changeable target characteristics and environmental conditions in autonomous driving scenarios, thereby improving target detection performance.
[0091] Based on the theory of deep learning, the embodiments of the present application propose a more robust three-dimensional object detection method, and lightweight the model to achieve the ultimate practical application; in view of the problem of difficulty in processing and acquiring three-dimensional point cloud data, a more suitable autonomous driving database is constructed, and a more reasonable data sampling and training method is proposed to ensure processing efficiency and improve the effect of the deep learning detection model; in view of the current problem of difficulty in representing three-dimensional data, a network feature fusion method is proposed to enhance the ability to extract detailed features of the target and improve the overall detection effect; in view of the practical limitations of three-dimensional detection algorithms based on deep learning, a fast detection framework based on candidate boxes is proposed, and the speed is improved while ensuring detection accuracy.
[0092] The above steps are not to be performed in a strict order as described in the numbers, but should be understood as an overall solution.
[0093] In the second aspect, based on the above embodiments, Figure 4 This is a schematic diagram of the structure of a target detection device based on a driving scene provided in an embodiment of the present application. Figure 4 The target detection device based on driving scenes provided in this embodiment specifically includes: a data acquisition module 401, a data processing module 402, a model training module 403 and a target detection module 404.
[0094] Among them, the data acquisition module 401 is used to acquire the three-dimensional point cloud data of the driving scene and pre-process the three-dimensional point cloud data; the data processing module 402 is used to establish a training data set based on the pre-processed three-dimensional point cloud data in combination with the automatic driving scene target detection database; the model training module 403 is used to train and verify the training data set to obtain a feature extraction network model; the target detection module 404 is used to perform target detection and feature extraction on the three-dimensional point cloud data through the fusion of multiple feature extraction network models combined with multi-scale feature fusion.
[0095] As described above, the embodiment of the present application obtains three-dimensional point cloud data of the driving scene and preprocesses the three-dimensional point cloud data; establishes a training data set based on the preprocessed three-dimensional point cloud data in combination with an autonomous driving scene target detection database; trains and verifies the training data set to obtain a feature extraction network model; performs target detection and feature extraction on the three-dimensional point cloud data through the fusion of multiple feature extraction network models and in combination with multi-scale feature fusion; proposes a network feature fusion method to enhance the ability to extract detailed features of the target, which helps the network learn more discriminative features, and can refine the position prediction accuracy and improve the overall detection effect.
[0096] The driving scene-based target detection device provided in the embodiment of the present application can be used to execute the driving scene-based target detection method provided in the above embodiment, and has corresponding functions and beneficial effects.
[0097] In a third aspect, an embodiment of the present application further provides an electronic device, which can integrate the target detection device based on driving scenarios provided in an embodiment of the present application. Figure 5 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Figure 5 The electronic device includes: an input device 43, an output device 44, a memory 42 and one or more processors 41; the memory 42 is used to store one or more programs; when the one or more programs are executed by the one or more processors 41, the one or more processors 41 implement the target detection method based on the driving scene provided in the above embodiment. The input device 43, the output device 44, the memory 42 and the processor 41 can be connected by a bus or other means. Figure 5 The example of connecting through bus is taken in the following.
[0098] The processor 41 executes various functional applications and data processing of the device by running the software programs, instructions and modules stored in the memory 42, that is, realizes the above-mentioned target detection method based on driving scenarios.
[0099] The electronic device provided above can be used to execute the target detection method based on driving scenarios provided in the above embodiments, and has corresponding functions and beneficial effects.
[0100] In a fourth aspect, an embodiment of the present application also provides a computer-readable storage medium, which includes a stored computer program; wherein, when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the target detection method based on the driving scene as described above, and can achieve the same beneficial effects.
[0101] Of course, the storage medium containing computer executable instructions provided in an embodiment of the present application is not limited to the target detection method based on the driving scene as described above, and can also execute related operations in the target detection method based on the driving scene provided in any embodiment of the present application.
[0102] In a fifth aspect, the embodiments of the present application also provide a computer program product, and the methods described in the various embodiments of the present application can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instruction is loaded and executed on a computer, the processes or functions described in the various embodiments of the present application are executed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device, a core network device, an OAM (OpenAplicationModel, open application model) or other programmable device.
[0103] The computer program or instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer program or instructions may be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired or wireless means. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it may also be an optical medium, such as a digital video disk; it may also be a semiconductor medium, such as a solid-state drive. The computer-readable storage medium may be a volatile or non-volatile storage medium, or may include both volatile and non-volatile types of storage media.
[0104] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.
[0105] In addition, the functional modules in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.
[0106] If the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions for an electronic device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0107] The above description is only an embodiment of the present application and is not intended to limit the scope of protection of the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application. It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings.
[0108] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0109] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
Claims
1. A target detection method based on driving scenes, characterized in that: include: Acquire three-dimensional point cloud data of the driving scene and pre-process the three-dimensional point cloud data; Based on the pre-processed 3D point cloud data and the autonomous driving scene object detection database, a training data set is established; Train and verify the training data set to obtain a feature extraction network model; By fusing multiple feature extraction network models and combining multi-scale feature fusion, target detection and feature extraction are performed on three-dimensional point cloud data.
2. The target detection method based on driving scene according to claim 1, characterized in that: The step of obtaining three-dimensional point cloud data of a driving scene and preprocessing the three-dimensional point cloud data includes: The three-dimensional point cloud data of the driving scene is obtained, and the farthest point sampling technology is used to perform preliminary sampling on the three-dimensional point cloud data.
3. The target detection method based on driving scene according to claim 1, characterized in that: The method of performing target detection and feature extraction on three-dimensional point cloud data by fusing multiple feature extraction network models and combining multi-scale feature fusion includes: The voxelized convolutional network model is fused with the point cloud feature learning network model to extract key point features of 3D point cloud data, and feature fusion is performed by combining multi-scale feature fusion.
4. The target detection method based on driving scene according to claim 3, characterized in that: The method uses a voxelized convolutional network model to fuse a point cloud feature learning network model, extracts key point features from three-dimensional point cloud data, and combines multi-scale feature fusion to perform feature fusion, including: Processing the three-dimensional point cloud data to obtain voxelized point cloud data; Train the voxelized point cloud data to obtain a voxelized convolutional network model; The convolution features trained by the voxelized convolutional network model are processed by the point cloud feature learning network model and combined with multi-scale feature fusion to extract key point features; The key point features are fused based on the voxelized convolutional network model, and the shallow feature information is fused with the deep features of the deep network.
5. The target detection method based on driving scene according to claim 1, characterized in that: The training data set is trained and verified to obtain a feature extraction network model, including: Using the single-stage detection technology as the basic framework, the feature extraction network model is obtained by improving the position frame generation method, using the prior 3D object position information as an auxiliary, and adopting the multi-scale target prediction method to train the training data set; The training dataset is validated to generate candidate location boxes at multiple scales.
6. The target detection method based on driving scene according to claim 1, characterized in that: Also includes: Conduct field scenario tests on the feature extraction network model, receive feedback results, and adjust the feature extraction network model based on the feedback results.
7. A target detection device based on a driving scene, characterized in that: include: A data acquisition module, used to acquire three-dimensional point cloud data of the driving scene and pre-process the three-dimensional point cloud data; A data processing module is used to establish a training data set based on the pre-processed 3D point cloud data and the autonomous driving scene object detection database; Model training module, used to train and verify the training data set to obtain a feature extraction network model; The target detection module is used to perform target detection and feature extraction on 3D point cloud data by fusing multiple feature extraction network models and combining multi-scale feature fusion.
8. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the processor is connected to the memory via the bus, and the memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the method for target detection based on a driving scene is implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a server, implements the target detection method based on driving scenarios as described in any one of claims 1 to 6.
10. A computer program product, characterized in that The computer program product comprises instructions, which, when executed by a computer, enable the computer to implement the target detection method based on driving scenarios according to any one of claims 1 to 6.
Citation Information
Patent Citations
Three-dimensional dynamic target detection method and device based on voxel point cloud fusion
CN113989797A
Point cloud target detection method and device, equipment and storage medium
CN115082885A
Semantic scene completion method based on image and point cloud fusion in automatic driving scene
CN116503825A
Far small target point cloud data feature enhancement method based on attention mechanism
CN116824533A
Automatic driving data fusion method and device, equipment and medium
CN117152574A
Cited By
Performance test system and equipment of fan grating sensor and storage medium
CN120740652A