Detection Method and Device for Road Object, Storage Medium and Electronic Device

By using the target point cloud model to identify the first recognition result of the target training sample, the initial camera model is obtained, and the target camera model is solved, and the problem of low road object detection accuracy in autonomous driving technology is achieved, achieving higher detection accuracy and robustness.

CN115527186BActive Publication Date: 2025-06-13INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211202848.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-29
Publication Date
2025-06-13
Estimated Expiration
2042-09-29

AI Technical Summary

Technical Problem

In the existing autonomous driving technology, the detection accuracy of road objects is low, especially in severe weather and darker light scenes, which affects the accuracy of 3D object detection.

Method used

By obtaining the first recognition result obtained by identifying the target training sample by the target point cloud model, the initial imaging model is trained using the target training sample marked with the first recognition result, and the target imaging model is obtained. Compared with the initial camera model, the target imaging model has a higher accuracy on the target road object that is displayed in the target image data that carries the target spatial information.

Benefits of technology

Improve the detection accuracy of road objects and enhance the detection stability and robustness in severe weather and dark light scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115527186B_ABST
    Figure CN115527186B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a method and apparatus for detecting road objects, a storage medium, and an electronic device. The method for detecting road objects includes: obtaining a first recognition result obtained by identifying a target training sample from a target point cloud model; training an initial camera model using the target training sample labeled with the first recognition result to obtain a target camera model; and detecting a target road object carrying target spatial information shown in target image data through the target camera model. Through the present application, the problem of low detection accuracy of road objects is solved, and thus the technical effect of improving the detection accuracy of road objects is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of autonomous driving technology. Specifically, the embodiments relate to a method and apparatus for detecting road objects, a storage medium, and an electronic device. Background Art

[0002] In the past few years, autonomous driving technology has developed rapidly. However, due to the complex and dynamic driving environment, achieving fully autonomous driving remains a daunting task. To understand the surrounding driving environment, autonomous vehicles are equipped with a set of sensors for powerful and accurate environmental perception. This set of sensor devices and their supporting processing algorithms are called the perception system. In autonomous driving technology, the perception system takes the data from a group of sensors as input, and after a series of processing steps, outputs information about the environment, other objects (such as cars), and the autonomous vehicle itself. 3D object detection is an important task of the perception system, whose purpose is to identify all objects of interest in the sensor data and determine their positions and categories (such as vehicles, bicycles, pedestrians, etc.). In the 3D object detection task, output parameters are required to specify the 3D-oriented bounding boxes around the objects. Therefore, the perception system has the following basic requirements: First, it needs to be accurate and give an accurate description of the driving environment. Second, it should be robust and ensure the stability and safety of AV (Audio Video) even in bad weather or when some sensors degrade or even fail.

[0003] In a real autonomous driving scenario, when only using a camera as a sensor, the sensor data collected by the camera is easily interfered by the external environment. For example, when an object is occluded, the sensor data collected by the camera will affect the subsequent 3D object detection. In addition, the light conditions in the environment will also affect the sensor data collected by the camera. Especially in a dimly lit scene, it will have a great impact on the work of the sensor data collected by the camera, and the accuracy of the subsequent 3D object detection will also be reduced. At the same time, since the sensor data collected by the camera is only two-dimensional picture data, using two-dimensional picture data for 3D object detection may result in inaccurate spatial depth prediction in the detection results. Summary of the Invention

[0004] The embodiments of the present application provide a method and apparatus for detecting road objects, a storage medium, and an electronic device, so as to at least solve the problem of low detection accuracy of road objects in the related art.

[0005] According to an embodiment of the present application, a method for detecting road objects is provided, including:

[0006] Obtain a first recognition result obtained by recognizing a target training sample using a target point cloud model, where the target point cloud model is used to recognize a road object carrying spatial information from input point cloud data, and the first recognition result is used to represent a reference road object carrying reference spatial information shown in the target training sample, and the reference spatial information is used to indicate the position and motion state of the reference road object in space;

[0007] Use the target training sample annotated with the first recognition result to train an initial camera model to obtain a target camera model;

[0008] Detect a target road object carrying target spatial information shown in the target image data through the target camera model.

[0009] Optionally, the using the target training sample annotated with the first recognition result to train an initial camera model to obtain a target camera model includes:

[0010] Obtain a second recognition result obtained by the initial camera model recognizing the target training sample;

[0011] Adjust the model parameters of the target network layer in the initial camera model according to the difference degree between the first recognition result and the second recognition result until the difference degree between the second recognition result and the first recognition result is less than a difference degree threshold to obtain the target camera model.

[0012] Optionally, the obtaining a second recognition result obtained by the initial camera model recognizing the target training sample includes:

[0013] Take pictures of three-dimensional scene data to obtain an image data sample, where the target training sample includes the three-dimensional scene data;

[0014] Input the image data sample into the initial camera model to obtain the second recognition result output by the initial camera model.

[0015] Optionally, the adjusting the model parameters of the target network layer in the initial camera model according to the difference degree between the first recognition result and the second recognition result includes:

[0016] Calculate the KL divergence between the second recognition result and the first recognition result, where the difference degree includes the KL divergence;

[0017] Adjust the model parameters of the target network layer in the initial camera model according to the KL divergence, where when the KL divergence is less than a divergence threshold, it is determined that the difference degree between the second recognition result and the first recognition result is less than the difference degree threshold.

[0018] Optionally, adjusting the model parameters of the target network layer in the initial camera model according to the difference degree between the first recognition result and the second recognition result includes:

[0019] Obtain the target network layer from one or more network layers included in the initial camera model, where the target network layer includes at least one of the following: an aerial view feature extraction layer, an aerial view feature decoding layer, and an output layer. The aerial view feature extraction layer is used to extract features from the input aerial view data, the aerial view feature decoding layer is used to decode the input aerial view features, and the output layer is used to obtain a recognition result according to the decoding result of the aerial view features;

[0020] Calculate the difference degree between the first recognition result and the second recognition result;

[0021] Adjust the model parameters of the target network layer according to the difference degree.

[0022] Optionally, obtaining the first recognition result obtained by the target point cloud model for recognizing the target training sample includes:

[0023] Scan the three-dimensional scene data to obtain a first point cloud data sample, where the target training sample includes the three-dimensional scene data;

[0024] Input the first point cloud data sample into the target point cloud model to obtain the first recognition result output by the target point cloud model.

[0025] Optionally, before obtaining the first recognition result obtained by the target point cloud model for recognizing the target training sample, the method further includes:

[0026] Obtain a second point cloud data sample labeled with a sample label, where the sample label includes a sample road object carrying sample space information, and the second point cloud data sample is obtained by scanning with a radar device;

[0027] Use the second point cloud data sample to train the initial point cloud model to obtain the target point cloud model.

[0028] According to another embodiment of the present application, there is provided a detection device for road objects, including:

[0029] A first acquisition module, configured to acquire a first recognition result obtained by using a target point cloud model to recognize a target training sample, where the target point cloud model is used to recognize a road object carrying spatial information from input point cloud data, the first recognition result is used to represent a reference road object carrying reference spatial information shown in the target training sample, and the reference spatial information is used to indicate the position and motion state of the reference road object in space;

[0030] A first training module, configured to use the target training sample annotated with the first recognition result to train an initial camera model to obtain a target camera model;

[0031] A detection module, configured to detect a target road object carrying target spatial information shown in target image data through the target camera model.

[0032] According to another embodiment of the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, where the computer program is configured to execute the steps in any one of the above method embodiments when running.

[0033] According to another embodiment of the present application, there is also provided an electronic device, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0034] Through the present application, since the first recognition result is obtained by using a target point cloud model to recognize a target training sample, and because the target point cloud model can recognize a road object carrying spatial information from input point cloud data, the first recognition result can represent a reference road object carrying reference spatial information shown in the target training sample, and the reference spatial information can indicate the position and motion state of the reference road object in space. Then, using the target training sample annotated with the first recognition result to train an initial camera model to obtain a target camera model, the target camera model has a higher accuracy in detecting a target road object carrying target spatial information shown in target image data compared to the initial camera model. Therefore, the problem of low detection accuracy of road objects can be solved, and the technical effect of improving the detection accuracy of road objects can be achieved. Description of the Drawings

[0035] Figure 1 is a hardware structure block diagram of a mobile terminal for a method for detecting a road object according to an embodiment of the present application;

[0036] Figure 2 is a flowchart of a method for detecting a road object according to an embodiment of the present application;

[0037] Figure 3 It is a schematic diagram of training a target camera model according to an embodiment of the present application;

[0038] Figure 4 It is a schematic diagram of a detection process of a road object according to an embodiment of the present application;

[0039] Figure 5 It is a structural block diagram of a detection device for road objects according to an embodiment of the present application. Detailed implementation manners

[0040] In the following, embodiments of the present application will be described in detail with reference to the drawings and in combination with embodiments.

[0041] It should be noted that the terms "first", "second", etc. in the description and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence.

[0042] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 It is a hardware structural block diagram of a mobile terminal for a detection method of a road object according to an embodiment of the present application. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1 only one is shown in Figure 1 a processor 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above-mentioned mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown in Figure 1 is only schematic and does not limit the structure of the above-mentioned mobile terminal. For example, the mobile terminal may further include more or fewer components than those shown in

[0043] The memory 104 can be used to store computer programs, such as software programs and modules of application software, like the computer program corresponding to the detection method of road objects in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, to implement the above method. The memory 104 may include a high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the mobile terminal through a network. Examples of the above network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.

[0044] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include the wireless network provided by the communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0045] In this embodiment, a detection method of road objects is provided, which is applied to the above computer terminal. Figure 2 It is a flowchart of a detection method of road objects according to an embodiment of the present application, as Figure 2 shown. The process includes the following steps:

[0046] Step S202, obtaining a first recognition result obtained by identifying a target training sample from a target point cloud model, where the target point cloud model is used to identify road objects carrying spatial information from the input point cloud data, and the first recognition result is used to characterize the reference road objects carrying reference spatial information shown in the target training sample, and the reference spatial information is used to indicate the position and motion state of the reference road objects in space;

[0047] Step S204, training an initial camera model with the target training sample labeled with the first recognition result to obtain a target camera model;

[0048] Step S206, detecting a target road object carrying target spatial information shown in the target image data through the target camera model.

[0049] Through the above steps, the first recognition result is obtained by recognizing the target training sample through the target point cloud model. Since the target point cloud model can identify road objects carrying spatial information from the input point cloud data, the first recognition result can represent the reference road object carrying the reference spatial information shown in the target training sample. The reference spatial information can indicate the position and motion state of the reference road object in space. Then, the initial camera model is trained using the target training sample labeled with the first recognition result to obtain the target camera model. The target camera model has higher accuracy in detecting the target road object carrying the target spatial information shown in the target image data compared to the initial camera model. By adopting the above technical solution, the problems such as low detection accuracy of road objects in the related art are solved, and the technical effect of improving the detection accuracy of road objects is achieved.

[0050] In the technical solution provided in the above step S202, the target point cloud model is a model that allows identifying road objects carrying spatial information from the input point cloud data, where the point cloud data can be, but is not limited to, the scanned point set data obtained by lidar scanning.

[0051] Optionally, in this embodiment, the spatial information is used to indicate the position and motion state of the reference road object in space, where the position in space can include, but is not limited to, depth information and height information; the motion state can include, but is not limited to, velocity information and rotation angle information.

[0052] Optionally, in this embodiment, the road object can include, but is not limited to, all types of objects that may appear on the road, such as, but not limited to, vehicle types, pedestrian types, pet types, and building types, etc. Among them, each type can be further divided into more precise sub - categories, which can be divided according to actual application requirements without limitation.

[0053] In the technical solution provided in the above step S204, Figure 3 is a schematic diagram of training a target camera model according to an embodiment of the present application, as Figure 3As shown in the figure, the training of the target camera model is divided into two branches, including the target point cloud model branch and the target camera model branch. The target point cloud model in the lower branch of the figure obtains Lidar Point (the above-mentioned point cloud data). First, Lidar Point is voxelized by Voxelization, and then the data output by Voxelization is feature-encoded by Encoder to obtain encoded features. The encoded features are input into the Bird's Eye View Feature (BEV Feature) layer to obtain the bird's eye view feature output by BEV Feature. Further, the bird's eye view feature is input into the BEV Decoder of the feature decoding layer to obtain the decoded features output by BEV Decoder. Finally, the decoded features are input into the 3D detection head Head for recognition to obtain the first recognition result. Further, in the target point cloud model branch, the image features output by the corresponding layer of the initial camera model are supervised and trained at several places such as the BEV Feature layer, the BEV Decoder of the feature decoding layer, and the 3D detection head Head, respectively, to obtain the target camera model. That is to say, the target point cloud model branch only participates in the training and does not participate in the inference.

[0054] Optionally, in this embodiment, as Figure 3 shown, in the target camera model branch, the target feature encoder Encoder is used to extract features from the image data (Image) corresponding to the target training samples. The result output by Encoder enters the feature geometric transformation (Transformer), and the two-dimensional planar features are obtained through feature geometric transformation (Transformer) according to the internal and external camera parameters. Further, the two-dimensional planar features are input into the Bird's Eye View Feature (BEV Feature) layer to obtain the bird's eye view feature output by BEV Feature.

[0055] Optionally, in this embodiment, after the bird's eye view feature output by BEV Feature, the bird's eye view feature is input into the BEV Decoder of the feature decoding layer. The BEV Decoder of the feature decoding layer is used to decode the bird's eye view feature to obtain the decoded features. The above-mentioned decoded features can be regarded as a further extraction of the bird's eye view feature and are the embodiment of the bird's eye view feature. The bird's eye view feature may be fragmented features, and the decoded features are further integrated on the basis of the bird's eye view feature. Finally, the decoded features are recognized by the 3D detection head Head to obtain the classification result.

[0056] In an exemplary embodiment, the target training sample labeled with the first recognition result can be used, but not limited to, the following method to train the initial camera model to obtain the target camera model: obtain the second recognition result obtained by the initial camera model for the target training sample; adjust the model parameters of the target network layer in the initial camera model according to the difference degree between the first recognition result and the second recognition result until the difference degree between the second recognition result and the first recognition result is less than the difference degree threshold, and obtain the target camera model.

[0057] Optionally, in this embodiment, the initial camera model can be, but not limited to, a multi-camera fusion network model architecture, such as Figure 3 As shown, at the beginning stage, the initial camera model extracts features from the image data (Image) corresponding to the target training sample through the feature encoding Encoder, and then performs feature geometric transformation (Transformer) according to the internal and external camera parameters to obtain the bird's-eye view features. Instead of the original single-camera single-image feature extraction. On the one hand, multi-camera fusion can splice the features of the truncated targets in the edge area of the picture to form complete target features; on the other hand, through the fusion of multi-angle features, it is beneficial to more accurate feature expression. Then design a BEV Decoder based on the ResNet residual network architecture, and finally perform further feature parsing through a 3D detection head (Head), and the output result is used for classification and detection loss calculation.

[0058] In an exemplary embodiment, the second recognition result obtained by the initial camera model for the target training sample can be obtained, but not limited to, the following method: shoot the three-dimensional scene data to obtain an image data sample, where the target training sample includes the three-dimensional scene data; input the image data sample into the initial camera model to obtain the second recognition result output by the initial camera model.

[0059] Optionally, in this embodiment, the three-dimensional scene data can be, but not limited to, scene data containing three-dimensional features. Shooting the three-dimensional scene data to obtain an image data sample can be, but not limited to, using multiple cameras to shoot the three-dimensional scene data, replacing the original single-camera single-image feature extraction. Multi-camera fusion can splice the features of the truncated targets in the edge area of the picture to form a complete image data sample.

[0060] In an exemplary embodiment, the model parameters of the target network layer in the initial camera model can be adjusted according to the difference degree between the first recognition result and the second recognition result in the following ways, but not limited to: calculating the Kullback-Leibler (KL) divergence between the second recognition result and the first recognition result, where the difference degree includes the KL divergence; adjusting the model parameters of the target network layer in the initial camera model according to the KL divergence, where when the KL divergence is less than the divergence threshold, it is determined that the difference degree between the second recognition result and the first recognition result is less than the difference degree threshold.

[0061] Optionally, in this embodiment, as Figure 3 shown, the target point cloud model branch can adopt point cloud voxelization, 3D sparse convolution, and then dimensionality reduction to generate 2D bird's-eye view features. Through the point cloud feature supervision architecture of the target point cloud model branch, feature supervision is performed on the target camera model branch at three places: the bird's-eye view feature layer BEV Feature, the feature decoding layer BEV Decoder, and the 3D detection head Head. The Loss design mainly considers: KL divergence. Further, in the BEV feature, the BEV feature (1*256*128*128) of the target point cloud model branch is used to supervise the BEV feature (1*256*128*128) of the target camera model branch to make their feature distributions consistent, so that the high-precision features of the Lidar Point (target point cloud model) can guide the target camera model to improve the feature extraction accuracy. In addition to improving the depth information of traditional algorithms, there are also significant improvements in the recognition accuracy of height information, velocity information, rotation angle information, etc.; similarly, the BEV decoder feature (1*256*128*128), Head feature (1*1*128*128), etc. can also be supervised, guided, and corrected. Among them, the KL divergence can be calculated using the following formula, but not limited to:

[0062] L(y cam , y lidar ) = y cam *(logy lidar - logy cam )

[0063] where L(y cam , y lidar ) is the KL divergence value, y cam is the output result probability value corresponding to the target camera model, and y lidar is the output result probability value corresponding to the target point cloud model.

[0064] In an exemplary embodiment, the model parameters of the target network layer in the initial camera model can be adjusted according to the degree of difference between the first recognition result and the second recognition result in the following ways, but not limited thereto: obtain the target network layer from one or more network layers included in the initial camera model, wherein the target network layer includes at least one of the following: a bird's-eye view feature extraction layer, a bird's-eye view feature decoding layer, and an output layer. The bird's-eye view feature extraction layer is used to extract features from the input bird's-eye view data, the bird's-eye view feature decoding layer is used to decode the input bird's-eye view features, and the output layer is used to obtain a recognition result according to the decoding result of the bird's-eye view features; calculate the degree of difference between the first recognition result and the second recognition result; adjust the model parameters of the target network layer according to the degree of difference.

[0065] Optionally, in this embodiment, as Figure 3 shown, the bird's-eye view feature extraction layer can be, but not limited to, the above-mentioned bird's-eye view feature layer BEV Feature, the bird's-eye view feature decoding layer can be, but not limited to, the above-mentioned feature decoding layer BEV Decoder, the output layer can be, but not limited to, the above-mentioned 3D detection head Head, and the above-mentioned degree of difference can include, but not limited to, the bird's-eye view feature difference degree BEV Feature loss, the feature decoding difference degree BEV Decoder loss, and the 3D detection head difference degree BEV Headloss, etc.

[0066] Optionally, in this embodiment, the degree of difference of the corresponding network layer is indicated by parameters such as BEV Feature loss, BEV Decoder loss, and BEV Headloss, etc., and the model parameters of the target camera model are further adjusted, that is, the initial camera model is trained to obtain the target camera model. The degree of difference between the output results of the corresponding network layers of the target camera model and the target point cloud model needs to be less than a preset value, that is, to ensure that the output results of the corresponding network layers of the target camera model and the target point cloud model are similar.

[0067] In an exemplary embodiment, the first recognition result obtained by the target point cloud model for recognizing the target training sample can be obtained in the following ways, but not limited thereto: scan the three-dimensional scene data to obtain a first point cloud data sample, wherein the target training sample includes the three-dimensional scene data; input the first point cloud data sample into the target point cloud model to obtain the first recognition result output by the target point cloud model.

[0068] Optionally, in this embodiment, the three-dimensional scene data can be scanned by, but not limited to, using a lidar to obtain a set of scan points corresponding to the three-dimensional scene data as the first point cloud data sample. For example, in a road scene, there may be vehicle objects, pedestrian objects, pet objects of pedestrian objects, building objects, etc. All the above objects constitute the road scene. A large number of detection rays are emitted by the lidar, and a large number of detection rays hit various positions in the road scene to form a set of scan points, and the corresponding position information and motion information are returned. The set of scan points can be used as the first point cloud data sample.

[0069] In the technical solution provided in the above step S206, before obtaining the first recognition result by using the target point cloud model to recognize the target training sample, it may also include, but not limited to, the following methods: obtaining a second point cloud data sample labeled with a sample label, where the sample label includes a sample road object carrying sample space information, and the second point cloud data sample is obtained by scanning with a radar device; using the second point cloud data sample to train an initial point cloud model to obtain the target point cloud model.

[0070] Optionally, in this embodiment, the process of training the initial point cloud model to obtain the target point cloud model may include, but not limited to, training with the second point cloud data sample for 20 epochs (the process of sending data into the network and completing one forward calculation + backward propagation process), and then freezing the target point cloud model, only performing forward inference and not participating in backward propagation.

[0071] To better understand the above process of detecting road objects, the following further describes the detection process of road objects in combination with optional embodiments, but it is not used to limit the technical solutions of the embodiments of the present application.

[0072] In this embodiment, a method for detecting road objects is provided. Figure 4 It is a schematic diagram of a detection process of road objects according to an embodiment of the present application, as Figure 4 shown, mainly including the following steps:

[0073] Step S401: Design the fusion network architecture of the point cloud supervised camera;

[0074] Step S402: Design the optimization of the feature supervised loss;

[0075] Step S403: Perform joint data training of Lidar (target point cloud model) and Camera (target camera model) for 20 epochs, and then freeze the Lidar part of the model and continue training for 5 epochs;

[0076] In the inference stage, laser point cloud data does not need to be introduced, and the test accuracy is greatly improved.

[0077] Through the above embodiments, the first recognition result is obtained by recognizing the target training sample through the target point cloud model. Since the target point cloud model can identify the road object carrying spatial information from the input point cloud data, the first recognition result can represent the reference road object carrying the reference spatial information shown in the target training sample. The reference spatial information can indicate the position and motion state of the reference road object in space. Then, the initial camera model is trained using the target training sample annotated with the first recognition result to obtain the target camera model. The target camera model has a higher accuracy in detecting the target road object carrying the target spatial information shown in the target image data compared to the initial camera model. By adopting the above technical solution, the problems of low detection accuracy of road objects in the related art are solved, and the technical effect of improving the detection accuracy of road objects is achieved.

[0078] Through the above implementation manners, there are innovations in the overall model architecture training part. Generally, it is divided into two steps. The first step: overall iterative training, so that Lidar can quickly reach a stable state with high accuracy, and the image part can also reach a certain accuracy. The second step: freeze the model of the Lidar point cloud processing part, only perform forward inference, and do not participate in backpropagation; continue to train the image part for 3 - 5 epochs to achieve a higher detection accuracy.

[0079] In the embodiment of the present application, first, the initial point cloud model is trained until the output result of the initial point cloud model converges and is approximated to the label of the training sample to obtain the target point cloud model. Then, the target point cloud model is used to recognize the target training sample to obtain the first recognition result output by the target point cloud model. The first recognition result is used as a benchmark to train the initial camera model, that is, feature supervision is performed on the target camera model branch at three places: the Bird's Eye View Feature layer (BEV Feature), the BEV Decoder layer, and the 3D detection head (Head) until the result output by the initial camera model approaches the first recognition result, which can be regarded as the completion of the training of the initial camera model to obtain the target camera model. Then, the target camera model is used to detect the target road object carrying the target spatial information shown in the target image data. The detection result of the target camera model has a greater improvement in the recognition accuracy of height information, velocity information, rotation angle information, etc., in addition to the improvement in depth information.

[0080] The embodiment of the present application proposes a fusion network feature extraction architecture for supervising the features of a target camera model through the laser point cloud features of a target point cloud model; it improves the problem of inaccurate depth prediction in the previous target camera model, and directly through feature supervision, significantly improves the camera feature recognition accuracy of the target camera model. It has a certain degree of innovation in the field of 3D target detection of single-class camera sensors for autonomous driving. The embodiment of the present application mainly innovates and optimizes the 3D target detection algorithm of the autonomous driving camera (corresponding to the above-mentioned target camera model), and greatly improves the detection accuracy of the target camera model for 3D targets. That is, first, the image features are encoded, a multi-camera feature fusion architecture is designed, and then the bird's-eye view is generated; then, through the laser point cloud features, the image features output by the corresponding layers of the initial camera model are supervised and trained respectively at several places such as the bird's-eye view feature layer BEV Feature, the feature decoding layer BEV Decoder, and the 3D detection head Head, etc., to obtain the target camera model, which has higher detection accuracy and better robustness for 3D targets.

[0081] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application.

[0082] In this embodiment, a detection device for road objects is also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0083] Figure 5 is a structural block diagram of the detection device for road objects according to the embodiment of the present application. As Figure 5 shown, the device includes:

[0084] The first acquisition module 502 is configured to acquire a first recognition result obtained by using a target point cloud model to recognize a target training sample. The target point cloud model is used to recognize a road object carrying spatial information from the input point cloud data. The first recognition result is used to represent a reference road object carrying reference spatial information shown in the target training sample. The reference spatial information is used to indicate the position and motion state of the reference road object in space.

[0085] The first training module 504 is configured to train an initial camera model by using the target training sample annotated with the first recognition result to obtain a target camera model.

[0086] The detection module 506 is configured to detect a target road object carrying target spatial information shown in the target image data by using the target camera model.

[0087] Through the above embodiments, the first recognition result is obtained by using a target point cloud model to recognize a target training sample. Since the target point cloud model can recognize a road object carrying spatial information from the input point cloud data, the first recognition result can represent a reference road object carrying reference spatial information shown in the target training sample. The reference spatial information can indicate the position and motion state of the reference road object in space. Then, the initial camera model is trained by using the target training sample annotated with the first recognition result to obtain a target camera model. The target camera model has a higher accuracy in detecting a target road object carrying target spatial information shown in the target image data than the initial camera model. By adopting the above technical solution, the problem of low detection accuracy of road objects in the related art is solved, and the technical effect of improving the detection accuracy of road objects is achieved.

[0088] It should be noted that the above-mentioned modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to this: all the above-mentioned modules are located in the same processor; or, the above-mentioned modules are respectively located in different processors in any combination form.

[0089] In an exemplary embodiment, the first training module includes:

[0090] An acquisition unit configured to acquire a second recognition result obtained by using the initial camera model to recognize the target training sample;

[0091] An adjustment unit configured to adjust the model parameters of the target network layer in the initial camera model according to the difference degree between the first recognition result and the second recognition result until the difference degree between the second recognition result and the first recognition result is less than a difference degree threshold, so as to obtain the target camera model.

[0092] In an exemplary embodiment, the obtaining unit is further configured to:

[0093] Take pictures of the three-dimensional scene data to obtain an image data sample, where the target training sample includes the three-dimensional scene data;

[0094] Input the image data sample into the initial camera model to obtain the second recognition result output by the initial camera model.

[0095] In an exemplary embodiment, the adjusting unit is further configured to:

[0096] Calculate the KL divergence between the second recognition result and the first recognition result, where the difference degree includes the KL divergence;

[0097] Adjust the model parameters of the target network layer in the initial camera model according to the KL divergence. Wherein, when the KL divergence is less than the divergence threshold, it is determined that the difference degree between the second recognition result and the first recognition result is less than the difference degree threshold.

[0098] In an exemplary embodiment, the adjusting unit is further configured to:

[0099] Obtain the target network layer from one or more network layers included in the initial camera model, where the target network layer includes at least one of the following: an aerial view feature extraction layer, an aerial view feature decoding layer, and an output layer. The aerial view feature extraction layer is used to extract features from the input aerial view data, the aerial view feature decoding layer is used to decode the input aerial view features, and the output layer is used to obtain a recognition result according to the decoding result of the aerial view features;

[0100] Calculate the difference degree between the first recognition result and the second recognition result;

[0101] Adjust the model parameters of the target network layer according to the difference degree.

[0102] In an exemplary embodiment, the first obtaining module includes:

[0103] A scanning unit, configured to scan the three-dimensional scene data to obtain a first point cloud data sample, where the target training sample includes the three-dimensional scene data;

[0104] An input unit, configured to input the first point cloud data sample into the target point cloud model to obtain the first recognition result output by the target point cloud model.

[0105] In an exemplary embodiment, the device further includes:

[0106] A second acquisition module, configured to acquire a second point cloud data sample labeled with a sample label before the first recognition result obtained by recognizing a target training sample from the acquired target point cloud model, where the sample label includes a sample road object carrying sample space information, and the second point cloud data sample is obtained by scanning with a radar device;

[0107] A second training module, configured to train an initial point cloud model using the second point cloud data sample to obtain the target point cloud model.

[0108] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, where the computer program is configured to execute the steps in any one of the above method embodiments when running.

[0109] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disc that can store a computer program.

[0110] An embodiment of the present application further provides an electronic device, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0111] In an exemplary embodiment, the above electronic device may further include a transmission device and an input / output device, where the transmission device is connected to the above processor, and the input / output device is connected to the above processor.

[0112] Specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be repeated here.

[0113] Obviously, those skilled in the art should understand that the above modules or steps of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described may be executed in a different order than here, or they may be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them may be fabricated into a single integrated circuit module to implement. Thus, the present application is not limited to any specific combination of hardware and software.

[0114] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the principle of the present application shall be included within the protection scope of the present application.

Claims

1. A method for detecting road objects, characterized in that, it includes: Obtaining a first recognition result obtained by identifying a target training sample using a target point cloud model, wherein the target point cloud model is used to identify road objects carrying spatial information from the input point cloud data, and the first recognition result is used to represent a reference road object carrying reference spatial information shown in the target training sample, and the reference spatial information is used to indicate the position and motion state of the reference road object in space; Training an initial camera model using the target training sample annotated with the first recognition result to obtain a target camera model; Detecting a target road object carrying target spatial information shown in the target image data through the target camera model; Among them, the step of training the initial camera model using the target training sample annotated with the first recognition result to obtain a target camera model includes: obtaining a second recognition result obtained by the initial camera model identifying the target training sample; adjusting the model parameters of the target network layer in the initial camera model according to the difference degree between the first recognition result and the second recognition result until the difference degree between the second recognition result and the first recognition result is less than the difference degree threshold to obtain the target camera model.

2. The method according to claim 1, characterized in that, the step of obtaining the second recognition result obtained by the initial camera model identifying the target training sample includes: Taking pictures of three-dimensional scene data to obtain an image data sample, wherein the target training sample includes the three-dimensional scene data; Inputting the image data sample into the initial camera model to obtain the second recognition result output by the initial camera model.

3. The method according to claim 1, characterized in that, the step of adjusting the model parameters of the target network layer in the initial camera model according to the difference degree between the first recognition result and the second recognition result includes: Calculating the KL divergence between the second recognition result and the first recognition result, wherein the difference degree includes the KL divergence; Adjusting the model parameters of the target network layer in the initial camera model according to the KL divergence, wherein when the KL divergence is less than the divergence threshold, it is determined that the difference degree between the second recognition result and the first recognition result is less than the difference degree threshold.

4. The method according to claim 1, characterized in that, the step of adjusting the model parameters of the target network layer in the initial camera model according to the difference degree between the first recognition result and the second recognition result includes: Obtaining the target network layer from one or more network layers included in the initial camera model, wherein the target network layer includes at least one of the following: a bird's-eye view feature extraction layer, a bird's-eye view feature decoding layer, and an output layer. The bird's-eye view feature extraction layer is used to extract features from the input bird's-eye view data, the bird's-eye view feature decoding layer is used to decode the input bird's-eye view features, and the output layer is used to obtain a recognition result according to the decoding result of the bird's-eye view features; Calculate the difference degree between the first recognition result and the second recognition result; Adjust the model parameters of the target network layer according to the difference degree.

5. The method according to claim 1, wherein, the obtaining the first recognition result obtained by the target point cloud model recognizing the target training sample includes: scanning three-dimensional scene data to obtain a first point cloud data sample, wherein the target training sample includes the three-dimensional scene data; inputting the first point cloud data sample into the target point cloud model to obtain the first recognition result output by the target point cloud model.

6. The method according to claim 1, wherein, before obtaining the first recognition result obtained by the target point cloud model recognizing the target training sample, the method further includes: obtaining a second point cloud data sample labeled with a sample label, wherein the sample label includes a sample road object carrying sample space information, and the second point cloud data sample is obtained by scanning with a radar device; training an initial point cloud model with the second point cloud data sample to obtain the target point cloud model.

7. A detection device for road objects, wherein, comprising: a first acquisition module, configured to acquire a first recognition result obtained by a target point cloud model recognizing a target training sample, wherein the target point cloud model is configured to recognize a road object carrying spatial information from input point cloud data, and the first recognition result is used to represent a reference road object carrying reference spatial information shown in the target training sample, and the reference spatial information is used to indicate the position and motion state of the reference road object in space; a first training module, configured to train an initial camera model with the target training sample labeled with the first recognition result to obtain a target camera model; a detection module, configured to detect a target road object carrying target spatial information shown in target image data through the target camera model; wherein, the first training module includes: an acquisition unit, configured to acquire a second recognition result obtained by the initial camera model recognizing the target training sample; an adjustment unit, configured to adjust the model parameters of a target network layer in the initial camera model according to the difference degree between the first recognition result and the second recognition result until the difference degree between the second recognition result and the first recognition result is less than a difference degree threshold, to obtain the target camera model.

8. A computer-readable storage medium, wherein, the computer-readable storage medium includes a stored program, wherein the program, when running, executes the method according to any one of claims 1 to 6.

9. An electronic device, including a memory and a processor, wherein, a computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 6 through the computer program.

Citation Information

Patent Citations

  • Training method of image recognition model, and image recognition method and device

    CN113255444A

  • Automatic labeling, detection model training and target recognition methods, and electronic devices

    CN114937177A