Training method and detection method of target detection model, device, equipment and medium

By adjusting the loss weights based on the magnitude of the target object's influence on vehicle behavior during the target detection model training process, the problem of insufficient applicability of target detection models in intelligent driving scenarios is solved, thereby improving the accuracy and safety of vehicle environmental perception.

CN122116305APending Publication Date: 2026-05-29BEIJING VOYAGER TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING VOYAGER TECH CO LTD
Filing Date
2024-11-29
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies struggle to provide target detection models suitable for intelligent driving scenarios, resulting in inaccurate perception of the vehicle's surroundings and impacting driving safety.

Method used

By determining the loss weight of each target object during training based on the relative state information of the target object relative to the specified vehicle and the magnitude of its influence on the vehicle behavior, the model parameters are adjusted to improve the recall rate of target objects with a greater impact.

Benefits of technology

This improves the applicability of the target detection model in intelligent driving scenarios and enhances vehicle driving safety, ensuring that target objects that have a significant impact on the vehicle can be effectively recalled and detected.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122116305A_ABST
    Figure CN122116305A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a target detection model training method, a target detection model detection method, a device, equipment and a medium. Label information of a target object in a training sample is determined. Relative state information of the target object relative to a specified vehicle in the training sample is determined. The target detection model to be trained is used to perform target detection on the training sample to obtain a detection result of the target object. Based on the relative state information of each target object relative to the specified vehicle and the influence size relationship of the target objects in different relative states on the behavior of the specified vehicle, a loss weight corresponding to each target object is determined. Based on the position label and position detection value, the type label and type detection value of each target object, and the loss weight, a sub-loss corresponding to each target object is determined. The network parameters are updated based on the sub-loss corresponding to each target object. In response to satisfying a preset training completion condition, a target detection model is obtained, thereby improving the applicability of the target detection model in an intelligent driving scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to intelligent driving technology, and in particular to a training method, detection method, apparatus, device, and medium for a target detection model. Background Technology

[0002] In the field of intelligent driving, such as autonomous driving and assisted driving, it is necessary to perceive the environment around the vehicle, such as detecting and identifying target objects around the vehicle. The perception range can reach more than 100 meters. The detection of target objects is usually achieved based on target detection models, which need to be trained. How to obtain target detection models that are more suitable for intelligent driving scenarios has become an urgent technical problem to be solved. Summary of the Invention

[0003] The embodiments of this disclosure provide a training method, detection method, apparatus, device, and medium for an object detection model to improve the applicability of the object detection model to intelligent driving scenarios.

[0004] A first aspect of this disclosure provides a method for training a target detection model, comprising: determining annotation information of target objects in at least one training sample; each training sample including at least one of image samples and radar point cloud samples collected by sensors on a designated vehicle, wherein the annotation information of the target objects includes a position label and a type label of the target objects; determining, for each training sample, the relative state information of each target object in the training sample relative to the designated vehicle; performing target detection on the at least one training sample using a detection model to be trained, to obtain a detection result of at least one target object, wherein the detection result of any target object includes a position detection value and a type detection value of the target object; and based on the relative state information of each target object relative to the designated vehicle... Based on the state information and the influence of target objects in different relative states on the behavior of the specified vehicle, a loss weight is determined for each target object. Sub-losses are determined for each target object based on its position label and position detection value, type label and type detection value, and the corresponding loss weights. The network parameters of the detection model to be trained are updated based on the sub-losses. The process iteratively executes the steps from determining the annotation information of target objects in at least one training sample to updating the network parameters of the detection model to be trained based on the sub-losses, until a preset training completion condition is met, and a target detection model is obtained from the detection model to be trained.

[0005] A second aspect of this disclosure provides a method for detecting a target object, comprising: acquiring sensor data collected by sensors on a vehicle; the sensor data including at least one of an image and a radar point cloud; and determining a target object detection result based on the sensor data using a pre-trained target detection model; wherein the target detection model is obtained based on the training method for the target detection model provided in any of the above embodiments.

[0006] A third aspect of this disclosure provides a training apparatus for a target detection model, comprising: a first processing module, configured to determine annotation information of target objects in at least one training sample; each training sample includes at least one of image samples and radar point cloud samples collected by sensors on a designated vehicle, and the annotation information of the target objects includes a position label and a type label of the target objects; a second processing module, configured to determine, for each training sample, the relative state information of each target object in the training sample relative to the designated vehicle; a third processing module, configured to perform target detection on the at least one training sample using a detection model to be trained, to obtain detection results of at least one target object, wherein the detection result of any target object includes a position detection value and a type detection value of the target object; and a fourth processing module, configured to, based on the relative state information of each target object relative to the designated vehicle... The first processing module determines the loss weight corresponding to each target object based on the state information and the influence of target objects in different relative states on the behavior of the specified vehicle. The second processing module determines the sub-loss corresponding to each target object based on the position label and position detection value, type label and type detection value, and the loss weight corresponding to each target object. The third processing module updates the network parameters of the detection model to be trained based on the sub-loss corresponding to each target object. The first processing module to the sixth processing module iteratively execute the operation of determining the annotation information of target objects in at least one training sample and updating the network parameters of the detection model to be trained based on the sub-loss corresponding to each target object until the preset training completion condition is met, and the target detection model is obtained from the detection model to be trained.

[0007] A fourth aspect of this disclosure provides a target object detection device, comprising: an acquisition module for acquiring sensor data collected by sensors on a vehicle; the sensor data including at least one of an image and a radar point cloud; and a detection processing module for determining a target object detection result based on the sensor data and a pre-trained target detection model; wherein the target detection model is obtained based on the training method of the target detection model provided in any of the above embodiments.

[0008] A third aspect of this disclosure is to provide a computer-readable storage medium storing a computer program for executing a training method for a target detection model as described in any of the above embodiments of this disclosure; or the computer program for executing a target object detection method as described in any of the above embodiments of this disclosure.

[0009] A fourth aspect of this disclosure provides an electronic device, the electronic device comprising: a processor; a memory for storing executable instructions of the processor; the processor being configured to read the executable instructions from the memory and execute the instructions to implement a training method for a target detection model or a target object detection method according to any of the above embodiments of this disclosure; or, the electronic device comprising: a training device for a target detection model or a target object detection device according to any of the above embodiments.

[0010] A fifth aspect of this disclosure provides a computer program product that, when instructions in the computer program product are executed by a processor, performs a training method for a target detection model or a target object detection method provided in any of the above embodiments of this disclosure.

[0011] Based on the training method, detection method, apparatus, device, and medium of the target detection model provided in the above embodiments of this disclosure, during the training process, based on the relative state information of each target object relative to a specified vehicle in the training samples, and the magnitude of the influence of target objects in different relative states on the behavior of the specified vehicle, the loss weight corresponding to each target object is determined. This results in a higher loss weight for target objects that have a greater impact on the behavior of the specified vehicle, and a relatively lower loss weight for target objects that have a smaller impact on the behavior of the specified vehicle. This improves the recall rate of the model for target objects that have a greater impact on the vehicle's behavior. For example, the impact of target objects close to the vehicle on the vehicle's behavior is higher than that of target objects far away, and the impact of target objects in front of the vehicle on the vehicle's behavior is higher than that of target objects behind the vehicle. Thus, when the model is deployed to a vehicle for target object detection, it can effectively recall target objects that have a greater impact on the vehicle, which can be used as a basis for vehicle planning and control, thereby effectively improving vehicle driving safety. Therefore, the training method of the target detection model in the embodiments of this disclosure can effectively improve the applicability of the target detection model in intelligent driving scenarios. Attached Figure Description

[0012] Figure 1 This is an exemplary application scenario of the training method for the target detection model provided in the embodiments of this disclosure;

[0013] Figure 2This is a flowchart illustrating a training method for an object detection model provided in an exemplary embodiment of this disclosure;

[0014] Figure 3 This is a flowchart illustrating a training method for an object detection model provided in another exemplary embodiment of this disclosure;

[0015] Figure 4 This is a flowchart illustrating a training method for an object detection model provided in yet another exemplary embodiment of this disclosure;

[0016] Figure 5 This is a flowchart illustrating a training method for an object detection model provided in yet another exemplary embodiment of this disclosure;

[0017] Figure 6 This is a flowchart illustrating a training method for an object detection model provided in yet another exemplary embodiment of this disclosure;

[0018] Figure 7 This is a visual schematic diagram illustrating the relative state between a target object and a vehicle provided in an exemplary embodiment of this disclosure;

[0019] Figure 8 This is a schematic flowchart of a target object detection method provided in an example embodiment of this disclosure;

[0020] Figure 9 This is a schematic diagram of the structure of a training device for an object detection model provided in an exemplary embodiment of this disclosure;

[0021] Figure 10 This is a schematic diagram of the structure of a target object detection device provided in an exemplary embodiment of this disclosure;

[0022] Figure 11 This is a structural diagram of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0023] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.

[0024] Those skilled in the art will understand that the terms "first," "second," etc., in the embodiments of this disclosure are only used to distinguish different steps, devices, or modules, and do not represent any specific technical meaning, nor do they indicate a necessary logical order between them.

[0025] It should also be understood that in the embodiments disclosed herein, "a plurality of" may refer to two or more, and "at least one" may refer to one, two or more.

[0026] It should also be understood that any component, data or structure mentioned in the embodiments of this disclosure can generally be understood as one or more unless expressly defined or given to the contrary in the context.

[0027] Furthermore, the term "and / or" in this disclosure is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this disclosure generally indicates that the preceding and following related objects have an "or" relationship.

[0028] It should also be understood that the description of the various embodiments in this disclosure emphasizes the differences between the various embodiments, and the similarities or similarities can be referred to each other. For the sake of brevity, they will not be described in detail.

[0029] At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn according to actual scale.

[0030] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit this disclosure or its application or use.

[0031] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.

[0032] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.

[0033] This disclosure outlines

[0034] In the process of developing this disclosure, the inventors discovered that in the field of intelligent driving, such as autonomous driving and assisted driving, it is necessary to perceive the environment around the vehicle, such as detecting and identifying target objects around the vehicle. The perception range can reach over 100 meters. Target object detection is usually achieved based on target detection models, which need to be trained. How to obtain a target detection model that is more suitable for intelligent driving scenarios has become an urgent technical problem to be solved.

[0035] Exemplary Overview

[0036] Figure 1 This is an exemplary application scenario of the training method for the object detection model provided in the embodiments of this disclosure. For example... Figure 1As shown, the target detection model can be trained on the server 11 using the training method of the target detection model provided in this embodiment, resulting in a trained target detection model. The target detection model is deployed to a vehicle (referred to as a self-driving vehicle) 12. While driving on the road, the vehicle 12 can collect sensor data of the surrounding environment based on sensor 13, and then, based on the sensor data, use the target detection model to detect target objects around the vehicle 12 (e.g., target object 14 in the figure), for downstream planning and control, enabling the vehicle 12 to continuously drive on the road (divided into one or more lanes by lane lines 15). Sensor 13 may include at least one of cameras, various radars, etc. Various radars may include, for example, radar, lidar, millimeter-wave radar, ultrasonic radar, etc. Specifically, the server 11 can determine the annotation information of target objects in at least one training sample. Each training sample includes at least one of image samples collected by sensors on the specified vehicle and radar point cloud samples. The radar point cloud samples are point cloud data collected based on one or more of various radars. The annotation information of the target object includes the target object's location label and type label. The designated vehicle (distinguished from vehicle 12 in the figure) can be a designated vehicle actually driving on the road (e.g., a user-authorized vehicle) or a dedicated data collection vehicle; this embodiment of the disclosure is not limited to this. Then, for each training sample, the relative state information of each target object in the training sample relative to the designated vehicle is determined. Subsequently, the detection model to be trained can be used to perform target detection on at least one training sample to obtain the detection result of at least one target object. The detection result of any target object includes the position detection value and type detection value of the target object. Then, based on the relative state information of each target object relative to the designated vehicle, and the influence of target objects in different relative states on the behavior of the designated vehicle, the loss weights corresponding to each target object can be determined. Then, based on the position label and position detection value, type label and type detection value, and the loss weights corresponding to each target object, the sub-loss corresponding to each target object can be determined; based on the sub-loss corresponding to each target object, the network parameters of the detection model to be trained are updated; the above process is iteratively executed until the preset training completion conditions are met, and the target detection model is obtained from the detection model to be trained. The detection model after the last network parameter update after training is completed is determined as the target detection model.During training, based on the relative state information of each target object relative to the specified vehicle in the training samples, and the influence of target objects in different relative states on the behavior of the specified vehicle, the loss weights corresponding to each target object are determined. This results in target objects with a greater impact on the behavior of the specified vehicle having a larger loss weight, while target objects with a smaller impact on the behavior of the specified vehicle have a relatively smaller loss weight. This improves the recall rate of the model for target objects that have a greater impact on the vehicle's behavior. For example, target objects close to the vehicle have a greater impact on the vehicle's behavior than target objects far away, and target objects in front of the vehicle have a greater impact on the vehicle's behavior than target objects behind the vehicle. Therefore, when the model is deployed to the vehicle for target object detection, it can effectively recall target objects that have a greater impact on the vehicle, which can be used as a basis for vehicle planning and control, thus effectively improving vehicle driving safety. Therefore, the training method of the target detection model in this embodiment can effectively improve the applicability of the target detection model in intelligent driving scenarios.

[0037] Exemplary methods

[0038] Figure 2 This is a flowchart illustrating a training method for an object detection model provided in an exemplary embodiment of this disclosure. Embodiments of this disclosure can be applied to electronic devices, specifically servers, terminal devices, and other electronic equipment. Figure 2 As shown, the training method for the target detection model provided in the embodiments of this disclosure may include the following steps:

[0039] Step 210: Determine the annotation information of the target object in each training sample in at least one training sample.

[0040] Each training sample includes at least one of image samples collected by sensors on a designated vehicle and radar point cloud samples. The annotation information of the target object includes the target object's location label and type label.

[0041] In some optional embodiments, the designated vehicle may be a designated vehicle actually driving on the road (e.g., a user vehicle authorized by the user) or a dedicated data collection vehicle, and this disclosure does not limit the scope of the vehicle.

[0042] In some alternative embodiments, the sensors on the designated vehicle may include at least one of cameras, various radars, etc. Various radars may include, for example, radar (Radar), lidar (Lidar), millimeter-wave radar, ultrasonic radar, etc. Accordingly, the radar point cloud samples are based on point cloud data collected by one or more of these various radars.

[0043] In some optional embodiments, the target object may include various dynamic and static objects surrounding the specified vehicle. Dynamic objects may include, for example, other vehicles, pedestrians, cyclists, and other dynamic objects in the vicinity. Static objects may include, for example, surrounding static obstacles such as traffic cones, water-filled barriers, and traffic warning signs. Optionally, static objects may also include static elements such as lane lines, zebra crossings, and stop lines. The type label of the target object indicates whether the target object belongs to a certain type. The type label can include two states: 1 indicates that it belongs to that type, and 0 indicates that it does not belong to that type. For a single-task detection model, such as a target vehicle detection model, the type label of each target object includes a label value. For example, if the target object belongs to a vehicle, the type label is 1; if the target object does not belong to a vehicle, the type label of the target object is 0. For a multi-task detection model, the type label of each target object includes a label value indicating whether the target object belongs to each type. For example, if the model simultaneously detects three types of target objects: target vehicles, target pedestrians, and target cyclists, then for each target object in the training samples, the type label of the target object includes three label values. For instance, if the target object belongs to the target vehicle category, then the type label of the target object can be represented as (whether it is a target vehicle, whether it is a target pedestrian, whether it is a target cyclist) = (1, 0, 0), with the label values ​​of each type arranged in order. This is only an exemplary representation of the type label. In practical applications, the type label can also adopt any other feasible representation, and this disclosure does not limit it.

[0044] In some optional embodiments, when the training samples are image samples, the location label of the target object is the location label of the target object in the corresponding image sample (which can be referred to as the first location label). When the training samples are radar point cloud samples, the location label of the target object is the location label of the target object in the vehicle coordinate system of the specified vehicle (which can be referred to as the second location label) or the location label of the target object in the radar coordinate system. There is a pre-defined transformation relationship between the radar coordinate system and the vehicle coordinate system.

[0045] In some optional embodiments, when the training samples include image samples and radar point cloud samples, the location label of the target object may include at least one of a first location label of the target object in the corresponding image sample and a second location label in the vehicle coordinate system (or radar coordinate system). The specific labeling information can be set according to the detection task. For example, if the detection task is to detect a target object in the image coordinate system based on images and radar point clouds, then the location label of the target object is the first location label. If the detection task is to detect a target object in the vehicle coordinate system based on images and radar point clouds, then the location label of the target object is the second location label. If the detection task is multi-task, that is, the detection task is to detect a target object in both the image coordinate system and the vehicle coordinate system based on images and radar point clouds, then the location label of the target object includes both the first location label and the second location label.

[0046] In some optional embodiments, the number of designated vehicles can be one or more, and the specific number is not limited. For each training sample, when acquiring the training sample, the pose information of the designated vehicle corresponding to the training sample, the extrinsic and intrinsic parameters of the camera, the extrinsic parameters of the radar, etc., can be recorded to achieve mutual transformation between the image coordinate system, camera coordinate system, radar coordinate system, vehicle coordinate system, and global coordinate system (e.g., world coordinate system, vehicle coordinate system of the vehicle's initial position). The pose information of the designated vehicle can be used to determine the extrinsic parameters of the vehicle coordinate system of the designated vehicle, that is, the transformation parameters between the vehicle coordinate system and the global coordinate system.

[0047] In some optional embodiments, each training sample may include one or more frames of sensor data (including at least one of image samples and radar point cloud samples). That is, the model can perform target object detection based on a single frame of sensor data or based on a sequence of multiple frames of sensor data. For example, the training samples may include image samples at time t (frame t), image samples at time t-1, image samples at time t-2, ..., image samples at time tN, totaling N+1 frames of image samples. The detection model's task is to detect the target object in the image sample at time t based on the N+1 frames of image samples. Multiple frames of images can reflect the dynamic information of the target object relative to the specified vehicle, such as changes in relative distance, relative speed, and relative direction, allowing the detection model to learn the complex correlation between the target object detection result and various states of the target object, which helps improve model performance and the accuracy of the target object detection results. Multiple frames of radar point cloud samples are similar to multiple frames of image samples and will not be described in detail here.

[0048] In some optional embodiments, the specific number of training samples in at least one training sample can be set according to actual needs, and this disclosure does not limit it.

[0049] In some optional embodiments, the annotation information of the target object can be obtained through any of the annotation methods such as manual annotation, semi-automatic annotation, and automatic annotation.

[0050] Step 220: For each training sample, determine the relative state information of each target object relative to the specified vehicle in the training sample.

[0051] For each training sample, the training sample may include one or more target objects. The relative state information of each target object relative to the specified vehicle (the specified vehicle for collecting the training sample) may include at least one of the following: relative distance, relative angle (or relative direction), relative speed, and relative direction of motion (or relative trend of motion) of the target object relative to the specified vehicle.

[0052] In some optional embodiments, the relative state information of the target object relative to the designated vehicle can be predetermined and stored in a designated space corresponding to the training samples. Then, the relative state information of each target object relative to the designated vehicle in the training samples can be obtained from the designated space. Optionally, the relative state information of the target object relative to the designated vehicle can be determined in advance based on the target object's annotation information. For example, the relative state information of the target object can be determined in advance based on the target object's position label. Alternatively, the annotation information of the target object can include the target object's pose (or orientation) in addition to its position label and type label. Then, the relative state information of the target object can be determined based on its position label and orientation.

[0053] In some optional embodiments, the relative state information of the target object with respect to the specified vehicle can be determined in real time during the training process. The determination principle is similar to the pre-determined principle and will not be described in detail here.

[0054] Step 230: Perform target detection on at least one training sample using the detection model to be trained, and obtain the detection result of at least one target object.

[0055] The detection result for any target object includes the target object's position detection value and type detection value.

[0056] In some optional embodiments, the detection model to be trained can be an initialized model or a model updated in the previous iteration. For example, in the first iteration, the network parameters of the detection model to be trained are the initialized parameters. Therefore, the detection model to be trained is an initialized model. After the i-th (i is a positive integer) iteration, the network parameters of the model change due to the update. In the next (i+1) iteration, the detection model to be trained is the model updated in the i-th iteration.

[0057] In some optional embodiments, the coordinate system of the target object's position detection value is consistent with the coordinate system of the target object's position label. That is, if the position label is a label in the image coordinate system, then the position detection value is the detection value in the image coordinate system. If the position label is a label in the vehicle coordinate system, then the position detection value is the detection value in the vehicle coordinate system. If the position label includes labels in both the image coordinate system and the vehicle coordinate system, then the position detection value includes the detection value in both the image coordinate system and the vehicle coordinate system. The specific settings can be configured according to the actual detection task.

[0058] In some optional embodiments, the detection model to be trained can be any implementable detection model, and this disclosure does not limit it. For example, the detection model to be trained can be an image-based detection model, a point cloud-based detection model, a multi-sensor fusion-based detection model, and so on. The network structure of the detection model to be trained can be set to any implementable network structure according to the actual detection task. For example, a series of detection models based on convolutional neural networks and their variations, detection models based on the YOLO series, a series of detection models based on Transformer and its variations, and so on.

[0059] In some optional embodiments, the detection model to be trained can be used to perform object detection on each of at least one training sample to obtain the detection result of the target object in each training sample. Each training sample may include one or more target objects, and the detection result of the target object in each training sample may include the detection results of one or more target objects in that training sample.

[0060] It should be noted that the execution of steps 230 and 220 is not in any particular order.

[0061] Step 240: Based on the relative state information of each target object relative to the designated vehicle, and the relationship between the influence of target objects in different relative states on the behavior of the designated vehicle, determine the loss weight corresponding to each target object.

[0062] The influence of target objects in different relative states on the behavior of a specified vehicle can be pre-set based on experience. That is, based on the relative state information of each target object relative to the specified vehicle and the pre-configured influence relationships of target objects in different relative states on the behavior of the specified vehicle, the loss weight corresponding to each target object is determined. The influence relationship of target objects in different relative states on the behavior of a specified vehicle refers to the relationship between the influence of objects in different relative states around the vehicle and the vehicle's behavior during its movement, with the specified vehicle as the "self-vehicle". For example, the influence of a target object in front of the vehicle is greater than that of a target object behind the vehicle; the influence of a target object close to the vehicle is greater than that of a target object far away; the influence of a target object whose relative speed causes it to continuously approach the vehicle is greater than that of a target object continuously moving away from the vehicle; the influence of a target object in front of the vehicle that has a tendency to cross or change lanes is greater than that of a target object traveling in the same direction as the vehicle, and so on. For object detection models in intelligent driving scenarios, it is desirable to accurately detect objects that have a significant impact on the vehicle's behavior, while objects with less impact on the vehicle's behavior should have slightly lower accuracy but still have a relatively small impact. Based on this, the loss weights for each object can be determined by considering the relative state information of each object to the specified vehicle and the magnitude of the impact of objects in different relative states on the specified vehicle's behavior. This ensures that the loss weights for objects with a significant impact on the vehicle's behavior are relatively greater than those for objects with less impact, thus guiding the model's learning and enabling the model to have a higher recall rate for objects with a greater impact. This allows the trained object detection model to better meet the detection requirements of intelligent driving scenarios, thereby improving the model's adaptability to intelligent driving environments.

[0063] It should be noted that the execution order of steps 240 and 230 is not important.

[0064] Step 250: Based on the location label and location detection value, type label and type detection value of each target object, and the loss weight corresponding to each target object, determine the sub-loss corresponding to each target object.

[0065] For each target object, the position detection error is determined based on the position label and position detection value, and the type detection error is determined based on the type label and type detection value. The sum of the position detection error and the type detection error is the total error of the target object. Multiplying the loss weight corresponding to the target object by the total error of the target object yields the sub-loss of the target object. Alternatively, the loss weight of the target object is multiplied by the position detection error and the type detection error separately to obtain the adjusted position detection error and the adjusted type detection error. Adding the adjusted position detection error and the type detection error together yields the sub-loss of the target object.

[0066] Step 260: Update the network parameters of the detection model to be trained based on the sub-loss corresponding to each target object.

[0067] In this process, the network parameters of the detection model to be trained can be updated using a pre-configured optimization algorithm, which can be combined with the sub-loss corresponding to each target object. The specific optimization algorithm can be set according to actual needs, and this embodiment does not limit it. For example, the optimization algorithm may include stochastic gradient descent, gradient descent with adaptive learning rate, Newton's method, Gauss-Newton method, etc., which will not be described in detail in this embodiment.

[0068] The process iteratively executes step 210, which involves determining the annotation information of target objects in at least one training sample, and step 260, which involves updating the network parameters of the detection model to be trained based on the sub-loss corresponding to each target object, until the preset training completion condition is met. The training ends, and the updated detection model is determined as the target detection model, thus obtaining the target detection model from the detection model to be trained.

[0069] Step 270: In response to meeting the preset training completion conditions, the updated detection model is determined as the target detection model.

[0070] The updated detection model is the one updated after the final iteration of training. Preset training completion conditions may include model convergence, reaching a preset threshold number of iterations, etc., but are not specifically limited.

[0071] The target detection model training method provided in this disclosure determines the loss weight for each target object during training based on the relative state information of each target object relative to a specified vehicle in the training samples, and the influence of target objects in different relative states on the behavior of the specified vehicle. This results in a higher loss weight for target objects that have a greater impact on the behavior of the specified vehicle, and a relatively lower loss weight for target objects that have a smaller impact. This improves the recall rate of the model for target objects that significantly influence vehicle behavior. For example, target objects close to the vehicle have a higher impact on vehicle behavior than target objects at a greater distance, and target objects in front of the vehicle have a higher impact than target objects behind the vehicle. Therefore, when the model is deployed to a vehicle for target object detection, it can effectively recall target objects that significantly impact the vehicle, which can then be used as a basis for vehicle planning and control, thus effectively improving driving safety. Therefore, the target detection model training method in this disclosure can effectively improve the applicability of target detection models in intelligent driving scenarios.

[0072] Figure 3 This is a flowchart illustrating a training method for an object detection model provided in another exemplary embodiment of this disclosure.

[0073] In some alternative embodiments, in the above... Figure 2 Based on the illustrated embodiments, as Figure 3 As shown, step 220, which determines the relative state information of each target object relative to the specified vehicle in each training sample, may include:

[0074] Step 2210: For each training sample, determine the first pose information of each target object in the training sample in the vehicle coordinate system of the specified vehicle.

[0075] In some optional embodiments, for each training sample, if the annotation information of the target object includes the first pose information of the target object in the vehicle coordinate system of the specified vehicle, the first pose information of each target object in the training sample in the vehicle coordinate system of the specified vehicle can be obtained from the annotation information of the target object. If the annotation information of the target object does not include the first pose information of the target object, the first pose information of the target object can be obtained by other means. The first pose information includes the position and attitude of the target object (i.e., heading angle, yaw angle, and orientation).

[0076] Step 2220: Determine the relative state information of each target object based on the first pose information of each target object.

[0077] In some optional embodiments, for each target object, the relative state information of the target object may include at least one of the following: relative distance (referred to as the first relative distance), relative direction (referred to as the first relative direction), relative velocity (referred to as the first relative velocity), and relative motion direction (referred to as the first relative motion direction) between the target object and the designated vehicle. For the relative distance, the distance between the target object and the designated vehicle can be calculated based on the position in the first pose information of the target object. If the first pose information is the pose information in the vehicle coordinate system of the designated vehicle, the relative distance between the target object and the designated vehicle can be calculated using the distance calculation formula between the position in the first pose information (i.e., the three-axis coordinates (x, y, z)) and the origin (0, 0, 0) of the vehicle coordinate system. The relative direction can be determined based on the attitude in the first pose information. The relative velocity and relative motion direction can be determined by combining the pose information of the target object across multiple frames (if the pose information of multiple frames is available).

[0078] In some optional embodiments, the relative direction may include a direction based on a preset angle granularity. The preset angle granularity can be set to any angle, such as 30 degrees, 60 degrees, 90 degrees, 180 degrees, etc. Taking 180 degrees as an example, the relative direction may include forward and backward. Taking 60 degrees as an example, the relative direction may include forward, backward, left front, right front, left rear, and right rear. The specific division of relative direction is not limited in this disclosure embodiment. Relative speed may include the forward speed and reverse speed relative to the specified vehicle. The forward speed causes the target object to gradually move away from the specified vehicle, and the reverse speed causes the target object to gradually move closer to the specified vehicle. Relative speed may include relative lateral speed and relative longitudinal speed. Optionally, the forward speed and reverse speed can also be finely divided. For example, the forward speed can be divided into multiple speed levels according to speed magnitude, and the reverse speed can be divided into multiple speed levels according to speed magnitude, etc. The relative motion direction can be determined by the relative speed, which may include relative lateral speed and relative longitudinal speed. The relative lateral speed represents the lateral movement of the target object relative to the specified vehicle, and the relative longitudinal speed represents the longitudinal movement of the target object relative to the specified vehicle. If the relative lateral velocity of a target object exceeds a preset lateral velocity threshold, it indicates that the target object has a tendency to cross or change lanes, which may significantly affect the behavior of the designated vehicle. For example, a target object crossing in front of a designated vehicle has a significant impact on its behavior. In practical applications, the relative motion direction can be considered in conjunction with the target object's relative direction to the designated vehicle to determine whether its relative motion direction should be taken into account in determining the loss weight. For example, if the target object is behind the designated vehicle, its lateral movement has little impact on the vehicle's behavior, and its relative motion direction may not be considered when determining its loss weight.

[0079] In the embodiments of this disclosure, the relative state information of the target object with respect to the specified vehicle can be effectively determined by using the pose information of the target object in the vehicle coordinate system of the specified vehicle, thus providing an accurate and effective reference for determining the loss weight of the target object.

[0080] Figure 4 This is a flowchart illustrating a training method for an object detection model provided in another exemplary embodiment of this disclosure.

[0081] In some optional embodiments, based on any of the above embodiments, step 240, which determines the loss weight corresponding to each target object based on the relative state information of each target object relative to the designated vehicle and the magnitude of the influence of target objects in different relative states on the behavior of the designated vehicle, may include:

[0082] Step 2410: Based on the relative state information of each target object, determine the first relative distance, first relative direction, first relative speed and first relative motion direction of each target object relative to the designated vehicle.

[0083] For each target object, the first relative distance, first relative direction, first relative speed, and first relative motion direction relative to the designated vehicle can be found in the aforementioned embodiments and will not be repeated here.

[0084] Step 2420: Based on the first relative distance, first relative direction, first relative velocity, first relative motion direction of each target object, and the mapping relationship between the pre-configured loss weight and the relative distance, relative direction, relative velocity and relative motion direction, determine the loss weight corresponding to each target object.

[0085] The pre-configured mapping relationship between loss weights and relative distance, relative direction, relative speed, and relative motion direction (which can be referred to as the first mapping relationship) can be set according to the influence of target objects in different relative states on the behavior of a specified vehicle. This mapping relationship is used to increase the loss weight of target objects with greater influence and decrease the loss weight of target objects with less influence.

[0086] In some optional embodiments, the first mapping relationship may include mapping portions (referred to as sub-relationships) corresponding to influencing factors such as relative distance, relative direction, relative speed, and relative motion direction. Each sub-relationship is used to adjust the contribution of the corresponding influencing factor to the loss weight according to the actual situation of the target object's corresponding influencing factor. For example, the sub-relationship corresponding to relative distance states that the contribution of relative distance to the loss weight decreases as the distance increases. The sub-relationship corresponding to relative direction states that different relative directions contribute differently to the loss weight; for example, forward direction increases the loss weight, while backward direction decreases it. The sub-relationship corresponding to relative speed is similar to that of relative direction, that is, different relative speeds contribute differently to the loss weight; for example, forward speed decreases the loss weight, while reverse speed increases it. The sub-relationship corresponding to relative motion direction states that different relative motion directions contribute differently to the loss weight. For example, a relative motion direction with a traversing tendency increases the loss weight, while a relative motion direction without a traversing tendency decreases it. And so on. The relative direction, relative speed, and relative motion direction can be finely divided so that different situations contribute differently to the loss weight.

[0087] In some optional embodiments, the first mapping relationship can be represented as follows:

[0088] Lossweight=f1(d)+f2(dir)+f3(v)+f4(r)

[0089] Where f1(d) represents the sub-relationship function corresponding to the relative distance d, f2(dir) represents the sub-relationship function corresponding to the relative direction dir, f3(v) represents the sub-relationship function corresponding to the relative velocity v, and f4(r) represents the sub-relationship function corresponding to the relative motion direction r. Each sub-relationship function can be set according to the specific circumstances of each influencing factor.

[0090] In practical applications, the first mapping relationship can be set to any one of the four sub-relation functions mentioned above, or any other arbitrary combination, depending on actual needs. For example, any combination of two functions (such as f1(d) + f2(dir)), any combination of three functions, etc. Sub-relation functions for other influencing factors can also be added to the first mapping relationship described above.

[0091] For example, the relative distance is represented by the variable d, and the relative direction is represented by the variable is. front Relative velocity is represented by the variable is neg The direction of relative motion is represented by the variable is. cross , is front The values ​​of is include 1 (representing the front) and -1 (representing the back). neg The value of is includes 1 (representing reverse velocity) and -1 (representing forward velocity).cross The value of includes 1 (indicating a crossover trend) and -1 (indicating no crossover trend), so the first mapping relationship can be expressed as follows:

[0092] Lossweight = f1(d) + f2(is) front )+f3(is neg )+f4(is cross )

[0093] Where f1(d) represents the sub-relation function corresponding to the relative distance, and f2(is front f3(is) represents the sub-relation function corresponding to the relative direction. neg f4(is) represents the sub-relation function corresponding to the relative velocity. cross ) represents the sub-relation function corresponding to the direction of relative motion.

[0094] In some optional embodiments, the sub-relationships can be represented as follows:

[0095]

[0096] f2(is front )=β*is front

[0097] f3(is neg )=ρ*is neg

[0098]

[0099] Among them, α, γ, β, ρ, These are preset parameters, all positive numbers. γ can be set to 100 or other values, and α can be set according to the threshold between near and far distances. For example, to improve the recall rate of targets within 30 meters, α = 1.3 and γ = 100, so that the loss weight of targets within 30 meters (in front) is greater than 1, and the loss weight of targets greater than 30 meters is less than 1. Similarly, to improve the recall rate of targets within 50 meters, α = 1.5. β, ρ, These are the direction coefficient, velocity coefficient, and motion direction coefficient, respectively, β, ρ, It can be set to 0.1, 0.2, etc. β, ρ, They can be the same or different, for example, β = 0.1, ρ = 0.05, No specific restrictions are imposed.

[0100] In some alternative embodiments, f1(d) can be set as a sub-relation function in exponential form, for example...

[0101] The sub-relation functions described above are merely illustrative. In practical applications, the sub-relationships are not limited to the functions described above and can be set to other forms of functions according to actual needs, as long as they can effectively express the contribution of factors such as relative distance, relative direction, relative speed, and relative motion direction to the loss weights. For example, the relative direction dir can be finely granularized to values ​​between 0 and 1, and between -1 and 0. For example, taking the X-axis of the vehicle coordinate system as 0 degrees, clockwise as positive, and counterclockwise as negative, the relative direction dir of the target object in front (-30 degrees to 30 degrees) is represented as 1, the relative direction dir of the target object to the left front (-90 degrees to -30 degrees) and right front (30 degrees to 90 degrees) is represented as 0.5, the relative direction dir of the target object behind (150 degrees to 180 degrees, -180 degrees to -150 degrees) is represented as -1, and the relative direction dir of the target object to the left rear (-150 degrees to -90 degrees) and right rear (90 degrees to 150 degrees) is represented as -0.5. If the relative velocity v and the relative direction of motion r are similar, then the loss weights are expressed as follows:

[0102]

[0103] The meanings of each symbol are explained above.

[0104] In some optional embodiments, for target detection of single-frame sensor data, if the relative velocity and relative motion direction of the target object cannot be calculated during the training process, the first mapping relationship may include at least one of f1(d) and f2(dir) mentioned above.

[0105] In the embodiments of this disclosure, by pre-configuring the mapping relationship between loss weights and relative distance, relative direction, relative speed, and relative motion direction, the loss weights of each target object can be dynamically determined in real time during model training by combining the first relative distance, first relative direction, first relative speed, and first relative motion direction of each target object relative to the specified vehicle. This increases the contribution of target objects with greater influence to the model loss, enabling the model to pay more attention to the features of target objects with greater influence during the training process, thereby improving the recall rate of the model for target objects with greater influence.

[0106] In some alternative embodiments, based on any of the above embodiments, the location label of the target object includes a second location label of the target object in the vehicle coordinate system of the specified vehicle.

[0107] Determining the relative state information of each target object in the training sample relative to the specified vehicle in step 220 may include: determining the relative state information of each target object relative to the specified vehicle based on the second position label of each target object in the training sample.

[0108] The vehicle coordinate system in the second position label of the target object in the vehicle coordinate system of the specified vehicle is the vehicle coordinate system at the time when the specified vehicle collects the training sample corresponding to the target object. This vehicle coordinate system can be determined based on the recorded pose information of the specified vehicle when collecting the training sample. The pose information of the specified vehicle is its pose information in the global coordinate system, including its position and orientation. The second position label of the target object can include the three-axis coordinates of the target object in the corresponding vehicle coordinate system, which can be represented as (x, y, z). Alternatively, the second position label can include the label of the overall position of the target object's bounding box in the vehicle coordinate system, such as including the coordinates and size of the bounding box's center point, or including the coordinates of the bounding box's vertices. For example, it could include the coordinates of the eight vertices of the bounding box. The specific representation of the second position label can be set according to actual needs. The relative distance of the target object can be calculated based on the coordinates of the target object's center point and the origin of the vehicle coordinate system. For the relative direction of the target object, the line connecting the target object's center point and the origin of the vehicle coordinate system can be determined, and the angle between this line and the specified coordinate axis of the vehicle coordinate system can be used to determine the direction. The specified coordinate axis can be, for example, the horizontal axis (Y-axis) or the vertical axis (X-axis) of the vehicle coordinate system, without limitation. Taking the vertical axis (X-axis) as 0 degrees and clockwise as positive as an example, if the angle between the line connecting the target object and the origin and the X-axis is within the range of -90 degrees to 90 degrees (this is only an example range; in actual applications, it can be set to other ranges), it means that the target object is in front of the specified vehicle. If the angle between the line connecting the target object and the X-axis is within the range of 90 degrees to 270 degrees, it means that the target object is behind the specified vehicle. The relative directions of other angular granularities are similar and will not be elaborated here.

[0109] In some optional embodiments, the relative distance and relative direction of the target object relative to the specified vehicle can be determined based on the second position label of the target object and the origin coordinates of the vehicle coordinate system, which serves as the relative state information of the target object.

[0110] In embodiments of this disclosure, when the annotation information of the target object includes a second position label of the target object in the vehicle coordinate system of the specified vehicle, the relative state information of the target object relative to the specified vehicle can be determined based on the second position label of the target object, thereby providing an effective relative state information reference for determining the loss weight of the target object.

[0111] In some optional embodiments, the relative state information of the target object relative to the designated vehicle includes a second relative distance and a second relative direction of the target object relative to the designated vehicle.

[0112] Step 240, based on the relative state information of each target object relative to the designated vehicle and the magnitude of the influence of target objects in different relative states on the behavior of the designated vehicle, determines the loss weight corresponding to each target object, which may include:

[0113] Based on the second relative distance, the second relative direction, and the first mapping relationship between the loss weight and the relative distance and the relative direction of each target object, the loss weight corresponding to each target object is determined. The first mapping relationship includes a first sub-relationship corresponding to the relative distance and a second sub-relationship corresponding to the relative direction. The first sub-relationship is a relationship that decreases as the relative distance increases.

[0114] The first sub-relation can be found in the sub-relation function f1(d) of the above embodiment, and the second sub-relation can be found in the sub-relation function f2(dir) of the above embodiment; further details will not be provided. The first sub-relation decreases with increasing relative distance, making the contribution of nearby target objects to the model loss greater than that of distant target objects, thereby helping to improve the recall rate of nearby target objects.

[0115] In some optional embodiments, the second sub-relation is to increase the loss weight of the target object when the second relative direction is the first preset direction, and to decrease the loss weight of the target object when the second relative direction is the second preset direction.

[0116] Here, the first preset direction is forward, and the second preset direction is backward. That is, if the target object is in front of the specified vehicle, the loss weight for that target object increases based on the first sub-relation; if the target object is behind the specified vehicle, the loss weight for that target object decreases based on the first sub-relation. For example, f2(is...) front )=β*is front The target object in front is front The is value is 1, indicating the target object behind is. tront The value is -1. Therefore, if the relative direction of the target object is forward, the loss weight increases by β, and if the relative direction of the target object is backward, the loss weight decreases by β.

[0117] In the embodiments of this disclosure, by setting the second sub-relationship to increase the loss weight of the target object when the second relative direction is the first preset direction, and to decrease the loss weight of the target object when the second relative direction is the second preset direction, the contribution of the target object in the first preset direction to the model loss is greater than the contribution of the target object in the second preset direction. This makes the model pay more attention to the target object in the first preset direction that has a greater impact on the vehicle's behavior, thereby improving the recall rate of the target object in the first preset direction.

[0118] In some optional embodiments, based on any of the above embodiments, step 240, which determines the loss weight corresponding to each target object based on the relative state information of each target object relative to the designated vehicle and the magnitude of the influence of target objects in different relative states on the behavior of the designated vehicle, may include:

[0119] Based on the relative state information of each target object, the type label of each target object, and the relationship between the influence of different relative states and different types of target objects on the behavior of the specified vehicle, the loss weight corresponding to each target object is determined.

[0120] In addition to the above embodiments, the loss weight of the target object is further determined by combining the type of the target object. Based on the influence of different types of target objects on the behavior of the specified vehicle, and combined with the influence of target objects in different relative states on the behavior of the specified vehicle, the loss weight of the target object is comprehensively determined, which further improves the recall rate of the model for target objects with greater influence, and thus improves the adaptability of the model to intelligent driving scenarios.

[0121] In some optional embodiments, a sub-relationship related to the type of the target object can be added based on the first mapping relationship of the loss weight described above. Different coefficients are set for different types to represent different contributions to the loss weight. The specifics will not be elaborated further.

[0122] Figure 5 This is a flowchart illustrating a training method for an object detection model provided in yet another exemplary embodiment of this disclosure.

[0123] In some optional embodiments, based on any of the above embodiments, the training samples are image samples; the location label of the target object includes the first location label of the target object in the corresponding image sample; the location detection value of the target object includes the first location detection value of the target object in the corresponding image sample.

[0124] The first location label of the target object in the corresponding image sample may include a region bounding box label of the target object in the image sample, such as the center point and size of the region bounding box, or the coordinates of the four corner points of the region bounding box. Correspondingly, the first location detection value has the same format as the first location label; for example, the first location detection value includes the detection box of the detected target location in the image sample, which may specifically include the center point and size of the detection box, or the coordinates of the four corner points of the detection box. The specific location representation method is not limited.

[0125] Step 250, which determines the sub-loss for each target object based on its location label and location detection value, type label and type detection value, and the loss weight corresponding to each target object, may include:

[0126] Step 2510: Determine the position detection error for each target object based on its first position label and first position detection value.

[0127] Specifically, the first position label can be compared with the first position detection value to calculate the position detection error.

[0128] In some alternative embodiments, the location detection error can be determined based on the intersection over union (IoU) ratio of the detection box and the region box label. For example, the IoU ratio represents the degree of overlap between the detection box and the region box label. The location detection error can be obtained by subtracting the IoU ratio value from 1.

[0129] In some optional embodiments, the position detection error can be determined based on the positional error between the four corner points of the detection box and the four corner points of the region box label. The specific method for calculating the position detection error is not limited. The position detection error can be calculated using Mean Absolute Error (MAE), Mean Squared Error (MSE), etc., and is not specifically limited.

[0130] Step 2520: Determine the type detection error for each target object based on its type label and type detection value.

[0131] The target object's type label includes two labels: 0 and 1. The type detection value is a probability value between 0 and 1, representing the confidence level that the model detects the target object as belonging to a certain type. The type detection error is calculated by comparing the type label with the type detection value. For example, if target object A has a label of 1 for vehicle type, 0 for pedestrian type, and 0 for cyclist type, and the model detects a probability of 0.7 for target object A as vehicle type, 0.4 for pedestrian type, and 0.3 for cyclist type, then the type detection error is determined by the absolute values ​​of 0.7-1, 0.4-0, and 0.3-0.

[0132] In some alternative embodiments, the type detection error can be the mean absolute error, mean square error, etc.

[0133] It should be noted that the execution order of steps 2520 and 2510 is not important.

[0134] Step 2530: Based on the loss weight, position detection error and type detection error corresponding to each target object, determine the sub-loss corresponding to each target object.

[0135] For each target object, the sub-loss Lob corresponding to that target object can be determined using the following formula:

[0136] Lob = Lossweight * (E p+E c )

[0137] Among them, E p E represents the position detection error of the target object. c This represents the type detection error of the target object, and Lossweight represents the loss weight of the target object.

[0138] After determining the sub-loss corresponding to each target object, the sub-losses corresponding to each target object can be added together to obtain the model loss, which is used to calculate the iteration step size of the network parameters. The network parameters are then updated based on the iteration step size to obtain the updated detection model.

[0139] In some optional embodiments, after each iteration of updating the network parameters, the updated detection model can be tested based on test samples to determine whether the model meets the preset training completion conditions. For example, if the model's loss is less than a preset loss threshold after multiple consecutive iterations, it indicates that the model has converged and the preset training completion conditions are met.

[0140] The embodiments of this disclosure train the detection model through image samples to obtain an image-based target detection model. This allows the target detection model to perceive target objects around the vehicle based on real-time acquired images, improve the recall rate of nearby target objects, and ensure the driving safety of the vehicle.

[0141] In some optional embodiments, for each training sample, the loss can be determined based on the location label and location detection value, type label and type detection value, and the loss weight corresponding to each target object in the training sample. Then, based on the losses corresponding to each training sample, the network parameters of the detection model to be trained are updated. The loss of each training sample can be determined based on a pre-configured loss function. The loss function is a function of the detection error of the target object and the loss weight. The specific type of loss function is not limited, such as L1 loss (i.e., absolute error loss), L2 loss (i.e., mean squared error loss), cross-entropy loss, etc.

[0142] Figure 6 This is a flowchart illustrating a training method for an object detection model provided in another exemplary embodiment of this disclosure.

[0143] In some optional embodiments, the training samples are radar point cloud samples; the position label of the target object includes a second position label of the target object in the vehicle coordinate system of the specified vehicle; the position detection value of the target object includes the second position detection value of the target object in the vehicle coordinate system.

[0144] The second location label can be found in the aforementioned embodiments, and the data structure of the second location detection value is consistent with that of the second location label.

[0145] Step 250, which determines the sub-loss for each target object based on its location label and location detection value, type label and type detection value, and the loss weight corresponding to each target object, may include:

[0146] Step 25a0: Determine the position detection error for each target object based on the second position label and the second position detection value of each target object.

[0147] Among them, radar point cloud samples are used as training samples, and the corresponding target detection model is a three-dimensional (3D) target detection model. The position detection error can be determined by any implementable position loss related function in the training of the 3D target detection model, such as the smoothing L1 loss function, the cross-union ratio loss function, etc., without any specific limitation.

[0148] Step 25b0: Determine the type detection error for each target object based on its type label and type detection value.

[0149] The type detection error is similar to step 2520 above, and will not be described in detail here.

[0150] It should be noted that the execution order of steps 25b0 and 25a0 is not important.

[0151] Step 25c0: Based on the loss weights, position detection errors, and type detection errors corresponding to each target object, determine the sub-loss corresponding to each target object.

[0152] The specific operation of step 25c0 is similar to that of step 2530, and will not be described in detail here.

[0153] The embodiments of this disclosure train the detection model using radar point cloud samples, thereby obtaining a 3D target detection model based on radar point clouds. This enables the target detection model to perceive target objects around the vehicle based on real-time collected radar point clouds, obtain the position and type of the target objects in three-dimensional space, improve the recall rate of nearby target objects, and ensure the driving safety of the vehicle.

[0154] In some optional embodiments, the training samples include image samples and radar point cloud samples, meaning the detection model is a multi-sensor detection model capable of detecting at least one of the two-dimensional position of a target object in an image and its three-dimensional position in three-dimensional space. In this case, during model training, the sub-loss of the target object can be determined by combining the position detection error, type detection error, and loss weights in both the image sample and radar point cloud sample cases. Specifically, the first position detection error for each target object can be determined based on its first position label and first position detection value; the second position detection error can be determined based on its second position label and second position detection value; the type detection error can be determined based on its type label and type detection value; and finally, the sub-loss for each target object can be determined based on its respective loss weights, first position detection error, second position detection error, and type detection error. Further details are omitted here.

[0155] In some alternative embodiments, Figure 7 This is a visual schematic diagram illustrating the relative state between a target object and a vehicle, provided in an exemplary embodiment of this disclosure. For example... Figure 7 As shown, the gray rectangles represent the vehicle 12, and each white rectangle represents a target object 14, which can be other vehicles, pedestrians, cyclists, etc. The arrows within the rectangles indicate the direction. The vehicle 12 is surrounded by multiple target objects 14. Taking target objects A, B, and C as examples, the distance between A and the vehicle 12 is denoted as d1, and the distance between C and the vehicle 12 is denoted as d2, where d2 is greater than d1. During the vehicle 12's movement, it is clear that A has a greater impact on the vehicle 12's behavior than C. Therefore, in the environmental perception phase, it is expected that A can be accurately detected, while C, even if temporarily undetectable, has little impact on the vehicle 12's behavior. Thus, in the model training phase, the loss weights for nearby target objects are relatively larger than those for distant target objects to guide the model to achieve higher detection accuracy for nearby target objects. Regarding the direction of the target object, A is in front of vehicle 12 and B is behind vehicle 12. The influence of A on the behavior of vehicle 12 is greater than that of B. Therefore, in the perception phase, it is expected that the detection accuracy and precision of A are greater than that of B to ensure the safety of the behavior of vehicle 12.

[0156] The target detection model training method provided in the embodiments of this disclosure dynamically calculates the loss weight of each target object by combining the relative state information of each target object in the training samples with that of the specified vehicle during the training process. This improves the recognition accuracy and recall rate of target objects that have a significant impact on the behavior of the specified vehicle, such as objects at close range and objects in front, in intelligent driving scenarios. Compared with the loss function design that uses the same loss weight for all target objects, the method provided in the embodiments of this disclosure has stronger applicability to intelligent driving scenarios and can better meet the perception needs of intelligent driving scenarios.

[0157] The embodiments described above can be implemented individually or in any combination without conflict. The specific implementation can be set according to actual needs, and this disclosure does not limit them.

[0158] The training method for any object detection model provided in this disclosure can be executed by any suitable electronic device with data processing capabilities, including but not limited to: terminal devices and servers. Alternatively, the training method for any object detection model provided in this disclosure can be executed by a processor, such as by a processor executing the training method for any object detection model mentioned in this disclosure by calling corresponding instructions stored in memory. Further details will not be elaborated below.

[0159] Figure 8 This is a schematic flowchart illustrating a target object detection method provided in an example embodiment of this disclosure. Embodiments of this disclosure can be applied to electronic devices, specifically, for example, to in-vehicle computing platforms (or in-vehicle terminals). Figure 8 As shown, the target object detection method provided in the embodiments of this disclosure may include the following steps:

[0160] Step 410: Obtain sensor data collected by the sensors on the vehicle.

[0161] The sensor data includes at least one of images and radar point clouds.

[0162] In some optional embodiments, the sensor data includes data types consistent with the training samples during the model training phase. For example, if the training samples are image samples, then the sensor data used in the real-time detection process is images; if the training samples are radar point cloud samples, then the sensor data used in the real-time detection process is radar point cloud.

[0163] Step 420: Based on sensor data, determine the target object detection result using a pre-trained target detection model.

[0164] The target detection model is obtained based on the training method provided in any of the above embodiments. The specific training process can be found in the corresponding embodiments described above, and will not be repeated here. The target object detection result may include the position and type of the target object.

[0165] In some optional embodiments, the target detection model performs target detection on sensor data, outputting a predicted position value and a predicted type value for the target object. The position prediction value determines the target object's location, and the predicted type value is a probability value. Post-processing can be performed based on these probability values ​​to obtain the target object's type. For example, the predicted type value includes the probability that the target object belongs to each of several preset types (e.g., vehicle, pedestrian, cyclist, etc.). The target object's type is determined based on the probabilities corresponding to each preset type. For instance, the preset type with the highest probability is taken as the target object's type.

[0166] The target object detection method provided in the embodiments of this disclosure, since the target object detection model used is obtained based on the training method provided in the above embodiments, can improve the recall rate of target objects around the vehicle that have a significant impact on the vehicle's behavior, effectively improve the driving safety of the vehicle, and the target object detection results are more suitable for the needs of intelligent driving scenarios.

[0167] The embodiments described above can be implemented individually or in any combination without conflict. The specific implementation can be set according to actual needs, and this disclosure does not limit them.

[0168] The target object detection method provided in this disclosure can be executed by any suitable electronic device with data processing capabilities, including but not limited to: terminal devices and servers. Alternatively, the target object detection method provided in this disclosure can be executed by a processor, such as by a processor executing any of the target object detection methods mentioned in this disclosure by calling corresponding instructions stored in memory. Further details will not be elaborated below.

[0169] Exemplary device

[0170] Figure 9 This is a schematic diagram of the structure of a training apparatus for an object detection model provided in an exemplary embodiment of this disclosure. The apparatus of this embodiment can be used to implement corresponding training method embodiments of the object detection model of this disclosure, such as... Figure 9 The apparatus shown may include: a first processing module 51, a second processing module 52, a third processing module 53, a fourth processing module 54, a fifth processing module 55, a sixth processing module 56, and a seventh processing module 57.

[0171] The first processing module 51 is used to determine the annotation information of the target object in each training sample in at least one training sample; each training sample includes at least one of image samples collected by sensors on a specified vehicle and radar point cloud samples, and the annotation information of the target object includes the location label and type label of the target object.

[0172] The second processing module 52 is used to determine the relative state information of each target object in the training sample relative to the specified vehicle for each training sample.

[0173] The third processing module 53 is used to perform target detection on at least one training sample through the detection model to be trained, and obtain the detection result of at least one target object. The detection result of any target object includes the position detection value and type detection value of the target object.

[0174] The fourth processing module 54 is used to determine the loss weight corresponding to each target object based on the relative state information of each target object relative to the specified vehicle and the influence of target objects in different relative states on the behavior of the specified vehicle.

[0175] The fifth processing module 55 is used to determine the sub-loss corresponding to each target object based on the position label and position detection value, type label and type detection value, and loss weight corresponding to each target object.

[0176] The sixth processing module 56 is used to update the network parameters of the detection model to be trained based on the sub-loss corresponding to each target object.

[0177] The first to sixth processing modules iteratively execute operations to determine the annotation information of target objects in each training sample in at least one training sample, and update the network parameters of the detection model to be trained based on the sub-loss corresponding to each target object, until the preset training completion conditions are met.

[0178] The seventh processing module 57 is used to determine the updated detection model as the target detection model in response to the satisfaction of the preset training completion conditions.

[0179] In some alternative embodiments, in the above... Figure 9 Based on the illustrated embodiment, the second processing module 52 can specifically be used to: determine the first pose information of each target object in the training sample in the vehicle coordinate system of the specified vehicle for each training sample; and determine the relative state information of each target object based on the first pose information of each target object.

[0180] In some optional embodiments, based on any of the above embodiments, the fourth processing module 54 may specifically be used to: determine the first relative distance, first relative direction, first relative speed, and first relative motion direction of each target object relative to the designated vehicle, based on the relative state information of each target object. Based on the first relative distance, first relative direction, first relative speed, first relative motion direction of each target object, and the pre-configured mapping relationship between the loss weights and the relative distance, relative direction, relative speed, and relative motion direction, determine the loss weight corresponding to each target object.

[0181] In some alternative embodiments, based on any of the above embodiments, the location label of the target object includes a second location label of the target object in the vehicle coordinate system of the specified vehicle.

[0182] The second processing module 52 can be specifically used to: determine the relative state information of each target object relative to the specified vehicle based on the second position label of each target object in the training samples.

[0183] In some optional embodiments, the relative state information of the target object relative to the designated vehicle includes a second relative distance and a second relative direction of the target object relative to the designated vehicle.

[0184] The fourth processing module 54 can be specifically used to: determine the loss weight corresponding to each target object based on the second relative distance, the second relative direction, and the first mapping relationship between the loss weight and the relative distance and the relative direction; the first mapping relationship includes a first sub-relationship corresponding to the relative distance and a second sub-relationship corresponding to the relative direction; the first sub-relationship is a relationship that decreases as the relative distance increases.

[0185] In some optional embodiments, the second sub-relation is to increase the loss weight of the target object when the second relative direction is the first preset direction, and to decrease the loss weight of the target object when the second relative direction is the second preset direction.

[0186] In some optional embodiments, based on any of the above embodiments, the fourth processing module 54 may specifically be used for:

[0187] Based on the relative state information of each target object, the type label of each target object, and the relationship between the influence of different relative states and different types of target objects on the behavior of the specified vehicle, the loss weight corresponding to each target object is determined.

[0188] In some optional embodiments, based on any of the above embodiments, the training samples are image samples; the location label of the target object includes the first location label of the target object in the corresponding image sample; the location detection value of the target object includes the first location detection value of the target object in the corresponding image sample.

[0189] The fifth processing module 55 can be specifically used to: determine the position detection error corresponding to each target object based on the first position label and the first position detection value of each target object; determine the type detection error corresponding to each target object based on the type label and the type detection value of each target object; and determine the sub-loss corresponding to each target object based on the loss weight, position detection error, and type detection error corresponding to each target object.

[0190] In some optional embodiments, based on any of the above embodiments, the training samples are radar point cloud samples; the position label of the target object includes a second position label of the target object in the vehicle coordinate system of the specified vehicle; the position detection value of the target object includes the second position detection value of the target object in the vehicle coordinate system.

[0191] The fifth processing module 55 can be specifically used to: determine the position detection error corresponding to each target object based on the second position label and the second position detection value of each target object; determine the type detection error corresponding to each target object based on the type label and type detection value of each target object; and determine the sub-loss corresponding to each target object based on the loss weight, position detection error, and type detection error corresponding to each target object.

[0192] The embodiments described above can be implemented individually or in any combination without conflict. The specific implementation can be set according to actual needs, and this disclosure does not limit them.

[0193] The beneficial technical effects corresponding to the exemplary embodiments of this device can be found in the corresponding beneficial technical effects in the exemplary method section above, and will not be repeated here.

[0194] Figure 10 This is a schematic diagram of a target object detection apparatus provided in an exemplary embodiment of this disclosure. The apparatus of this embodiment can be used to implement corresponding target object detection method embodiments of this disclosure, such as... Figure 10 The apparatus shown may include an acquisition module 61 and a detection processing module 62.

[0195] The acquisition module 61 is used to acquire sensor data collected by sensors on the vehicle; the sensor data includes at least one of images and radar point clouds.

[0196] The detection processing module 62 is used to determine the detection result of the target object based on sensor data and a pre-trained target detection model.

[0197] The target detection model is obtained based on the training method of the target detection model provided in any of the above embodiments.

[0198] The beneficial technical effects corresponding to the exemplary embodiments of this device can be found in the corresponding beneficial technical effects in the exemplary method section above, and will not be repeated here.

[0199] Exemplary electronic devices

[0200] Figure 11 This is a structural diagram of an electronic device provided in an embodiment of the present disclosure, including at least one processor 91 and a memory 92.

[0201] The processor 91 may be a central processing unit (CPU) or other form of processing unit with data processing and / or instruction execution capabilities, and may control other components in the electronic device 90 to perform desired functions.

[0202] The memory 92 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 91 may execute one or more computer program instructions to implement the methods and / or other desired functions of the various embodiments of this disclosure described above.

[0203] In one example, the electronic device 90 may also include an input device 93 and an output device 94, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).

[0204] The input device 93 may also include, for example, a keyboard, mouse, touchscreen, microphone, various sensors, etc. Sensors may include, for example, image sensors (e.g., cameras, webcams), LiDAR, millimeter-wave radar, ultrasonic radar, positioning sensors, pressure sensors, air quality sensors, temperature sensors, etc. Image sensors, LiDAR, millimeter-wave radar, ultrasonic radar, etc., can be used for environmental perception, i.e., detecting moving and static objects in the surrounding environment. Moving and static objects may include, for example, static objects such as lane lines, curbs, arrows, signs, trees, and buildings, as well as dynamic objects such as surrounding vehicles, pedestrians, and cyclists. Positioning sensors are used to locate the mobile device (e.g., a bicycle, a robot, etc.) where the electronic device is located. Positioning sensors may include, for example, an Inertial Measurement Unit (IMU) and a Global Positioning System (GPS). Pressure sensors can be used to detect seat pressure. Temperature sensors can be used to detect the temperature inside the vehicle cabin. Air quality sensors can be used to detect the air quality inside the vehicle cabin.

[0205] The output device 94 can output various information to the outside, including, for example, a display, a speaker, a communication network and its connected remote output devices, etc.

[0206] Of course, for the sake of simplicity, Figure 11 Only some of the components of the electronic device 90 relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device 90 may include any other suitable components depending on the specific application.

[0207] Exemplary computer program products and computer-readable storage media

[0208] In addition to the methods and apparatus described above, embodiments of this disclosure may also provide a computer program product, including computer program instructions that, when executed by a processor, cause the processor to perform the steps of the methods in the various embodiments of this disclosure described in the "Exemplary Methods" section above.

[0209] Computer program products can be written in any combination of one or more programming languages ​​to perform the operations of embodiments of this disclosure. These programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0210] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps of the methods in the various embodiments of this disclosure described in the "Exemplary Methods" section above.

[0211] Computer-readable storage media may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may include, but is not limited to, systems, apparatuses, or devices that are electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0212] The basic principles of this disclosure have been described above with reference to specific embodiments. However, the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.

[0213] Various modifications and variations can be made to this disclosure without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this disclosure and their equivalents, this disclosure is also intended to include such modifications and variations.

Claims

1. A method for training an object detection model, comprising: Determine the annotation information of the target object in each of the training samples in at least one training sample; Each of the training samples includes at least one of image samples collected by sensors on a specified vehicle and radar point cloud samples, and the annotation information of the target object includes the location label and type label of the target object; For each training sample, determine the relative state information of each target object in the training sample relative to the specified vehicle; The detection model to be trained is used to perform target detection on the at least one training sample to obtain the detection result of at least one target object. The detection result of any target object includes the position detection value and type detection value of the target object. Based on the relative state information of each target object relative to the designated vehicle, and the influence of target objects in different relative states on the behavior of the designated vehicle, the loss weight corresponding to each target object is determined. Based on the location label and location detection value, type label and type detection value of each target object, and the loss weight corresponding to each target object, a sub-loss corresponding to each target object is determined. The network parameters of the detection model to be trained are updated based on the sub-loss corresponding to each of the target objects. The process iteratively executes the steps of determining the annotation information of target objects in each of the training samples in at least one training sample and updating the network parameters of the detection model to be trained based on the sub-loss corresponding to each target object, until the preset training completion condition is met, and the target detection model is obtained from the detection model to be trained.

2. The method according to claim 1, wherein, Determining the relative state information of each target object in the training sample relative to the designated vehicle includes: Determine the first pose information of each target object in the training sample in the vehicle coordinate system of the specified vehicle; Based on the first pose information of each of the target objects, the relative state information of each of the target objects is determined.

3. The method according to claim 1, wherein, The step of determining the loss weight corresponding to each target object based on the relative state information of each target object relative to the designated vehicle, and the influence relationship between target objects in different relative states on the behavior of the designated vehicle, includes: Based on the relative state information of each target object, the first relative distance, first relative direction, first relative speed and first relative motion direction of each target object relative to the designated vehicle are determined respectively; Based on the first relative distance, first relative direction, first relative velocity, first relative motion direction of each target object, and the pre-configured mapping relationship between the loss weight and the relative distance, relative direction, relative velocity, and relative motion direction, the loss weight corresponding to each target object is determined.

4. The method according to claim 1, wherein, The training samples are image samples; the location label of the target object includes the first location label of the target object in the corresponding image sample; The position detection value of the target object includes the first position detection value of the target object in the corresponding image sample; The step of determining the sub-loss corresponding to each target object based on the location label and location detection value, type label and type detection value, and the loss weight corresponding to each target object includes: Based on the first position label and the first position detection value of each target object, the position detection error corresponding to each target object is determined respectively; Based on the type label and type detection value of each target object, the type detection error corresponding to each target object is determined respectively; Based on the loss weight, position detection error, and type detection error corresponding to each target object, a sub-loss corresponding to each target object is determined.

5. The method according to claim 1, wherein, The training samples are radar point cloud samples; the position label of the target object includes the second position label of the target object in the vehicle coordinate system of the specified vehicle; the position detection value of the target object includes the second position detection value of the target object in the vehicle coordinate system. The step of determining the sub-loss corresponding to each target object based on the location label and location detection value, type label and type detection value, and the loss weight corresponding to each target object includes: Based on the second position label and the second position detection value of each target object, the position detection error corresponding to each target object is determined respectively; Based on the type label and type detection value of each target object, the type detection error corresponding to each target object is determined respectively; Based on the loss weight, position detection error, and type detection error corresponding to each target object, a sub-loss corresponding to each target object is determined.

6. The method according to any one of claims 1-5, wherein, The location label of the target object includes a second location label of the target object in the vehicle coordinate system of the specified vehicle; Determining the relative state information of each target object in the training sample relative to the designated vehicle includes: Based on the second location label of each target object in the training sample, the relative state information of each target object relative to the designated vehicle is determined.

7. The method according to claim 6, wherein, The relative state information of the target object relative to the designated vehicle includes a second relative distance and a second relative direction of the target object relative to the designated vehicle; The step of determining the loss weight corresponding to each target object based on the relative state information of each target object relative to the designated vehicle, and the influence relationship between target objects in different relative states on the behavior of the designated vehicle, includes: Based on the second relative distance, the second relative direction, and the first mapping relationship between the loss weight and the relative distance and relative direction of each target object, the loss weight corresponding to each target object is determined respectively; the first mapping relationship includes a first sub-relationship corresponding to the relative distance and a second sub-relationship corresponding to the relative direction; the first sub-relationship is a relationship that decreases as the relative distance increases.

8. The method according to claim 7, wherein, The second sub-relation is to increase the loss weight of the target object when the second relative direction is the first preset direction, and to decrease the loss weight of the target object when the second relative direction is the second preset direction.

9. The method according to any one of claims 1-5, wherein, The step of determining the loss weight corresponding to each target object based on the relative state information of each target object relative to the designated vehicle, and the influence relationship between target objects in different relative states on the behavior of the designated vehicle, includes: Based on the relative state information of each target object, the type label of each target object, and the influence of different relative states and different types of target objects on the behavior of the specified vehicle, the loss weight corresponding to each target object is determined.

10. A method for detecting a target object, comprising: Acquire sensor data collected by sensors on the vehicle; The sensor data includes at least one of images and radar point clouds; Based on the sensor data, the target object detection result is determined using a pre-trained target detection model; The target detection model is obtained based on the training method of the target detection model according to any one of claims 1-9.

11. A training device for an object detection model, comprising: The first processing module is used to determine the annotation information of the target object in each of the training samples in at least one training sample; Each of the training samples includes at least one of image samples collected by sensors on a specified vehicle and radar point cloud samples, and the annotation information of the target object includes the location label and type label of the target object; The second processing module is used to determine the relative state information of each target object in the training sample relative to the specified vehicle for each training sample. The third processing module is used to perform target detection on the at least one training sample through the detection model to be trained, and obtain the detection result of at least one target object. The detection result of any target object includes the position detection value and type detection value of the target object. The fourth processing module is used to determine the loss weight corresponding to each of the target objects based on the relative state information of each target object relative to the designated vehicle and the influence of target objects in different relative states on the behavior of the designated vehicle. The fifth processing module is used to determine the sub-loss corresponding to each target object based on the position label and position detection value, type label and type detection value, and the loss weight corresponding to each target object respectively; The sixth processing module is used to update the network parameters of the detection model to be trained based on the sub-loss corresponding to each of the target objects. The first to the sixth processing modules iteratively execute the operations of determining the annotation information of target objects in each of the training samples in at least one training sample and updating the network parameters of the detection model to be trained based on the sub-loss corresponding to each target object, until the preset training completion conditions are met, and the target detection model is obtained from the detection model to be trained.

12. A device for detecting a target object, comprising: The acquisition module is used to acquire sensor data collected by sensors on the vehicle. The sensor data includes at least one of images and radar point clouds; The detection processing module is used to determine the target object detection result based on the sensor data and a pre-trained target detection model. The target detection model is obtained based on the training method of the target detection model according to any one of claims 1-9.

13. An electronic device, the electronic device comprising: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method of any one of claims 1-9, or to execute the instructions to implement the method of claim 10. or, The electronic device includes: a training device for the target detection model as described in claim 11 or a target object detection device as described in claim 12.

14. A computer-readable storage medium storing a computer program for performing the method of any one of claims 1-9, or for performing the method of claim 10.

15. A computer program product, wherein when instructions in the computer program product are executed by a processor, it performs the method described in any one of claims 1-9 of the present disclosure, or performs the method described in claim 10 of the present disclosure.