Method and device for collecting and processing output data of autonomous driving perception algorithm model
Through data fusion and scorer judgment methods, we collected and utilized supplementary training sets for the autonomous driving perception algorithm model, solved the problem of abnormal output of the model under limited training sets, and improved the performance and adaptability of the model.
Patent Information
- Application Number
- CN202210556973.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-19
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2042-05-19
AI Technical Summary
Existing autonomous driving perception algorithm models are prone to abnormal output results when the data covered by the training set is limited, and lack the use of data in scenarios of interest to target detection personnel, resulting in limited improvement in model performance.
Through data fusion, data from multiple vehicle-mounted sensors are input into the perception algorithm model, the target perception results are time-series, and a scorer is used to determine whether they meet the preset scenarios. Data that meets the scenarios are collected as a supplementary training set to retrain the model.
It effectively improves the training effect of the perception algorithm model for preset scenarios, optimizes model performance, and improves the accuracy and adaptability of target perception.
Smart Images

Figure CN114972911B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving technology, and specifically provides a method for collecting and processing output data of an autonomous driving perception algorithm model, an electronic device, a storage medium, and a vehicle. Background Art
[0002] In autonomous driving scenarios, perception algorithms typically use data from cameras, radars, and various sensors to predict and then fuse features to achieve perception of surrounding targets. Perception algorithms are typically implemented using models based on deep learning and neural networks, which are trained using training sets to achieve their functionality. However, due to the limited data coverage of the training set, there may be some scenarios that do not appear in the training set. This can cause the model to produce abnormal output results, known as bad cases, when making predictions. These abnormal output results are often very valuable in improving model performance. Furthermore, data on scenes of interest to target detectors output during training is also valuable in improving model performance. How to effectively utilize these output results to improve model performance is a challenge that needs to be addressed in this field. Summary of the Invention
[0003] In order to overcome the above-mentioned defects, the present invention is proposed to provide a solution or at least partially solve the problem of how to effectively collect supplementary training sets of preset scenarios to improve the performance of the perception algorithm model.
[0004] In a first aspect, the present invention provides a method for collecting and processing output data of an autonomous driving perception algorithm model, comprising:
[0005] fusing data collected by multiple vehicle-mounted sensors, inputting the fusion results into the perception algorithm model, and obtaining the target perception results output by the perception algorithm model;
[0006] Sequencing the target perception result into a plurality of perception frame data;
[0007] Determine whether the current perception frame data conforms to the preset scenario;
[0008] Collecting data collected by the vehicle-mounted sensor within a time window of a preset length where the current frame that meets the preset scenario is located as a supplementary training set for the perception algorithm model;
[0009] The supplementary training set is applied to retraining of the perception algorithm model to optimize the performance of the perception algorithm model.
[0010] In one technical solution of the above method, the determining whether the current perception frame data conforms to a preset scenario includes:
[0011] Different scorers are set for different preset scenarios, wherein each scorer sets its own scoring criteria and its own scoring weight based on the corresponding preset scenario;
[0012] Inputting the plurality of perception frame data into the different scorers respectively, and obtaining a scoring result of each scorer on the current perception frame data based on the scoring criteria;
[0013] Based on the scoring weight of each scorer, the scoring results of all scorers are weighted averaged to obtain the evaluation score of the current perception frame data; and
[0014] If the evaluation score exceeds a predetermined threshold, it is determined that the current perception frame data meets the preset scenario.
[0015] In one technical solution of the above method, the preset scenario includes an abnormal scenario that occurs during the target perception process;
[0016] The step of inputting the plurality of perception frame data into the different scorers respectively, and obtaining a scoring result of each scorer on the current perception frame data based on the scoring criteria, comprises:
[0017] Based on the scoring criteria, a scoring result for the current perception frame data is obtained by comparing the current perception frame data with perception frame data before the current frame.
[0018] In a technical solution of the above method, the preset scene includes a scene of interest that appears during the target perception process;
[0019] The step of inputting the plurality of perception frame data into the different scorers respectively and obtaining the scoring result of each scorer on the current perception frame data based on the scoring criteria includes: obtaining the scoring result of the current perception frame data by analyzing the current perception frame data based on the scoring criteria.
[0020] In one technical solution of the above method, different scorers are set for different preset scenarios, wherein each scorer sets its own scoring criteria and its own scoring weight based on the corresponding preset scenario, including: the scoring weight set for the scorer corresponding to the scene of interest is greater than the scoring weight set for the scorer corresponding to the abnormal scene.
[0021] In a technical solution of the above method, the method further includes:
[0022] The data collected by the vehicle-mounted sensor is subjected to a time series conversion after being subjected to another perception algorithm model different from the perception algorithm model to obtain a plurality of sensor frame data;
[0023] Determining whether the current sensor frame data conforms to the preset scenario;
[0024] The data collected by the vehicle-mounted sensor within a time window of a preset length where the current frame that meets the preset scenario is located is collected as a supplementary training set for the perception algorithm model.
[0025] In one technical solution of the above method, collecting data acquired by the vehicle-mounted sensor within a time window of a preset length where the current frame that meets the preset scenario is located as a supplementary training set for the perception algorithm model includes:
[0026] When the current perception frame data meets the preset scenario, the data collected by the on-board sensor within the time window of the preset length where the perception frame data is located is transmitted back as a supplementary training set for the perception algorithm model.
[0027] In one technical solution of the above method, the vehicle-mounted sensor includes a vehicle-mounted camera and a vehicle-mounted laser radar;
[0028] The step of fusing data collected by multiple vehicle-mounted sensors, inputting the fusion result into the perception algorithm model, and obtaining the target perception result output by the perception algorithm model includes:
[0029] Acquire 2D visual data from the vehicle’s onboard camera;
[0030] Obtain 3D point cloud data from vehicle-mounted lidar;
[0031] Projecting the two-dimensional visual data from the image coordinate system to the camera coordinate system to obtain a first projection result;
[0032] Projecting the first projection result into the world coordinate system according to a conversion relationship between the camera coordinate system and the world coordinate system to obtain a second projection result;
[0033] Dedistorting the second projection result to obtain a third projection result;
[0034] Projecting the three-dimensional point cloud data into a world coordinate system to obtain a fourth projection result;
[0035] Performing data fusion on the third projection result and the fourth projection result, and projecting the data into a two-dimensional space to obtain a fusion result;
[0036] The fusion result is input into the perception algorithm model to obtain the target perception result.
[0037] In a second aspect, an electronic device is provided, which includes a processor and a storage device, wherein the storage device is suitable for storing multiple program codes, and the program codes are suitable for being loaded and run by the processor to execute the output data collection and processing method of the autonomous driving perception algorithm model described in any one of the technical solutions of the above-mentioned output data collection and processing method of the autonomous driving perception algorithm model.
[0038] In a third aspect, a computer-readable storage medium is provided, which stores a plurality of program codes, wherein the program codes are suitable for being loaded and run by a processor to execute the output data collection and processing method of the autonomous driving perception algorithm model described in any one of the technical solutions of the above-mentioned output data collection and processing method of the autonomous driving perception algorithm model.
[0039] In a fourth aspect, a vehicle is provided, comprising the electronic device described in the above-mentioned electronic device technical solution.
[0040] The above one or more technical solutions of the present invention have at least one or more of the following beneficial effects:
[0041] In implementing the technical solution of the present invention, the data collected by the on-board sensors can be fused and input into the perception algorithm model to obtain the target perception result, the target perception result can be time-series to obtain multiple perception frame data, and a judgment can be made as to whether each frame of the perception frame data conforms to the preset scene. The data collected by the on-board sensors within the time window where the perception frame data conforms to the preset scene is located is used as a supplementary training set of the perception algorithm model. The perception algorithm model can be retrained by using the supplementary training set, which can achieve more effective training of the perception algorithm model for the preset scene and better optimize and improve the performance of the perception algorithm model.
[0042] Solution 1. A method for collecting and processing output data of an autonomous driving perception algorithm model, comprising:
[0043] fusing data collected by multiple vehicle-mounted sensors, inputting the fusion results into the perception algorithm model, and obtaining the target perception results output by the perception algorithm model;
[0044] Sequencing the target perception result into a plurality of perception frame data;
[0045] Determine whether the current perception frame data conforms to the preset scenario;
[0046] Collecting data collected by the vehicle-mounted sensor within a time window of a preset length where the current frame that meets the preset scenario is located as a supplementary training set for the perception algorithm model;
[0047] The supplementary training set is applied to retraining of the perception algorithm model to optimize the performance of the perception algorithm model.
[0048] Solution 2. The method according to Solution 1, wherein determining whether the current perception frame data conforms to a preset scenario comprises:
[0049] Different scorers are set for different preset scenarios, wherein each scorer sets its own scoring criteria and its own scoring weight based on the corresponding preset scenario;
[0050] Inputting the plurality of perception frame data into the different scorers respectively, and obtaining a scoring result of each scorer on the current perception frame data based on the scoring criteria;
[0051] Based on the scoring weight of each scorer, the scoring results of all scorers are weighted averaged to obtain the evaluation score of the current perception frame data; and
[0052] If the evaluation score exceeds a predetermined threshold, it is determined that the current perception frame data meets the preset scenario.
[0053] Solution 3. The method according to Solution 2, characterized in that:
[0054] The preset scenarios include abnormal scenarios that occur during target perception;
[0055] The step of inputting the plurality of perception frame data into the different scorers respectively, and obtaining a scoring result of each scorer on the current perception frame data based on the scoring criteria, comprises:
[0056] Based on the scoring criteria, a scoring result for the current perception frame data is obtained by comparing the current perception frame data with perception frame data before the current frame.
[0057] Solution 4. The method according to Solution 3, characterized in that:
[0058] The preset scenes include scenes of interest that appear during target perception;
[0059] The step of inputting the plurality of perception frame data into the different scorers respectively, and obtaining a scoring result of each scorer on the current perception frame data based on the scoring criteria, comprises:
[0060] Based on the scoring criteria, a scoring result of the current perception frame data is obtained by analyzing the current perception frame data.
[0061] Solution 5. The method according to Solution 4 is characterized in that different scorers are set for different preset scenarios, wherein each scorer sets its own scoring criteria and its own scoring weight based on the corresponding preset scenario, including: the scoring weight set for the scorer corresponding to the scene of interest is greater than the scoring weight set for the scorer corresponding to the abnormal scene.
[0062] Solution 6. The method according to Solution 1, further comprising:
[0063] The data collected by the vehicle-mounted sensor is subjected to a time series conversion after being subjected to another perception algorithm model different from the perception algorithm model to obtain a plurality of sensor frame data;
[0064] Determining whether the current sensor frame data conforms to the preset scenario;
[0065] The data collected by the vehicle-mounted sensor within a time window of a preset length where the current frame that meets the preset scenario is located is collected as a supplementary training set for the perception algorithm model.
[0066] Solution 7. The method according to Solution 1, characterized in that
[0067] The collecting of data collected by the vehicle-mounted sensor within a time window of a preset length where the current frame that meets the preset scenario is located as a supplementary training set for the perception algorithm model includes:
[0068] When the current perception frame data meets the preset scenario, the data collected by the on-board sensor within the time window of the preset length where the perception frame data is located is transmitted back as a supplementary training set for the perception algorithm model.
[0069] Solution 8. The method according to any one of Solutions 1-6, wherein the vehicle-mounted sensor includes a vehicle-mounted camera and a vehicle-mounted laser radar;
[0070] The step of fusing data collected by multiple vehicle-mounted sensors, inputting the fusion result into the perception algorithm model, and obtaining the target perception result output by the perception algorithm model includes:
[0071] Acquire 2D visual data from the vehicle’s onboard camera;
[0072] Obtain 3D point cloud data from vehicle-mounted lidar;
[0073] Projecting the two-dimensional visual data from the image coordinate system to the camera coordinate system to obtain a first projection result;
[0074] Projecting the first projection result into the world coordinate system according to a conversion relationship between the camera coordinate system and the world coordinate system to obtain a second projection result;
[0075] Dedistorting the second projection result to obtain a third projection result;
[0076] Projecting the three-dimensional point cloud data into a world coordinate system to obtain a fourth projection result;
[0077] Performing data fusion on the third projection result and the fourth projection result, and projecting the data into a two-dimensional space to obtain a fusion result;
[0078] The fusion result is input into the perception algorithm model to obtain the target perception result.
[0079] Solution 9. An electronic device comprising a processor and a storage device, wherein the storage device is suitable for storing multiple program codes, and is characterized in that the program codes are suitable for being loaded and run by the processor to execute any one of the methods of Solutions 1 to 8.
[0080] Solution 10. A computer-readable storage medium storing a plurality of program codes, wherein the program codes are suitable for being loaded and executed by a processor to execute the method according to any one of Solutions 1 to 8.
[0081] Solution 11. A vehicle, characterized in that the vehicle includes the electronic device described in Solution 9. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] The disclosure of the present invention will become more easily understood with reference to the accompanying drawings. Those skilled in the art will readily appreciate that these drawings are for illustrative purposes only and are not intended to limit the scope of protection of the present invention. Among them:
[0083] Figure 1 This is a flowchart of the main steps of a method for collecting and processing output data of an autonomous driving perception algorithm model according to one embodiment of the present invention;
[0084] Figure 2 is a schematic diagram of target perception results according to an example of an embodiment of the present invention;
[0085] Figure 3 This is a flowchart of the main steps of obtaining target perception results according to an embodiment of the present invention;
[0086] Figure 4 This is a flowchart of main steps for determining whether current perception frame data conforms to a preset scenario according to an embodiment of the present invention;
[0087] Figure 5 1 is a flowchart illustrating the main steps of obtaining a supplementary training set for a perception algorithm model based on sensor frame data according to an embodiment of the present invention;
[0088] Figure 6 1 is a flow chart of the main steps of a method for training an autonomous driving perception algorithm model according to an embodiment of the present invention. DETAILED DESCRIPTION
[0089] Some embodiments of the present invention are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0090] In the description of the present invention, "module" and "processor" may include hardware, software, or a combination of both. A module may include hardware circuitry, various suitable sensors, communication ports, and memory. It may also include software components, such as program code, or a combination of software and hardware. A processor may be a central processing unit, a microprocessor, an image processor, a digital signal processor, or any other suitable processor. A processor has data and / or signal processing capabilities. A processor may be implemented in software, hardware, or a combination of both. Non-transitory computer-readable storage media include any suitable medium capable of storing program code, such as magnetic disks, hard disks, optical disks, flash memory, read-only memory, random access memory, and the like. The term "A and / or B" refers to all possible combinations of A and B, such as only A, only B, or both A and B. The terms "at least one of A or B" or "at least one of A and B" have similar meanings to "A and / or B" and may include only A, only B, or both A and B. The singular forms "one" and "the" may also include the plural forms.
[0091] In the field of autonomous driving, perception algorithms are often used to perceive the surrounding environment and assist vehicles in achieving autonomous driving. Therefore, the accuracy of the perception results output by these algorithms is crucial to the vehicle's autonomous driving capabilities. However, existing technologies often suffer from a certain degree of anomalies in the perception results output by these algorithms, necessitating further improvements in their performance.
[0092] See attached Figure 1 , Figure 1 1 is a flow chart of the main steps of a method for collecting and processing output data of an autonomous driving perception algorithm model according to an embodiment of the present invention.
[0093] like Figure 1 As shown, the output data collection and processing method of the autonomous driving perception algorithm model in the embodiment of the present invention mainly includes the following steps S101 to S105.
[0094] Step S101: Fusing data collected by multiple vehicle-mounted sensors, inputting the fusion results into a perception algorithm model, and obtaining target perception results output by the perception algorithm model.
[0095] In this embodiment, data collected by multiple vehicle-mounted sensors can be fused, and the fusion results can be input into the perception algorithm model. The output of the perception algorithm model is the target perception result. Among them, the vehicle-mounted sensors are, for example, vehicle-mounted cameras and lidars. Correspondingly, the collected data can be image data and point cloud data. The target perception result can include the classification and location information of other vehicles, obstacles, lane lines, traffic lights and other targets in the vehicle's surrounding environment. Please refer to the attached Figure 2 , Figure 2 FIG. 1 is a schematic diagram of target perception results according to an example of an embodiment of the present invention. Figure 2 As shown, different vehicles can be classified in the target perception results, including the vehicle 1 and the other vehicle 2 in front.
[0096] Step S102: Sequencing the target perception result into a plurality of perception frame data.
[0097] In this embodiment, a time sequence operation may be performed on the target perception result to obtain a plurality of perception frame data.
[0098] In one embodiment, the target perception result can be framed according to a preset time interval and time-sequenced to obtain multiple perception frame data. In a specific example, 200 frames of data are collected.
[0099] Step S103: Determine whether the current perception frame data conforms to a preset scenario.
[0100] In this embodiment, for each perception frame data, a judgment may be made on the current perception frame data to determine whether the current perception frame data conforms to a preset scenario.
[0101] In one embodiment, the preset scenario may be an abnormal scenario that occurs during the target perception process, that is, a situation in which the target perception result obtained by the perception algorithm model is erroneous, that is, a bad case.
[0102] For example, in the previous frame data of the current perception frame data (such as the 5th frame), there is a car within 10-15 meters in front of the vehicle on the straight road, but in the current frame data, the car disappears, which is regarded as an abnormal scene. This type of abnormal scene is called "inter-frame result target disappearance."
[0103] Furthermore, when there are a large number of targets, there may be multiple targets overlapping with each other. In this case, only individual targets disappear in different frames, and it is not necessarily defined as "target disappearance between frames". This rule can be set according to the experience of the target detection personnel (such as the number of targets, degree of overlap, etc.).
[0104] For another example, during the perception process, each identified target will be assigned a different ID. For targets with the same ID, the ID value remains unchanged in different frames, so that there will be no repeated numbers within a certain period of time. However, due to the influence of the model itself or environmental factors, when a target disappears for a large number of frames (the first type of abnormal scenario), when the target reappears, the perception algorithm model will consider it to be a new target and assign it a new ID. There will be a situation where the ID jumps between different frames, which is considered an abnormal scenario. This type of abnormal scenario is called "inter-frame same-target track_id jump". In this abnormal scenario, it is preferred to set the current frame to be compared with as many previous frames as possible.
[0105] For example, in one frame, the target is 5 meters ahead of the vehicle. In the next frame, it's 5.5 meters ahead, and in the next frame, it's 6 meters ahead. This is a normal scenario. However, if the perception result shows the target is 10 meters ahead in the next frame, this is also considered an abnormal scenario, known as a "sudden change in the position of a close-range target." In this abnormal scenario, because the target stops while the vehicle continues to move, the 10-meter position may be normal. Comparing only two adjacent frames may lead to misjudgment. Therefore, in this case, it is preferable to set the current frame to be compared with as many previous frames as possible.
[0106] For example, in a certain frame, two targets are detected, one 4 meters ahead and the other 5 meters ahead. However, considering the length of the car itself (for example, 4 meters), the overlap between the two targets is very high, which is considered an abnormal scene. This type of abnormal scene is called "excessive intersection area between targets."
[0107] The above abnormal scenarios generally reflect abnormal perception results.
[0108] In another embodiment, the preset scene can also be a scene of interest that appears during the target perception process. For example, the perception algorithm model used does not have a good perception result for the data of some scenes (such as scenes with dense vehicles). Then the target detection personnel are interested in scenes with a large number of targets (called scenes of interest). Then they can set quantity rules and conduct targeted recovery of data that meets the rules. Similarly, if you are interested in a certain type of target that appears in the target perception results, you can set rules to recover such data. Those skilled in the art can set scenes of interest according to the needs of actual applications.
[0109] Step S104: Collect data collected by the vehicle-mounted sensor within a time window of a preset length where the current frame that meets the preset scenario is located as a supplementary training set for the perception algorithm model.
[0110] In this embodiment, when the current perception frame data meets the preset scenario, the data collected by the vehicle-mounted sensor within the time window where the current perception frame data is located can be used as a supplementary training set for the perception algorithm model.
[0111] For example, the data collected by the sensor within a time window of 5 seconds before and after the current frame is used as a supplementary training set for the perception algorithm model. Those skilled in the art can set the value of the time window according to the needs of actual applications.
[0112] Step S105: Apply the supplementary training set to retrain the perception algorithm model to optimize the performance of the perception algorithm model.
[0113] In this embodiment, the supplementary training set can be used to retrain the perception algorithm model to optimize the performance of the perception algorithm model.
[0114] Based on the above steps S101 to S105, an embodiment of the present invention can fuse the data collected by the on-board sensors and input the data into the perception algorithm model to obtain a target perception result, time-series the target perception result to obtain multiple perception frame data, and judge whether each frame of perception frame data conforms to the preset scene. The data collected by the on-board sensors within the time window where the perception frame data that conforms to the preset scene is located is used as a supplementary training set for the perception algorithm model. The perception algorithm model is retrained by using the supplementary training set, which can achieve more effective training of the perception algorithm model for the preset scene and better optimize and improve the performance of the perception algorithm model.
[0115] Step S101, step S103 and step S104 are further described below.
[0116] In one embodiment of the present invention, see the attached Figure 3 , Figure 3 The figure is a flowchart of the main steps of obtaining target perception results according to an embodiment of the present invention.
[0117] like Figure 3 As shown, step S101 may include the following steps S1011 to S1018:
[0118] Step S1011: Acquire two-dimensional visual data from the vehicle-mounted camera.
[0119] Step S1012: Acquire three-dimensional point cloud data from the vehicle-mounted laser radar.
[0120] Step S1013: Project the two-dimensional visual data from the image coordinate system to the camera coordinate system to obtain a first projection result.
[0121] Step S1014: projecting the first projection result into the world coordinate system according to the conversion relationship between the camera coordinate system and the world coordinate system to obtain a second projection result.
[0122] Step S1015: Dedistort the second projection result to obtain a third projection result.
[0123] Step S1016: Project the three-dimensional point cloud data into the world coordinate system to obtain a fourth projection result.
[0124] Step S1017: fusing the third projection result and the fourth projection result, and projecting them into a two-dimensional space to obtain a fusion result.
[0125] Step S1018: Input the fusion result into the perception algorithm model to obtain the target perception result.
[0126] In this embodiment, the onboard sensors are a camera and a LiDAR. The 2D visual data acquired by the camera is projected from the image coordinate system to the camera coordinate system, then from the camera coordinate system to the world coordinate system and dedistorted to obtain a third projection result. The 3D point cloud data acquired by the LiDAR is projected into the world coordinate system to obtain a fourth projection result. With both the third and fourth projection results in the world coordinate system, data fusion is performed on the third and fourth projection results and projected into 2D space to obtain a fusion result of the 2D visual data and the 3D point cloud data. This fusion result is then input into a perception algorithm model to obtain target perception results.
[0127] In one embodiment of the present invention, see the attached Figure 4 , Figure 4 FIG. 1 is a flow chart showing the main steps of determining whether the current perception frame data conforms to a preset scene according to an embodiment of the present invention. Figure 4 As shown, step S103 may include the following steps S1031 to S1034:
[0128] Step S1031: different scorers are set for different preset scenarios, wherein each scorer sets its own scoring criteria and its own scoring weight based on the corresponding preset scenario.
[0129] In this embodiment, different scorers can be set according to different preset scenarios based on actual application needs. The scorers can be scoring models trained according to different evaluation criteria. For example, the preset scenarios can include abnormal scenarios that occur during target perception and interesting scenarios that occur during target perception.
[0130] Those skilled in the art will understand that the evaluation criteria for various preset scenarios can be set according to actual needs / experience. The present invention does not focus on how to set the evaluation criteria, but rather uses the preset evaluation criteria to screen out abnormal scenario data or data of scenarios of interest for recovery.
[0131] In one embodiment, step S1031 may further include: setting a scoring weight for the scorer corresponding to the scene of interest greater than a scoring weight for the scorer corresponding to the abnormal scene.
[0132] In this embodiment, when the data collected by the on-board sensors corresponding to the scene of interest is needed as a supplementary training set for the perception algorithm model, since the probability of abnormal scenes occurring is relatively high, the scorer corresponding to the scene of interest can be set with a higher scoring weight, and the scorer corresponding to the abnormal scene can be set with a lower weight, so as to facilitate the collection of data collected by the on-board sensors corresponding to the scene of interest. The target detection personnel can flexibly configure the weight according to the type of data that needs to be recovered. For example, if only the data corresponding to certain special scenes needs to be recovered, the weight of the scorer for other scenes can be set to 0.
[0133] Step S1032: inputting the plurality of perception frame data into different scorers respectively, and obtaining the scoring result of each scorer on the current perception frame data based on the scoring criteria.
[0134] In this embodiment, the perception frame data obtained in step S102 may be input into different scorers respectively, and each frame of perception frame data may be scored by the scorer to obtain a scoring result for each frame of perception frame data.
[0135] In one embodiment, the preset scenario is an abnormal scenario that occurs during the target perception process. Step S1032 may include the following steps:
[0136] Based on the scoring criteria, a scoring result for the current perception frame data is obtained by comparing the current perception frame data with the perception frame data before the current frame.
[0137] In this embodiment, when the preset scenario is an abnormal scenario that occurs during the target perception process, the current perception frame data may be compared with the perception frame data before the current perception frame data to obtain a scoring result of the current perception frame data.
[0138] In one embodiment, the preset scene is a scene of interest that occurs during the target perception process.
[0139] Step S1032 may include the following steps:
[0140] Based on the scoring criteria, the scoring result of the current perception frame data is obtained by analyzing the current perception frame data.
[0141] In this embodiment, when the preset scene is a scene of interest that appears during the target perception process, the current perception frame data can be analyzed based on the scoring criteria to obtain a scoring result.
[0142] The following describes step S1032 by taking an example:
[0143] When the preset scenario is an abnormal scenario that occurs during target perception, the scoring criteria may include scoring rules for "target disappearance between frames," "track_id jumps for the same target between frames," "sudden position changes of close-range targets," and "excessive intersection area between targets." Each scoring criterion corresponds to a scorer, and the perception frame data is fed into different scorers to obtain a scoring result for the perception frame data.
[0144] Specifically, the perception frame data is fed into the scorer corresponding to "Inter-frame Target Disappearance," which evaluates the current perception frame data based on the total number of targets in the perception frame data and the number of targets that have disappeared. For example, if the total number of targets in the current perception frame data is large, and considering the occlusion between targets, if one target disappears, this can be considered a normal situation, and the scorer will give a low score. If the total number of targets in the current perception frame data is small and there are targets that have disappeared, the scorer will give a high score, indicating that the current perception frame data is an abnormal scene.
[0145] The perception frame data is sent to the scorer corresponding to the "track_id jump of the same target between frames". The track_id of the target in the current perception frame data is compared with the track_id of the target in the perception frame data before the current perception frame data. When the track_id of the same target changes, the scorer will give a higher score result, that is, the current perception frame data is an abnormal scene.
[0146] The perception frame data is sent to the scorer corresponding to "sudden change in position of close-range targets". The distance between targets in the current perception frame data is compared with the distance between targets in the perception frame data before the current perception frame data. When the distance between targets in adjacent frames changes greatly, the scorer will give a higher score result, that is, the current perception frame data is an abnormal scene.
[0147] The perception frame data is fed into the scorer for "Excessively Large Intersection Area Between Objects." The scorer determines whether the current perception frame data is an abnormal scene based on the intersection area between the objects in the current perception frame data. For example, if objects 1 and 2 in the current perception frame data are both vehicles, and the intersection area between objects 1 and 2 is greater than a preset threshold, the scorer will assign a high score, indicating that the current perception frame data is an abnormal scene.
[0148] When the preset scene is a scene of interest that occurs during target perception, the scoring criterion can be "vehicle density." Specifically, the perception frame data can be fed into a scorer corresponding to "vehicle density." If the number of vehicles in the current perception frame exceeds a preset threshold, the scorer will assign a higher score, indicating that the current perception frame data is a scene of interest.
[0149] Those skilled in the art may set different scoring criteria according to the needs of actual applications, and set a corresponding scorer and the scoring result calculation logic within the scorer according to the scoring criteria.
[0150] Step S1033: Based on the scoring weight of each scorer, the scoring results of all scorers are weighted averaged to obtain an evaluation score for the current perception frame data.
[0151] In this embodiment, the scoring results of each scorer may be weighted averaged to obtain the evaluation score of the current perception frame data.
[0152] In one embodiment, the evaluation score may range from 0 to 1. A higher evaluation score indicates that the current perception frame data is more consistent with the preset scenario.
[0153] Step S1034: If the evaluation score exceeds a predetermined threshold, it is determined that the current perception frame data meets the preset scenario.
[0154] In this embodiment, the evaluation score obtained in step S1033 can be compared with a predetermined threshold to determine whether the current perception frame data conforms to the preset scene. Those skilled in the art can set the value of the predetermined threshold according to the needs of actual application.
[0155] In one implementation of the embodiment of the present invention, step S104 may further include the following steps:
[0156] When the current perception frame data meets the preset scenario, the data collected by the on-board sensor within the preset length time window where the perception frame data is located is transmitted back as a supplementary training set for the perception algorithm model.
[0157] In this embodiment, when the current perception frame data meets the preset scenario, the data collected by the on-board sensor within the time window of the preset length where the perception frame data is located can be triggered to be transmitted back, so that these returned data can be used as a supplementary training set for the perception algorithm model.
[0158] Generally speaking, in order to ensure the real-time and speed of calculations, the vehicle perception algorithm model needs to make trade-offs in terms of real-time and speed when building it. For example, building a lightweight model may lead to limited calculation accuracy, resulting in abnormal target perception results due to the model itself.
[0159] In order to screen out this situation, in one embodiment of the present invention, in addition to the above steps S101 to S104, the present invention may also include the following steps S105 to S107. Figure 5 Steps S105 to S107 are explained. Figure 5 FIG. 1 is a flow chart of the main steps of obtaining a supplementary training set for a perception algorithm model based on sensor frame data according to an embodiment of the present invention. Figure 5 As shown, the present invention may also include:
[0160] Step S105: The data collected by the vehicle-mounted sensor is subjected to a time series after being subjected to other perception algorithm models different from the perception algorithm model to obtain a plurality of sensor frame data.
[0161] Step S106: Determine whether the current sensor frame data conforms to a preset scenario.
[0162] Step S107: The data collected by the vehicle-mounted sensor within the time window of the preset length where the current frame that meets the preset scene is located is used as a supplementary training set for the perception algorithm model.
[0163] In this embodiment, a different perception algorithm model from the perception algorithm model in step S101 is reset. The other perception algorithm model is only used to collect a supplementary training set, so there is no need to consider real-time performance (for example, the perception algorithm model in step S101 requires real-time calculation, while the other perception algorithm model does not, and it can calculate data at intervals). For this reason, it can use a model with higher accuracy and more complex algorithm logic than the perception algorithm model in step S101 to determine whether there is a problem with the perception algorithm model with low accuracy. For example, if the scorer scores the target perception result abnormally after the perception algorithm model in step S101, and the scorer scores the target perception result normal after the other perception algorithm model, it can be determined that there is a problem with the perception algorithm model in step S101 itself.
[0164] The data collected by the vehicle's sensors is fed into other perception algorithm models and then time-sequenced. The output is multiple sensor frames. This sensor frame data can be judged to determine if it matches a pre-defined scenario. If so, the data collected by the vehicle's sensors within the time window of the current perception frame data can be used as a supplementary training set for the perception algorithm model.
[0165] In one embodiment, multiple sensor frame data can be sequentially input into different scorers set according to different preset scenarios, and each frame of sensor frame data can be scored based on the scoring criteria to obtain a scoring result. The scoring results of all scorers are weighted averaged to obtain an evaluation score for the current sensor frame data. If the evaluation score is higher than a predetermined threshold, it is determined that the current sensor frame data meets the preset scenario, and the data collected by the on-board sensor within the preset length time window of the current frame that meets the preset scenario is used as a supplementary training set for the perception algorithm model.
[0166] Furthermore, the present invention also provides a training method for an autonomous driving perception algorithm model.
[0167] See attached Figure 6 , Figure 6 FIG. 1 is a flow chart showing the main steps of a method for training an autonomous driving perception algorithm model according to an embodiment of the present invention. Figure 6 As shown, the training method of the autonomous driving perception algorithm model in the embodiment of the present invention mainly includes the following steps S301 and S302.
[0168] Step S301: Obtain a supplementary training set through the output data collection and processing method embodiment of the above-mentioned autonomous driving perception algorithm model.
[0169] In this embodiment, a supplementary training set of the model can be obtained by the method described in the above-mentioned embodiment of the method for collecting and processing the output data of the autonomous driving perception algorithm model.
[0170] Step S302: Use the training set to train the perception algorithm model, wherein the supplementary training set is at least a part of the training set.
[0171] In this embodiment, the perception algorithm model may be trained using a training set, where the training set includes the supplementary training set obtained in step S301.
[0172] In one embodiment, the perception algorithm model may be trained using only the supplementary training set.
[0173] In one embodiment, a supplementary training set may be added to the training set, and the training set may be used to train the perception algorithm model.
[0174] In one embodiment, the perception algorithm model can be trained using a training set using a model training method commonly used in the art.
[0175] It should be pointed out that although the various steps in the above embodiments are described in a specific order, those skilled in the art will understand that in order to achieve the effects of the present invention, different steps do not have to be performed in such an order. They can be performed simultaneously (in parallel) or in other orders. These changes are within the scope of protection of the present invention.
[0176] Those skilled in the art will appreciate that all or part of the processes in the method for implementing the above-mentioned embodiment of the present invention may also be accomplished by instructing the relevant hardware through a computer program. The computer program may be stored in a computer-readable storage medium. When the computer program is executed by a processor, it may implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable storage medium may include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electric carrier signal, telecommunication signal, and software distribution medium capable of carrying the computer program code. It should be noted that the content contained in the computer-readable storage medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable storage media do not include electric carrier signals and telecommunication signals.
[0177] Furthermore, the present invention also provides an electronic device. In an electronic device embodiment according to the present invention, the electronic device includes a processor and a storage device. The storage device can be configured to store a program for executing the method for collecting and processing the output data of the autonomous driving perception algorithm model according to the above-mentioned method embodiment. The processor can be configured to execute the program in the storage device, which includes but is not limited to a program for executing the method for collecting and processing the output data of the autonomous driving perception algorithm model according to the above-mentioned method embodiment. For ease of explanation, only the parts related to the embodiment of the present invention are shown. For specific technical details not disclosed, please refer to the method section of the embodiment of the present invention.
[0178] Furthermore, the present invention also provides a computer-readable storage medium. In a computer-readable storage medium embodiment according to the present invention, the computer-readable storage medium can be configured to store a program for executing the output data collection and processing method of the autonomous driving perception algorithm model of the above-mentioned method embodiment. The program can be loaded and run by the processor to implement the output data collection and processing method of the above-mentioned autonomous driving perception algorithm model. For ease of explanation, only the parts related to the embodiment of the present invention are shown. For specific technical details not disclosed, please refer to the method part of the embodiment of the present invention. The computer-readable storage medium can be a storage device formed by various electronic devices. Optionally, the computer-readable storage medium in the embodiment of the present invention is a non-transitory computer-readable storage medium.
[0179] Furthermore, the present invention also provides a vehicle. In one embodiment of a vehicle according to the present invention, the vehicle may include the electronic device in the above electronic device embodiment.
[0180] Furthermore, it should be understood that since the configuration of each module is merely for the purpose of illustrating the functional units of the apparatus of the present invention, the physical devices corresponding to these modules may be the processor itself, or a portion of the software in the processor, a portion of the hardware, or a combination of software and hardware. Therefore, the number of modules in the figure is merely illustrative.
[0181] Those skilled in the art will appreciate that the various modules in the device can be adaptively split or merged. Such splitting or merging of specific modules does not cause the technical solution to deviate from the principles of the present invention. Therefore, the technical solutions after splitting or merging will fall within the scope of protection of the present invention.
[0182] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.
Claims
1. A method for collecting and processing output data of an autonomous driving perception algorithm model, characterized in that: include: fusing data collected by multiple vehicle-mounted sensors, inputting the fusion results into the perception algorithm model, and obtaining the target perception results output by the perception algorithm model; Sequencing the target perception result into a plurality of perception frame data; Different scorers are set for different preset scenarios, wherein each scorer sets its own scoring criteria and its own scoring weight based on the corresponding preset scenario; Inputting the plurality of perception frame data into the different scorers respectively, and obtaining a scoring result of each scorer on the current perception frame data based on the scoring criteria; According to the scoring weight of each scorer, the scoring results of all scorers are weighted averaged to obtain the evaluation score of the current perception frame data; as well as If the evaluation score exceeds a predetermined threshold, it is determined that the current perception frame data meets the preset scenario; Collecting data collected by the vehicle-mounted sensor within a time window of a preset length where the current frame that meets the preset scenario is located as a supplementary training set for the perception algorithm model; Applying the supplementary training set to retrain the perception algorithm model to optimize the performance of the perception algorithm model; The preset scenes include abnormal scenes that appear between different frames during target perception; The step of inputting the plurality of perception frame data into the different scorers respectively, and obtaining a scoring result of each scorer on the current perception frame data based on the scoring criteria, comprises: Based on the scoring criteria, a scoring result for the current perception frame data is obtained by comparing the current perception frame data with perception frame data before the current frame.
2. The method according to claim 1, characterized in that The preset scenes include scenes of interest that appear during target perception; The step of inputting the plurality of perception frame data into the different scorers respectively, and obtaining a scoring result of each scorer on the current perception frame data based on the scoring criteria, comprises: Based on the scoring criteria, a scoring result of the current perception frame data is obtained by analyzing the current perception frame data.
3. The method according to claim 2, characterized in that The method of setting different scorers for different preset scenarios, wherein each scorer sets its own scoring criteria and its own scoring weight based on the corresponding preset scenario, includes: setting a scoring weight for the scorer corresponding to the scenario of interest greater than the scoring weight for the scorer corresponding to the abnormal scenario.
4. The method according to claim 1, wherein Also includes: The data collected by the vehicle-mounted sensor is subjected to a time series conversion after being subjected to another perception algorithm model different from the perception algorithm model to obtain a plurality of sensor frame data; Determining whether the current sensor frame data conforms to the preset scenario; The data collected by the vehicle-mounted sensor within a time window of a preset length where the current frame that meets the preset scenario is located is collected as a supplementary training set for the perception algorithm model.
5. The method according to claim 1, wherein The collecting of data collected by the vehicle-mounted sensor within a time window of a preset length where the current frame that meets the preset scenario is located as a supplementary training set for the perception algorithm model includes: When the current perception frame data meets the preset scenario, the data collected by the on-board sensor within the time window of the preset length where the perception frame data is located is transmitted back as a supplementary training set for the perception algorithm model.
6. The method according to any one of claims 1 to 4, characterized in that The vehicle-mounted sensors include a vehicle-mounted camera and a vehicle-mounted laser radar; The step of fusing data collected by multiple vehicle-mounted sensors, inputting the fusion result into the perception algorithm model, and obtaining the target perception result output by the perception algorithm model includes: Acquire 2D visual data from the vehicle’s onboard camera; Obtain 3D point cloud data from vehicle-mounted lidar; Projecting the two-dimensional visual data from the image coordinate system to the camera coordinate system to obtain a first projection result; Projecting the first projection result into the world coordinate system according to a conversion relationship between the camera coordinate system and the world coordinate system to obtain a second projection result; Dedistorting the second projection result to obtain a third projection result; Projecting the three-dimensional point cloud data into a world coordinate system to obtain a fourth projection result; Performing data fusion on the third projection result and the fourth projection result, and projecting the data into a two-dimensional space to obtain a fusion result; The fusion result is input into the perception algorithm model to obtain the target perception result.
7. An electronic device comprising a processor and a storage device, wherein the storage device is adapted to store a plurality of program codes, wherein: The program code is suitable for being loaded and executed by the processor to perform the method according to any one of claims 1 to 6.
8. A computer-readable storage medium storing a plurality of program codes, characterized in that: The program code is suitable for being loaded and executed by a processor to perform the method according to any one of claims 1 to 6.
9. A vehicle, characterized in that: The vehicle includes the electronic device according to claim 7.
Citation Information
Patent Citations
Method, device, apparatus and medium for extracting obstacle perception error data
CN109255341A
Vision and radar perception algorithm optimization method and system based on fusion perception and automobile
CN113449632A
Road surface condition sensing method and system based on multi-MEMS sensor data fusion
CN114282430A