Evaluation method and device of automatic driving algorithm and computer readable storage medium
By rendering the autonomous driving data packets into scene playback videos, and using visual language models to automatically identify driving behavior problems, the existing autonomous driving algorithms are solved, and efficient and automated algorithm evaluation is achieved.
Patent Information
- Application Number
- CN202510119433.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-06
AI Technical Summary
Existing autonomous driving algorithm testing relies on manual execution, resulting in low testing efficiency and high cost, and difficult to unify the test standards, so it is impossible to quickly complete the test within the algorithm iteration cycle.
By obtaining autonomous driving data packets, the scene playback video of the vehicle's autonomous driving process is rendered, the video is recognized using a visual language model, the driving behavior problem is automatically identified, and the autonomous driving algorithm is evaluated based on the recognition results.
The automated algorithm evaluation process is realized, which improves testing efficiency, reduces testing costs, and can achieve rapid and comprehensive evaluation of autonomous driving algorithms.
Smart Images

Figure CN119937521A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of autonomous driving technology, and in particular to an evaluation method, device and computer-readable storage medium for an autonomous driving algorithm. Background Art
[0002] In the process of rapid development of autonomous driving technology, how to efficiently and comprehensively test and verify the effectiveness and safety of autonomous driving algorithms has become a core issue of common concern in the academic field and the industry. At present, the combination of virtual simulation technology and large-scale real road testing has been widely used in this field due to its advantage of accelerating algorithm iteration and optimization.
[0003] However, there are many problems in the actual application of this method. At present, the testing process relies on manual execution, which requires personnel to spend a lot of time and energy to learn relevant knowledge and skills, and the learning cost remains high. At the same time, different testers have differences in professional level, experience and subjective judgment, which makes it difficult to achieve complete unification of testing standards.
[0004] From the perspective of efficiency, algorithm iteration will generate massive amounts of data. The speed of manual processing cannot keep up with the pace of algorithm iteration, and testing cannot be completed quickly within the iteration cycle. This leads to a large backlog of data, which undoubtedly forms a bottleneck restricting the rapid development of autonomous driving technology.
[0005] Therefore, it is urgent to explore an efficient evaluation method to solve current problems and promote the stable development of autonomous driving technology. Summary of the invention
[0006] In order to overcome the problems existing in the related art, this specification provides an evaluation method, device and computer-readable storage medium for an autonomous driving algorithm.
[0007] According to a first aspect of an embodiment of this specification, a method for evaluating an autonomous driving algorithm is provided, the method comprising:
[0008] Acquire the autonomous driving data packets collected during the vehicle's autonomous driving process;
[0009] Based on the autonomous driving data packet, rendering a scene playback video showing the autonomous driving process of the vehicle;
[0010] According to a plurality of set driving behavior problems and a plurality of corresponding set driving scene data, the scene playback video is identified, and when the driving scene data of the scene playback video meets the set driving scene data, the corresponding identified driving behavior problem is obtained;
[0011] Based on the identified driving behavior problems and the driving behavior problems obtained by the vehicle based on the current autonomous driving algorithm, the autonomous driving algorithm is evaluated.
[0012] According to an evaluation method for an autonomous driving algorithm provided in this application,
[0013] The scene playback video is identified according to a plurality of set driving behavior problems and a corresponding plurality of set driving scene data, and when the driving scene data of the scene playback video is obtained to conform to the set driving scene data, the corresponding identified driving behavior problem includes:
[0014] Based on the target visual language big model, the scene playback video is identified, and when the driving scene data of the scene playback video conforms to the set driving scene data, the corresponding identified driving behavior problem is obtained;
[0015] The target visual language large model is trained by a plurality of the set driving behavior problems and a corresponding plurality of set driving scene data.
[0016] According to an evaluation method for an autonomous driving algorithm provided in this application, the training process of the visual language large model includes:
[0017] Build a prompt word library based on the set driving behavior problems;
[0018] Determining scenario data corresponding to the set driving behavior problem in a historical autonomous driving scenario as set driving scenario data;
[0019] Based on the set driving scenario data, construct a problem scenario library;
[0020] The prompt word library and the problem scenario library are used as input to train the basic visual language model until the training end condition is reached, and the target visual language model is obtained for understanding data in the field of autonomous driving.
[0021] According to an evaluation method for an autonomous driving algorithm provided by the present application, the scene playback video is identified based on the target visual language large model, and when the driving scene data of the scene playback video is obtained to meet the set driving scene data, the corresponding identified driving behavior problems include:
[0022] Traversing the set driving behavior problems based on the target visual language big model, and evaluating the scene data of the scene playback video according to the set driving scene data corresponding to the traversed set driving behavior problems;
[0023] When the driving scene data outputted from the scene playback video is consistent with the set driving scene data, the driving behavior problem corresponding to the set driving scene data.
[0024] According to an evaluation method for an autonomous driving algorithm provided by the present application, when the driving scene data of the scene playback video obtained meets the set driving scene data, after the corresponding driving behavior problem is identified, the method further includes:
[0025] Obtain identification information of an object associated with the identified driving behavior problem for verifying an output result of the target visual language large model.
[0026] According to an evaluation method for an autonomous driving algorithm provided by the present application, after rendering a scene playback video showing the autonomous driving process of the vehicle based on the autonomous driving data packet, the method further includes:
[0027] The scene playback video is lightweight processed so that the adjusted scene playback video meets the input requirements of the target visual language large model, and the target visual language large model is trained by multiple set driving behavior problems and corresponding multiple set driving scene data.
[0028] According to an evaluation method for an autonomous driving algorithm provided by the present application, before rendering a scene playback video showing the autonomous driving process of the vehicle based on the autonomous driving data packet, the method further includes:
[0029] Parsing the autonomous driving data packet to extract key data related to the autonomous driving function;
[0030] The rendering of a scene playback video showing the vehicle's autonomous driving process based on the autonomous driving data packet includes:
[0031] The key data are fused to render a scene playback video showing the vehicle's autonomous driving process.
[0032] According to an evaluation method for an autonomous driving algorithm provided in the present application, the driving scene data includes images, videos or data records of vehicle sensors.
[0033] The present application also provides an evaluation device for an autonomous driving algorithm, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, an evaluation method for an autonomous driving algorithm as described above is implemented.
[0034] The present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a method for evaluating an autonomous driving algorithm as described in any one of the above.
[0035] The evaluation method, device and computer-readable storage medium of the automatic driving algorithm in the embodiment of this specification, compared with the low efficiency and high testing cost of the current manual test, obtains the collected automatic driving data packet during the automatic driving process of the vehicle; after integrating the data in the automatic driving data packet, renders a scene playback video showing the automatic driving process of the vehicle, thereby converting the abstract data into an intuitive visual scene, which is convenient for the subsequent identification of driving behavior problems during the vehicle driving process. Thereafter, according to the multiple set driving behavior problems existing in the automatic driving process and the corresponding multiple set driving scene data, the scene playback video is identified, and the driving scene data of the scene playback video is obtained when it meets the set driving scene data. The process can use the understanding ability of the visual language large model in the field of automatic driving to realize the automatic analysis of the driving behavior problems existing in the driving scene. Finally, based on the identified driving behavior problems and the driving behavior problems obtained by the vehicle based on the current automatic driving algorithm, the accuracy of the automatic driving algorithm is evaluated, and the automatic evaluation process is realized, the test verification efficiency of the algorithm is improved, and the test cost is reduced.
[0036] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the specification and, together with the description, serve to explain the principles of the specification.
[0038] Figure 1 is a flowchart of an evaluation method of an autonomous driving algorithm according to an exemplary embodiment of this specification;
[0039] Figure 2 is another flowchart of an evaluation method of an autonomous driving algorithm according to an exemplary embodiment of this specification;
[0040] Figure 3 is a module architecture diagram of an evaluation algorithm for an autonomous driving algorithm according to an exemplary embodiment of this specification;
[0041] Figure 4 is a schematic diagram of an evaluation device for an autonomous driving algorithm according to an exemplary embodiment of this specification;
[0042] Figure 5 This is a schematic block diagram of an evaluation device for an autonomous driving algorithm shown in this specification according to an exemplary embodiment. DETAILED DESCRIPTION
[0043] Here, the technical solutions in the embodiments (or "implementations") of the present application will be described clearly and completely in conjunction with the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements.
[0044] If there are terms involving directional indications or positional relationships in the embodiments of the present application (such as up, down, left, right, front, back, inside, outside, top, bottom, center, vertical, horizontal, longitudinal, transverse, length, width, counterclockwise, clockwise, axial, radial, circumferential, etc.), such terms are only used to explain the relative positional relationship, movement, etc. between the components under a certain specific posture (as shown in the accompanying drawings); if the specific posture changes, the directional indication or positional relationship will also change accordingly. In addition, the terms "first", "second", etc. involved in the embodiments of the present application are only used for the purpose of convenience of description and cannot be understood as indicating or implying relative importance.
[0045] The present application provides an evaluation method, device and computer-readable storage medium for an autonomous driving algorithm. The present application is described in detail below in conjunction with the accompanying drawings. The features of the following embodiments and implementations may be combined with each other unless they conflict.
[0046] Since autonomous driving algorithms are complex and involve multiple links such as perception, decision-making, and control, and autonomous driving vehicles are directly involved in traffic operations, once an algorithm fails or errors occur, it may lead to serious traffic accidents, causing casualties and property losses. Therefore, how to efficiently and comprehensively test and verify the effectiveness and safety of autonomous driving algorithms has become a core issue of common concern in the academic field and industry. For example, through comprehensive testing, potential problems and loopholes in the algorithm can be discovered in different scenarios, environments, and working conditions, and then optimized and improved to improve the stability and reliability of the algorithm and ensure that the vehicle can operate correctly in various situations.
[0047] The current process of testing and verifying autonomous driving algorithms relies on manual execution. Large amounts of test data often require a large number of testers, which not only increases labor costs, but also makes it difficult to fully unify test standards. At the same time, the algorithm will accumulate tens of thousands of data during the rapid iteration process, and the manual processing speed is difficult to keep up with the algorithm iteration rhythm, so it is impossible to quickly complete the test within the iteration cycle.
[0048] In order to solve the above technical problems, this specification provides an evaluation method for an autonomous driving algorithm.
[0049] It aims to utilize the common sense reasoning capabilities of large visual language models to automatically evaluate problems in autonomous driving scenarios, thereby improving testing efficiency and reducing testing costs.
[0050] like Figure 1 As shown, Figure 1 1 is a flow chart of an evaluation method of an autonomous driving algorithm provided in an embodiment of this specification, including the following steps S100-S400:
[0051] Step S100, obtaining an autonomous driving data packet collected during the autonomous driving process of the vehicle;
[0052] Step S200, based on the autonomous driving data packet, rendering a scene playback video showing the autonomous driving process of the vehicle;
[0053] Step S300, identifying the scene playback video according to a plurality of set driving behavior problems and a plurality of corresponding set driving scene data, and obtaining a corresponding identified driving behavior problem when the driving scene data of the scene playback video meets the set driving scene data;
[0054] Step S400, evaluating the autonomous driving algorithm based on the identified driving behavior problems and the driving behavior problems obtained by the vehicle based on the current autonomous driving algorithm.
[0055] In this embodiment, the abstract autonomous driving data packet is converted into an intuitive visualization scene. The rich common sense and semantic knowledge learned by the visual language large model is used to understand and analyze the input autonomous driving scene image or video, quickly extract key information from the scene, identify the driving behavior of the vehicle during driving, automatically analyze the problems in the scene, and improve test efficiency.
[0056] The method specifically comprises the following steps:
[0057] like Figure 2 , Figure 3 As shown, Figure 2 This is another flowchart of an evaluation method for an autonomous driving algorithm provided in an embodiment of this specification. Figure 3 This is a module architecture diagram of an evaluation algorithm for an autonomous driving algorithm provided in an embodiment of this specification.
[0058] In step S100, an autonomous driving data packet collected during the vehicle's autonomous driving process is obtained.
[0059] As an example, an autonomous driving data packet obtains a large amount of autonomous driving data through autonomous driving road test scenarios or simulation test scenarios.
[0060] In the actual road test of autonomous driving (hereinafter referred to as road test), the vehicle will be equipped with various sensors, such as cameras, lidar, millimeter-wave radar, inertial measurement unit (IMU), etc. These sensors will collect real-time environmental information around the vehicle and the vehicle's own status information. The vehicle body status information can be, but is not limited to, the vehicle's own speed, acceleration, steering angle and other status information.
[0061] These data are used for subsequent evaluation of autonomous driving algorithms.
[0062] The autonomous driving simulation test simulates the vehicle's driving process in a virtual environment. The simulation data also contains parts similar to the road test data, which can be but not limited to virtual visual data, virtual radar data, virtual vehicle status data, etc. Among them, virtual visual data: images or videos from the camera perspective simulated in the simulation environment. These visual data can reflect the road conditions and traffic participants in the virtual scene. Virtual radar data: simulated laser radar and millimeter wave radar detection data, used to represent the relative position, speed and other relationships between the vehicle and objects in the virtual environment. Virtual vehicle status data: simulated vehicle speed, acceleration, steering angle and other status information in the virtual scene.
[0063] All data from road testing or simulation is recorded to form a data packet containing various information during the autonomous driving process.
[0064] In order to facilitate subsequent analysis, the autonomous driving data packet is stored in a set data storage format. The set data storage format can be, but is not limited to, bag file format, cvs format, json format, etc.
[0065] As an example, when conducting road tests or simulations, all sensor data, vehicle status data, etc. are collected and stored in bag files in a certain time series. In this way, the bag data becomes a data packet containing various information during the autonomous driving process. A complete test process can be recorded in a bag file, from vehicle startup, driving on the road, encountering various traffic conditions (such as overtaking, parking), until the vehicle stops. All data collected by the sensors during this period are stored in this file, which is convenient for the subsequent complete playback, analysis and processing of the entire driving process, providing basic data support for manual testing.
[0066] As an example, the rosbag tool of ROS (Robot Operating System) is used to record the data of the vehicle's autonomous driving system during operation, and all data is stored as a bag data packet.
[0067] In this embodiment, it is ensured that all kinds of data obtained from different sensors can be recorded completely and accurately, providing a comprehensive data basis for subsequent analysis and processing. The data from different sensors are strictly aligned in time, which is conducive to reproducing the real driving scene.
[0068] In step S200, the autonomous driving data packets during the vehicle's driving process are presented in a visual form using graphics rendering technology to generate a scene playback video showing the vehicle's autonomous driving process. During the playback process, the information corresponding to various sensor data is synchronously displayed according to the time series in the bag file. For example, the detected objects are drawn in the virtual scene based on the perception data, and the vehicle's driving trajectory and real-time position are displayed based on the positioning and trajectory data. At the same time, some auxiliary information such as vehicle speed, time, etc. can be added to make the playback more intuitive.
[0069] In this embodiment, abstract data is converted into intuitive visualization scenes, which makes it easier for the subsequent large visual language model to quickly understand the behavior of the autonomous driving vehicle and the surrounding environment during driving, and can more efficiently identify driving behavior problems.
[0070] In some embodiments, key information is extracted from a huge bag data packet and irrelevant data is filtered out, making subsequent processing and analysis more targeted and reducing the workload of data processing and computing resource consumption.
[0071] Specifically, before step S200, the method further includes:
[0072] Step a11, parse the autonomous driving data packet and extract key data related to the autonomous driving function.
[0073] The bag file stores a large amount of data of different types. Through specific parsing tools (such as the parsing library provided by ROS), topic data related to the key functions of autonomous driving can be extracted.
[0074] As an example, topic data includes but is not limited to positioning data, perception data, trajectory data, etc. Among them, positioning data is used to determine the precise position of the vehicle in the map; perception data contains identification information of the surrounding environment, such as the category and location of the detected objects; trajectory data records the driving path of the vehicle over a period of time. These data are key information for understanding the operating status of the autonomous driving system and conducting subsequent analysis.
[0075] After extracting the key data from the bag data packet, the step 200 comprises:
[0076] Step a12, fusing the key data and rendering a scene playback video showing the vehicle's autonomous driving process.
[0077] By utilizing the extracted key data (such as positioning and perception data) and combining it with graphics rendering technology, the environment and vehicle status information during vehicle driving are presented in a visual form to generate a scene playback video.
[0078] In other embodiments, in order to make the rendered visual driving scene playback video data more suitable for subsequent model processing, a series of preprocessing operations are required.
[0079] As an example, after step S200, the method further includes:
[0080] Step b11, performing lightweight processing on the scene playback video so that the adjusted scene playback video meets the input requirements of the target visual language large model, and the target visual language large model is trained by multiple set driving behavior problems and corresponding multiple set driving scene data.
[0081] Since the frame rate of the rendered scene playback video is relatively high, it adds a certain burden to the recognition and understanding of the subsequent visual language large model. Therefore, the scene playback video is lightweight processed. The adjustment includes but is not limited to adjusting the resolution and frame rate of the scene playback video, and the lightweight processing includes but is not limited to frame extraction, compression and other processing.
[0082] For example, adjusting the resolution can adjust the video to a suitable size according to the input requirements of the model to avoid wasting computing resources due to too high a resolution or affecting the model's recognition of details due to too low a resolution. Adjusting the frame rate to match the video playback speed with the model's processing capabilities can also help remove some redundant frames and reduce the amount of data.
[0083] In this embodiment, by reasonably adjusting the resolution and frame rate, the data processing amount is reduced, the consumption of computing resources is reduced, and the operating efficiency of the entire evaluation process is improved while ensuring the performance of the model.
[0084] In combination with the above embodiments, the data for the target visual language big model recognition used for automatic testing is the preprocessed scene data. This data is the scene playback data after fusion of sensor data and other topic data, rather than the data directly collected by the sensor on the vehicle. The fused and rendered scene is easier to be understood by the big model.
[0085] In step S300, driving behavior problems that meet the set conditions are accurately identified from the scene playback video, providing a key basis for subsequent evaluation of the autonomous driving algorithm based on these problems.
[0086] As an example, the driving scene data includes images, videos, or data records from vehicle sensors. Such scene data including images, videos, and vehicle sensor data records provide a rich source of information for the visual language model. Compared with traditional language models, such multimodal data input enables the visual language model to more comprehensively understand the actual situation of the autonomous driving scene.
[0087] First, clearly define a series of driving behavior problems that may occur during autonomous driving. These driving behavior problems cover problems that exist in autonomous driving scenarios, such as "vehicle running a red light", "self-vehicle crossing a solid line", "self-vehicle collision", "self-vehicle not maintaining a safe distance from the vehicle in front", etc. Each driving behavior problem corresponds to a specific traffic rule or safety standard.
[0088] Furthermore, the corresponding driving scenario data is constructed. For each driving behavior problem, the corresponding set driving scenario data is collected or generated. These data are used to describe the characteristics that the autonomous driving scenario should present when a specific driving behavior problem occurs. For example, for the problem of "vehicle running a red light", the set driving scenario data may include but is not limited to: the color status of the traffic light in the video screen (red light on), the position of the vehicle (crossing the stop line), the direction of the vehicle, and the status of other vehicles and pedestrians in the surrounding area.
[0089] For another example, for the problem of "the vehicle runs over a solid line", the set driving scene data may include, but is not limited to: the position of the vehicle in the video, the position of the solid line, the vehicle's driving trajectory and other information, and these data may exist in various forms, such as specific image feature descriptions, video clip examples, threshold ranges of vehicle sensor data, etc. Exemplarily, the driving scene data set for "the vehicle runs over a solid line" includes a white or yellow solid line on the road clearly presented in the video, the body of the vehicle overlaps with the solid line, and this overlap persists in multiple consecutive frames of images.
[0090] Parse the rendered scene playback video, decompose the video into a series of image frames, and extract key features from each frame and related sensor data records. For example, use the edge detection algorithm to extract the contour information of the road lines from each frame to identify the lane lines. Use the target recognition algorithm to determine the position and contour of the vehicle in the image. Extract key features from the identified lane lines and vehicle information. For lane lines, extract their color, shape, position in the image and other features; for the vehicle, extract its position, driving direction, body posture and other features. At the same time, extract the vehicle's speed, acceleration and other information from the vehicle sensor data records, and combine the position and direction of the vehicle in the image to comprehensively judge the vehicle's driving status.
[0091] Match the extracted video features with the features in the set driving scene data. Taking "self-vehicle pressing solid line" as an example, match the extracted video features with the features in the set "self-vehicle pressing solid line" driving scene data. Check whether there is a lane solid line in the image, and determine whether the position of the self-vehicle overlaps with the solid line. Combined with the vehicle's driving direction, speed, steering wheel steering angle and other information, further confirm whether it is a "self-vehicle pressing solid line" scene.
[0092] When the scene features in the video successfully match the features of a set driving scene data, it can be determined that the corresponding driving behavior problem has been identified. Exemplarily, when the scene features in the video successfully match the features of the set "own vehicle pressing the solid line" driving scene data, and the confidence of the match reaches a pre-set threshold, it can be determined that the driving behavior problem of "own vehicle pressing the solid line" has been identified. For example, after a series of feature matching and analysis, the system's confidence in judging the current scene as "own vehicle pressing the solid line" reaches 92%, which exceeds the threshold, then it is determined that the driving behavior problem has been identified.
[0093] It should be noted that in the actual scene playback video analysis, the identification of "the vehicle crossing the solid line" and other set driving behavior problems (such as running a red light, etc.) are carried out in parallel. The system will simultaneously detect and judge multiple aspects of the video to comprehensively evaluate various violations that may occur in the autonomous driving scene.
[0094] In some embodiments, the visual language large model's ability in visual understanding and reasoning is used to quickly identify and analyze key objects in the scene and their impact on the vehicle, thereby improving evaluation efficiency, reducing the workload of manual labeling and analysis, and lowering costs.
[0095] As an example, step S300 includes:
[0096] Step c11, based on the target visual language large model, the scene playback video is identified to obtain the driving behavior problem corresponding to the driving scene data of the scene playback video when the driving scene data conforms to the set driving scene data.
[0097] The target visual language large model is trained by a plurality of the set driving behavior problems and a corresponding plurality of set driving scene data.
[0098] In this embodiment, the features of the visual language big model are introduced, and the trained visual language big model is used to automatically identify problems in the preprocessed scene data and display the problem results, thereby eliminating the need for a large amount of manual testing.
[0099] As an example, the training process of the visual language large model includes:
[0100] The first step is to build a prompt vocabulary based on the set driving behavior problems.
[0101] In the autonomous driving scenario, the set driving behavior problems include various problems that may occur during the autonomous driving process, such as "the vehicle crosses the solid line", "the vehicle collides", "runs a red light", etc. These problems cover different situations that the autonomous driving vehicle may encounter during driving, including violations of traffic rules, safety accident risks, and inaccurate perception of the surrounding environment.
[0102] For each driving behavior problem, construct the corresponding prompt words. For example, the prompt words corresponding to "the vehicle crossed the solid line" can be "the vehicle crossed the solid line of the road while driving", and the prompt words corresponding to "the vehicle collided" can be "the vehicle collided with a pedestrian crossing the road". Of course, other different ways of expression can also be used to increase the richness and flexibility of the prompt words.
[0103] A prompt word library is constructed based on prompt words corresponding to multiple driving behavior problems.
[0104] The second step is to determine the scene data corresponding to the set driving behavior problem in the historical automatic driving scene as the set driving scene data.
[0105] The driving scene data of different types of autonomous driving behavior problems refer to various types of scene data (such as truncated video data) corresponding to the driving behavior problems in the above-mentioned prompt vocabulary, and the scene data is obtained by manually annotating known autonomous driving data.
[0106] The labelers need to accurately identify whether there are specific driving behavior problems in each scene based on the video and other relevant data, and make detailed annotations on the specific circumstances of the driving behavior problems. For example, when labeling the scene of "the vehicle crosses the solid line", it is necessary to label the specific location, time point, and road conditions of the vehicle crossing the solid line; for the scene of "the vehicle collides", it is necessary to label the time of the collision, the location and direction of movement of the pedestrians, and the speed of the vehicle.
[0107] The third step is to build a problem scenario library based on the set driving scenario data.
[0108] After the labeling is completed, the labeled data are classified and sorted according to different question types to form various scenario data corresponding to the prompt vocabulary library and build a question scenario library.
[0109] The fourth step is to use the prompt word library and the problem scenario library as input to train the basic visual language model until the training end conditions are met, thereby obtaining the target visual language model for understanding data in the field of autonomous driving.
[0110] Input the constructed prompt word library and the corresponding scene data into the basic multimodal visual language model. The model will process the input text prompt words and the corresponding visual scene data to extract the feature information. For example, for the prompt word "the car presses the solid line" and the corresponding solid line scene video, the model will extract the semantic features in the text and the visual features in the video, such as the position of the vehicle, the position of the solid line, and the relative relationship between the vehicle and the solid line.
[0111] During the training process, the model continuously adjusts its own parameters to learn how to accurately understand and judge based on the input prompt words and scene data. For example, when the prompt word "the vehicle crosses the line" and the corresponding scene data are input, the model will try to predict whether the vehicle crosses the line and the specific circumstances of the cross-line; when the prompt word "the vehicle collides" and the scene data are input, the model will predict whether the collision is about to occur or has already occurred, and give the corresponding judgment and explanation.
[0112] After multiple rounds of training, the model is fine-tuned. During the fine-tuning process, the model will adjust its own parameters based on these specific autonomous driving data, and learn to associate visual scenes with corresponding text descriptions (prompt words), so that it can better understand the characteristics and patterns of common problems in the field of autonomous driving, so as to further improve the performance and accuracy of the model in autonomous driving scenarios.
[0113] Among them, fine-tuning can be carried out according to specific application requirements and actual test results, such as adjusting the model's sensitivity to different problems, optimizing the model's ability to handle complex scenarios, etc.
[0114] Through continuous fine-tuning, the model can better adapt to various situations in autonomous driving scenarios, thereby obtaining a fine-tuned multimodal target visual language large model, which enables it to perform better in the detection and processing of autonomous driving problems.
[0115] In this embodiment, fine-tuning of the autonomous driving data enables the model to more accurately identify and understand various problems in the autonomous driving scene. Compared with direct processing using a general model, the model's judgment accuracy and analysis capabilities for autonomous driving problems are greatly improved. The trained visual language model is used to automatically identify driving behavior problems in the pre-processed driving scene data and display the problem results, eliminating the need for a large amount of manual testing.
[0116] Based on the target visual language large model obtained after the above training, the step c11 includes:
[0117] Step c111, traversing the set driving behavior problems based on the target visual language large model, and evaluating the scene data of the scene playback video according to the set driving scene data corresponding to the traversed set driving behavior problems;
[0118] Step c112, when the driving scene data of the scene playback video is consistent with the set driving scene data, the driving behavior problem corresponding to the set driving scene data is output.
[0119] Using the fine-tuned target visual language model, the pre-processed scene data in the scene playback video is evaluated for each prompt word in the prompt vocabulary built during the training process of the visual language model.
[0120] The target visual language model will analyze whether the scene data matches the problem type described by a prompt based on the learned visual and language association patterns. For example, for the prompt word "the vehicle hits the solid line", the model will analyze the position and distance relationship between the vehicle and the solid line in the scene data and determine whether there is a problem of hitting the solid line.
[0121] After traversing every prompt word in the prompt word library, when the driving scene data of the output scene playback video is consistent with the set driving scene data, the corresponding driving behavior problem.
[0122] In this embodiment, a prompt word library of question types and corresponding question scenario library data dedicated to autonomous driving are constructed to align the understanding ability of the visual language large model in the field of autonomous driving, and applied to testing and verification for automated testing and verification.
[0123] In other examples, the results of model processing are recorded in a certain format, such as recording the type of question, judgment result, and possible confidence level corresponding to each scene. Exemplarily, the result output by the model is a clear "yes" or "no" judgment result. For example, the target visual language model is used to accurately judge whether there are problems such as collision and compaction of the vehicle's behavior. If it is detected that the vehicle has collided during the autonomous driving process, the target visual language model outputs "yes".
[0124] After step c112, the method further includes:
[0125] Step c113, obtaining identification information of objects associated with the identified driving behavior problem, for verifying the output result of the target visual language large model.
[0126] The target visual language large model outputs driving behavior issues and more detailed information to facilitate subsequent verification and improve evaluation accuracy.
[0127] As an example, if a collision is detected during the autonomous driving process, the target visual language model outputs "yes". At the same time, since the perception module has identified the collision obstacle and assigned a unique ID during the road test or simulation, the target visual language model will obtain and output the obstacle ID from the relevant data storage. The output results (including collision judgment and obstacle ID) provide key information for subsequent verification.
[0128] In this embodiment, by traversing the prompt library, the scene data can be fully tested for problems, ensuring that no predefined problem type is missed, thereby improving the comprehensiveness and accuracy of the evaluation. In addition, the target visual language large model is used to achieve automated evaluation of multiple types of problems in autonomous driving scenarios, greatly improving evaluation efficiency and reducing the workload and subjectivity of manual evaluation.
[0129] In step S400, the autonomous driving algorithm is evaluated based on the identified driving behavior problems and the driving behavior problems obtained by the vehicle based on the current autonomous driving algorithm.
[0130] The results of the driving behavior problems given by the vehicle's current autonomous driving algorithm are compared with the results obtained based on the above-mentioned target visual language large model, and the autonomous driving algorithm is evaluated based on this. If the vehicle's current autonomous driving algorithm determines that "the vehicle has collided", and the target visual language large model predicts that the vehicle and the obstacle will overlap at a certain moment in the future based on the kinematic model, and the distance is less than the set collision threshold, then it can be considered that the judgment of the vehicle's current autonomous driving algorithm is correct in this scenario; conversely, if the target visual language large model predicts that there will be no collision, and the vehicle's current autonomous driving algorithm determines that it is "collision", it indicates that the vehicle's current autonomous driving algorithm has made an incorrect judgment.
[0131] In some embodiments, the consistency of the judgment results is not just a simple "right" or "wrong" judgment, but also requires the introduction of confidence assessment. For example, if the target visual language large model calculates that the collision probability is 95%, and the vehicle's current autonomous driving algorithm judges that "the vehicle has collided", then the confidence of this judgment is higher; if the target visual language large model calculates that the collision probability is only 55%, even if the vehicle's current autonomous driving algorithm judges that "the vehicle has collided", its accuracy needs to be questioned. In this way, a more comprehensive and scientific evaluation of the accuracy of the judgment results of the vehicle's current autonomous driving algorithm can be conducted.
[0132] The present application provides an evaluation method, device and computer-readable storage medium for an autonomous driving algorithm. Compared with the current manual testing with low efficiency and high testing cost, the method obtains the collected autonomous driving data packets during the vehicle autonomous driving process; after fusing the data in the autonomous driving data packets, a scene playback video showing the vehicle autonomous driving process is rendered, thereby converting the abstract data into an intuitive visual scene, which is convenient for the subsequent identification of driving behavior problems during the vehicle driving process. Thereafter, according to the multiple set driving behavior problems existing in the autonomous driving process and the corresponding multiple set driving scene data, the scene playback video is identified, and the driving behavior problems corresponding to the driving scene data of the scene playback video are obtained when the driving scene data of the scene playback video meets the set driving scene data. This process can utilize the understanding ability of the visual language large model in the field of autonomous driving to realize the automated analysis of driving behavior problems existing in the driving scene. Finally, based on the identified driving behavior problems and the driving behavior problems obtained by the vehicle based on the current autonomous driving algorithm, the accuracy of the autonomous driving algorithm is evaluated, an automated evaluation process is realized, the test verification efficiency of the algorithm is improved, and the test cost is reduced.
[0133] Based on the same application concept as the above method, the present application embodiment also proposes an evaluation device for an autonomous driving algorithm, such as Figure 4 shown.
[0134] The device comprises:
[0135] The data acquisition module 402 is used to acquire the autonomous driving data packets collected during the autonomous driving process of the vehicle;
[0136] A scene playback module 404, configured to render a scene playback video showing the vehicle's autonomous driving process based on the autonomous driving data packet;
[0137] The automatic testing module 406 is used to identify the scene playback video according to a plurality of set driving behavior problems and a corresponding plurality of set driving scene data, and obtain the driving behavior problem corresponding to the identification when the driving scene data of the scene playback video meets the set driving scene data;
[0138] The algorithm evaluation module 408 is used to evaluate the autonomous driving algorithm based on the identified driving behavior problems and the driving behavior problems obtained by the vehicle based on the current autonomous driving algorithm.
[0139] The implementation process of the functions and effects of each module / submodule / unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and the same technical effect can be achieved, so it will not be repeated here.
[0140] Figure 5The following is a schematic diagram of the physical structure of an evaluation device for an autonomous driving algorithm. Figure 5 As shown, the evaluation device of the autonomous driving algorithm may include: a processor (processor) 810, a communication interface (CommunicationsInterface) 820, a memory (memory) 830 and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call the logic instructions in the memory 830 to execute the evaluation method of the autonomous driving algorithm.
[0141] In addition, the logic instructions in the above-mentioned memory 830 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art, and the computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0142] On the other hand, the present application also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the evaluation method of the autonomous driving algorithm provided by the above-mentioned methods.
[0143] On the other hand, the present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the evaluation method of the autonomous driving algorithm provided by the above-mentioned methods.
[0144] It should be noted that the technical solutions or technical features described in the above embodiments can be combined or supplemented with each other without causing conflicts. The scope of protection of this application is not limited to the precise structures described in the above embodiments and shown in the drawings; all modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of this application should be included in the scope of protection of this application.
Claims
1. A method for evaluating an autonomous driving algorithm, characterized in that: The method comprises: Acquire the autonomous driving data packets collected during the vehicle's autonomous driving process; Based on the autonomous driving data packet, rendering a scene playback video showing the autonomous driving process of the vehicle; According to a plurality of set driving behavior problems and a plurality of corresponding set driving scene data, the scene playback video is identified, and when the driving scene data of the scene playback video meets the set driving scene data, the corresponding identified driving behavior problem is obtained; Based on the identified driving behavior problems and the driving behavior problems obtained by the vehicle based on the current autonomous driving algorithm, the autonomous driving algorithm is evaluated.
2. The method for evaluating an autonomous driving algorithm according to claim 1, wherein: The scene playback video is identified according to a plurality of set driving behavior problems and a corresponding plurality of set driving scene data, and when the driving scene data of the scene playback video is obtained to conform to the set driving scene data, the corresponding identified driving behavior problem includes: Based on the target visual language big model, the scene playback video is identified, and when the driving scene data of the scene playback video conforms to the set driving scene data, the corresponding identified driving behavior problem is obtained; The target visual language large model is trained by a plurality of the set driving behavior problems and a corresponding plurality of set driving scene data.
3. The method for evaluating an autonomous driving algorithm according to claim 2, wherein: The training process of the visual language large model includes: Build a prompt word library based on the set driving behavior problems; Determining scenario data corresponding to the set driving behavior problem in a historical autonomous driving scenario as set driving scenario data; Based on the set driving scenario data, construct a problem scenario library; The prompt word library and the problem scenario library are used as input to train the basic visual language model until the training end condition is reached, and the target visual language model is obtained for understanding data in the field of autonomous driving.
4. The method for evaluating an autonomous driving algorithm according to claim 3, wherein: The scene playback video is identified based on the target visual language large model, and when the driving scene data of the scene playback video is obtained to meet the set driving scene data, the corresponding identified driving behavior problem includes: Traversing the set driving behavior problems based on the target visual language big model, and evaluating the scene data of the scene playback video according to the set driving scene data corresponding to the traversed set driving behavior problems; When the driving scene data outputted from the scene playback video is consistent with the set driving scene data, the driving behavior problem corresponding to the set driving scene data.
5. The method for evaluating an autonomous driving algorithm according to claim 4, wherein: When the driving scene data of the scene playback video is obtained to meet the set driving scene data, after the corresponding driving behavior problem is identified, the method further includes: Obtain identification information of an object associated with the identified driving behavior problem for verifying an output result of the target visual language large model.
6. The method for evaluating an autonomous driving algorithm according to claim 1, wherein: After rendering a scene playback video showing the vehicle's autonomous driving process based on the autonomous driving data packet, the method further includes: The scene playback video is lightweight processed so that the adjusted scene playback video meets the input requirements of the target visual language large model, and the target visual language large model is trained by multiple set driving behavior problems and corresponding multiple set driving scene data.
7. The method for evaluating an autonomous driving algorithm according to claim 1, wherein: Before rendering a scene playback video showing the vehicle's autonomous driving process based on the autonomous driving data packet, the method further includes: Parsing the autonomous driving data packet to extract key data related to the autonomous driving function; The rendering of a scene playback video showing the vehicle's autonomous driving process based on the autonomous driving data packet includes: The key data are fused to render a scene playback video showing the vehicle's autonomous driving process.
8. The method for evaluating an autonomous driving algorithm according to any one of claims 1 to 7, characterized in that: The driving scene data includes images, videos or data records of vehicle sensors.
9. An evaluation device for an autonomous driving algorithm, characterized in that: The method comprises a memory, a processor and an evaluation program of an autonomous driving algorithm stored in the memory and executable on the processor, wherein the processor implements the steps of the evaluation method of the autonomous driving algorithm as described in any one of claims 1 to 8 when executing the evaluation program of the autonomous driving algorithm.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores an evaluation program for an autonomous driving algorithm, and when the evaluation program for the autonomous driving algorithm is executed, the steps of the evaluation method for an autonomous driving algorithm as described in any one of claims 1-8 are implemented.
Citation Information
Cited By
Road offset distance measurement and verification method and system for lane departure early warning system, and storage medium
CN121661863A