Automatic driving test case self-evolution repairing method based on large language model

The deep semantic mining and closed-loop autoevolution process of the performance log of the autonomous driving system through a large language model solve the problem of failure mode recognition and optimization of the autonomous driving system in complex scenarios, and achieve the improvement of the system's safety and generalization capabilities.

CN120430169APending Publication Date: 2025-08-05JILIN UNIVERSITY
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510525760.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

When facing complex and uncertain real road scenarios, existing autonomous driving systems lack effective failure pattern recognition and semantic analysis capabilities, and decoupling of testing and optimization processes makes it difficult to achieve systematic continuous evolution.

Method used

Use large language models to conduct in-depth semantic mining of performance logs, and build a closed-loop self-evolution process from failure mode to test scenarios, including performance log collection, failure mode semantic extraction, scene semantic retrieval, initial recommendation construction, reflective optimization and small sample fine-tuning to form a closed-loop self-evolution mechanism.

Benefits of technology

It significantly improves the safety and generalization capabilities of the autonomous driving system, realizes rapid identification and repair of key failure modes, and improves test maintainability and continuous optimization efficiency of models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention discloses an automatic driving test case self-evolution repairing method based on a large language model, and belongs to the technical field of testing and performance optimization of an automatic driving system. According to the automatic driving test case self-evolution repairing method based on the large language model, deep semantic mining is conducted on a performance log of the model by means of semantic comprehension and recommendation capacity of the large language model, and a high-quality repairing scene is recommended from a structured test scene library. The method comprises the steps of performance log collection and pre-evaluation execution, failure mode semantic extraction, a scene semantic retrieval mechanism, initial recommendation construction, reflection type semantic optimization, few-sample fine adjustment adaptation and a closed-loop self-evolution mechanism. The method has obvious advantages in the aspects of automatic driving test case recommendation, failure mode recognition, rapid model capability adaptation and system closed-loop self-evolution, and the safety, generalization capability and test maintainability of the automatic driving system can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of testing and performance optimization of automatic driving systems. Background Art

[0002] In recent years, with the continuous development of autonomous driving systems, their core algorithmic architectures have gradually evolved from traditional modular perception-decision-control processes to end-to-end deep learning models. While these models demonstrate good performance in standard traffic scenarios, the uncertainty, scene diversity, and behavioral complexity encountered on real roads still pose significant challenges to the robustness and generalization capabilities of these systems. This is particularly true in "long-tail" scenarios characterized by frequent obstructions, complex interactions, uncertain rules, or a high concentration of unexpected events. Autonomous driving models are prone to prediction bias or strategic failure, seriously impacting driving safety.

[0003] To improve the robustness and adaptability of models, researchers have proposed scenario-driven testing frameworks in recent years, which cover potential risks by constructing typical test cases. However, existing test scenarios mainly come from the following approaches: First, scripted scenario design based on accident cases or rule templates is highly subjective and has sparse scenarios. Second, scenarios are synthesized through search or optimization algorithms, such as optimizing interference factors to generate collision scenarios based on genetic algorithms or reinforcement learning. However, such methods often rely on high-cost simulations and single optimization objectives, making it difficult to cover abnormal behaviors in high-dimensional semantic spaces. Third, existing scenarios are retrieved from datasets, but existing retrieval is mostly based on surface features and cannot perform deep matching on specific failure semantics.

[0004] Furthermore, current testing systems generally lack an effective closed-loop feedback mechanism. Problems exposed during testing often cannot be efficiently converted into training data, nor can they guide the model to targeted capability improvements. Even with manual intervention, additional data collection remains costly, update cycles are slow, and the disconnect between testing and optimization makes it difficult to achieve systematic and continuous evolution.

[0005] In recent years, large language models (LLMs) have demonstrated strong generalization and context modeling capabilities in tasks such as natural language understanding, knowledge reasoning, and semantic retrieval, offering new opportunities for testing and repairing autonomous driving systems. Preliminary research has explored the application of LLMs in task instruction generation, scene textualization, or language control interfaces, but most focus on the single-stage task of "natural language → autonomous driving scenario" and have yet to establish a complete chain of "performance failure → semantic analysis → scenario repair."

[0006] In summary, existing technologies still have obvious shortcomings in the following aspects: lack of the ability to automatically identify failure modes of autonomous driving systems, and low efficiency in log data utilization; lack of semantic-level test scenario retrieval and recommendation mechanism, and the recommendation results and failure causes are often semantically mismatched; the testing process is decoupled from the model optimization process, and lacks a closed-loop self-evolution mechanism; the capabilities of existing large language models have not yet been systematically integrated into the autonomous driving test recommendation and repair process. Summary of the Invention

[0007] The purpose of this invention is to utilize the semantic understanding and recommendation capabilities of a large language model to conduct deep semantic mining on the model's performance logs, and to recommend high-quality repair scenarios from a structured test scenario library using a large language model-based self-evolutionary repair method for autonomous driving test cases.

[0008] The steps of the present invention are: Step 1: Performance log collection and pre-assessment execution: Deploy the autonomous driving model to be tested on the simulation platform, execute standardized test tasks, collect driving behavior logs, and generate structured performance records containing information on events such as collisions, deviations, and violations. Step 2: Failure pattern semantic extraction: Use a large language model to perform semantic analysis on log data, extract behavioral anomalies and semantic elements, identify typical failure patterns, and generate structured semantic labels. Step 3: Scenario semantic retrieval mechanism: Failure semantics are embedded as query conditions into the structured test scenario library. Through semantic matching calculations using a large language model, a set of candidate scenarios with similar semantics is screened out. Step 4: Initial recommendation construction: Based on the candidate set, construct an initial set of recommended test cases according to multi-dimensional similarity and behavior criticality scores; Step 5: Reflective Semantic Optimization: Introduce a large language model reflection mechanism to perform semantic redundancy analysis and behavior coverage assessment on the initial recommendation set, output optimization suggestions, and update the recommended use case set. Step 6: Few-sample fine-tuning and adaptation: Use the final recommended test cases as training data to fine-tune the original model with a few samples to achieve targeted enhancement of key capabilities. Step 7, closed-loop self-evolution mechanism: Deploy the repaired model for testing again, collect a new round of logs, and re-enter the failure analysis and repair process to form a closed-loop system with continuous optimization capabilities.

[0009] The system structure of the present invention is as follows: (1) Log collection module: performs simulation tests and collects autonomous driving behavior logs; (2) Failure analysis module: Identifies behavioral failures and abnormal patterns in performance logs based on a large language model; (3) Scenario retrieval module: used to filter the candidate test case sets related to failure semantics from the scenario library; (4) Initial recommendation module: sorts and screens the candidate set to form the first round of recommended test case set; (5) Reflection optimization module: performs semantic reflection, behavioral coverage analysis, and redundancy removal on the initial recommendation set to generate the optimized final recommendation set; (6) Model adaptation module: Enhance the capabilities of the autonomous driving model through a few-sample training strategy; (7) Closed-loop control module: responsible for retesting and re-identifying failures of the repaired model, forming an automatic iterative optimization process.

[0010] The present invention expresses the extracted failure semantic set as P = {p1, p2, ..., p n}, the scenes to be retrieved in the scene library are recorded as set B = {s1, s2, ..., s m}, overall matching score: Where: T s The natural language description text representing the candidate scene s, Represents scene T s The semantic similarity scoring function between the candidate scene s and the failure semantics P, r(s,P) represents the overall matching score between the candidate scene s and the failure set P; Select the top K test scenarios with the highest scores to form the initial recommendation set C: C=Top K (argmax s∈B r(s,P)) (2) Where: C represents the initial recommended scene set.

[0011] The present invention further adjusts the structure and enhances the semantics of the generated initial recommendation scene set C. The main process is as follows: (1) Calculate the semantic redundancy between scenes within the initial recommendation set C and identify scene pairs with highly similar descriptions, repeated triggering behaviors, or overly similar traffic layouts (s i ,s j ); (2) Determine whether there is a subset of failed labels that is not covered by the recommended set If there are any missing behavior types, guide the LLM to complete such scenarios in the prompts; (3) The optimization suggestion set is expressed as: R={Replace(s i ,s j ),Add(s k ),Remove(sl )} (3) Where: Add(s k ) to add a new scene k ;Remove(s k ) is to remove redundant scenes s i ;Replace(s i ,s j ) is to replace or reconstruct an existing scene; (4) Based on the optimized recommendation set R, the initial recommendation set C is updated to generate the final set of recommended test cases: C′=RefineLLM(C,R) (4) Where: RefineLLM(.) represents the recommendation set update function based on the language model reflection suggestion execution, and C′ is the final recommendation result input to the model adaptation module.

[0012] Based on the output of the final recommended scenario set C′, the present invention performs a few-sample fine-tuning on the original autonomous driving model. The process is as follows: (1) Training sample construction: Extract the structured training sample set D from the recommended scene set = {(o t ,a t )}, where: t is the observation input of the t-th frame, a t Label actions or optimal strategy outputs for corresponding experts; (2) Fine-tuning method configuration: According to the structural differences of the original autonomous driving model, choose efficient parameter fine-tuning, full parameter fine-tuning or policy network alternative adaptation; (3) Loss function and optimization target design: Construct a training objective: Where: π θ (o t ) represents the current autonomous driving model strategy for observation o under parameter θ t Given the action output, L fail represents the directional loss function defined under the failure scenario, represents the expected value calculation under the recommended sample distribution, θ * Represents the model parameters obtained after optimization; In order to ensure the stability of the training process, the overall optimized loss function is set as follows: min θ L=L fail +λ·Ω(θ) (6) Where: Ω(θ) is the regularization term, λ is the regularization term loss weight hyperparameter; (4) Output the updated model parameters θ * , forming a new version of the strategic network with targeted repair capabilities

[0013] Closed-loop update of the present invention: convergence judgment mechanism The repaired model is put back into the test process, and the system enters the next round of self-evolution recommendation and repair until the expected performance indicators are achieved or there are no new failure samples: Where: t is the current iteration number; ε is the set failure threshold.

[0014] The present invention has obvious advantages in autonomous driving test case recommendation, failure pattern recognition, rapid adaptation of model capabilities and system closed-loop self-evolution, and can significantly improve the safety, generalization capability and test maintainability of the autonomous driving system. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a diagram of the overall system architecture of the self-evolutionary repair method for autonomous driving test cases based on a large language model according to the present invention; Figure 2 This is a diagram of the functional modules of the system described in the present invention, including a performance log collection module, a failure analysis module, a semantic recommendation module, a reflection optimization module, a model adaptation module, and a closed-loop control module; Figure 3 This is a flowchart of the failure mode extraction and semantic recommendation process described in the present invention, showing the processing path from performance logs to recommended scenarios; Figure 4 A flowchart of optimizing recommendation results using the reflection mechanism of the present invention, including diversity assessment, redundancy filtering, and semantic completion steps; Figure 5 This is a control logic diagram for the model's few-sample fine-tuning and closed-loop self-evolution process in the method of the present invention; Figure 6 A comparison chart of the behavior repair of the method of the present invention in a high-speed lane change scenario; Figure 7 This is a comparison diagram of the interactive repair method of the present invention in the right-turn scene at the intersection. DETAILED DESCRIPTION

[0016] Due to existing problems, there is an urgent need for a failure-driven self-evolutionary repair method for autonomous driving systems. This method can leverage the semantic understanding and recommendation capabilities of large language models to conduct deep semantic mining of the model's performance logs, recommend high-quality repair scenarios from a structured test scenario library, and achieve continuous evolution of model capabilities through small-sample fine-tuning and closed-loop mechanisms.

[0017] The present invention has the ability of automatic fault identification and semantic modeling: This invention introduces a large language model to perform semantic analysis and attribution modeling on performance logs, which can automatically identify the failure behavior of the system in specific scenarios, replacing the traditional method of relying on manual rules or label clustering, and effectively improving the recognition accuracy and processing efficiency of "semantic-level failures".

[0018] The scenario recommendations of this invention have high semantic alignment and behavioral diversity: This paper proposes a scenario retrieval and recommendation mechanism driven by failure semantics, achieving a precise mapping from system failure semantics to test case semantics. By introducing a reflective optimization module, the breadth and representativeness of recommended test cases in terms of behavioral trigger conditions and risk coverage can be further improved, overcoming the limitations of traditional methods in terms of limited recommendation results and weak scenario matching.

[0019] Supporting a small number of samples and rapid adaptation mechanism for key capability points: This method uses recommended test cases as high-value training samples for the model input, combined with a small sample fine-tuning strategy to achieve rapid repair and adaptation of specific capability points. This mechanism has fast training convergence and high data utilization, outperforming traditional coarse-grained methods based on full data retraining.

[0020] Build a self-feedback closed-loop capability evolution system: This paper proposes a closed-loop optimization mechanism that covers the entire process of "failure analysis—recommended repairs—and verification of results." The results of these repairs are directly fed back into the next round of testing, enabling continuous iterative updates of the system in real or simulated environments. This mechanism significantly improves the efficiency and verifiability of the model optimization process, resolving the issue of traditional decoupling of training and testing.

[0021] Possessing good versatility and engineering deployment compatibility: The method of the present invention can be seamlessly integrated into existing mainstream autonomous driving test platforms (such as CARLA, Waymo Open Sim, etc.), does not rely on specific model structures or interfaces, and is suitable for regression testing, capability verification, and risk reinforcement scenarios of various types of autonomous driving systems. It has good engineering portability and promotion value.

[0022] This invention aims to solve the following key technical problems in test case recommendation and capability repair of existing autonomous driving systems: 1. Lack of automatic failure mode identification and semantic analysis capabilities leads to low efficiency in performance log utilization and difficulty in accurately diagnosing system weaknesses. 2. The lack of semantic alignment between scenario matching and recommendation results in the inability of recommended test cases to effectively cover key failure behaviors. 3. The test data and model training process are separated from each other, making it impossible to form a closed-loop system for continuous optimization and verification; 4. The lack of an efficient and transferable small-sample adaptation mechanism limits the rapid iteration and generalization capabilities of the model in key scenarios.

[0023] To this end, the present invention proposes a self-evolution repair method and system for autonomous driving test cases based on a large language model, constructing a seven-stage integrated process of "log collection - failure analysis - scenario retrieval - recommendation construction - semantic optimization - capability adaptation - closed-loop control", and has the ability to adaptively recommend and quickly repair key failure modes.

[0024] The technical solution of the present invention comprises the following steps: 1. Performance log collection and pre-assessment execution: Deploy the autonomous driving model to be tested on the simulation platform, execute standardized test tasks, collect driving behavior logs, and generate structured performance records containing information on events such as collisions, deviations, and violations; 2. Failure pattern semantic extraction: Use a large language model to perform semantic analysis on log data, extract abnormal behavior segments and semantic elements, identify typical failure patterns, and generate structured semantic labels. 3. Scenario semantic retrieval mechanism: Failure semantics are embedded as query conditions in the structured test scenario library. Through semantic matching calculations using a large language model, a set of candidate scenarios with similar semantics is screened out. 4. Initial recommendation construction: Based on the candidate set, the initial recommendation test case set is constructed according to multi-dimensional similarity and behavior criticality scores; 5. Reflective Semantic Optimization: Introducing a large language model reflection mechanism to perform semantic redundancy analysis and behavior coverage assessment on the initial recommendation set, output optimization suggestions, and update the recommended use case set; 6. Few-sample fine-tuning and adaptation: Use the final recommended test cases as training data to fine-tune the original model with a few samples to achieve targeted enhancement of key capabilities. 7. Closed-loop self-evolution mechanism: The repaired model is redeployed and tested, a new round of logs are collected, and the failure analysis and repair process is re-entered to form a closed-loop system with continuous optimization capabilities.

[0025] To implement the above method, the present invention also provides a supporting system structure, including the following modules: 1. Log collection module: performs simulation tests and collects autonomous driving behavior logs; 2. Failure Analysis Module: Identifies behavioral failures and abnormal patterns in performance logs based on a large language model; 3. Scenario retrieval module: used to filter test case candidate sets related to failure semantics from the scenario library; 4. Initial recommendation module: sorts and screens the candidate set to form the first round of recommended test case sets; 5. Reflection and Optimization Module: This module performs semantic reflection, behavioral coverage analysis, and redundancy removal on the initial recommendation set to generate an optimized final recommendation set. 6. Model Adaptation Module: Enhances the capabilities of autonomous driving models through few-sample training strategies; 7. Closed-loop control module: responsible for retesting the repaired model and re-identifying failures, forming an automatic iterative optimization process.

[0026] The method and system proposed in the present invention can be deployed on existing mainstream autonomous driving simulation platforms (such as CARLA, Waymo Open Sim, etc.), seamlessly integrated with test servers and training platforms, and have good scalability and engineering practicality.

[0027] The following is a detailed description of the implementation of the self-evolution repair method and system for autonomous driving test cases based on a large language model proposed in the present invention in conjunction with the accompanying drawings.

[0028] like Figure 1 As shown, the present invention provides an autonomous driving test case self-evolution and repair system, which includes the following functional modules: Log collection module; failure analysis module; scene retrieval module; initial recommendation module; reflection optimization module; model adaptation module; closed-loop control module.

[0029] The above modules can exchange information through data channels, function interfaces or intermediate cache systems, forming a complete closed-loop self-evolution process of "testing-identification-recommendation-repair-verification".

[0030] like Figure 2 As shown, the execution process of the system includes the following steps: Step S1: Performance log collection The autonomous driving model to be tested is deployed in a simulation platform (such as CARLA), and automated driving tasks are performed sequentially according to a set of preset routes. After each round of tasks is completed, a structured performance log record is automatically generated by the log collection module.

[0031] The log data is stored in a standardized JSON format and includes but is not limited to: Task execution identification information: such as route_id, index, etc.; Task execution status: For example, status = "Failed-Agent deviated from the route"; Violation statistics: including whether the driver ran a red light, deviated from the route, failed to stop and yield, etc., recording the specific violation type, violation description text and violation location; Key scoring indicators: score_route: indicates the overall route completion of the current test task; score_penalty: indicates the score deduction caused by system behavior; score_composed: final combined score; Execution metadata: route_length: total length of the route; duration_game: total time of the task; duration_system: controls the actual occupied time of the system.

[0032] This structured log provides raw input data for subsequent failure analysis modules, supporting operations such as semantic modeling, behavior clustering, and feature attribution.

[0033] Step S2: Failure mode identification This step is completed by the failure analysis module. Its purpose is to extract semantic failure modes that can reflect system behavior defects from structured performance logs, providing semantic input for subsequent recommendations and repairs.

[0034] Specifically, it includes the following sub-processes: Field-level anomaly aggregation analysis: This process reads the status field and infractions dictionary items in the performance log and combines them with scenario context information (such as route number, violation time, and location) to construct an "abnormal behavior set" for each test task. Semantic label construction: This process converts the aforementioned structured abnormal behaviors into human-interpretable semantic labels. This process mainly includes: Behavior type (e.g., “route deviation,” “red light violation,” “target missed”); Environmental conditions (e.g., “nighttime,” “intersection,” “presence of obstructions”); Trigger mechanism (such as "misjudgment of vehicle speed", "failure to detect red light signal", "delayed braking"); the label construction process can be implemented by prompting the large language model (LLM) or the rule-based label induction module to generate a unified failure label string, such as: "Failure to slow down at an intersection during a red light violation triggers a point deduction for crossing a red light"; "The route deviation system deviated from the route by 3.2 meters at the bend"; Semantic tag structured output: The above tags are standardized and output into a structured form, which is used as retrieval conditions in the subsequent scene matching stage.

[0035] The structured tag contains the following fields: { "fail_type":"Red Light Violation", "trigger_point":"(x=341.25,y=209.1)", "context":"Intersection / Daytime / Static Occlusion", "description":"Failed to complete deceleration, triggering a traffic light violation" } The above structured failure semantic label set will be used as the semantic matching input for the scene retrieval in step S3 to construct highly relevant test recommendation scenes.

[0036] Step S3: Scene semantic retrieval and initial recommendation construction This step is completed collaboratively by the "scene retrieval module" and the "initial recommendation module". The goal is to filter out test cases that are highly semantically relevant from the structured scene library based on the set of failed semantic labels generated in step S2, and build an initial recommendation set.

[0037] The main process is as follows Figure 3 shown In the structured scenario library constructed by the present invention, each test scenario contains the following components: Natural language scene description text: such as "a right-turn intersection in the rain at night, with an obstruction ahead"; Structural metadata: including time, weather, road type, traffic flow density, traffic rule configuration, etc.; Participant behavior logic definition: including trajectory arrangement and interaction action definition of other vehicles, pedestrians, bicycles, etc.

[0038] First, the failure semantic set extracted in step S2 is expressed as P = {p1, p2, ..., p n}, the scenes to be retrieved in the scene library are recorded as set B = {s1, s2, ..., s m For each candidate scene s∈B, its natural language description text is recorded as T s , generate its semantic embedding through a large language model, combined with the semantic similarity scoring function between failure labels p∈P Calculate the overall matching score: Where T s The natural language description text representing the candidate scene s; Represents scene T s The semantic similarity scoring function between the candidate scenario s and the failure semantics P can be calculated based on the embedding distance or the autoregressive confidence of the large language model; r(s,P) represents the overall matching score between the candidate scenario s and the failure set P, which is used for recommendation ranking.

[0039] All candidate scenarios are sorted according to the above scoring function, and the top K test scenarios are selected to form the initial recommendation set C: C=Top K (argmax s∈B r(s,P)) (2) Where C represents the initial recommended scene set (Top-K semantic matching results).

[0040] Step S4: Reflection Optimization This step is completed by the reflective optimization module, which aims to further structurally adjust and semantically enhance the initial recommendation scenario set C generated in step S3, thereby improving the behavioral diversity, semantic coverage, and response quality to failure modes of the recommended test case set.

[0041] The optimization process uses a semantic hint mechanism based on the Large Language Model (LLM), which uses hint templates and behavioral constraints to guide the large language model to reflect on the content of the recommendation set and generate targeted modification suggestions. Figure 4 .

[0042] The reflective optimization process includes the following sub-steps: 1) Diversity Assessment Calculate the semantic redundancy between scenes in the initial recommendation set C and identify scene pairs with highly similar descriptions, repeated triggering behaviors or too similar traffic layouts (s i ,s j ). Diversity calculation methods include but are not limited to: Sentence embedding distance between scene texts; Similarity of behavior triggering patterns between behavior tags; The duplication between scene metadata (such as weather, road type, interactive characters). For candidate pairs whose redundancy exceeds the set threshold, they are recorded as redundant scenarios.

[0043] 2) Coverage analysis Align the recommendation set C with the failure semantic set P to determine whether there is a failure label subset not covered by the recommendation set. If there are any missed behavior types (such as missed detection of "merge failure" related scenarios), the LLM will be guided to complete such scenarios in the prompt.

[0044] 3) Reflection and suggestion generation After completing the diversity and coverage analysis, a prompt instruction template is constructed and LLM is called to generate a set of recommended actions, including: Add(s k )Add new scenes k ; Remove(sk ) Eliminate redundant scenes i ; Replace(s i ,s j )Replace or reconstruct existing scenes The optimization suggestion set is expressed as: R={Replace(s i ,s j ),Add(s k ),Remove(s l )} (3)

[0045] 4) Recommended episode updates According to the above optimization recommendation set R, the initial recommendation set C is updated to generate the final recommended test case set. C′=RefineLLM(C,R) (4) RefineLLM(.) represents the recommendation set update function based on language model reflection suggestions; C′ is the final recommendation result input to the model adaptation module, which has stronger behavioral coverage and robustness.

[0046] Ultimately, the optimized scenario set C′ will serve as the core input for subsequent model fine-tuning training to promote targeted adaptation to failure modes.

[0047] Step S5: Model adaptation This step is completed by the model adaptation module. The goal is to perform a few-sample fine-tuning on the original autonomous driving model based on the final recommended scene set C′ output in step S4, so as to achieve targeted repair and performance improvement of key capability points. The main process is shown in Figure 5 .

[0048] 1) Training sample construction Extract the structured training sample set D from the recommended scene set = {(o t ,a t )}, where: t is the observation input of the t-th frame (such as the sensor state after fusion, historical trajectory, etc.); a t Label the actions or optimal strategy outputs (such as reference paths and vehicle control instructions) for the corresponding experts; the sample sources can include data from different time periods and different agent roles in the scene, and support the matching relationship screening between "target failure label → sample subset".

[0049] 2) Fine-tuning method configuration Depending on the structural differences of the original autonomous driving model, different fine-tuning strategies can be selected for capability adaptation, including: Efficient parameter fine-tuning (such as LoRA / Adapter): only modify the inserted module or low-rank matrix part; Full parameter fine-tuning: perform backpropagation training on the entire network; Policy network alternative adaptation: Train a lightweight residual policy based on the original policy to avoid direct interference with the backbone network.

[0050] 3) Loss function and optimization target design: With the failure behavior repair as the goal, define the directional loss function L fail , we can combine behavior label weights, time window filtering, or multi-task strategy embedding to construct the following training objectives: where π θ (o t ) represents the current autonomous driving model strategy for observation o under parameter θ t The action output given; L fail represents the directional loss function defined under the failure scenario, which is used to measure the error between the model output and the expert behavior; represents the expected value calculation under the recommended sample distribution; θ * Represents the model parameters obtained after optimization.

[0051] In order to ensure the stability of the training process, the overall optimized loss function is set as follows: min θ L=L fail +λ·Ω(θ) (6) Ω(θ) is a regularization term, which is often used to constrain the complexity or range of model parameters; λ is a regularization term loss weight hyperparameter.

[0052] 4) Output the adapted model After fine-tuning is completed, the updated model parameters θ are output * , forming a new version of the strategic network with targeted repair capabilities The model will be redeployed in the next phase (S6) for testing and verification to build a closed-loop evolution mechanism.

[0053] Step S6: Closed-loop update 1) Convergence determination mechanism The repaired model is put back into the test process, and S1-S5 are repeated. The system enters the next round of self-evolution recommendation and repair until the expected performance indicators are achieved or there are no new failed samples. t: represents the current iteration number; ε: is the set failure threshold. When the number of failure modes in the new round does not exceed the threshold, the system is considered to have converged.

[0054] 2) Restoration effect evaluation indicators In addition to the number of failures, the present invention supports the expansion of the following performance indicators for determining the effectiveness of model repair: Driving Score increment (such as route completion rate, comfort, etc.); the decline in the number of violations; The number of key behavioral tasks (such as obstacle avoidance and yielding) successfully completed; Recommended sample usage efficiency (training effect / sample size ratio); Behavioral robustness growth (how stable the model is under uncertainty).

[0055] The evaluation results can be automatically written back to the control module for subsequent policy adjustments or exception tracking.

[0056] Simulation verification process To verify the feasibility and effectiveness of the method in the actual autonomous driving test process, the present invention constructed a closed-loop verification framework based on the CARLA simulation platform, and conducted before-and-after comparative experiments on multiple autonomous driving models on a set of standard test routes. The experimental steps are as follows:

[0057] 1. Deploy the model and perform pre-evaluation testing: Select a typical driving model (such as Rule-Based, End-to-End strategy, etc.), perform a complete test on the route set predefined by the CARLA platform, and record performance logs;

[0058] 2. Identify failure modes: Extract typical failure behaviors (such as red light violations, route deviations, and sudden braking delays) based on structured logs.

[0059] 3. Generate recommended use cases and optimize: Use the method of the present invention to recommend test scenarios that are highly matched with the failure mode, and further remove redundancy gains through the reflection module.

[0060] 4. Execution fine-tuning and capability adaptation: Input the recommended test cases into the original model for few-sample fine-tuning.

[0061] 5. Re-run the test and conduct comparative analysis: Compare the performance of the model before and after the repair on the same test route, including multiple indicators such as completion rate, number of violations, comfort, reaction time, etc.

[0062] Experimental results Experimental results show that the proposed method is significantly superior to the original model or the training method without recommended samples in multiple capability dimensions, mainly reflected in: 1. Reduced failure rate: Failure scenarios such as red light crossing and route deviation have been significantly reduced; 2. Driving score improvement: Driving Score increased by an average of 4%; 3. High sample efficiency: Targeted capability enhancement can be achieved using only <10 recommended use cases; 4. Strong closed-loop stability: Target behavior repair can be achieved within an average of 2 to 3 iterations.

[0063] These results verify the technical effectiveness of the present invention in the "recommendation-repair-evolution" process, and its good deployability and engineering adaptability.

[0064] Figure 6 This is a comparison chart of the effects of the present invention in a typical scenario of "high-speed lane change", showing the behavioral performance of the original autonomous driving model (first row above) and the model repaired by the present invention (second row below) under the same time sequence. When the original model faced a vehicle changing lanes on the right in the main lane of the highway, it did not slow down and avoid it in time, resulting in insufficient vehicle distance and even interactive danger. After identifying the failure of "lane change decision response delay" in the previous round of testing, the method of the present invention fine-tuned and repaired the model by recommending test cases with similar semantic conditions, so that the model can judge the intentions of surrounding vehicles in advance in this scenario, adjust its own speed appropriately, and safely complete the avoidance, showing better response robustness.

[0065] Figure 2 This is a comparison chart of the effects of the present invention in a typical interaction scenario of "turning right at an intersection", showing the difference in response under complex interaction conditions between the original autonomous driving model (first row above) and the model repaired by the method of the present invention (second row below). It can be observed from the figure that the original model failed to correctly identify the intention of the vehicle crossing laterally during the turning process, resulting in slow decision-making and even conflict with other vehicles. The method of the present invention identifies failure labels such as "failure to avoid turning right at an intersection" in the failure analysis stage, and recommends test cases containing similar semantic elements for fine-tuning and repair, so that the optimized model can actively judge the intention of the lateral vehicle, slow down and avoid it in advance, and safely complete the right turn operation.

Claims

1. A self-evolutionary repair method for autonomous driving test cases based on a large language model, characterized by: Step 1: Performance log collection and pre-assessment execution: Deploy the autonomous driving model to be tested on the simulation platform, execute standardized test tasks, collect driving behavior logs, and generate structured performance records containing information on events such as collisions, deviations, and violations. Step 2: Failure pattern semantic extraction: Use a large language model to perform semantic analysis on log data, extract behavioral anomalies and semantic elements, identify typical failure patterns, and generate structured semantic labels. Step 3: Scenario semantic retrieval mechanism: Failure semantics are embedded as query conditions into the structured test scenario library. Through semantic matching calculations using a large language model, a set of candidate scenarios with similar semantics is screened out. Step 4: Initial recommendation construction: Based on the candidate set, construct an initial set of recommended test cases according to multi-dimensional similarity and behavior criticality scores; Step 5: Reflective Semantic Optimization: Introduce a large language model reflection mechanism to perform semantic redundancy analysis and behavior coverage assessment on the initial recommendation set, output optimization suggestions, and update the recommended use case set. Step 6: Few-sample fine-tuning and adaptation: Use the final recommended test cases as training data to fine-tune the original model with a few samples to achieve targeted enhancement of key capabilities. Step 7, closed-loop self-evolution mechanism: Deploy the repaired model for testing again, collect a new round of logs, and re-enter the failure analysis and repair process to form a closed-loop system with continuous optimization capabilities.

2. The self-evolution repair method for autonomous driving test cases based on a large language model according to claim 1, characterized in that: Its supporting system structure: (1) Log collection module: performs simulation tests and collects autonomous driving behavior logs; (2) Failure analysis module: Identifies behavioral failures and abnormal patterns in performance logs based on a large language model; (3) Scenario retrieval module: used to filter the candidate test case sets related to failure semantics from the scenario library; (4) Initial recommendation module: sorts and screens the candidate set to form the first round of recommended test case set; (5) Reflection optimization module: performs semantic reflection, behavioral coverage analysis, and redundancy removal on the initial recommendation set to generate the optimized final recommendation set; (6) Model adaptation module: Enhance the capabilities of the autonomous driving model through a few-sample training strategy; (7) Closed-loop control module: responsible for retesting and re-identifying failures of the repaired model, forming an automatic iterative optimization process.

3. The self-evolution repair method for autonomous driving test cases based on a large language model according to claim 1 or 2, characterized in that: The extracted failure semantic set is represented as P = {p1, p2, ..., p n }, record the scene to be retrieved in the scene library as a set Combined B = {s1,s2,...,s m }, overall matching score: Where: T s The natural language description text representing the candidate scene s, Represents scene T s The semantic similarity scoring function between the candidate scene s and the failure semantics P, r(s,P) represents the overall matching score between the candidate scene s and the failure set P; Select the top K test scenarios with the highest scores to form the initial recommendation set C: C=Top K (argmax s∈B r(s,P)) (2) Where: C represents the initial recommended scene set.

4. The self-evolution repair method for autonomous driving test cases based on a large language model according to claim 1 or 2, characterized in that: The generated initial recommendation scene set C is further structurally adjusted and semantically enhanced. The main process is as follows: (1) Calculate the semantic redundancy between scenes within the initial recommendation set C and identify scene pairs with highly similar descriptions, repeated triggering behaviors, or overly similar traffic layouts (s i ,s j ); (2) Determine whether there is a subset of failed labels that is not covered by the recommended set If there are any missing behavior types, guide the LLM to complete such scenarios in the prompts; (3) The optimization suggestion set is expressed as: R={Replace(s i ,s j ),Add(s k ),Remove(s l )} (3) Where: Add(s k ) to add a new scene k ;Remove(s k ) is to remove redundant scenes s i ;Replace(s i ,s j ) is to replace or reconstruct an existing scene; (4) Based on the optimized recommendation set R, the initial recommendation set C is updated to generate the final set of recommended test cases: C′=RefineLLM(C,R) (4) Where: RefineLLM(.) represents the recommendation set update function based on the language model reflection suggestion execution, and C′ is the final recommendation result input to the model adaptation module.

5. The self-evolution repair method for autonomous driving test cases based on a large language model according to claim 1 or 2, characterized in that: Based on the output of the final recommended scene set C′, the original autonomous driving model is fine-tuned with a few samples. The process is as follows: (1) Training sample construction: Extract the structured training sample set D from the recommended scene set = {(o t ,a t )}, where: t is the observation input of the tth frame, a t Label actions or optimal strategy outputs for corresponding experts; (2) Fine-tuning method configuration: Based on the structural differences of the original autonomous driving model, choose efficient parameter fine-tuning, full parameter fine-tuning, or policy network alternative adaptation; (3) Loss function and optimization target design: Construct a training objective: Where: π θ (o t ) represents the current autonomous driving model strategy for observation o under parameter θ t Given the action output, L fail represents the directional loss function defined under the failure scenario, represents the expected value calculation under the recommended sample distribution, θ * Represents the model parameters obtained after optimization; In order to ensure the stability of the training process, the overall optimized loss function is set as follows: minutes θ L=L fail +λ·Ω(θ) (6) Where: Ω(θ) is the regularization term, λ is the regularization term loss weight hyperparameter; (4) Output the updated model parameters θ * , forming a new version of the strategic network with targeted repair capabilities 6. The self-evolution repair method for autonomous driving test cases based on a large language model according to claim 1 or 2, characterized in that: Closed-loop update: convergence determination mechanism The repaired model is put back into the test process, and the system enters the next round of self-evolution recommendation and repair until the expected performance indicators are achieved or there are no new failure samples: And|P (t+1) |≤ε (7) Where: t is the current iteration number; ε is the set failure threshold.

Citation Information

Cited By

  • Integrated training test method for automatic driving algorithm

    CN121029622A

  • Site test scene construction method based on feature matching and trajectory optimization

    CN121031143A

  • Method for constructing a test site scene based on feature matching and trajectory optimization

    CN121031143B

  • Multi-dimensional evaluation scale construction and test feedback method and system for electromagnetic tasks

    CN122526955A

  • Multi-dimensional evaluation scale construction and test feedback method and system for electromagnetic tasks

    CN122526955B