Automatic driving test case generation method based on large model retrieval enhancement technology
Through the large-scale retrieval enhancement technology, efficient and automated generation of autonomous driving test cases is achieved, solving the problems of low efficiency and insufficient diversity in the existing technology, and improving the authenticity and controllability of the test scenarios.
Patent Information
- Application Number
- CN202510503431.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-25
AI Technical Summary
The existing autonomous driving test scenario generation methods are difficult to cover complex realities. The generated scenarios lack diversity and innovation, and rely on a large amount of manual participation, which is inefficient.
The autonomous driving test case generation method based on large-scale retrieval enhancement technology is adopted, and multi-source data is processed through format analysis and vectorization, combined with an improved hybrid search algorithm and a BERT cross encoder, and the large language model is fine-tuned by LoRA strategy, and the prompt word template is designed to generate test cases to achieve efficient and automated generation.
It improves the generation efficiency and diversity of autonomous driving test scenarios, reduces the cost of manual design and data handling, significantly improves the authenticity and controllability of the scenarios, and supports the coverage of complex traffic situations.
Smart Images

Figure CN120371708A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving, and particularly to a method for generating test cases for autonomous driving based on large model retrieval enhancement technology. Background Art
[0002] In the field of autonomous driving, tasks such as autonomous driving tests of vehicles and training of prediction and decision-making models rely on a large amount of driving scenario data. In the collected driving scenario dataset, most scenarios exhibit relatively safe driving behaviors, while only a few extreme scenarios have higher risks. Therefore, it is very necessary to generate and expand more edge scenarios, which should be able to better simulate extreme situations in natural driving. By increasing the quantity and diversity of these scenarios, it helps to improve the performance and robustness of the model in various driving environments. The core of simulation testing is the automatic generation of scenarios, and mining safety-critical scenarios is of profound significance for autonomous driving tests.
[0003] Many domestic and foreign research scholars have proposed several solutions to the problem of scenario generation. Existing test scenario generation methods can be roughly divided into three categories: data-driven scenario generation methods, optimization search-based scenario generation methods, and neural network-based scenario generation methods. The data-driven generation method identifies and derives key scenarios from a large amount of real data, thereby increasing the authenticity and danger of the scenarios. However, the collection of natural driving data takes a long time and is costly, and it is often safe driving data; the core idea of the optimization search-based scenario generation method is to transform the scenario generation problem into an objective optimization problem, design an objective function to guide the search direction, and use an optimization algorithm to find the optimal solution. However, the method requires high computer resources, and the final effect depends on the quality of the designed objective function; the neural network-based scenario generation method uses a neural network model to generate test scenarios, which may lead to the lack of diversity and innovation in the generated scenarios.
[0004] Previous autonomous driving test scenario generation methods are difficult to cover complex real situations and generate sufficient edge scenarios, and cannot be easily extended to large-scale scenarios. In addition, a large amount of manual participation in scenario design is required, resulting in low efficiency. Traditional methods restrict the effectiveness and practicality in dealing with increasingly complex autonomous driving test requirements. Therefore, it is of great significance to study a method for improving the reliability and usability of autonomous driving test scenario generation. Summary of the Invention
[0005] In order to solve the above problems, the present invention proposes a method for generating test cases for autonomous driving based on large model retrieval enhancement technology to solve the problems of low efficiency, long time consumption in scenario design, and lack of diversity and innovation in the generated scenarios.
[0006] The present invention adopts the following technical solutions:
[0007] A method for generating autonomous driving test cases based on large model retrieval enhancement technology includes the following contents:
[0008] Step S1: Collect multi-source information data of autonomous driving test scenarios, and use format parsing and format conversion technologies to preprocess and convert the multi-source information data into plain text data; perform chunking processing on long texts in the plain text data, and split them into small fragments with concentrated content and complete semantics to ensure that the texts returned during retrieval match the query highly and have stronger information relevance; then use an embedding model to perform vectorization operations on the chunked texts to construct a text vector index, and store the index in a vector database.
[0009] Step S2: For the relevant scenario requirements of the user's autonomous driving test task, use an improved hybrid retrieval algorithm to quickly retrieve relevant text chunks from the vector database, and select TOP-N as the preliminary retrieval result; introduce a re-ranking mechanism based on the BERT cross-encoder to further optimize the ranking of the preliminary recall results to ensure that the most valuable information is displayed first; after re-ranking, select TOP-K text chunks as the final retrieval result.
[0010] Step S3: Construct a scenario model fine-tuning dataset for guiding the model to generate autonomous driving professional term-compliant ones; use the dataset to fine-tune the large language model using the LoRA (Low-Rank Adaptation) strategy, and adopt low-rank matrix adaptation technology to only update the newly added low-rank matrix parameters, thereby achieving efficient and flexible update of model parameters, while ensuring the effective retention of pre-trained knowledge and improving the model's adaptability in the autonomous driving test scenario generation task.
[0011] Step S4: Design a prompt template dedicated to adapting to the test case generation task, combine the TOP-K retrieval results generated in Step S2 with the prompt template, and input them into the large language model fine-tuned in Step S3, and use its reasoning ability to generate test cases that meet the autonomous driving test requirements.
[0012] Step S5: The Agent performs task planning and dynamic decision-making according to the user's needs and context, including complex problem decomposition and multi-round dialogue management, etc., so as to achieve a more customized and interactive service experience. The RouterAgent is responsible for distributing the user's needs to different sub-task Agents, and each sub-task Agent retrieves relevant information in its respective knowledge base. The system will comprehensively integrate the outputs of each sub-task Agent, and continuously optimize and improve the final answer through multi-round interaction and feedback iteration to provide a more accurate and personalized service.
[0013] Furthermore, the specific steps of Step S1 are as follows:
[0014] Step S11: Preprocess the data in the autonomous driving test scenario file: Parse and convert documents in various formats such as PDF, Word, txt, and CSV according to the corresponding document loaders into plain text format. For example, PyPDFLoader for PDF, Docx2txtLoader for Word documents, and CSVLoader for parsing CSV data. Then construct the converted text into a local knowledge base L consisting of n documents:
[0015] L = {l1, l2, l3,..., l n}
[0016] where li i represents the i-th document.
[0017] Step S12: After completing the format conversion, the text data chunking process includes two parts: coarse-grained splitting and fine-grained optimization:
[0018] (1) Coarse-grained splitting uses the CharacterTextSplitter to split large-scale text. The text is split according to a fixed size, setting the parameter chunk_size to 2000, that is, the maximum length of each text chunk is 2000 characters. To reduce semantic fragmentation caused by text truncation, a sliding window mechanism is introduced between adjacent text chunks, setting chunk_overlap to 100, so that each text chunk shares 100 characters with adjacent text chunks, thus retaining necessary context information.
[0019] (2) Fine-grained optimization, after completing the preliminary segmentation based on a fixed length, uses the RecursiveCharacterTextSplitter, which is an improved version of the ordinary CharacterTextSplitter, to further split. At this stage, the text is recursively split in the order of natural language levels: paragraphs, sentences, and phrases. separators=["\n\n", "\n", "。", ","] uses paragraph separators, line breaks, full stops, and commas as the splitting basis. If the segments after higher-level splitting are still too long, they are further recursively split at smaller units until the length requirements are met, so as to better retain the text sentence structure while ensuring that the length of each segment is appropriate.
[0020] Implement text splitting into text chunks, and split L into segments suitable for vectorized storage and information retrieval:
[0021] D = {d1, d2,..., d m}
[0022] where dj iRepresents the i-th segment in document l.
[0023] Step S13. The specific method of using the Embedding Model to convert text blocks into vector representations and storing them in the vector library to build an index is as follows: Select bge-base-zh as the Embedding Model to map text segments to a vector space of a fixed dimension, so that texts with similar semantics are closer in the vector space, facilitating subsequent similarity calculations, thereby improving the accuracy and relevance of retrieval. The vectorized text blocks are stored in the FAISS index to support subsequent efficient semantic retrieval and enhance the accuracy of text matching and information recall. The text block D is converted into an embedding vector V for semantic search through the Embedding Model E d , and the formula is:
[0024] V D = E(D)
[0025] V d is stored in the vector library V together with D, and (V D , D) represents the correspondence between the embedding vector and the text block. The formula is:
[0026] V = {(V D , D)}
[0027] Furthermore, the specific steps of step S2 are as follows:
[0028] Step S21. Based on the improved hybrid retrieval algorithm, make full use of the high precision of keyword retrieval and the semantic understanding ability of vector retrieval, and significantly improve the relevance of retrieval results. Sort the text blocks according to the hybrid retrieval score, and select TOP-N as the preliminary retrieval results.
[0029] (1) For keyword retrieval, a sparse embedding algorithm BM25 is adopted. The sparse embedding algorithm BM25 is an improved term frequency-inverse document frequency (TF-IDF) model, mainly used to calculate the correlation score between the query Q and the text block D. The calculation formula is:
[0030]
[0031] In the formula, q i is the i-th word in the query, f(q i , D) is the frequency of occurrence of q i in D, |D| is the length of D, avgdl is the average length of all documents, k1 and b are hyperparameters, generally taking k1 = 1.2 - 2.0, b = 0.75, which are used to control the influence of term frequency and the influence of document length normalization respectively; N is the total number of documents, and n(q i ) is the number of documents containing the word q iThe number of documents, IDF(q i ) is the inverse document frequency.
[0032] (2) Vector retrieval is based on the text embedding method, which converts the text into high-dimensional vectors and measures their semantic relevance by calculating the similarity between vectors. d(v1, v2) measures the spatial distance between vectors. The smaller the distance, the more similar the two vectors are. The data objects are sorted according to the similarity scores to ensure that the most relevant search results can be obtained. The Euclidean distance similarity calculation formula in FAISS vector retrieval is:
[0033]
[0034] To avoid a certain scoring method having too much impact on the final score, the BM25 and vector retrieval scores are normalized. The final retrieval score is a mixture of the BM25 score and the vector retrieval score, and its expression is:
[0035] hybrid(α) = α·BM25(D, Q) + (1 - α)·faiss_score(D, Q)
[0036] In the formula, α is the weighting coefficient, which controls the contribution ratio of the BM25 and vector retrieval scores in the final score, and is usually adjusted between [0, 1] to optimize the retrieval effect. BM(D, Q) is the BM25 similarity score. faiss_score(D, Q) is the similarity score between the query vector and the document vector calculated by vector retrieval.
[0037] Step S22: Introduce a re-ranking algorithm to optimize the ranking of the retrieved results. The re-ranking algorithm selects BERT as the cross-encoder, takes the query and the document as input pairs and inputs them into the model together, and uses BERT's powerful semantic understanding ability to calculate the relevance between the query and the document. After re-ranking, the TOP-K text blocks are selected as the final retrieval results.
[0038] First, the i-th query Q i and one of its corresponding N hybrid retrieval results d i (d i = d1, d2,..., d N ) The combined input sequence original text is converted into a format suitable for BERT, and the WordPiece tokenizer is used to process Q i , d iThe text is segmented into multiple sub - word units. During the training phase of the BERT model, a fixed - size vocabulary is constructed, where each sub - word unit corresponds to a unique integer ID. The tokenizer maps each sub - word unit to the corresponding entry in the vocabulary. To ensure that all input sequences have the same length, shorter sequences are padded with the [PAD] token. The query and retrieval results are merged into a continuous input sequence, with the [CLS] token added at the beginning of the sequence as the starting point, and [SEP] is used as a separator to distinguish between the query and the document part. The final input sequence is:
[0039] x = [CLS]Q i1 ,...,Q ij [PAD]...[PAD][SEP]d i1 ,...,d ij
[0040] where x represents the input sequence, Q ij represents the j - th sub - word unit in the query Q i , and d ij represents the j - th sub - word unit in the retrieval result d i .
[0041] Then, the input sequence is fed into BERT and optimized using the cross - entropy loss function during training. The specific calculation method is as follows:
[0042]
[0043] where y i is the actual label, is the predicted probability distribution of the model, and M is the number of samples.
[0044] Furthermore, the specific steps of step S3 are as follows:
[0045] Step S31: Construct a scenario model fine - tuning dataset for guiding the model to generate terms conforming to autonomous driving. This dataset is constructed in the form of an artificially annotated JSON file and randomly divided into a training set and a test set in a ratio of 8:2. The dataset is organized and stored in the Alpaca format, and the annotation format is <instruction, input, output>. Among them, instruction describes the task that the model needs to complete; Input is the optional input providing additional input information; output is the answer or result generated by the model.
[0046] Step S32: Use the fine - tuning dataset and adopt the LoRA method to fine - tune the large - language model. Two low - rank matrices A and B are introduced for the original weight parameter matrix W in the self - attention layer of the model. Where W ∈ R d×k, the matrices A ∈ R r ×k and B ∈ R d×r introduced by LoRA have a rank r that is much smaller than min(d, k), and are thus used to approximately represent the update of the original weights ΔW = BA. The update of the model parameters during LoRA fine-tuning can be expressed as:
[0047] h = Wx + ΔWx = Wx + BAx
[0048] At the start of training, matrix A is initialized with a random Gaussian distribution, and matrix B is initialized to zero. By calculating the gradients of the loss function with respect to all parameters and only passing the gradients to the low-rank matrices A and B during backpropagation, it is ensured that the original model parameters remain frozen and only the newly added low-rank modules are updated. The AdamW optimizer continuously updates the low-rank matrices according to the set hyperparameters in multiple iterations, gradually adjusting their parameters to gradually improve the performance of the model on the generation task.
[0049] Further, the specific steps of step S4 are as follows:
[0050] Design a prompt template dedicated to the task of generating test cases, combine the TOP-K retrieval results generated in step S2 with the prompt template, and input them into the large language model fine-tuned in step S3 to use its generation ability to generate test cases that meet the requirements of autonomous driving tests. The content of the designed prompt template mainly includes role-playing prompts, chain-of-thought prompts, format constraint prompts, few-shot prompts, and problem definition prompts. Designing the prompt template can effectively control and optimize the model output to make it more in line with the requirements of the scenario generation task.
[0051] Role-playing prompt: You are an expert in the field of autonomous driving scenarios. Limiting the output of the large model with a certain perspective or role can help the model focus on providing more professional and accurate answers in line with the autonomous driving field; Chain-of-thought prompt: Provide instruction prompts for the task to be completed by the model, embed logical steps or reasoning chains in the input prompt to guide the model to reason step by step to improve its performance in the scenario test case task, and prompt the model to imitate the human thinking process to generate more accurate and reasonable outputs; Format constraint prompt: Specify the format of the output content to ensure that the generated results meet the requirements of subsequent processing; Few-shot prompt: Provide several formatted examples to illustrate the generation scenario; Problem definition prompt: Clearly describe the core problems, goals, and constraints of the current scenario generation task.
[0052] Further, the specific steps of step S5 are as follows:
[0053] The Agent performs task planning and dynamic decision-making based on user needs and context, including complex problem decomposition and multi-turn dialogue management, thereby achieving more customized and interactive services. The Router Agent is responsible for the overall analysis and decomposition of task requirements, allocating different subtasks to each specialized Agent, and integrating their outputs in subsequent stages. Coordinate and monitor the system workflow, including multi-turn iteration, feedback processing, etc. Each subtask Agent corresponds to a specific type of retrieval knowledge base. After receiving the task, it will automatically generate retrieval keywords and query statements according to the specific requirements of the subtask, perform dynamic retrieval, local reasoning, and preliminary generation. It can conduct multi-turn negotiation with the Router Agent when necessary to further refine the retrieval scope or adjust the output format. The system comprehensively integrates the outputs of each subtask Agent and continuously optimizes and improves the final answer through multi-turn interaction and feedback iteration.
[0054] The beneficial effects of the present invention are as follows: Through the deep integration of the large model fine-tuning strategy and retrieval enhancement technology, making full use of the extensive knowledge reserve and powerful reasoning ability of the large language model, the efficient automatic generation and diverse coverage of autonomous driving test cases are realized. This technology can not only cover complex traffic scenarios in both conventional and edge cases, but also significantly reduce the costs of manual design and data handling, and remarkably improve the authenticity and controllability of the scenarios. Compared with the prior art, the present invention has significant advantages in terms of high efficiency and low cost, and has important theoretical significance and practical application value for improving the level of autonomous driving test evaluation and accelerating the implementation of autonomous driving technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 It is a flowchart of the autonomous driving test case generation system of the present invention.
[0056] Figure 2 It is a schematic diagram of the text vectorization indexing library entry of the present invention.
[0057] Figure 3 It is a schematic diagram of the hybrid retrieval re-ranking of the present invention.
[0058] Figure 4 It is a schematic diagram of the BERT cross-encoder principle of the present invention.
[0059] Figure 5 It is a structural diagram of the LoRA fine-tuning model of the present invention.
[0060] Figure 6 It is a schematic diagram of the multi-agent RAG of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] The present invention will be further clarified below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.
[0062] As Figure 1 shown, the present invention provides a method for generating autonomous driving test cases based on large model retrieval enhancement technology, including the following steps:
[0063] Step S1: Collect multi-source information data of autonomous driving test scenarios, and use format parsing and format conversion technologies to preprocess and convert the multi-source information data into pure text data; perform chunking processing on the long text in the pure text data, and split it into small fragments with concentrated content and complete semantics to ensure that the text returned during retrieval highly matches the query and has stronger information relevance; then use an embedding model to perform vectorization operations on the chunked text to construct a text vector index, and store the index in a vector database.
[0064] Step S2: For the relevant scenario requirements of the user's autonomous driving test task, use an improved hybrid retrieval algorithm to quickly retrieve relevant text blocks from the vector database, and select TOP-N as the preliminary retrieval result; introduce a re-ranking mechanism based on the BERT cross-encoder to further optimize and rank the preliminary retrieval result to ensure that the most valuable information is displayed first; after re-ranking, select TOP-K text blocks as the final retrieval result.
[0065] Step S3: Construct a scenario model fine-tuning dataset for guiding the model to generate terms conforming to autonomous driving; use the dataset to fine-tune the large language model using the LoRA (Low-Rank Adaptation) strategy, and adopt low-rank matrix adaptation technology to only update the newly added low-rank matrix parameters, thereby realizing efficient and flexible update of model parameters, while ensuring the effective retention of pre-trained knowledge and improving the adaptability of the model in the task of generating autonomous driving test scenarios.
[0066] Step S4: Design a prompt template dedicated to adapting to the test case generation task, combine the TOP-K retrieval results generated in Step S2 with the prompt template, input them into the large language model fine-tuned in Step S3, and use its reasoning ability to generate test cases that meet the autonomous driving test requirements.
[0067] Step S5: The Agent can actively determine the retrieval strategy and dynamically adjust the retrieval and generation methods according to the context. The greatest advantage is that it supports the intelligent decomposition and progressive solution of complex problems. The RouterAgent will allocate the research requirements to three dedicated sub-task Agents. After they retrieve information from their respective responsible knowledge bases, the system will perform comprehensive information integration and adjust and improve the answer according to the user's feedback.
[0068] In this embodiment, the process of indexing the text vectorization library in step S1 is as follows Figure 2 shown, specifically:
[0069] After preprocessing the data of the autonomous driving test scenario file, it is converted into a plain text format. After coarse-grained splitting and fine-grained optimization, text blocks are formed. Select bge-base-zh as the Embedding model to map the text fragments to a vector space of a fixed dimension, so that texts with similar semantics are closer in the vector space, thereby improving the accuracy and relevance of retrieval. The vectorized text blocks are stored in the FAISS index to support subsequent efficient semantic retrieval and improve the accuracy of text matching and information recall. The text block D is converted into an embedding vector V for semantic search through the Embedding model E d , and the formula is:
[0070] V D = E(D)
[0071] V d is stored in the vector library V together with D, and (V D , D) represents the corresponding relationship between the embedding vector and the text block. The formula is:
[0072] V = {(V D , D)}
[0073] In this embodiment, the hybrid retrieval and re-ranking process in step S2 is as follows Figure 3 shown, and the principle of the BERT cross-encoder is as follows Figure 4 shown, specifically:
[0074] First, the present invention sorts the text blocks according to the hybrid retrieval score based on the improved hybrid retrieval algorithm, and selects TOP-N as the preliminary retrieval result.
[0075] (1) Keyword retrieval uses the sparse embedding algorithm BM25. The sparse embedding algorithm BM25 is an improved term frequency-inverse document frequency (TF-IDF) model, which is mainly used to calculate the correlation score between the query Q and the text block D. The calculation formula is:
[0076]
[0077] (2) Vector retrieval is based on the text embedding method. The text is converted into a high-dimensional vector, and the semantic correlation is measured by calculating the similarity between the vectors. d(v1, v2) measures the spatial distance of the vectors. The smaller the distance, the more similar the two vectors are. The data objects are sorted according to the similarity score to ensure that the most relevant search results can be obtained. The Euclidean distance similarity calculation formula in FAISS vector retrieval is:
[0078]
[0079] To avoid a certain scoring method having too much impact on the final score, the BM25 and vector retrieval scores are normalized. The final retrieval score is a mixture of the BM25 score and the vector retrieval score, and its expression is:
[0080] hybrid(α) = α·BM25(D,Q) + (1 - α)·faiss_score(D,Q)
[0081] Then, the present invention introduces a re-ranking algorithm to optimize the ranking of the retrieved results. The re-ranking algorithm selects BERT as the cross-encoder, inputs the query and the document as an input pair into the model together, and utilizes the powerful semantic understanding ability of BERT to calculate the relevance between the query and the document. After re-ranking, the TOP-K text chunks are selected as the final retrieval results.
[0082] The i-th query Q i and one of its corresponding N hybrid retrieval results d i (d i = d1, d2,..., d N ) are combined into an input sequence original text and converted into an input format suitable for BERT. The WordPiece tokenizer is used to tokenize the Q i , d i texts into multiple sub-word units. During the training stage of the BERT model, a fixed-size vocabulary is constructed, where each sub-word unit corresponds to a unique integer ID. The tokenizer maps each sub-word unit to the corresponding entry in the vocabulary. To ensure that all input sequences have the same length, shorter sequences are padded with the [PAD] token. The query and the retrieval result are combined into a continuous input sequence, and the [CLS] token is added at the beginning of the sequence as the starting point, and [SEP] is used as a separator to distinguish the query and the document parts. The final input sequence is:
[0083] x = [CLS]Q i1 ,..., Q ij [PAD]...[PAD][SEP]d i1 ,..., d ij
[0084] where x represents the input sequence, Q ij represents the j-th sub-word unit in the query Q i , and d ij represents the j-th sub-word unit in the retrieval result d i .
[0085] The input sequence is fed into BERT and optimized using the cross-entropy loss function during training. The specific calculation method is as follows:
[0086]
[0087] Among them, y i is the actual label, is the predicted probability distribution of the model, and M is the number of samples.
[0088] In this embodiment, the LoRA fine-tuning model structure in step S3 is as Figure 5 shown, specifically:
[0089] First, a scenario model fine-tuning dataset is constructed to guide the model to generate terms conforming to the professional terms of autonomous driving. This dataset is constructed in the form of an artificially annotated JSON file and randomly divided into a training set and a test set according to a ratio of 8:2. The dataset is organized and stored in the Alpaca format, and the annotation format is <instruction, input, output>. Among them, instruction is the instruction describing the task that the model needs to complete; Input is the optional input providing additional input information; output is the answer or result generated by the model.
[0090] Then, the large language model is fine-tuned using the fine-tuning dataset and the LoRA method. Two low-rank matrices A and B are introduced for the original weight parameter matrix W in the self-attention layer of the model. Among them, W ∈ R d×k , the matrix A introduced by LoRA ∈ R r×k , B ∈ R d ×r , and its rank r is much smaller than min(d, k), so as to approximately represent the update of the original weight ΔW = BA. The model parameter update of LoRA fine-tuning can be expressed as:
[0091] h = Wx + ΔWx = Wx + BAx
[0092] At the beginning of training, the matrix A is initialized with a random Gaussian distribution, and the matrix B is initialized to zero. By calculating the gradients of the loss function with respect to all parameters and only passing the gradients to the low-rank matrices A and B during backpropagation, it is ensured that the original model parameters remain frozen and only the newly added low-rank modules are updated. The AdamW optimizer continuously updates the low-rank matrices according to the set hyperparameters in multiple iterations, gradually adjusting its parameters to gradually improve the performance of the model in the generation task.
[0093] In this embodiment, step S4 is specifically as follows:
[0094] Design a prompt template specifically for the task of generating test cases, combine the TOP-K retrieval results generated in step S2 with the prompt template, and input them into the fine-tuned large language model in step S3. Utilize its inference ability to generate test cases that meet the requirements of autonomous driving tests. The content of the designed prompt template mainly includes role-playing prompts, chain-of-thought prompts, format constraint prompts, few-shot prompts, and problem definition prompts. The designed prompt template can effectively control and optimize the model output to make it more in line with the requirements of the scenario generation task.
[0095] Role-playing prompt: You are an expert in the field of autonomous driving scenarios. Limiting the output of the large model with a certain view or role can help the model focus on providing more professional and accurate answers in line with the autonomous driving field; Chain-of-thought prompt: Provide instruction prompts for the task to be completed by the model, embed logical steps or reasoning chains in the input prompts, and guide the model to reason step by step to improve its performance in the scenario test case task, prompting the model to imitate the human thinking process and thus generate more accurate and reasonable outputs; Format constraint prompt: Specify the format of the output content to ensure that the generated results meet the requirements of subsequent processing; Few-shot prompt: Provide several formatted examples to illustrate the generated scenarios; Problem definition prompt: Clearly describe the core issues, goals, and constraints of the current scenario generation task.
[0096] In this embodiment, the multi-agent RAG process of step S5 is as Figure 6 shown, specifically:
[0097] The Agent performs task planning and dynamic decision-making based on user needs and context, including complex problem decomposition and multi-round dialogue management, so as to achieve a more customized and interactive service. The RouterAgent is responsible for overall analysis and decomposition of task requirements, allocating different subtasks to each specialized Agent, and integrating their outputs in subsequent stages. Coordinate and monitor the system workflow, including multi-round iteration, feedback processing, etc. Each subtask Agent corresponds to a specific type of retrieval knowledge base. After receiving the task, it will automatically generate retrieval keywords and query statements according to the specific requirements of the subtask, perform dynamic retrieval, local reasoning, and preliminary generation. It can conduct multi-round negotiation with the RouterAgent when necessary to further refine the retrieval scope or adjust the output format. The system will conduct comprehensive information integration and adjust and improve the answer through multi-round interaction and iteration.
[0098] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements all fall within the protection scope of the claims of the present invention.
Claims
1. A method for generating autonomous driving test cases based on large model retrieval enhancement technology, characterized in that: It includes the following steps: Step S1: Collect multi-source information data of autonomous driving test scenarios, and use format parsing and format conversion technologies to preprocess and convert the multi-source information data into plain text data; perform chunking on long texts in the plain text data, and split them into small fragments with concentrated content and complete semantics to ensure that the texts returned during retrieval highly match the query and have stronger information relevance; use an embedding model to perform vectorization operations on the chunked texts to construct a text vector index, and store the index in a vector database; Step S2: For the relevant scenario requirements of the user's autonomous driving test task, use an improved hybrid retrieval algorithm to quickly retrieve relevant text blocks from the vector database, and select TOP-N as the preliminary retrieval results; introduce a re-ranking mechanism based on a BERT cross-encoder to further optimize and rank the preliminary retrieval results to ensure that the most valuable information is displayed first; select TOP-K text blocks as the final retrieval results after re-ranking; Step S3: Construct a scenario model fine-tuning dataset for guiding the model to generate terms conforming to autonomous driving; use the dataset to fine-tune the large language model with the LoRA strategy, and adopt low-rank matrix adaptation technology to only update the newly added low-rank matrix parameters, thereby achieving efficient and flexible update of model parameters, while ensuring the effective retention of pre-trained knowledge and improving the model's adaptability in the autonomous driving test scenario generation task; Step S4: Design a prompt template dedicated to adapting to the test case generation task, combine the TOP-K retrieval results generated in Step S2 with the prompt template, and input them into the large language model fine-tuned in Step S3, and use its reasoning ability to generate test cases that meet the autonomous driving test requirements; Step S5: The Agent performs task planning and dynamic decision-making according to the user's needs and context, including complex problem decomposition and multi-turn dialogue management, so as to achieve a more customized and interactive service experience; the RouterAgent is responsible for allocating the user's needs to different sub-task Agents, and each sub-task Agent retrieves relevant information in its own knowledge base; the system comprehensively integrates all parties' information and adjusts and improves the answer according to the user's feedback to provide more accurate and personalized services.
2. The method for generating an autonomous driving test case based on the large model retrieval enhancement technology according to claim 1, wherein: The specific steps of Step S1 are as follows: Step S11: Preprocess the data of the autonomous driving test scenario file: Parse the file formats of the collected PDF, Word, txt, and CSV documents according to the corresponding document loaders and convert them into plain text format; Step S12: After completing the format conversion, the text data chunking process includes two parts: coarse-grained splitting and fine-grained optimization: (1) Coarse-grained splitting uses the CharacterTextSplitter to split large-scale texts; split the texts according to a fixed size, and introduce a sliding window mechanism between adjacent text blocks to reduce semantic fragmentation caused by text truncation; (2) Fine-grained optimization: After the initial segmentation based on a fixed length is completed, the RecursiveCharacterTextSplitter, which is an improved version of the ordinary character text splitter, is used for further segmentation. At this stage, the text is recursively segmented in the order of natural language levels, namely paragraphs, sentences, and phrases. If the segments after a higher-level segmentation are still too long, further recursive segmentation is performed at a smaller unit until the length limit is met, so as to better preserve the semantic structure of the text and ensure that the length of each segment is appropriate. Step S13: Use an embedding model to perform text vectorization on the text blocks. First, map the text segments to vectors for subsequent similarity calculation, and then store the generated vectors in a vector database and construct an efficient index structure to quickly retrieve similar texts in a large-scale vector set.
3. The method for generating an autonomous driving test case based on the large model retrieval enhancement technology according to claim 1, characterized in that: The specific steps of step S2 are as follows: Step S21: Sort the text blocks based on an improved hybrid retrieval algorithm. This method combines the high precision of keyword retrieval and the semantic understanding ability of vector retrieval, significantly improving the relevance of retrieval results. (1) Keyword retrieval uses the sparse embedding algorithm BM25. The sparse embedding algorithm BM25 is an improved term frequency-inverse document frequency model, which is used to calculate the relevance score between the query Q and the text block D. The calculation formula is: where q i is the i-th word in the query, f(q i , D) is the frequency of q i appearing in D, |D| is the length of D, avgdl is the average length of all documents, k1 and b are hyperparameters, k1 = 1.2 to 2.0, b = 0.75, which are used to control the influence of term frequency and the influence of document length normalization respectively; N is the total number of documents, n(q i ) is the number of documents containing the word q i , IDF(q i ) is the inverse document frequency; (2) Vector retrieval is based on the text embedding method. The text is converted into high-dimensional vectors, and the semantic relevance is measured by calculating the similarity between the vectors. d(v1, v2) measures the spatial distance between the vectors. The smaller the distance, the more similar the two vectors are. The data objects are sorted according to the similarity score to ensure that the most relevant search results can be obtained. The Euclidean distance similarity calculation formula in FAISS vector retrieval is: To avoid a certain scoring method having too much influence on the final score, normalization processing is performed on the BM25 and vector retrieval scores. The final retrieval score is a mixture of the BM25 score and the vector retrieval score, and its expression is: hybrid(α) = α·BM25(D, Q) + (1 - α)·faiss_score(D, Q) In the formula, α is the weighting coefficient, which controls the contribution ratio of the BM25 and vector retrieval scores in the final score and is adjusted between [0, 1]; BM(D, Q) is the BM25 similarity score; faiss_score(D, Q) is the similarity score between the query vector and the document vector calculated by vector retrieval. Step S22: Introduce a re-ranking algorithm to optimize the ranking of the retrieved results. The re-ranking algorithm selects BERT as the cross-encoder, inputs the query and the document as a pair together into the model, and uses BERT's powerful semantic understanding ability to calculate the relevance between the query and the document. After re-ranking, select the TOP-K text blocks as the final retrieval results. First, take the i-th query Q i and one of its corresponding hybrid retrieval results d i (d i = d1, d2,..., d N ) Combine the original input sequence into an input format suitable for BERT. Use the WordPiece tokenizer to tokenize the Q i , d i texts into multiple sub-word units; During the BERT model training phase, a fixed-size vocabulary is constructed, where each sub-word unit corresponds to a unique integer ID; The tokenizer maps each sub-word unit to the corresponding entry in the vocabulary; To ensure that all input sequences have the same length, shorter sequences are padded with the [PAD] token; The query and retrieval results are combined into a continuous input sequence, with the [CLS] token added at the beginning of the sequence as the starting point, and [SEP] is used as a separator to distinguish the query and document parts; The final input sequence is: x = [CLS]Q i1 ,..., Q ij [PAD]...[PAD][SEP]d i1 ,..., d ij where x represents the input sequence, and Q ij represents the j-th sub-word unit in the query Q i , and d ij represents the j-th sub-word unit in the retrieval result d i ; Then, the input sequence is input into BERT and optimized using the cross-entropy loss function during training. The specific calculation method is: where y i is the actual label, is the predicted probability distribution of the model, and M is the number of samples.
4. The method for generating an autonomous driving test case based on the large model retrieval enhancement technology according to claim 3, wherein: The specific steps of step S3 are as follows: Step S31: Construct a fine-tuning dataset for the scenario model that guides the model to generate terms conforming to autonomous driving terminology. This dataset is constructed in the form of an artificially annotated JSON file and randomly divided into a training set and a test set in a ratio of 8:
2. The dataset is organized and stored in the Alpaca format, and the annotation format is <instruction, input, output>. Among them, instruction is the instruction describing the task that the model needs to complete; Input is the optional input to provide additional input information; output is the answer or result generated by the model. Step S32: Use the fine-tuning dataset and adopt the LoRA method to fine-tune the large language model. Introduce two low-rank matrices A and B for the original weight parameter matrix W in the self-attention layer of the model; where W ∈ R d×k , the matrix A ∈ R r×k introduced by LoRA, and B ∈ R d×r , whose rank r is much smaller than min(d, k), so as to approximately represent the update of the original weight ΔW = BA; The update of the model parameters fine-tuned by LoRA can be expressed as: h = Wx + ΔWx = Wx + BAx At the beginning of training, matrix A is initialized with a random Gaussian distribution, and matrix B is initialized to zero. By calculating the gradients of the loss function with respect to all parameters and only passing the gradients to the low-rank matrices A and B during backpropagation, it is ensured that the original model parameters remain frozen, and only the newly added low-rank module is updated. The AdamW optimizer continuously updates the low-rank matrix according to the set hyperparameters in multiple iterations, gradually adjusting its parameters to gradually improve the performance of the model in the generation task.
5. The method for generating an automatic driving test case based on the large model retrieval enhancement technology according to claim 1, characterized in that: The specific steps of step S4 are as follows: Design a prompt template dedicated to the task of generating test cases, combine the TOP-K retrieval results generated in step S2 with the prompt template, and input them into the large language model fine-tuned in step S3. Use its generation ability to generate test cases that meet the autonomous driving test requirements. The designed prompt template content includes role-playing prompts, chain-of-thought prompts, format constraint prompts, few-shot prompts, and problem definition prompts. Designing the prompt template can effectively control and optimize the model output to make it more in line with the requirements of the scenario generation task. Role-playing prompt: You are an expert in the field of autonomous driving scenarios; the output of the large model with certain view or role restrictions can help the model focus on providing more professional and accurate answers in line with the autonomous driving field. Chain-of-thought prompt: Provide instruction prompts for the task that the model needs to complete, embed logical steps or reasoning chains in the input prompts, guide the model to reason step by step to improve its performance in the scenario test case task, and prompt the model to imitate the human thinking process to generate more accurate and reasonable outputs. Format constraint prompt: Specify the format of the output content to ensure that the generated results meet the requirements of subsequent processing; few-shot Prompt: Provide several formatted examples to illustrate the generated scenario. Problem definition prompt: Clearly describe the core problems, goals, and constraints of the current scenario generation task.
6. The method for generating an autonomous driving test case based on the large model retrieval enhancement technology according to claim 1, wherein: The specific steps of step S5 are as follows: The Agent performs task planning and dynamic decision-making based on user needs and context, including complex problem decomposition and multi-turn dialogue management, so as to achieve more customized and interactive services; the routing agent RouterAgent is responsible for the overall analysis and decomposition of task requirements, allocating different subtasks to each specialized Agent, and integrating their outputs in subsequent stages; coordinating and monitoring the system workflow, including multi-turn iteration and feedback processing; each subtask Agent corresponds to a specific type of retrieval knowledge base, and after receiving the task, it will automatically generate retrieval keywords and query statements according to the specific requirements of the subtask, perform dynamic retrieval, local reasoning and preliminary generation; it can conduct multi-turn negotiation with the routing RouterAgent when necessary to further refine the retrieval scope or adjust the output format; the system will comprehensively integrate the outputs of each subtask Agent and continuously optimize and improve the final answer through multi-turn interaction and feedback iteration.
Citation Information
Cited By
Electricity marketing work order reply generation method and system based on generative large model
CN120670562A
Industrial quality data management system based on dialogue type large language model
CN120910401A
Automatic driving cloud simulation platform test method and device
CN121187943A
Automatic driving cloud simulation platform test method and device
CN121187943B
Automatic driving-oriented driving behavior semantic library construction method and scene retrieval method and system
CN122173408A