Model distillation method and system based on large language model automatic driving system

By adopting the model distillation method in the large language model autonomous driving system, an offline data set is constructed and a robust regularization process is generated, the problem of excessive inference time and insufficient adaptability of the large language model in the real-time autonomous driving environment is solved, and the real-time inference ability with efficient, robust and highly adaptable is achieved.

CN120146176APending Publication Date: 2025-06-13INST OF SOFTWARE - CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510133207.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing large language model has problems such as long inference time and difficulty in dealing with continuous data acquisition and learning in dynamic environments in real-time autonomous driving, which limits its application in real-time decision-making.

Method used

A model distillation method based on a large language model autonomous driving system is proposed. By constructing offline data sets, generating a distillation strategy with robust regularization processing, and achieving efficient and robustness of the model through interactive fine-tuning joint strategy with the online environment.

Benefits of technology

It realizes more efficient, robust and highly adaptable real-time inference capabilities, ensuring adaptability and stability in various driving scenarios, and improving the real-time decision-making capabilities of large-language models in autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146176A_ABST
    Figure CN120146176A_ABST
Patent Text Reader

Abstract

The invention discloses a model distillation method and system based on a large language model automatic driving system, and belongs to the technical field of artificial intelligence. The method comprises the following steps: constructing an offline data set DLLM; based on the offline data set, generating a distillation strategy of robustness regularization processing; fixing the distillation strategy and finely tuning the joint strategy through interaction with an online environment to generate a trained joint strategy; wherein the joint strategy comprises an adapter strategy and the distillation strategy. According to the invention, adaptability and robustness in various driving scenes can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a model distillation method and system for an autonomous driving system based on a large language model. Background Art

[0002] With its excellent common sense reasoning ability, the large language model shows great application potential in the field of autonomous driving. However, existing large language models have problems such as too long inference time and difficulty in processing continuous data collection and learning in dynamic environments in real-time autonomous driving scenarios, which limit their application in real-time decision-making. Summary of the Invention

[0003] The present invention proposes a model distillation method and system for an autonomous driving system based on a large language model, which can ensure adaptability and robustness in various driving scenarios.

[0004] To achieve the above object, the technical solution of the present invention includes the following content.

[0005] A model distillation method for an autonomous driving system based on a large language model, the method comprising:

[0006] Constructing an offline dataset D LLM , each set of offline data in the offline dataset D LLM includes: the current state s at each decision, the selected action a * , the obtained feedback r, and the updated state s ′ ;

[0007] Generating a distillation strategy with robustness regularization based on the offline dataset;

[0008] Fixing the distillation strategy and fine-tuning the joint strategy through interaction with the online environment to generate a trained joint strategy; wherein, the joint strategy includes: an adapter strategy and the distillation strategy.

[0009] Further, the constructing of the offline dataset includes:

[0010] Conducting a closed-loop driving experiment in the HighwayEnv environment and obtaining the current state s and historical information H;

[0011] Generating a prefix prompt for the GPT-3.5 model according to the current state s and historical information H;

[0012] The GPT-3.5 model makes a decision inference based on the prefix prompt to generate an action set A = {a 1 , a 2 , …, a n}, and for each action a iAllocate the probability distribution p(a i |s,H);

[0013] Based on the probability distribution p(a i |s,H), obtain the selected action a * ;

[0014] Execute the selected action a * , and obtain the feedback r and the updated state s ′ .

[0015] Furthermore, the selected action a * includes: changing to the left lane, changing to the right lane, accelerating, decelerating, and idling.

[0016] Furthermore, based on the offline dataset D LLM , generate a distillation strategy with robustness regularization, including:

[0017] Construct the objective function J robust (Q,π distill ,D LLM ); where the objective function J robust (Q,π distill ,D LLM ) represents the performance of the current Q-value function Q and the distillation strategy π distill on the offline dataset D LLM ;

[0018] On the offline dataset D LLM , based on the objective function J robust (Q,π distill ,D LLM ) with robustness regularization, train the large language model to obtain a distillation strategy with robustness regularization.

[0019] Furthermore, the objective function where J(Q,π distill ,D LLM ) represents the objective function to be minimized in the Q-value function, α represents the first adjustment parameter, β represents the second adjustment parameter, represents the Q-value expectation of action selection under the adversarial strategy, a~μ(·|s) represents the probability distribution of selecting action a under the adversarial strategy, the standard dataset D is the offline dataset D LLM , represents the Q-value expectation under the standard dataset D, Q(s,a) represents the Q-value function, onehot(a) represents the one-hot encoding of action a, and σ(Q(s,a)) represents the probability distribution after the Q-value passes through the softmax function.

[0020] Further, the process of generating the combined policy includes:

[0021] Select the most important K sub-policies π from N sub-policies i , where the K sub-policies π i include: an adapter policy and a distillation policy;

[0022] Based on the observation of the current state s, obtain the weight ω of the sub-policy π i ; i ;

[0023] Combine the weight ω of the sub-policy i to generate a combined policy where θ d are the parameters of the action decoder, θ r are the parameters of the router, θ p are the parameters of the sub-policy, are the parameters of the i-th sub-policy, represents the action decoder based on the parameters θ d , represents the router based on the parameters θ r .

[0024] Further, after generating the trained combined policy, it further includes:

[0025] Based on the trained combined policy, generate an action in the current state

[0026] Use the action as the control command of the vehicle.

[0027] A model distillation system for an autonomous driving system based on a large language model, the system includes:

[0028] A dataset construction module for constructing an offline dataset D LLM , where each set of offline data in the offline dataset D LLM includes: the current state s at each decision, the selected action a * , the obtained feedback r, and the updated state s ′ ;

[0029] A distillation policy generation module for generating a distillation policy with robust regularization processing based on the offline dataset;

[0030] The joint policy training module is used to fix the distillation policy and fine-tune the joint policy through interaction with the online environment to generate the trained joint policy; wherein, the joint policy includes: an adapter policy and the distillation policy.

[0031] An electronic device, the electronic device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the model distillation method of the large language model-based autonomous driving system described in any one of the above is implemented.

[0032] A computer-readable storage medium, characterized in that computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by a processor, the model distillation method of the large language model-based autonomous driving system described in any one of the above is implemented.

[0033] Compared with the prior art, the present invention extracts knowledge from a large language model-driven autonomous driving agent and converts it into a lighter-weight reinforcement learning policy, achieving more efficient, robust, and adaptable real-time inference. This method not only inherits the powerful reasoning ability of the large language model but also ensures adaptability and robustness in various driving scenarios through the hybrid policy technique. Description of the Drawings

[0034] Figure 1 is a schematic diagram of the joint policy generation and system output adjustment process in the autonomous driving system framework based on large language models and reinforcement learning according to an embodiment of the present invention.

[0035] Figure 2 is a flowchart of the offline data collection and driving action generation process in the autonomous driving system framework based on large language models and reinforcement learning according to an embodiment of the present invention.

[0036] Figure 3 is a schematic diagram of the robust regularization distillation optimization process in the autonomous driving system framework based on large language models and reinforcement learning according to an embodiment of the present invention.

[0037] Figure 4 is a schematic diagram of the robust policy adaptation process of large language model knowledge fusion in the autonomous driving system framework based on large language models and reinforcement learning according to an embodiment of the present invention. Detailed Embodiments

[0038] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below through specific embodiments in combination with the accompanying drawings.

[0039] As Figure 1As shown in the figure, the present invention includes three stages: offline data collection, robust regularization distillation, and robust strategy and environment adaptation of large language model knowledge fusion.

[0040] Stage 1: Offline Data Collection.

[0041] In an autonomous driving system, the processing and preprocessing of sensor data are key steps, which provide basic information for subsequent policy learning and decision-making. The main objectives of processing and preprocessing are to integrate different types of sensor data into a unified state representation after collection, and remove noise and fill in missing data to improve the stability of the model and the accuracy of decision-making. First, the model conducts closed-loop driving experiments in the HighwayEnv environment, and uses the GPT-3.5 model to collect offline data. Since GPT-3.5 cannot directly interact with the HighwayEnv simulator, the present invention uses a perception tool and agent prompts to assist it in making decisions. As Figure 2 shown, the specific process is as follows:

[0042] Stage 1.1: Scene Perception and Information Extraction.

[0043] GPT-3.5 obtains the current state and historical information through prefix prompts. Assuming the current state is s and the historical information is H, the content input to the model is the prefix prompt P(s, H), where:

[0044] P(s, H) = Concat(current_scene(s), history(H))

[0045] Stage 1.2: Decision Inference.

[0046] Based on the current state s and historical information H, the system will call GPT-3.5 to perform decision inference using the ReAct framework, and generate a rationality judgment of driving operations. The model outputs a series of possible driving actions A = {a 1 , a 2 , …, a n}, and assigns a probability distribution p(a i | s, H) to each primitive action a i .

[0047] Stage 1.3: Primitive Action Output.

[0048] In this stage, the present invention designs the matrix input data for GPT-3.5 to decide to execute a specific primitive action as: (lane_left, lane_right, faster, slower, idle). Assuming the selected primitive action is a * , we get:

[0049]

[0050] During this data collection process, the system generates a dataset D LLM , recording the current state s, the selected action a * , the obtained feedback r, and the updated state s ′ :

[0051] D LLM ={(s, a * , r, s ′ ) | a * ~π LLM (a | s)

[0052] Here, π LLM represents the intelligent agent policy based on the large language model. This dataset provides the basis for subsequent policy learning, recording the system's decisions and their effects in different scenarios. The core task at this stage is to collect a large amount of data through simulated driving experiments, including the state information, decision-making choices, and results of the vehicle in different driving scenarios.

[0053] Phase II: Robust Regularization Distillation.

[0054] The model of the present invention not only needs to be able to effectively execute driving tasks but also maintain stability when facing various uncertainties and potential attacks. Therefore, based on the offline dataset, this model further improves its adaptability in complex driving environments by applying a distillation strategy with robust regularization processing. The core of this stage is to optimize the objective function to obtain a more robust distillation strategy π distill , as Figure 3 shown, which includes two key parts in the following stages: standard Q-learning distillation and robust regularization.

[0055] Phase 2.1: Standard Q-learning Distillation.

[0056] During the distillation process, it is first necessary to clarify the mechanism of action of standard Q-learning. Its goal is to learn the maximum cumulative reward that can be obtained by executing action a in the current state s by repeatedly updating the Q-value function Q(s, a). The more accurate the Q-value function, the greater the possibility that the policy π distill makes efficient decisions during execution.

[0057] In this step, based on the offline dataset D LLM , a standard offline RL objective function is defined to optimize the Q-value function, and the specific expression is as follows:

[0058]

[0059] Among them, it should be noted that:

[0060] · In the formula, k represents the number of iterations of model update, k + 1 represents the increment of the number of iterations, referring to the next iteration;

[0061] · J(Q, π distill , D LLM ) represents the objective function to be minimized, and optimizing makes the error close to zero, representing the current Q-value function Q and the distillation policy π distill on the dataset D LLM performance. By the optimization process, the Q-value function Q(s, a) is made more accurate, thereby enhancing the ability of the policy π distill where represents the target Q-value function, represents the target Q-value function at the (k + 1)-th iteration;

[0062] · r(s, a): The immediate reward obtained by the system in state s after executing action a;

[0063] · The expected Q-value of the distillation policy π distill in the next state s′.

[0064] The core idea of this formula is to optimize the Q-function by minimizing the error between the predicted Q-value and the actual Q-value. In other words, the model continuously adjusts the Q-function to more accurately estimate the long-term reward for each state-action pair, thereby improving the accuracy of decision-making.

[0065] Phase 2.2: Robust regularization.

[0066] Due to the possible uncertainties and potential adversarial attacks faced by the autonomous driving system. Using simple standard Q-learning may perform poorly in the face of malicious adversarial inputs. Therefore, in this step, the present invention introduces a robust regularization term to improve the anti-interference ability of the policy. The introduction of the robust regularization term aims to counter the possible impacts of adversarial data (such as noise, malicious attacks) and ensure the stability and robustness of the policy π distill The robust objective function is defined as follows:

[0067]

[0068] Among them, it should be noted that:

[0069] · J robust (Q, π distill , D LLM ) represents the objective function after robust regularization. It measures the error between the predicted value and the target value of the Q-value function, and the goal is to make this error as small as possible, thereby training the distillation model π distill to have a higher Q-value;

[0070] · α and β are adjustment parameters used to balance the influence between standard Q-learning and robustness regularization. The larger the values of parameters α and β, the higher the attention of the model to robustness during training;

[0071] · represents the expected Q value of action selection under the adversarial policy, and a ∼ μ(·|s) represents the probability distribution of selecting actions under the adversarial policy. This term is used to evaluate the performance of the model under adversarial inputs;

[0072] · represents the expected Q value under the standard dataset D, which is used to compare with the adversarial Q value expectation,

[0073] where the standard dataset is usually composed of the dataset D collected offline in the first stage LLM constituted;

[0074] · σ(·) is the softmax function used to map Q values to probability distributions for facilitating the understanding of the relative probabilities of action selection;

[0075] · "onehot(a)" is the onehot encoding of action a, which uses the binary method to identify the uniqueness of a specific action.

[0076] By introducing robustness regularization, the model of the present invention not only optimizes the standard Q value function during training but also can handle possible adversarial inputs, thereby ensuring that the distillation policy π distill In the real driving environment, even when facing uncertainties or malicious interferences, it can still make stable and reasonable driving decisions.

[0077] Stage three: Robust policy and environmental adaptation with large language model knowledge fusion.

[0078] As Figure 4 shown, in this stage, the present invention further utilizes the interaction of the online environment, combines the large language model distillation policy and the adapter policy (the combined optimization of the hybrid policy model), and generates a more stable and adaptable joint policy.

[0079] Stage 3.1: Observation feature extraction and hybrid policy.

[0080] In this stage, the system generates a hybrid policy through the feature extraction and dynamic weight allocation mechanism. The features are processed by the matrix G(s), and the matrix G(s) receives the observations of the current state s, such as the environmental layout, speed, historical information, etc., to obtain the weight ω for the routing policy:

[0081] ω = softmax(topK(G(s)))

[0082] Among them, the TopK operation only selects the current most important K sub-policies for weighting to reduce the computational complexity while ensuring the diversity of the policies. During the generation process of the joint policy, the system dynamically allocates the sub-policy weights ω through the matrix G(s) and fuses the outputs of different sub-policies into the final joint policy. The sub-policy parameters include the fixed distilled policy parameters generated by offline distillation of the large language model and the adapter policy parameters that are dynamically optimized and continuously adjusted through environmental feedback. Then, the outputs of N policies are weighted by weights to obtain the final joint policy Π MoP :

[0083]

[0084] where θ d , θ r , θ p are the parameters of the action decoder, router, and sub-policies respectively; i represents the number of sub-policies. In this embodiment, K = N = 2, and this hybrid policy has two sources, namely the Luban regularization knowledge policy distilled by the large language model in stage two and the adaptive policy of online environmental interaction.

[0085] Stage 3.2: Environmental adaptation of the joint policy.

[0086] In actual deployment, the system further fine-tunes the joint policy through interaction with the online environment to ensure that it can adapt to dynamic driving scenarios. As mentioned before, during this process, the large language model distilled policy is fixed, and the adapter policy is incrementally trained through environmental feedback to gradually improve the adaptability and robustness of the joint policy.

[0087] The incremental training of the adapter policy is achieved through the following mechanism:

[0088] 1. Due to the generality of this framework, the adapter policy can increase the number of sub-policies i; update the policy parameters using the feedback of different environments (such as rewards or state changes)

[0089] 2. Dynamically adjust the fusion ratio of the adapter policy and the distilled policy through the weight allocation mechanism to ensure the effectiveness of both in the joint policy. This dual mechanism enables the RAPID framework to always maintain stable decision-making capabilities in complex scenarios.

[0090] Stage 3.3: System output and adjustment.

[0091] In the final output stage, the action generated by the joint policy Π MoP ​It will be directly used as the control command of the vehicle. The system continuously adjusts the policy parameters according to the environmental feedback to ensure that the policy remains efficient, robust, and adaptable throughout the driving process. Through the tight combination of the above three stages, the RAPID framework can effectively integrate the knowledge of large language models into the reinforcement learning policy to generate an adaptable and robust autonomous driving policy.

[0092] In summary, through the application of a deep reinforcement learning policy with robust regularization processing, based on the offline dataset, by minimizing the error between the predicted Q-value and the actual Q-value, the present invention effectively enhances the decision-making accuracy of the policy, and introduces an adversarial policy to improve the robustness of the model, ensuring that it can maintain stable driving decisions in the face of uncertainties or potential attacks, achieving the effect of improving the reliability of the model in the real driving environment.

[0093] The above embodiments are provided only for the purpose of describing the present invention, and are not intended to limit the scope of the present invention. The scope of the present invention is defined by the appended claims. All equivalent substitutions and modifications made without departing from the spirit and principles of the present invention shall be covered within the scope of the present invention.

Claims

1. A model distillation method for an autonomous driving system based on a large language model, characterized in that: The method comprises: Construct offline dataset D LLM , the offline dataset D LLM Each set of offline data in includes: the current state s at each decision, the selected action a * , the feedback r obtained and the updated state s ′ ; Based on the offline data set, generating a distillation strategy for robust regularization processing; The distillation strategy is fixed, and the joint strategy is fine-tuned through interaction with the online environment to generate a trained joint strategy; wherein the joint strategy includes: an adapter strategy and the distillation strategy.

2. The method according to claim 1, characterized in that The constructing of the offline data set includes: Conduct a closed driving experiment in the HighwayEnv environment and obtain the current state s and historical information H; Generate prefix hints for the GPT-3.5 model based on the current state s and historical information H; The GPT-3.5 model makes decision inferences based on the prefix prompts and generates an action set A = {a1, a2, ..., a n }, and for each action a i Assign probability distribution p(a i |s,H); Based on the probability distribution p(a i |s,H), get the selected action a * ; Execute the selected action a * , get feedback r and updated state s ′ .

3. The method according to claim 2, characterized in that The selected action a * Includes: changing to the left lane, changing to the right lane, accelerating, decelerating and idling.

4. The method according to claim 1, characterized in that Based on the offline dataset D LLM , generating a distillation strategy for robust regularization, including: Construct the objective function J after robust regularization robust (Q,π distill ,D LLM );Wherein, the objective function J robust (Q,π distill ,D LLM ) represents the current Q-value function Q and the distillation strategy π distill In the offline dataset D LLM Performance on In the offline dataset D LLM Based on the objective function J after the robust regularization robust (Q,π distill ,D LLM ) to train a large language model to obtain a distillation strategy with robust regularization.

5. The method according to claim 4, characterized in that The objective function Among them, J(Q,π distill ,D LLM ) represents the objective function to be minimized in the Q value function, α represents the first adjustment parameter, β represents the second adjustment parameter, represents the Q value expectation of action selection under the adversarial strategy, a~μ(·|s) represents the probability distribution of selecting action a under the adversarial strategy, and the standard dataset D is the offline dataset D LLM , represents the Q value expectation under the standard data set D, Q(s,a) represents the Q value function, onehot(a) represents the one-hot encoding of action a, and σ(Q(s,a)) represents the probability distribution of Q value after the softmax function.

6. The method according to claim 1, characterized in that The process of generating the joint strategy includes: Select the most important K sub-strategies π from N sub-strategies i , the K sub-strategies π i Includes: adapter strategy and distillation strategy; Based on the observation value of the current state s, we get the sub-strategy π i The weight ω i ; Combine the weights ω of the sub-strategies i , generate a joint strategy Among them, θ d is the parameter of the action decoder, θ r is the parameter of the router, θ p are the parameters of the sub-strategy, is the parameter of the ith sub-strategy, Represents the parameter θ d The action decoder, Represents the parameter θ r router.

7. The method according to any one of claims 1 to 6, characterized in that: After generating the trained joint strategy, the method further includes: Based on the trained joint strategy, generate actions in the current state The action As the vehicle's control command.

8. A model distillation system based on a large language model autonomous driving system, characterized in that: The system comprises: Dataset construction module, used to construct offline dataset D LLM , the offline dataset D LLM Each set of offline data in includes: the current state s at each decision, the selected action a * , the feedback r obtained and the updated state s ′ ; A distillation strategy generation module, used to generate a distillation strategy for robust regularization processing based on the offline data set; The joint strategy training module is used to fix the distillation strategy and fine-tune the joint strategy through interaction with the online environment to generate a trained joint strategy; wherein the joint strategy includes: an adapter strategy and the distillation strategy.

9. An electronic device, characterized in that: The electronic device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the model distillation method based on the large language model autonomous driving system as described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the model distillation method based on a large language model autonomous driving system as described in any one of claims 1 to 7.