Interpretable Reinforcement Learning Decision System and Method Based on Large Language Model Enhancement
By employing a large language model-based interpretable reinforcement learning decision-making system with soft decision trees and a natural language interpretation module, the system addresses the challenges of strategy interpretation and performance optimization in UAV adversarial game scenarios. This enables transparent and real-time UAV decision-making, enhancing operator trust and decision-making efficiency.
Patent Information
- Application Number
- CN202511210997.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-08-27
AI Technical Summary
In existing drone adversarial game scenarios, the black-box nature of deep reinforcement learning models makes it difficult to interpret strategies, understand intentions, and trust decisions. Furthermore, it is difficult to optimize model performance in a targeted manner, which limits its application in civilian drone adversarial game scenarios.
An interpretable reinforcement learning decision system based on a large language model is adopted. The system uses a white-box policy module to construct upper and lower layer models with soft decision trees. Combined with a natural language interpretation module and a policy optimization module, the decision-making process is made transparent and real-time. The large language model is used to process policy parameters and trajectory data, outputting interpretation content that conforms to domain cognition, and generating optimization suggestions based on interaction data.
It improves the intelligence, explainability, and real-time performance of UAV countermeasures decision-making, enhances operator trust, solves the problems of strategy difficulty in interpretation and performance optimization, and achieves efficient task decomposition and instant response.
Smart Images

Figure CN120722758B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of unmanned aerial vehicle (UAV) technology, and in particular to an interpretable reinforcement learning decision system and method based on large language model enhancement. Background Technology
[0002] With the rapid development of drone technology, drone applications are becoming increasingly widespread. Scenarios involving adversarial game theory (such as route disputes in drone logistics delivery and resource allocation conflicts in emergency rescue) place extremely high demands on the intelligence and real-time performance of decision-making. Deep reinforcement learning, as a crucial technology for achieving intelligent decision-making in drones, demonstrates significant advantages in this field due to its ability to autonomously learn optimal strategies in complex and dynamic environments. Currently, existing methods, primarily through hierarchical reinforcement learning, have achieved certain results in drone adversarial game scenarios and have become a common approach for solving such problems.
[0003] However, current drone adversarial decision-making methods based on deep reinforcement learning suffer from problems such as difficulty in interpreting strategies, understanding intentions, and trusting decisions due to their black-box model characteristics. Furthermore, the black-box nature makes it difficult to specifically optimize model performance, hindering their practical application in civilian drone adversarial game scenarios.
[0004] Existing interpretable reinforcement learning methods are mainly divided into two types: self-interpretable and post-interpretable. Both methods share common problems in drone adversarial game scenarios: difficulty in combining domain knowledge to provide explanations consistent with operator experience, difficulty in gaining trust, and inability to use the explanation results to specifically improve model performance. The rapid development of LLM methods has provided a solution to these problems, but due to the highly dynamic nature of drone adversarial games, the output delay characteristics of LLM are difficult to directly use for drone control strategies, limiting its application. Summary of the Invention
[0005] Therefore, it is necessary to provide an interpretable reinforcement learning decision system and method based on large language model enhancement to address the above-mentioned technical problems.
[0006] An interpretable reinforcement learning decision system based on large language model enhancement, the system comprising:
[0007] The system comprises a white-box strategy module, a natural language interpretation module, a strategy optimization module, and an adversarial environment module; the adversarial environment module is used to provide adversarial situation data and interaction data.
[0008] The white-box strategy module includes an upper-level strategy model and a lower-level strategy model. Both the upper-level strategy model and the lower-level strategy model are constructed using soft decision trees. The upper-level strategy model is used to make decisions based on the adversarial situation data and a preset reward function, and outputs macro-level decision data. The macro-level decision data is used as the upper-level sub-target input to the lower-level strategy model. The lower-level strategy model is used to make decisions based on the upper-level sub-target and the adversarial situation data, output the UAV control quantity, and generate UAV trajectory data.
[0009] The natural language interpretation module uses a pre-set decision behavior interpretation big model to process the soft decision tree parameters, computation process data, the preset reward function and UAV trajectory data of the upper-level strategy model and the lower-level strategy model, and outputs behavioral interpretation content of single-step decision logic, overall trajectory intent and strategy behavior pattern.
[0010] The strategy optimization module uses a pre-set decision behavior optimization model to analyze and process the behavior explanation content and UAV trajectory data. Based on the interaction data, it provides the white-box strategy module with suggestions for modifying the reward function and a failure trajectory repair scheme.
[0011] An interpretable reinforcement learning decision-making method based on large language model enhancement, the method comprising:
[0012] An upper-level strategy model and a lower-level strategy model are constructed. Both the upper-level strategy model and the lower-level strategy model are constructed using soft decision trees. The upper-level strategy model is used to make decisions based on adversarial situation data and a preset reward function, and outputs macro-decision data. The macro-decision data is used as the input of the upper-level sub-targets into the lower-level strategy model. The lower-level strategy model is used to make decisions based on the upper-level sub-targets and the adversarial situation data, outputs UAV control variables, and generates UAV trajectory data.
[0013] The pre-set decision behavior interpretation big model is used to process the soft decision tree parameters, computation process data, preset reward function and UAV movement trajectory data of the upper-level strategy model and lower-level strategy model, and output the behavior interpretation content of single-step decision logic, overall trajectory intention and strategy behavior pattern.
[0014] The pre-set decision-making behavior optimization model is used to analyze and process the behavior interpretation content and UAV trajectory data. Based on the interaction data, suggestions for modifying the reward function and failure trajectory repair schemes are provided to the white-box strategy module.
[0015] The aforementioned interpretable reinforcement learning decision-making system and method based on a large language model utilizes a soft decision tree to construct upper and lower layer models through a white-box policy module, achieving transparency in the decision-making process. The upper layer outputs macro-level decision data as sub-objectives, while the lower layer directly outputs control variables. This hierarchical decision-making mechanism enables efficient task decomposition, reducing the computational time consumption of a single model. The natural language interpretation module uses the large language model to process policy parameters and trajectory data, outputting interpretations consistent with domain cognition. This domain knowledge enhances the understandability of the decision intent and increases operator trust. The policy optimization module generates optimization suggestions based on the interpretation content and interaction data, enabling targeted improvements to the policy and addressing the difficulty of using interpretation to improve performance. The white-box policy module makes online dynamic decisions, avoiding the direct impact of large language model output latency on real-time UAV control. Simultaneously, the rapid inference characteristics of the soft decision tree ensure the immediate response of upper and lower layer models to highly dynamic adversarial situations. The asynchronous collaboration mode with the large language model significantly mitigates the impact of output latency on highly dynamic scenarios while retaining interpretability and optimization capabilities, achieving a balance between real-time performance and interpretability in intelligent decision-making. The embodiments of the present invention can improve the intelligence, interpretability, and real-time performance of UAV adversarial decision-making. Attached Figure Description
[0016] Figure 1 This is a block diagram of an interpretable reinforcement learning decision system based on a large language model enhancement in one embodiment.
[0017] Figure 2 This is a schematic diagram of a suboptimal heuristic layered interpretable reinforcement learning process in one embodiment;
[0018] Figure 3 This is a schematic diagram of a soft decision tree structure in one embodiment;
[0019] Figure 4 This is a flowchart illustrating the natural language interpretation of the decision-making process in one embodiment;
[0020] Figure 5 This is a flowchart illustrating the closed-loop enhancement of the interpretable reinforcement learning process in one embodiment. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0022] In one embodiment, such as Figure 1 As shown, an interpretable reinforcement learning decision system based on large language model enhancement is provided, including:
[0023] The system comprises a white-box strategy module, a natural language interpretation module, a strategy optimization module, and an adversarial environment module; the adversarial environment module is used to provide adversarial situation data and interaction data.
[0024] The white-box strategy module includes an upper-level strategy model and a lower-level strategy model. Both the upper-level strategy model and the lower-level strategy model are constructed using soft decision trees. The upper-level strategy model is used to make decisions based on the adversarial situation data and a preset reward function, and outputs macro-level decision data. The macro-level decision data is used as the upper-level sub-target input to the lower-level strategy model. The lower-level strategy model is used to make decisions based on the upper-level sub-target and the adversarial situation data, output the UAV control quantity, and generate UAV trajectory data.
[0025] The natural language interpretation module uses a pre-set decision behavior interpretation big model to process the soft decision tree parameters, computation process data, the preset reward function and UAV trajectory data of the upper-level strategy model and the lower-level strategy model, and outputs behavioral interpretation content of single-step decision logic, overall trajectory intent and strategy behavior pattern.
[0026] The strategy optimization module uses a pre-set decision behavior optimization model to analyze and process the behavior explanation content and UAV trajectory data. Based on the interaction data, it provides the white-box strategy module with suggestions for modifying the reward function and a failure trajectory repair scheme.
[0027] exist Figure 1 The system comprises three modules: Reinforcement Learning-Based Decision Control (ControlxHXRL), Large Language Model (LLM)-Based Decision Explanation (ExplanationxLLM), and LLM-Based Policy Optimization (OptimizationxLLM). First, an interpretable reinforcement learning method based on hierarchical soft decision trees is constructed (ControlxHXRL module). White-box interpretable policies are used to control UAV actions, and the output model parameters and action data are stored in the dataset. Based on this, the large decision behavior explanation model is used to process the soft decision tree parameters and computational processes output by the ControlxHXRL module, enhancing the natural language interpretation of the decisions (ExplanationxLLM module). Simultaneously, the large decision behavior optimization model, based on the explanations and action data from the ExlanationxLLM module, performs policy analysis and improvement derivation to achieve decision optimization (OptimizationxLLM module). The optimization suggestions are then fed back to the ControlxHXRL module. Combined with a simulation environment and a trustworthy and understandable human-computer interaction method, this supports UAV adversarial decision-making applications.
[0028] In the aforementioned interpretable reinforcement learning decision-making method based on a large language model, a white-box policy module uses a soft decision tree to construct upper and lower layer models, achieving transparency in the decision-making process. The upper layer outputs macro-level decision data as sub-objectives, while the lower layer directly outputs control variables. This hierarchical decision-making mechanism enables efficient task decomposition, reducing the computational time of a single model. The natural language interpretation module utilizes the large language model to process policy parameters and trajectory data, outputting interpretations consistent with domain cognition. This domain knowledge enhances the understandability of the decision intent and increases operator trust. The policy optimization module generates optimization suggestions based on the interpretation content and interaction data, enabling targeted improvements to the policy and addressing the difficulty of improving performance through interpretation. The white-box policy module makes online dynamic decisions, avoiding the direct impact of large language model output latency on real-time UAV control. Simultaneously, the rapid inference characteristics of the soft decision tree ensure the immediate response of upper and lower layer models to highly dynamic adversarial situations. The asynchronous collaboration mode with the large language model significantly mitigates the impact of output latency on highly dynamic scenarios while retaining interpretability and optimization capabilities, achieving a balance between real-time performance and interpretability in intelligent decision-making.
[0029] In one embodiment, the upper-layer policy model and the lower-layer policy model have the same network structure, including a policy soft decision tree, a target policy soft decision tree, a value network, and a target value network. The policy soft decision tree and the target policy soft decision tree are soft decision trees with binary probabilistic decision boundaries, and the value network and the target value network are multi-layer neural networks with the same structure.
[0030] This invention proposes a suboptimal-inspired hierarchical interpretable reinforcement learning approach, a white-box interpretable reinforcement learning method for constructing adversarial strategies for unmanned aerial vehicles (UAVs). Figure 2 As shown, this method first employs a hierarchical reinforcement learning strategy. The upper-layer strategy determines the target speed, altitude, and orientation of the aircraft based on situational data from both sides. The lower-layer strategy, based on the sub-targets and situational data determined by the upper-layer strategy, determines control variables such as elevators, ailerons, rudder, and throttle, interacting within the UAV adversarial environment. Both the upper and lower layers use a white-box self-explanatory method, soft decision trees, to construct the policy model. The training process is adaptively modified based on the standard TD3 (TwinDelayedDeepDeterministicPolicyGradient) algorithm and heuristically trained using behavior sets generated from simple UAV adversarial rules. Simultaneously, a curriculum learning method is used, setting the adversary's ability from weak to strong to improve training effectiveness.
[0031] The training network consists of a policy soft decision tree, a target policy soft decision tree, a value network, and a target value network. The value network and the target value network are fitted using the same multi-layer neural network structure. The policy soft decision tree and the target soft decision tree use the same SDT (Soft Decision Tree) structure with binary probabilistic decision boundaries, which exhibits excellent transparency and performance.
[0032] In one embodiment, the soft decision tree includes a root node, intermediate nodes, and leaf nodes. Each leaf node parameter includes an output action reference value. The root node and intermediate nodes contain weight parameters for the input observation values, used to calculate the traversal probability of the soft decision tree decision branch. The input observation values are adversarial situation data. The output action reference values are macroscopic decision parameters or control parameters of the UAV. The macroscopic decision data includes target speed, target altitude, and target orientation. The UAV control parameters include the deflection of the elevator, ailerons, and rudder, as well as the throttle opening.
[0033] Unlike traditional hard decision boundary growth learning models, SDT focuses on all input features and calculates the probability of traversing to the left child node through each internal node, then probabilistically selects the path. In this invention, the parameters of each leaf node of the SDT are determined as output action reference values, which are learnable parameters during reinforcement learning training; the root node and intermediate nodes contain weights for the input observations, used to calculate the traversal probability of the SDT decision branch, and are also set as learnable parameters during reinforcement learning training. Furthermore, to improve SDT performance, this invention employs a decision forest approach, obtaining a large number of SDTs with the same structure through training. During training, each SDT votes on the results to produce the final result. The structure of each SDT is as follows: Figure 3 As shown. In Figure 3 In this example, a two-layer SDT is used. More layers of SDT involve adding intermediate nodes, while the root and leaf node structures remain unchanged. For each non-leaf node j, the probability of a left branch traversing that node can be calculated as follows:
[0034] ;
[0035] in It is the Sigmoid function. It is the first The weight corresponding to this node in the subtree. For a specific leaf node, its selection probability is obtained by multiplying the probabilities of all paths leading to this node. For example... Figure 3 The leftmost leaf node has a selection probability that is the product of the probabilities of the two paths:
[0036] ;
[0037] In this scenario, action values are continuous, and each leaf code represents a relative action parameter. This value is for the leaf node. Learnable parameters at each leaf node. After SDT, the relative action parameters of each leaf node are weighted and summed:
[0038] ;
[0039] Because of this step Reference values for each sub-action included At this point, the output result also needs to be mapped to the action value range to obtain the final output action value:
[0040] ;
[0041] in For action Half of the range of values, For action The average value.
[0042] In one embodiment, the training process of the upper-layer policy model and the lower-layer policy model includes: initializing the training replay dataset with behavioral data generated by a predefined policy; the training replay dataset includes a high-reward replay dataset and a low-reward replay dataset, and the application frequency of different datasets is controlled by selecting probabilities; during training, the upper-layer policy model optimizes node weights and path probability parameters based on samples sampled from the replay dataset using gradient descent; the lower-layer policy model uses the same optimization algorithm as the upper-layer policy model and constructs the target weight matrix of the obtained upper-layer sub-targets; the target weight matrix includes the weights of completed sub-targets set to the minimum value, the weights of the currently learned sub-targets that can be dynamically adjusted, and the weights of unlearned sub-targets set to 0; a sub-target counter is set to count the learning progress of the sub-targets, and when the agent completes a valid sub-target learning, the counter is automatically incremented. When the cumulative number of times the counter reaches the maximum number of times the current sub-target is set, it is determined that the current sub-target learning is completed and the next sub-target learning begins. If the minimum threshold of the sub-target is not reached for several consecutive times, the previous learning task is returned.
[0043] Specifically, to address the performance bottleneck of the white-box model, the upper-layer strategy employs a dual-delay deep deterministic policy gradient algorithm based on suboptimal heuristics and soft decision trees. First, to avoid ineffective exploration, a predefined UAV adversarial strategy is used for heuristic guidance, initializing the training replay dataset with the behavioral data generated by this optimal heuristic. Second, this method divides the training replay dataset into high-reward and low-reward replay datasets, controlling the application frequency of different replay datasets by selecting probabilities, thereby improving training efficiency while encouraging exploration.
[0044] The lower-level strategy adopts a soft update learning mechanism for multiple sub-objectives of the upper-level output, based on the double-delay deep deterministic strategy gradient algorithm of suboptimal heuristic and soft decision tree.
[0045] To ensure the agent retains its memory of the first j-1 objectives while completing the j-th sub-objective, the weights of the first j-1 sub-objectives are set to the minimum value of their weight coefficients. As the agent learns the j-th sub-objective and gradually achieves the objective, it should then be guided to learn the (j+1)-th objective. The (j+2)-Nth sub-objectives are currently not required for learning, so the weights after the j-th sub-objective can be set to 0. Therefore, the objective weight matrix is set as follows:
[0046] ;
[0047] In the formula: and The weights are currently being dynamically adjusted. - The minimum reward weights corresponding to the first k sub-goals;
[0048] ;
[0049] ;
[0050] In the formula: For the current sub-target's counter, This represents the maximum number of counts for the current sub-objective. After the agent completes one effective target learning iteration, Automatically increment by one. When At that point, the task agent has completed learning target j and is moving on to learning target j+1.
[0051] Because during the learning process of sub-objectives, there are instances where the expected goal cannot be achieved, or where training deteriorates as the difficulty of the sub-objectives increases, this algorithm also incorporates a goal backoff mechanism, namely:
[0052] ;
[0053] In the formula: The total reward for the current trajectory sub-target j. This represents the lowest threshold for the current sub-target. When consecutive When none of the learning trajectories have reached the minimum threshold The learning task is automatically decremented by one, and the agent reverts to the previous learning task.
[0054] In one embodiment, the white-box strategy module employs a course learning mechanism during training to improve the opponent's intelligence level in stages.
[0055] In this embodiment, line-of-sight (LAS) drone combat possesses complex adversarial game characteristics. If the non-friendly drones are too strong initially, successful experience training is difficult and the process is slow. Later, without strong non-friendly drone strategies, it is difficult to train effective adversarial strategies for the friendly drones. This invention designs a learning mechanism that initially employs a stationary strategy for non-friendly drones, then a rule-inspired strategy, and finally a self-game framework. During training, the adversarial intelligence level of non-friendly drones is continuously improved to enhance the adversarial capabilities of the friendly drones. Considering that the adversarial game in LAS mainly manifests in the upper-level drone adversarial decision-making and has less to do with the lower-level maneuver control, this invention fixes the lower-level maneuver control algorithm during training while continuously improving the upper-level drone adversarial decision-making capabilities.
[0056] The heuristic rules used by non-friendly drones in the mid-term are the same as those used for the upper and lower layer drone adversarial decision-making. The final self-game training method consists of two parts: a training phase and an evaluation phase. In the training phase, non-friendly drones select the latest generation of strategies from the drone adversarial strategy library to play against friendly drones. Friendly drones are trained using the upper-layer drone adversarial method based on the ControlxHXRL algorithm. The pseudocode of the ControlxHXRL algorithm is shown in Table 1. Initially, friendly drones initialize using the network parameters of the previous generation. When the probability of friendly drones losing to non-friendly drones is less than 30%, the evaluation phase begins. In the evaluation phase, non-friendly drones sequentially select all generations of strategies from the strategy library to play against friendly drones. If friendly drones can maintain a win rate of over 50% (not considering draws), their strategies are added to the drone adversarial strategy library as the latest generation. If friendly drones fail to achieve a win rate of over 50% when playing against the i-th generation strategy, the opponent is set as the i-th generation strategy, and the training phase resumes.
[0057] Table 1. Pseudocode of the ControlxHXRL algorithm
[0058]
[0059] In one embodiment, the soft decision tree parameters, computational process data, preset reward function, and UAV trajectory data of the upper-level strategy model and lower-level strategy model are processed to output behavioral explanations of single-step decision logic, overall trajectory intent, and strategy behavior patterns. This includes: understanding and simulating the decision-making process of the soft decision tree based on the soft decision tree parameters and computational process data generated by prompts and enhanced retrieval; reproducing the soft decision tree decision process based on UAV adversarial state data and the internal decision data of the white-box strategy module, explaining the single-step decision logic, and obtaining behavioral explanations of the single-step decision logic; analyzing the consistency between the overall trajectory direction and the upper-level sub-target based on the changing trends of UAV trajectory data and the computational process data of the soft decision tree in multi-step decision-making, explaining the overall trajectory intent, and obtaining behavioral explanations of the overall trajectory intent; and explaining the strategy behavior patterns based on the common characteristics of multiple sets of UAV trajectory data, soft decision tree parameters, and the weight allocation logic of the preset reward function, and obtaining behavioral explanations of the strategy behavior patterns.
[0060] like Figure 4 As shown, the goal of the ExplanationxLLM module is to teach the LLM to generate natural language explanations of the decision-making process of the ControlxHXRL module. More specifically, LLM-Exp (a large language model for explanation, which can use deepseek) first understands the soft decision tree decision-making process based on the soft decision tree parameters and internal calculation processes given by Prompts and RAG, and has the ability to simulate the soft decision tree decision-making process. On this basis, the LLM, based on the given UAV adversarial state data and the internal decision data of the ControlxHXRL module, explains the tactical significance and rationality of the single-step decision logic analysis by reproducing the soft decision-making process. Then, based on this, it continues to analyze the overall trajectory actions and intentions, and interpret the strategic behavior patterns of the ControlxHXRL module. Table 2 shows the pseudocode for the ExplanationxLLM algorithm. This invention employs a two-stage R1-zero training procedure. In the first stage, domain knowledge is distilled into a pre-trained LLM-Opt (a large language model for optimization, which can use Qwen3 1.7B and 8B models) based on UAV adversarial public knowledge, training environment knowledge, and white-box policy knowledge using the SFT (Supervised Fine-Tuning) method. In the second stage, the model is further optimized using DPO (Direct Preference Optimization) / PPO (Proximal Policy Optimization). This process is guided by a static decision dataset, and a reward function is designed to enhance decision accuracy and output structure.
[0061] Table 2. Pseudocode of the ExplanationxLLM algorithm
[0062]
[0063] This invention constructs explanation datasets for three types of explanations: single-step decision logic explanation, overall trajectory explanation, and explanation of the policy behavior pattern of the ControlxHXRL module. Each instance is represented as a binary classification task to determine whether the decision is reasonable, and the explanation is provided using natural language. Manual annotation is provided by experts in the UAV adversarial domain, combining trajectory rewards and the output of a large evaluation model (averaging the explanation results using another large model) to provide comprehensive annotations.
[0064] In one embodiment, the decision behavior explanation model of the natural language interpretation module adopts a two-stage training method. In the first stage, the knowledge of the UAV adversarial domain and the soft decision tree knowledge are distilled into the pre-trained model through supervised fine-tuning. In the second stage, reinforcement learning is used to optimize the model. The total reward function used in the optimization process includes accuracy reward and format reward.
[0065] This invention employs two types of rewards, defined as accuracy and format rewards, therefore the total reward can be described as: ;
[0066] in It is based on a comprehensive score given by domain experts on the accuracy of the output. It is a format reward designed to encourage models to structure their answers, dividing them into reasoning, explanation, and response parts (e.g., using...). <reasoning> and< / reasoning> Labels are used to mark the beginning and end of the reasoning section. While the correctness reward incentivizes accurate decision-making, the format reward serves a dual purpose: it encourages the model to articulate its reasoning and ensures that the output remains structured and easy to parse.
[0067] In one embodiment, the decision behavior optimization model of the strategy optimization module adopts a composite reward function, which includes UAV adversarial performance reward, format reward, and parameter extraction reward.
[0068] Compared to the OptimizationxLLM module, aircraft control adaptation in UAV adversarial scenarios cannot rely solely on static datasets. As with many UAV adversarial applications, effective behavior is generated through interaction with the environment, rather than solely through parameters learned via SFT. Analogously, this distinction reflects the difference between learning to operate an aircraft by reading a manual (SFT) and learning to fly through actual flight lessons (PPO / DPO). In this setting, the OptimizationxLLM module places LLM within a closed loop with the interpretable reinforcement learning (XRL) proposed in this invention, enabling it to influence control behavior based on interaction, thereby promoting improvements through experience-driven, suboptimal-inspired hierarchical interpretable reinforcement learning methods.
[0069] like Figure 5 As shown, LLM provides closed-loop feedback enhancement for UAV adversarial simulation. During training, the simulation environment provides feedback on the UAV's action reward value and win rate. In the process, LLM-Exp first acts as an expert, participating in the manual correction and enhancement of a selected dataset based on the explanation content and environmental feedback. In response, LLM-Opt generates suggestions for modifying the reward function relied upon by XRL training, and solutions for correcting failed trajectories, based on the explanation and environmental feedback generated by ExplanationxLLM. During training, these XRL reward functions and failure correction solutions are used to improve XRL training and are tested in closed-loop simulation, thereby calculating the OptimizationxLLM behavioral adaptation reward. Furthermore, to enhance and emphasize the generalization ability of GRPO-trained LLM.
[0070] Control Adaptive Reward Modeling: To train the OptimizationxLLM module, this invention uses three different rewards: UAV adversarial performance reward, format reward, and parameter extraction reward, which together constitute the total reward. As shown below:
[0071] ;
[0072] The goal is to reward the LLM's improvement suggestions for their actual effectiveness in a UAV adversarial simulation environment. In each training step, these suggestions are applied to the improved XRL training, and after training, the algorithm is used adversarially against a baseline XRL algorithm, with a reward value given based on the reward value and win rate during training. This is achieved using... and This indicates the effectiveness of the baseline XRL algorithm against drones compared to the improved XRL algorithm proposed by LLM. It is a formatted reward that encourages models to structure their responses and reason about their answers. Finally, This is a parameter extraction reward used to prevent LLM from producing hallucinatory or invalid suggestion parameters. This composite reward guides LLM to generate interpretable, effective, and behavior-aligned control adaptation parameters through PPO / DPO. Table 3 shows the pseudocode of the OptimizationxLLM algorithm.
[0073] Table 3. Pseudocode of the OptimizationxLLM algorithm
[0074]
[0075] In one embodiment, an interpretable reinforcement learning decision-making method based on large language model enhancement is provided, comprising the following steps:
[0076] S1. Construct an upper-level strategy model and a lower-level strategy model. Both the upper-level strategy model and the lower-level strategy model are constructed using soft decision trees. The upper-level strategy model is used to make decisions based on adversarial situation data and a preset reward function, and outputs macro-decision data. The macro-decision data is used as the input of the upper-level sub-targets into the lower-level strategy model. The lower-level strategy model is used to make decisions based on the upper-level sub-targets and the adversarial situation data, outputs UAV control variables, and generates UAV trajectory data.
[0077] S2. Using a pre-set decision behavior interpretation model, the soft decision tree parameters, computation process data, preset reward function, and UAV trajectory data of the upper-level strategy model and lower-level strategy model are processed to output the behavior interpretation content of single-step decision logic, overall trajectory intent, and strategy behavior pattern.
[0078] S3. Analyze and process the behavior explanation content and UAV trajectory data using a pre-set decision behavior optimization model, and provide suggestions for modifying the reward function and failure trajectory repair schemes for the white-box strategy module based on the interaction data.
[0079] Specific limitations regarding the interpretable reinforcement learning decision-making method based on large language model reinforcement can be found in the limitations of the interpretable reinforcement learning decision-making system based on large language model reinforcement mentioned above, and will not be repeated here. Each module in the aforementioned interpretable reinforcement learning decision-making system based on large language model reinforcement can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0080] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0081] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. An interpretable reinforcement learning decision-making system based on large language model enhancement, characterized in that, The interpretable reinforcement learning decision system includes a white-box policy module, a natural language interpretation module, a policy optimization module, and an adversarial environment module; the adversarial environment module is used to provide adversarial situation data and interaction data. The white-box strategy module includes an upper-level strategy model and a lower-level strategy model. Both the upper-level strategy model and the lower-level strategy model are constructed using soft decision trees. The upper-level strategy model is used to make decisions based on the adversarial situation data and a preset reward function, and outputs macro-level decision data. The macro-level decision data is used as the upper-level sub-target input to the lower-level strategy model. The lower-level strategy model is used to make decisions based on the upper-level sub-target and the adversarial situation data, output the UAV control quantity, and generate UAV trajectory data. The natural language interpretation module uses a pre-set decision behavior interpretation big language model to process the soft decision tree parameters, computation process data, the preset reward function and UAV trajectory data of the upper-level strategy model and lower-level strategy model, and outputs behavioral interpretation content of single-step decision logic, overall trajectory intent and strategy behavior pattern. The strategy optimization module uses a pre-set decision behavior optimization big language model to analyze and process the behavior explanation content and UAV movement trajectory data. Based on the interaction data, it provides the white-box strategy module with preset reward function modification suggestions and failure trajectory repair schemes. The training process for the upper-layer policy model and the lower-layer policy model includes: The training replay dataset is initialized with behavioral data generated using a predefined strategy; the training replay dataset includes a high-reward replay dataset and a low-reward replay dataset, and the application frequency of the high-reward replay dataset and the low-reward replay dataset is controlled by selecting the probability. During training, the upper-layer policy model optimizes node weights and path probability parameters through gradient descent based on samples taken from the training replay dataset. The lower-level strategy model uses the same optimization algorithm as the upper-level strategy model and constructs the target weight matrix of the obtained upper-level sub-objectives; the target weight matrix includes the weights of completed sub-objectives set to the minimum value, the weights of currently learned sub-objectives that can be dynamically adjusted, and the weights of unlearned sub-objectives set to 0; A sub-target counter is set to track the learning progress of sub-targets. When the agent completes a valid sub-target learning, the counter automatically increments. When the cumulative number of times the counter reaches the maximum number of times set for the current sub-target, it is determined that the current sub-target learning is completed and the agent moves on to the next sub-target learning. If the minimum threshold of the sub-target is not reached for several consecutive times, the agent is returned to the previous learning task.
2. The interpretable reinforcement learning decision system according to claim 1, characterized in that, The upper-layer policy model and the lower-layer policy model have the same network structure, including a policy soft decision tree, a target policy soft decision tree, a value network, and a target value network. The policy soft decision tree and the target policy soft decision tree are soft decision trees with binary probabilistic decision boundaries, and the value network and the target value network are multi-layer neural networks with the same structure.
3. The interpretable reinforcement learning decision system according to claim 1 or 2, characterized in that, The soft decision tree includes a root node, intermediate nodes, and leaf nodes. Each leaf node parameter includes an output action reference value. The root node and intermediate nodes contain weight parameters for the input observation values, which are used to calculate the probability of passing through the decision branch of the soft decision tree. The input observation values are adversarial situation data. The output action reference values are the macroscopic decision parameters or control parameters of the UAV.
4. The interpretable reinforcement learning decision system according to claim 1, characterized in that, The macro-decision data includes target speed, target altitude, and target orientation, while the UAV control variables include the deflection of the elevator, ailerons, and rudder, as well as the throttle opening.
5. The interpretable reinforcement learning decision system according to claim 1, characterized in that, The white-box strategy module employs a course learning mechanism during training to improve the opponent's intelligence level in stages.
6. The interpretable reinforcement learning decision system according to claim 1, characterized in that, The soft decision tree parameters, computational process data, preset reward function, and UAV trajectory data of the upper-level and lower-level policy models are processed to output behavioral explanations of single-step decision logic, overall trajectory intent, and policy behavior patterns, including: Based on the prompts and search enhancements, the soft decision tree parameters and computational data provided are used to understand and simulate the decision-making process of the soft decision tree. Based on the UAV combat status data and the internal decision data of the white-box strategy module, the soft decision tree decision process is reproduced, the single-step decision logic is explained, and the behavioral explanation of the single-step decision logic is obtained. Based on the trend of changes in UAV trajectory data and the computational process data of soft decision trees in multi-step decision-making, the consistency between the overall trajectory direction and the upper-level sub-targets is analyzed, the overall trajectory intent is explained, and the behavioral explanation of the overall trajectory intent is obtained. Based on the common characteristics of multiple sets of UAV trajectory data, soft decision tree parameters, and the weight allocation logic of the preset reward function, the strategy behavior pattern is explained, and the behavioral explanation of the strategy behavior pattern is obtained.
7. The interpretable reinforcement learning decision system according to claim 1, characterized in that, The natural language interpretation module's decision behavior interpretation big language model adopts a two-stage training method. In the first stage, supervised fine-tuning is used to distill UAV adversarial domain knowledge and soft decision tree knowledge into the pre-trained model. In the second stage, reinforcement learning is used to optimize the model. The total reward function used in the optimization process includes accuracy reward and format reward.
8. The interpretable reinforcement learning decision system according to claim 1, characterized in that, The decision behavior optimization model of the strategy optimization module adopts a composite reward function, which includes UAV adversarial performance reward, format reward and parameter extraction reward.
9. An interpretable reinforcement learning decision-making method based on large language model enhancement, characterized in that, The method includes: An upper-level strategy model and a lower-level strategy model are constructed, both of which are built using soft decision trees. The upper-level strategy model is used to make decisions based on adversarial situation data and a preset reward function, and outputs macro-level decision data. The macro-level decision data is used as input to the lower-level strategy model as an upper-level sub-target. The lower-level strategy model is used to make decisions based on the upper-level sub-target and the adversarial situation data, outputs UAV control variables, and generates UAV trajectory data. The pre-set decision behavior interpretation big language model is used to process the soft decision tree parameters, computation process data, preset reward function and UAV movement trajectory data of the upper-level strategy model and lower-level strategy model, and output the behavior interpretation content of single-step decision logic, overall trajectory intention and strategy behavior pattern. The pre-set decision behavior optimization big language model is used to analyze and process the behavior explanation content and UAV movement trajectory data. Based on the interaction data, the white-box strategy module is provided with suggestions for modifying the preset reward function and failure trajectory repair scheme. The training process for the upper-layer policy model and the lower-layer policy model includes: The training replay dataset is initialized with behavioral data generated using a predefined strategy; the training replay dataset includes a high-reward replay dataset and a low-reward replay dataset, and the application frequency of the high-reward replay dataset and the low-reward replay dataset is controlled by selecting the probability. During training, the upper-layer policy model optimizes node weights and path probability parameters through gradient descent based on samples taken from the training replay dataset. The lower-level strategy model uses the same optimization algorithm as the upper-level strategy model and constructs the target weight matrix of the obtained upper-level sub-objectives; the target weight matrix includes the weights of completed sub-objectives set to the minimum value, the weights of currently learned sub-objectives that can be dynamically adjusted, and the weights of unlearned sub-objectives set to 0; A sub-target counter is set to track the learning progress of sub-targets. When the agent completes a valid sub-target learning, the counter automatically increments. When the cumulative number of times the counter reaches the maximum number of times set for the current sub-target, it is determined that the current sub-target learning is completed and the agent moves on to the next sub-target learning. If the minimum threshold of the sub-target is not reached for several consecutive times, the agent is returned to the previous learning task.
Citation Information
Patent Citations
Automatic driving lane selection decision-making method and system based on inverse reinforcement learning
CN116890855A
Unmanned aerial vehicle autonomous decision-making method and device based on Transform neural network state prediction
CN117192998A