Reinforcement learning decision optimization method, system and equipment based on causal big language model
By constructing a causal model using a causal large language model and introducing a causal intervention mechanism, the agent's policy network is optimized, solving the problem of low learning efficiency of the agent in complex environments and achieving efficient adaptation and decision optimization.
Patent Information
- Application Number
- CN202510944141.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-11-07
AI Technical Summary
Existing reinforcement learning agents are inefficient in learning in complex environments, lack adaptability, and are deficient in effective reasoning ability, especially in dynamic environments where they struggle to adapt to changes efficiently.
We employ a reinforcement learning approach based on a causal large language model. By constructing a structural causal model, we optimize the policy network using a causal intervention mechanism, design a multimodal reward function, and combine causal knowledge with the policy update.
It improves the decision-making efficiency and adaptability of intelligent agents in dynamic environments, reduces the dependence on large amounts of interaction data, and enables them to efficiently adapt to new tasks in low-sample conditions.
Smart Images

Figure CN120911539A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and more particularly, to a reinforcement learning decision optimization method, system and device based on a causal large language model. BACKGROUND
[0002] In recent years, reinforcement learning (RL) agents have been increasingly applied in the field of artificial intelligence, especially in robot control, autonomous driving, game AI and complex task planning. Reinforcement learning is a typical sequential decision-making method, in which an agent interacts with the environment, selects appropriate actions based on the current state, and continuously optimizes its strategy through rewards or punishments from the environment feedback to achieve long-term goals. The continuous development of RL technology enables agents to have stronger adaptability in high-dimensional, dynamic and complex environments.
[0003] However, traditional RL methods often require a large amount of interaction data to learn effective strategies when facing complex structured environments, which leads to low learning efficiency, slow convergence speed, and difficulty in adapting to environmental changes.
[0004] To solve these problems, existing technologies focus on improving the decision-making efficiency and adaptability of agents, mainly through improving model learning paradigms, optimizing exploration strategies, and introducing auxiliary learning modules. Model-based RL attempts to learn the dynamic model of the environment, allowing the agent to simulate the environment internally, thereby reducing the need for interaction with the real environment and improving learning efficiency. In addition, the balance between exploration and exploitation has always been a core challenge in reinforcement learning. Recent research in this direction has proposed exploration strategies based on intrinsic motivation, allowing agents to guide their exploration behavior through rewards for novelty or information gain, thereby more efficiently learning optimal strategies. Meanwhile, hierarchical RL constructs high-level decision-making structures, allowing agents to learn complex tasks in a more abstract manner, reducing the search space during the learning process and improving decision-making efficiency. Meta-RL is also widely used in reinforcement learning tasks, allowing agents to quickly adapt to multiple tasks by optimizing initial strategy parameters to quickly adjust their strategies with limited interaction data, thereby improving generalization ability.
[0005] In recent years, large language models (LLMs) have been introduced into the field of reinforcement learning, further enhancing the understanding of the environment and the ability to generate strategies for agents. LLMs have powerful knowledge storage and reasoning capabilities, which can help RL agents analyze environmental states, understand task descriptions, and generate reasonable action recommendations, thereby accelerating policy optimization. However, existing LLMs mainly rely on token prediction-based reasoning, lacking structured understanding and adaptive ability of the environment, resulting in limitations in decision-making of agents in complex tasks. SUMMARY
[0006] To solve the problems of low learning efficiency, insufficient adaptability and lack of effective reasoning ability of reinforcement learning agents in complex environments in the prior art, the present application provides a reinforcement learning decision optimization method, system and device based on a causal large language model, which has high decision efficiency in dynamic environments.
[0007] To achieve the above-mentioned purposes of the present application, the technical solutions adopted are as follows:
[0008] A reinforcement learning decision optimization method based on a causal large language model, comprising the following steps:
[0009] Initializing an agent and its policy network in a dynamic environment;
[0010] Interacting the agent with the environment to obtain trajectory information of historical sequence decisions generated by the interaction;
[0011] Extracting causal variables and initial causal relationships from the trajectory information using a large language model to construct a structural causal model;
[0012] Obtaining an agent policy-driven causal intervention mechanism through the policy network, executing a specified action in the environment, and comparing the variable state changes before and after the intervention, and dynamically correcting the causal relationships in the structural causal model by comparing the conditional probability after the intervention with the original observation probability;
[0013] According to the task-related causal chain extracted from the corrected structural causal model, generating semantic sub-goals corresponding to the causal relationships through the large language model reasoning;
[0014] Designing a multi-modal reward function that integrates semantic similarity;
[0015] Updating the agent's policy network in the dynamic environment using the obtained sub-goals and reward function.
[0016] Preferably, the trajectory information of the historical sequence decisions generated by the interaction is obtained, specifically:
[0017] Enables the agent to interact with the environment, storing the environmental state, the agent's state, and the action sequence τ={s} of the agent's historical decision-making process. t ,φ t ,a t ,r t ,s t+1 ,φ t+1}; where t represents time s t Environmental states include, but are not limited to, entity attributes, spatial topological relationships, and quantitative relationships; φ t Representing the state of an agent, including but not limited to coordinates, energy, and behavioral constraints; a t This represents the action calculated by the agent based on the currently observed state information using a policy function; r t This indicates that the intelligent agent uses action a. t Interacting with the environment, the state transitions to s t+1 Subsequently, there are feedback on the effects of the decision-making actions and rewards for changes in the state of the environment itself.
[0018] Furthermore, a large language model is used to extract causal variables and initial causal relationships from the trajectory information, and a structural causal model is constructed. The specific steps are as follows:
[0019] The environmental state s in the trajectory information τ t s t+1 Agent state φ t φ t+1 and action sequence a t The textual description serves as a prompt template, which is then input into a large language model for causal inference.
[0020] Through several rounds of iterative and multi-sample prompting engineering, the large language model is guided to learn context, outputting a set of candidate causal variables:
[0021] V,E=LLM(s t ,s t+1 ,φ t ,φ t+1 ,a t )
[0022] Where E is a relational variable set LLM representing a large language model, E = {e i→j |v i ,v j ∈V}; the set of causal variables is V={v1,v2,...,v n};
[0023] Construct an initial structural causal model G(V,E); store the structural causal model in a causal matrix, where matrix element A ij ∈{0,1} represents the variable v i to vj The existence of causal relationships.
[0024] Furthermore, a policy-driven causal intervention mechanism for the agent is obtained through a policy network. The agent then executes specified actions in the environment and compares the changes in variable states before and after the intervention. By comparing the conditional probability after intervention with the original observed probability, the causal relationship in the structural causal model is dynamically corrected. The specific steps are as follows:
[0025] In an independent validation environment, the target causal variable v i The agent implements the action a calculated by the policy network. t :
[0026]
[0027] Perform action a t Then record the target variable v j State changes Calculate the conditional probability after intervention With the original observation probability P(v) j |v i If they are equal, then it means v i With v j The causal relationship between them is not valid, but the converse is valid;
[0028] Update edge e in the structural causal model G(V,E) i→j Values in matrix A:
[0029]
[0030] Furthermore, based on the task-related causal chains extracted from the modified structural causal model, semantic sub-targets corresponding to the causal relationships are generated through large language model reasoning. The specific steps are as follows:
[0031] Extracting causal chains related to agent tasks from the structural causal model G(V,E):
[0032] Ρ={v i →v i+1 →···→v j |v i ∈G};
[0033] The causal chain is textualized into a prompt template and input into a large language model for causal sub-objective inference:
[0034] LLM(prompt,P i );
[0035] Where P i v is the i-th causal relationship segment in the causal chain. i →vi+1 The prompt template includes a description of causal relationships and target generation instructions; through multi-round iterative and multi-sample prompting engineering, the large language model is guided to learn the context and output the set of sub-targets for the current task.
[0036]
[0037] Each sub-target g i The corresponding causal relationship v with a length of 1 in the causal chain P i →v i+1 .
[0038] Furthermore, a multimodal reward function integrating semantic similarity is designed. The specific steps are as follows: based on the standard reward function, a multimodal causal knowledge reward is introduced.
[0039] r t =r env +λr causal
[0040] Where r env The reward represents the state change of the environment itself, λ is a configurable fusion weight parameter, and r causal The additional reward, representing the inclusion of causal knowledge, is achieved by encoding the set of sub-goals g into a semantic vector {g}. i}, and the environment state s at the current time t. t Embedded vector {s t}, and calculate the cosine similarity between the two.
[0041] Furthermore, specifically, the aforementioned r causal The cosine similarity between the two images was calculated using the text-image multimodal model Clip.
[0042]
[0043] Furthermore, the obtained sub-objectives and rewards are used to update the agent's policy network in a dynamic environment. The specific steps are as follows:
[0044] The sub-goal-driven policy network is represented as:
[0045] π(a t |s t ,φ t ,g)
[0046] Where φ t The state of the agent at time t is represented by the policy π, which learns to perform the optimal action a under causal constraints by maximizing the long-term cumulative reward. t ;
[0047] The agent performs action a according to policy π.t The interaction state with the environment is transferred to s t+1 After that, the agent receives the reward signal of the environment feedback;
[0048] After obtaining a series of trajectory sequences τ, the advantage function is calculated to measure the value of the current action a t relative to the average strategy:
[0049]
[0050] Where γ is the discount factor, which controls the weight of future rewards, λ is the smoothing factor, which controls the smoothing degree of advantage estimation, T is the trajectory length, δ is the TD error, which measures the prediction error of the current state value function:
[0051] δ t =r t +λV(s t+1 ,φ t+1 ,g t+1 )-V(s t ,φ t ,g t )
[0052] The loss function of the policy network update is obtained:
[0053] L clip (θ)=E θ [min(r t ·A t ,clip(r t ,1-ε,1+ε)·A t )]
[0054] Where clip is the gradient clipping strategy to prevent the gradient update amplitude from being too large, and ε is the parameter to control the policy change range, is the probability ratio of the new and old strategies.
[0055] A reinforcement learning decision optimization system based on a causal large language model, comprising a learning module, an adaptation module, and a decision module;
[0056] The learning module is used to interact the agent with the environment to obtain the trajectory information of the historical sequence decision generated by the interaction:
[0057] The adaptation module is used to extract causal variables and initial causal relationships from the trajectory information using a large language model to construct a structural causal model;
[0058] The decision module is used for generating semantic sub-targets corresponding to the causal relationship through a large language model inference according to the task-related causal chain extracted from the revised structural causal model; a multi-modal reward function is designed by fusing semantic similarity; and the obtained sub-targets and the reward function are used to update the strategy network of the agent in the dynamic environment.
[0059] A computer device capable of learning from a specific environment and making decisions using causal information, a computer program running on a processor, wherein the processor implements the steps of the causal large language model-based reinforcement learning decision optimization method when executing the computer program.
[0060] The beneficial effects of the present application are as follows:
[0061] The present application discloses a policy-driven causal intervention to revise a causal graph, so that an agent can actively optimize the causal structure according to task requirements and improve the efficiency of policy learning. By introducing causal intervention in the decision-making process, the agent can dynamically adjust key causal relationships to avoid policy failure due to incorrect correlation inference; when the environment changes, the agent can identify new influencing factors through causal analysis and adjust the decision-making strategy accordingly. The causal intervention mechanism can also reduce the dependence on a large amount of environment interaction data, so that the agent can still adapt to new tasks efficiently under low sample learning conditions. BRIEF DESCRIPTION OF DRAWINGS
[0062] Figure 1 A flowchart of a causal large language model-based reinforcement learning decision optimization method and system provided for the examples of the present application.
[0063] Figure 2 A brief schematic diagram of the environment provided for the examples of the present application.
[0064] Figure 3 A schematic diagram of the extraction of causal variables and the construction of a structural causal model by LLMs from a specific environment provided for the examples of the present application.
[0065] Figure 4 A schematic diagram of the principle of causal intervention provided for the examples of the present application.
[0066] Figure 5 A schematic diagram of the principle of obtaining causal sub-targets and rewards provided for the examples of the present application.
[0067] Figure 6 A schematic diagram of the principle of updating the strategy network of the agent in the dynamic environment using causal information provided for the examples of the present application.
[0068] Figure 7 A structural schematic diagram of a causal large language model-based reinforcement learning decision optimization method and system provided for the examples of the present application. Detailed Implementation
[0069] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0070] Example 1
[0071] like Figure 1 As shown, a reinforcement learning decision optimization method based on a causal large language model includes the following steps:
[0072] Initialize the agent and its policy network in a dynamic environment;
[0073] The agent interacts with the environment to obtain trajectory information of historical decision-making generated by the interaction;
[0074] A large language model is used to extract causal variables and initial causal relationships from trajectory information, and a structural causal model is constructed.
[0075] The agent's policy-driven causal intervention mechanism is obtained through the policy network. The agent performs a specified action in the environment and compares the changes in variable states before and after the intervention. By comparing the conditional probability after intervention with the original observation probability, the causal relationship in the structural causal model is dynamically corrected.
[0076] Based on the task-related causal chains extracted from the modified structural causal model, semantic sub-targets corresponding to the causal relationships are generated through reasoning using a large language model.
[0077] Design a multimodal reward function that integrates semantic similarity;
[0078] The obtained sub-objectives and reward functions are used to update the agent's policy network in a dynamic environment.
[0079] Example 2
[0080] More specifically, in this embodiment, such as Figure 2 As shown, within the Crafter environment, agents complete a series of complex tasks in an open-world environment, such as collecting resources, building tools, and responding to environmental changes.
[0081] In this embodiment, as Figure 3 As shown, the trajectory information of historical sequence decisions generated by the interaction is obtained, specifically as follows:
[0082] Enables the agent to interact with the environment, storing the environmental state, the agent's state, and the action sequence τ={s} of the agent's historical decision-making process. t ,φ t ,a t ,r t ,s t+1 ,φ t+1}; where t represents time st representing the environment state includes but is not limited to entity attributes, spatial topological relations, quantity relations; φ t representing the agent state, including but not limited to coordinates, energy, behavior constraint conditions; a t representing the action calculated by the agent according to the current observed state information through the strategy function; r t representing the state of the agent after taking action a t interacting with the environment, the state transitions to s t+1 , and the effect feedback of the decision action and the state change reward of the environment itself.
[0083] In one specific embodiment, a large language model is used to extract causal variables and initial causal relationships from trajectory information, and a structural causal model is constructed. The specific steps are as follows:
[0084] The environment state s t , s t+1 , the agent state φ t , φ t+1 , and the action sequence a t in the trajectory information τ are textually described as prompt templates, and a large language model is input to perform causal reasoning;
[0085] Through several rounds of iterative and multi-sample prompt engineering, the large language model is guided to learn the context, and a candidate causal variable set is output:
[0086] V, E = LLM(s t , s t+1 , φ t , φ t+1 , a t )
[0087] Wherein, E is the relationship variable set LLM, E = {e i→j |v i , v j ∈V}; the causal variable set is V = {v1, v2,..., v n};
[0088] An initial structural causal model G(V, E) is constructed, and the structural causal model is stored in a causal matrix, wherein the matrix element A ij ∈{0, 1} represents the existence of causal relationship from variable v i to v j .
[0089] In this embodiment, the large language model analyzes the potential causal relationship between the environmental state and the agent's action based on the causal reasoning ability in a large amount of training data, infers which environmental variables have a key impact on decision-making, for example, in the Crafter environment, making a "wood axe" requires "wood", and "wood" comes from "trees". The agent first needs to identify whether it has the conditions to collect wood in its current state, including whether there are trees around, whether it needs to use tools, etc. Then, the agent performs the "felling" behavior to obtain "wood" as a resource. By recording this process and inputting the large language model V, G = LLM(s t ,s t+1 ,φ t ,φ t+1 ,a t ), the agent can infer that "wood" is a necessary prerequisite for "wood axe", so it will prioritize collecting related materials in future task planning. In addition, the agent can also discover a better resource collection strategy based on causal reasoning, such as whether it should prioritize making other tools to improve resource acquisition efficiency, or how to reasonably allocate resources to complete more complex task goals.
[0090] In one specific embodiment, as Figure 4 shown, due to the bias problem of the large language model's reasoning, the correctness of the structural causal model cannot be guaranteed, and the agent's strategy is used to drive causal intervention to correct the wrong causal relationship. The agent strategy drives the causal intervention mechanism through the strategy network, and performs the specified action in the environment and compares the variable state changes before and after the intervention. By comparing the conditional probability after the intervention with the original observation probability, the causal relationship in the structural causal model is dynamically corrected, and the specific steps are as follows:
[0091] In the Crafter environment, the action performed by the agent strategy has an impact on the environment, and the target causal variable v i The agent implements the action a t calculated by the strategy network:
[0092]
[0093] After performing the action a t , record the state change of the target variable v j Calculate the conditional probability after the intervention and the original observation probability P(v j |v i ), if they are equal, it means that the causal relationship between v i and v j does not exist, otherwise it exists;
[0094] In this embodiment, specifically, the agent can choose to exert an operation on a specific target in the environment, such as felling a tree to observe whether wood will be obtained, or trying to dig for ore to determine whether minerals can be collected. If an operation always leads to a specific state change, the agent can confirm that there is a causal relationship between the operation and the state change, and use this information to correct the causal graph.
[0095] In this embodiment, if the LLM predicts that "felling a tree" can obtain "stone", but the agent actually performs the operation and does not obtain "stone" but obtains "wood" after the operation, the relationship is corrected through causal intervention, so that the causal model is more consistent with the real dynamic changes of the environment. From the perspective of structural causal models, when the agent's action is exerted on "tree", it is equivalent to exerting intervention on the "tree" node in the causal graph, that is, setting the value of the variable independent of the influence of its parent nodes. By performing the "felling" operation, the agent can observe the changes in the "wood" variable in the environment. If "wood" does indeed increase as a result of the "felling a tree" action, it can be confirmed that "tree" is a direct causal predecessor of "wood". If the number of "wood" does not change after intervention, further adjustment of the causal relationship is needed, and the influence of other potential variables, such as the use of tools, environmental conditions, etc. may be considered. Through continuous intervention and adjustment, the agent can dynamically optimize the causal structure, thereby improving the accuracy and adaptability of strategic decision-making;
[0096] Update the edge e in the structural causal model G(V, E) i→j The value in the matrix A:
[0097]
[0098] In one specific embodiment, as shown in Figure 5 According to the task-related causal chain extracted from the corrected structural causal model, the semantic sub-goals corresponding to the causal relationship are generated by reasoning through a large language model, and the specific steps are as follows:
[0099] Extract the agent's task-related causal chain from the structural causal model G(V, E):
[0100] P = {v i → v i+1 → ··· → v j | v i ∈ G};
[0101] Textualize the causal chain P as a prompt template and input it into a large language model for causal sub-goal reasoning:
[0102] LLM(prompt, P i );
[0103] Where Pi is the i-th segment of the causal chain v i →v i+1 , the prompt template contains the causal segment description and the target generation instruction; through multiple rounds of iterative and multi-sample prompt engineering, the large language model is guided to learn the context, and a sub-target set of the current task is output:
[0104]
[0105] Each sub-target g i corresponds to a causal relationship v i →v i+1 of length 1 in the causal chain P.
[0106] In this embodiment, in the Crafter environment, the agent needs to complete a series of tasks, such as manufacturing tools, collecting resources, and building structures. Based on the causal graph obtained by causal reasoning, the agent can decompose complex tasks into a series of sub-targets with causal relationships. For example, in order to make a "wooden axe", the agent first needs to obtain "wood", and "wood" comes from "cutting down trees". Therefore, the agent can set "cutting down trees" as a causal sub-target to complete the task of "making a wooden axe", and optimize the strategy of each stage in turn.
[0107] In one specific embodiment, in order to alleviate the problem of sparse rewards, the present application proposes a multi-modal cosine similarity reward function, which designs a multi-modal reward function that integrates semantic similarity. The specific steps are as follows: on the basis of the standard reward function, a multi-modal causal knowledge reward is introduced:
[0108] r t =r env +λr causal
[0109] Where r env represents the state change reward of the environment itself, λ is a configurable fusion weight parameter, and r causal represents an additional reward containing causal knowledge. By encoding the sub-target set g into a semantic vector {g i}, and calculating the cosine similarity between the embedding vector of the environment state s t at the current time t {s t}, and the cosine similarity is obtained.
[0110] In one specific embodiment, specifically, the r causal adopts a text-image multi-modal model Clip to calculate the cosine similarity between the two:
[0111]
[0112] In this embodiment, the current environment state (image) S = {s0, s1,..., s n} and the sub-target (text) Cosine similarity in embedding space measures the proximity of the agent's current task progress to the expected sub-target. Specifically, using a multi-modal model such as CLIP, text and images can be embedded into the same space, extracting semantic representations of task sub-targets and combining with environment state embeddings to calculate the similarity of the agent's current state and target state. The higher the similarity, the greater the reward the agent receives, guiding it to adopt more causal logic-compliant strategies. Specifically, assuming the current sub-target of the agent is "cutting down trees", the agent will identify whether there are trees to be cut down through visual perception. Over time, the semantic distance between the agent and "trees" can be captured from image embeddings. If the distance increases, the agent deviates from the semantic sub-target, and the reward decreases, otherwise it increases.
[0113] In one specific embodiment, as shown in Figure 6 , the obtained sub-target and reward are used to update the agent's policy network in a dynamic environment, with the following specific steps:
[0114] The sub-target-driven policy network is represented as:
[0115] π(a t |s t ,φ t ,g)
[0116] where φ t represents the state of the agent at time t, and the policy π learns to perform the optimal action a t under causal constraints by maximizing the long-term cumulative reward.
[0117] The agent performs action a t with policy π and interacts with the environment, and the state transitions to s t+1 , after which the agent receives the reward signal from the environment;
[0118] After obtaining a series of trajectory sequences τ, the advantage function is calculated to measure the value of the current action a t relative to the average policy:
[0119]
[0120] where γ is the discount factor, used to control the weight of future rewards, λ is the smoothing factor, used to control the smoothing degree of advantage estimation, T is the trajectory length, and δ is the TD error, measuring the prediction error of the current state value function:
[0121] δ t = rt +λV(s t+1 ,φ t+1 ,g t+1 )-V(s t ,φ t ,g t )
[0122] The loss function for obtaining the policy network update is:
[0123] L clip (θ)=E θ [min(r t ·A t ,clip(r t ,1-ε,1+ε)·A t )]
[0124] Wherein clip is a gradient clipping strategy to prevent the gradient update from being too large, and ε is a parameter for controlling the range of policy change, is the probability ratio of the new and old policies.
[0125] In this embodiment, in order to further optimize the decision-making process of the agent, the present application integrates semantic sub-goals and reward values into the agent's policy learning framework, so that the agent can not only optimize the policy using the cumulative reward of traditional reinforcement learning, but also enhance the decision-making efficiency and accuracy through the sub-goals obtained by causal reasoning. Specifically, during the training process, the agent not only optimizes based on the sparse rewards obtained from environmental interaction, but also uses the semantic sub-goals generated by the causal graph to provide additional supervision signals. For example, in the Crafter environment, if the final goal of the agent is to make a "wooden axe", first search for the corresponding causal knowledge in the causal graph, such as "cutting down trees" → "collecting wood" → "manufacturing wooden axes". During the policy training process, the agent will receive additional rewards every time it completes a sub-goal, allowing it to converge more efficiently to the optimal policy. At the same time, when the agent updates its policy, it will also fully consider the continuously updated causal knowledge in the dynamic environment. Specifically, the agent not only relies on fixed causal knowledge, but also dynamically adjusts the order of sub-goals to adapt to different environmental states. For example, if there are no available trees in the environment, the agent can adjust its strategy to prioritize finding trees or manufacturing tools suitable for felling. When executing the policy, an adaptive learning mechanism is also used to iteratively optimize the causal graph. For example, if the agent discovers after multiple attempts that "making a wooden axe" requires additional steps (such as "synthesizing wooden sticks"), it will automatically update the causal graph to more accurately describe the task dependency relationship. This dynamic adjustment mechanism can improve the task completion rate of the agent and reduce the time consumption caused by blind exploration. Through this method, the agent not only constructs an efficient task execution path using causal reasoning, but also optimizes the policy by combining semantic sub-goals and reward signals, improving the explainability and stability of decision-making.
[0126] The application is based on a large language model, a structural causal model, causal intervention, and reinforcement learning, and proposes a method capable of effectively improving the decision-making ability of an agent, enabling it to more efficiently learn a strategy in a complex dynamic environment, improve the task completion rate, and enhance adaptability and generalization ability.
[0127] In this embodiment, the agent can use a large language model to analyze environmental information and build relationships between task objectives through causal reasoning, thereby effectively reducing random exploration and improving learning efficiency. Through causal intervention technology, the agent can adaptively correct the causal graph, making strategy optimization more accurate. Further, combined with a cosine similarity reward mechanism based on multi-modal embedding, the agent can adjust the strategy in a dynamic environment to ensure the stability and generalization ability of task execution.
[0128] Embodiment 3
[0129] As shown in Figure 7 A reinforcement learning decision optimization system based on a causal large language model includes a learning module, an adaptation module, and a decision module.
[0130] The learning module is used to enable the agent to interact with the environment and obtain trajectory information of historical sequence decisions generated by the interaction: in this embodiment, after updating the strategy network of the agent in the dynamic environment, the agent is also allowed to interact with the environment for iterative learning.
[0131] The adaptation module is used to extract causal variables and initial causal relationships from the trajectory information using a large language model and construct a structural causal model.
[0132] The decision module is used to generate semantic sub-goals corresponding to the causal relationships through a large language model based on the task-related causal chains extracted from the corrected structural causal model; a multi-modal reward function that integrates semantic similarity is designed; and the obtained sub-goals and reward function are used to update the strategy network of the agent in the dynamic environment.
[0133] Embodiment 4
[0134] A computer device capable of learning from a specific environment and making decisions using causal information, a computer program running on a processor, wherein the processor executes the computer program to implement the steps of the reinforcement learning decision optimization method based on a causal large language model.
[0135] Obviously, the above embodiments of the application are merely examples for clearly illustrating the application, and are not intended to limit the embodiments of the application. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the application shall be included in the protection scope of the claims of the application.
Claims
1. A reinforcement learning decision optimization method based on a causal large language model, characterized in that, The method comprises the following steps: initializing an agent and a policy network thereof in a dynamic environment; interacting the agent with the environment to obtain trajectory information of historical sequence decisions generated by the interaction; extracting causal variables and initial causal relationships from the trajectory information by using a large language model to construct a structural causal model; obtaining an agent policy-driven causal intervention mechanism through the policy network, executing a specified action in the environment, comparing variable state changes before and after the intervention, comparing conditional probabilities after the intervention with original observation probabilities, and dynamically modifying causal relationships in the structural causal model; extracting a task-related causal chain from the modified structural causal model, reasoning and generating semantic sub-goals corresponding to the causal relationships by using the large language model; designing a multi-modal reward function that fuses semantic similarity; updating the policy network of the agent in the dynamic environment by using the obtained sub-goals and the reward function.
2. The reinforcement learning decision optimization method based on the causal large language model according to claim 1, characterized in that, The trajectory information of the historical sequence decisions generated by the interaction is obtained, specifically as follows: Let the agent interact with the environment, store the environment state, the state of the agent and the action sequence τ = {s t , φ t , a t , r t , s t+1 , φ t+1} of the agent historical sequence decision process; wherein t represents the time, s t represents the environment state including but not limited to entity attribute, spatial topological relation, quantity relation; φ t represents the agent state, including but not limited to coordinate, energy, behavior constraint condition; a t represents the action calculated by the agent according to the current observed state information through the strategy function; r t represents the effect feedback of the decision action and the state change reward of the environment itself after the state transition of the agent interacting with the environment with the action a t to s t+1 .
3. The reinforcement learning decision optimization method based on the causal large language model according to claim 2, characterized in that, The structural causal model is constructed by extracting causal variables and initial causal relationships from the trajectory information by using a large language model, and the specific steps are as follows: The environment state s in the trajectory information τ t , s t+1 , the agent state φ t , φ t+1 , and the action sequence a t are textually described as prompt templates, and a large language model is input to perform causal reasoning; Through several rounds of iterative and multi-sample prompting engineering, the large language model is guided to learn the context, and a candidate causal variable set is output: V, E = LLM(s t , s t+1 , φ t , φ t+1 , a t ) Wherein, E is a relationship variable set LLM representing a large language model, E = {e i→j |v i ,v j ∈V};The causal variable set is V = {v1, v2,..., v n} constructing an initial structural causal model G(V, E); storing the structural causal model into a causal matrix, wherein the matrix elements A ij ∈ {0, 1} represent the existence of a causal relationship from variable v i to v j causal relationship existence.
4. The reinforcement learning decision optimization method based on the causal large language model according to claim 3, characterized in that, The agent policy-driven causal intervention mechanism is obtained through the policy network, a specified action is executed in the environment, variable state changes before and after the intervention are compared, conditional probabilities after the intervention are compared with original observation probabilities, and causal relationships in the structural causal model are dynamically modified, and the specific steps are as follows: In an independent validation environment, the target causal variable v i The agent implements the action a computed by the policy network t : Perform action a t Post-record target variable v j State change Compute conditional probability after intervention Whether equal to original observation probability P(v j |v i ), if equal, then causal relationship between v i and v j does not exist, otherwise exists; updating an edge e in a structural causal model G(V, E) i→i values in matrix A:
5. The reinforcement learning decision optimization method based on the causal large language model according to claim 4, characterized in that, The task-related causal chain is extracted from the modified structural causal model, and the semantic sub-goals corresponding to the causal relationships are reasoned and generated by using the large language model, and the specific steps are as follows: The agent task-related causal chain is extracted from the structural causal model G(V, E): P = {v i → v i+1 → ··· → v j | v i ∈ G}; The causal chain is described in text as a prompt template, and the large language model is input to reason the causal sub-goals: LLM (prompt, P i ); wherein P i is the i-th segment of the causal relationship piece v i →v i+1 , the prompt template contains the causal relationship piece description and the target generation instruction; through multiple rounds of iterative and multi-sample prompt engineering, the large language model is guided to learn the context, and a sub-target set of the current task is output: where each sub-goal g i Causal relationships v of length 1 in the corresponding causal chain P i → v i+1 .
6. The reinforcement learning decision optimization method based on the causal large language model according to claim 5, characterized in that, The multi-modal reward function that fuses semantic similarity is designed, and the specific steps are as follows: On the basis of the standard reward function, multi-modal causal knowledge reward is introduced: r t = r env + λr causal where r env represents the state change reward of the environment itself, λ is a configurable fusion weight parameter, r causal represents an additional reward containing causal knowledge, by encoding the sub-goal set into a semantic vector {g i}, and embedding the environment state s t at the current time t {s t}, and calculating the cosine similarity between the two.
7. The reinforcement learning decision optimization method based on the causal large language model according to claim 6, characterized in that, Specifically, the r causal The cosine similarity of both is calculated using the text-image multimodal model Clip to get:
8. The causal-based large language model reinforcement learning decision optimization method according to claim 7, characterized in that, The sub-goals and the reward are used to update the policy network of the agent in the dynamic environment, and the specific steps are as follows: The policy network driven by the sub-goals is represented as: where φ t represents the state of the agent at the current time t, and the policy π learns to perform the optimal action a t under the causal constraints by maximizing the long-term cumulative reward The agent performs an action a with policy π t The state transition to s with environment interaction t+1 After that, the agent receives a reward signal from the environment feedback; After obtaining a series of trajectory sequences τ, the advantage function is calculated to measure the current action a. t The value of the strategy relative to the average strategy: Wherein, γ is a discount factor, used to control the weight of future rewards, λ is a smoothing factor, used to control the smoothing degree of advantage estimation, T is the length of the trajectory, δ is the TD error, measuring the prediction error of the current state value function: delta t = r t + lambda * V(s t+1 , phi t+1 , g t+1 ) - V(s t , phi t , g t ) The loss function of the updated policy network is obtained: L clip (θ) = E θ [min(r t · A t , clip(r t , 1 - ε, 1 + ε) · A t )] where clip is a gradient clipping strategy to prevent the gradient update magnitude from being too large, and ε is a parameter to control the range of the policy change, is the probability ratio of the new and old policies; The loss function is used to update the policy network.
9. A reinforcement learning decision optimization system based on a causal large language model, characterized in that, It comprises a learning module, an adaptation module and a decision module. The learning module is used to interact the agent with the environment to obtain trajectory information of historical sequence decisions generated by the interaction; The adaptation module is used to extract causal variables and initial causal relationships from the trajectory information by using a large language model to construct a structural causal model; The decision module is used to extract a task-related causal chain from the modified structural causal model, reason and generate semantic sub-goals corresponding to the causal relationships by using the large language model, design a multi-modal reward function that fuses semantic similarity, and update the policy network of the agent in the dynamic environment by using the obtained sub-goals and the reward function.
10. A computer device, comprising: The computer device can learn from a specific environment and make decisions using causal information, a computer program running on a processor, wherein the processor implements the steps of the reinforcement learning decision optimization method based on the causal large language model as claimed in claims 1-8 when executing the computer program.
Citation Information
Cited By
Multi-agent task routing method and device based on causal atlas and related products
CN121900891A
Mobile robot path navigation and planning method suitable for complex scene
CN122041907A
Interactive simulation method and system for intelligent robot with body
CN122222048A
Soft prior fused causal DAG discovery method for online service system
CN122242707A