Strategy model training method and device, computer equipment and storage medium
Through large language model evaluation and adaptive security entropy mechanism, combined with Off-policy training, the problems of security constraint sparsity and reward sparsity in reinforcement learning are solved, and security and efficiency are improved, and strategy optimization is optimized to adapt to different task scenarios.
Patent Information
- Application Number
- CN202510782184.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-12
AI Technical Summary
During the training process, traditional reinforcement learning is difficult to effectively ensure that the agent does not violate environmental security constraints during the exploration process, resulting in increased training costs and system damage, and the sparsity of safety constraint signals and reward sparsity affect the learning effect of strategy.
By introducing a large language model to evaluate the security of environmental state and actions, a dense security indication signal is generated, and combined with the adaptive security entropy mechanism and Off-policy training method, the security and exploration of the strategy are dynamically adjusted, and the dual-action value network and target network balance training process is used.
It improves security and data utilization efficiency during the training process, ensures that the policy model maximizes cumulative rewards while meeting security constraints, adapts to different task scenarios, and provides accurate security feedback.
Smart Images

Figure CN120278215A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of deep learning technology. Specifically, it relates to a training method, device, computer device, and storage medium for a policy model. Background Art
[0002] Reinforcement Learning (RL) is a machine learning method that learns the optimal policy by an agent interacting with the environment. In recent years, reinforcement learning has made remarkable progress in fields such as games, robot control, and autonomous driving. However, traditional reinforcement learning faces many challenges in practical applications. Especially during the training process, since the reinforcement learning agent may take dangerous actions during exploration, it may violate the safety constraints of the environment, which not only increases the training cost but may also cause irreversible damage to the system where the agent is deployed.
[0003] Safe Reinforcement Learning (Safe RL) ensures that the agent does not violate the safety rules of the environment when exploring and optimizing the policy by introducing safety constraints during the training process. Summary of the Invention
[0004] In view of this, this application provides a training method, device, computer device, and storage medium for a policy model.
[0005] Specifically, this application is implemented through the following technical solutions: In a first aspect, an embodiment of the present disclosure provides a training method for a policy model, the method comprising: Obtain a first environmental state, and input the first environmental state into a policy model to be trained to obtain a first action corresponding to the first environmental state; Process the first environmental state and the first action using a pre-trained large language model to obtain a safety indication signal corresponding to the first action; the safety indication signal is used to indicate whether it is safe to execute the first action in the first environmental state; and Interact with the environment based on the first action to obtain a second environmental state and a reward; Construct interaction data based on the first environmental state, the first action, the safety indication signal, the second environmental state, and the reward, and train the policy model to be trained based on the interaction data to obtain a target policy model.
[0006] Optionally, the process of using a pre-trained large language model to process the first environmental state and the first action to obtain a safety indication signal corresponding to the first action includes: Convert the first environmental state and the first action into a natural language description; Input the first environmental state and the first action described in natural language into the large language model to obtain the safety indication signal.
[0007] Optionally, training the to-be-trained policy model based on the interaction data includes: Taking the interaction data generated in the current training cycle as the target interaction data, and training the to-be-trained policy model for the current training cycle based on the target interaction data; and / or, Storing the interaction data determined in the current training cycle into the experience replay pool; Sampling the interaction data stored in the experience replay pool to obtain the target interaction data corresponding to the current training cycle; and training the to-be-trained policy model for the current training cycle based on the target interaction data.
[0008] Optionally, training the to-be-trained policy model for the current training cycle based on the target interaction data includes: Determining the to-be-trained policy model, action value network, and target network for the current training cycle; wherein, the to-be-trained policy model, action value network, and target network for the current training cycle are respectively the initialized policy model, initialized action value network, and initialized target network; and the network architectures and network internal parameters of the initialized action value network and initialized target network are the same; or, they are the to-be-trained policy model updated in the previous training cycle of the current training cycle, the action value network updated in the previous training cycle, and the target network updated in the previous training cycle. Based on the target interaction data, updating the action value network for the current training cycle by minimizing the Bellman residual, updating the policy model for the current training cycle by maximizing the objective function, and updating the target network for the current training cycle through a slow update mechanism.
[0009] Optionally, the action value network includes: a first action value network and a second action value network; the target network includes: a first target network and a second target network; the network architectures of the first action value network and the second action value network are the same, and the network internal parameters are different; the network architectures and network internal parameters of the initialized first action value network and the initialized first target network are the same; the network internal parameters of the initialized second action value network and the initialized second target network are the same.
[0010] Optionally, based on the target interaction data, update the action-value network of the current training cycle by minimizing the Bellman residual, update the policy model of the current training cycle by maximizing the objective function, and update the target network of the current training cycle through a slow update mechanism, including: Perform a weighted sum processing on the network parameters of the target network in the current training cycle and the network parameters of the action-value network in the current training cycle based on a preset weight to obtain the target network parameters of the target network in the current training cycle; update the target network of the current training cycle based on the target network parameters; Use the action-value network of the current training cycle to process the first environmental state and the first action in the target interaction data to obtain a first cumulative reward expectation; and use the target network of the current training cycle or the updated target network to process the second environmental state and the second action corresponding to the second environmental state in the target interaction data to obtain a second cumulative reward expectation; substitute the first cumulative reward expectation, the second cumulative reward expectation, the safety indication information in the target interaction data, and the reward into a pre-constructed Bellman residual function, and update the action-value network of the current training cycle with the goal of minimizing the Bellman residual; Substitute the first cumulative reward expectation into a pre-constructed objective function, and update the policy model of the current training cycle with the goal of maximizing the objective function.
[0011] Optionally, the interaction data further includes a completion indication signal; the completion indication signal is used to indicate the task execution status corresponding to the policy network after performing the first action in the first environmental state; In the case where the task execution status indicates that no other actions can be performed, the updating of the action-value network of the current training cycle by minimizing the Bellman residual includes: Use the action-value network of the current training cycle to process the first environmental state and the first action in the target interaction data to obtain a first cumulative reward expectation; substitute the first cumulative reward expectation and the safety indication information in the target interaction data into a pre-constructed Bellman residual function, and update the action-value network of the current training cycle with the goal of minimizing the Bellman residual.
[0012] Optionally, the method further includes: Determine the safety entropy weight parameter of the current training cycle; the safety entropy weight parameter is used to adjust the weight of the safety information in the target interaction information when updating the action-value network; the safety entropy weight parameter of the current training cycle is a preset parameter, or is determined based on the safety information in the target interaction data determined in the previous training cycle; Updating the action value network of the current training cycle by minimizing the Bellman residual, including: Based on the safety entropy weight parameter, update the action value network of the current training cycle by minimizing the Bellman residual.
[0013] In a second aspect, an embodiment of the present disclosure further provides a training device for a policy model, the device includes: An acquisition module, configured to acquire a first environmental state, and input the first environmental state into a policy model to be trained, so as to obtain a first action corresponding to the first environmental state; A processing module, configured to use a pre-trained large language model to process the first environmental state and the first action to obtain a safety indication signal corresponding to the first action; the safety indication signal is used to indicate whether it is safe to execute the first action in the first environmental state; and Interact with the environment based on the first action to obtain a second environmental state and a reward; A training module, configured to form interaction data based on the first environmental state, the first action, the safety indication signal, the second environmental state, and the reward, and train the policy model to be trained based on the interaction data to obtain a target policy model.
[0014] Optionally, when the processing module uses a pre-trained large language model to process the first environmental state and the first action to obtain a safety indication signal corresponding to the first action, it is configured to: Convert the first environmental state and the first action into a natural language description; Input the first environmental state and the first action in natural language description into the large language model to obtain the safety indication signal.
[0015] Optionally, when the training module trains the policy model to be trained based on the interaction data, it is configured to: Use the interaction data generated in the current training cycle as target interaction data, and based on the target interaction data, train the policy model to be trained for the current training cycle; and / or, Store the interaction data determined in the current training cycle in an experience replay pool; Sample the interaction data stored in the experience replay pool to obtain target interaction data corresponding to the current training cycle; based on the target interaction data, train the policy model to be trained for the current training cycle.
[0016] Optionally, when the training module trains the policy model to be trained for the current training cycle based on the target interaction data, it is configured to: Determine the policy model, action-value network, and target network to be trained in the current training cycle; wherein, the policy model, action-value network, and target network to be trained in the current training cycle are respectively the initialized policy model, initialized action-value network, and initialized target network; and the network architectures and internal network parameters of the initialized action-value network and initialized target network are the same; or, they are the policy model to be trained updated in the previous training cycle of the current training cycle, the action-value network updated in the previous training cycle, and the target network updated in the previous training cycle. Based on the target interaction data, update the action-value network of the current training cycle by minimizing the Bellman residual, update the policy model of the current training cycle by maximizing the objective function, and update the target network of the current training cycle through a slow update mechanism.
[0017] Optionally, the action-value network includes: a first action-value network and a second action-value network; the target network includes: a first target network and a second target network; the network architectures of the first action-value network and the second action-value network are the same, and the internal network parameters are different; the network architectures and internal network parameters of the initialized first action-value network and the initialized first target network are the same; the internal network parameters of the initialized second action-value network and the initialized second target network are the same.
[0018] Optionally, when the training module updates the action-value network of the current training cycle by minimizing the Bellman residual, updates the policy model of the current training cycle by maximizing the objective function, and updates the target network of the current training cycle through a slow update mechanism based on the target interaction data, it is used for: Perform a weighted sum processing on the network parameters of the target network in the current training cycle and the network parameters of the action-value network in the current training cycle based on a preset weight to obtain the target network parameters of the target network in the current training cycle; update the target network of the current training cycle based on the target network parameters. Use the action-value network of the current training cycle to process the first environmental state and the first action in the target interaction data to obtain the first cumulative reward expectation; and use the target network or the updated target network of the current training cycle to process the second environmental state and the second action corresponding to the second environmental state in the target interaction data to obtain the second cumulative reward expectation; substitute the first cumulative reward expectation, the second cumulative reward expectation, the safety indication information in the target interaction data, and the reward into a pre-constructed Bellman residual function, and update the action-value network of the current training cycle with the goal of minimizing the Bellman residual. Substitute the first cumulative reward expectation into the pre-constructed objective function, and update the policy model for the current training cycle with the goal of maximizing the objective function.
[0019] Optionally, the interaction data further includes a completion indication signal; the completion indication signal is used to indicate the task execution status corresponding to the policy network after performing the first action in the first environmental state. In the case where the task execution status indicates that no other actions can be performed, the training module 53, when updating the action-value network for the current training cycle by minimizing the Bellman residual, is used for: Utilize the action-value network for the current training cycle to process the first environmental state and the first action in the target interaction data to obtain the first cumulative reward expectation; substitute the first cumulative reward expectation and the safety indication information in the target interaction data into the pre-constructed Bellman residual function, and update the action-value network for the current training cycle with the goal of minimizing the Bellman residual.
[0020] Optionally, the determination module is further used for: Determine the safety entropy weight parameter for the current training cycle; the safety entropy weight parameter is used to adjust the weight of the safety information in the target interaction information when updating the action-value network; the safety entropy weight parameter for the current training cycle is a preset parameter, or is determined based on the safety information in the target interaction data of the previous training cycle. The training module, when updating the action-value network for the current training cycle by minimizing the Bellman residual, is used for: Based on the safety entropy weight parameter, update the action-value network for the current training cycle by minimizing the Bellman residual.
[0021] In a third aspect, an optional implementation manner of the present disclosure further provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps in the above first aspect or any possible implementation manner in the first aspect are implemented.
[0022] In a fourth aspect, an optional implementation manner of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above first aspect or any possible implementation manner in the first aspect are implemented.
[0023] In a fifth aspect, an optional implementation manner of the present disclosure further provides a computer program product. The computer program product carries program code, and the instructions included in the program code can be used to execute the steps of the model training method as described in the first aspect or any item in the first aspect.
[0024] In the training method of the policy model provided by the embodiments of the present disclosure, after obtaining the first environmental state, the first environmental state is input into the policy model to be trained to obtain an action corresponding to the first environmental state. Then, the pre-trained large language model is used to process the first environmental state and the action to obtain a safety indication signal corresponding to the action, and an interaction is performed based on the action and the environment to obtain a second environmental state, a reward, and a completion signal. After that, the first environmental state, the action, the safety indication signal, the second environmental state, the reward, and the completion signal are used as interaction data, and the policy model to be trained is trained based on the interaction data. Thus, the large language model is used to evaluate the safety of the environmental state-action pair, and a densified safety indication signal is generated as a safety constraint during the training process. The large language model can quickly adapt to different task scenarios and provide accurate safety feedback by using its rich prior knowledge, thereby ensuring sufficient safety information is obtained during the training process.
[0025] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the technical solutions of the present disclosure.
[0026] To make the above objects, features, and advantages of the present disclosure more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. Description of the Drawings
[0027] Figure 1 Shows a flowchart of the training method of the policy model provided by some embodiments of the present disclosure; Figure 2 Shows a flowchart of a specific way to train the policy model to be trained in the current training cycle provided by some embodiments of the present disclosure; Figure 3 Shows a flowchart of a specific example of the training method of the policy model provided by some embodiments of the present disclosure; Figure 4 Shows a schematic structural diagram of a computer device provided by some embodiments of the present disclosure; Figure 5 Shows a schematic structural diagram of the training device of the policy model provided by some embodiments of the present disclosure. Detailed Embodiments
[0028] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are only examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0029] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms "a", "the", and "said" used in this application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0030] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".
[0031] In the related art, there are the following several problems in safety reinforcement learning: 1: In many practical scenarios, the safety constraint signals of the environment are sparse, that is, only when the agent takes dangerous actions will the corresponding feedback be triggered. This sparsity of safety constraint signals makes it difficult for the agent to obtain sufficient safety information in the initial stage of training, thus increasing the number of violations of safety constraints.
[0032] 2: The safety constraints in safety reinforcement learning are implemented by relying on manually designed constraint functions, which are difficult to adapt to the dynamically changing environment and task requirements, and have limited generalization ability.
[0033] 3: In current safety reinforcement learning, the rewards are also sparse, that is, in general, the agent needs to interact with the environment multiple times to obtain rewards. In an environment with sparse rewards, the safety reinforcement learning algorithm may overly focus on safety constraints, resulting in a too conservative policy and being unable to effectively complete the task. This problem is particularly prominent in complex task scenarios.
[0034] 4: Traditional safety reinforcement learning algorithms usually adopt the On-policy training method, with low data utilization efficiency, resulting in too long training time. In addition, frequent interaction with the environment also increases the risk of violating safety constraints.
[0035] To solve the above problems, the embodiments of the present disclosure provide a method for training a policy model. After obtaining the first environmental state, the first environmental state is input into the policy model to be trained, and an action corresponding to the first environmental state is obtained. Then, the pre-trained large language model is used to process the first environmental state and the action to obtain a safety indication signal corresponding to the action, and an interaction is performed based on the action and the environment to obtain a second environmental state and a reward. After that, the first environmental state, the action, the safety indication signal, the second environmental state, and the reward are used as interaction data, and the policy model to be trained is trained based on the interaction data. Thus, the large language model is used to evaluate the safety of the environmental state-action pair, and a densified safety indication signal is generated as a safety constraint during the training process. The large language model can quickly adapt to different task scenarios and provide accurate safety feedback by using its rich prior knowledge, thereby ensuring that sufficient safety information is obtained during the training process.
[0036] In addition, in the embodiments of the present disclosure, the Adaptive Safety Entropy mechanism is also used as a safety measurement index to dynamically adjust the safety of the policy in different environmental states. In the early stage of training, the Adaptive Safety Entropy reduces the access frequency to dangerous states by the action probability of the uniform distribution, and dynamically adjusts according to the learning of the policy in the later stage of training to increase the access frequency to safe states, thereby dynamically balancing exploration and safety, and ensuring that the policy optimization in the training process maximizes the cumulative reward under the premise of meeting the safety constraints.
[0037] The embodiments of the present disclosure also adopt an Off-policy training method, and store historical interaction data through an experience replay pool to improve the data utilization efficiency.
[0038] In addition, when training the policy model, the embodiments of the present disclosure use two action value networks and two target networks as assistance to dynamically balance exploration and safety, and ensure that the policy optimization in the training process maximizes the cumulative reward under the premise of meeting the safety constraints.
[0039] All the defects existing in the above solutions are the results obtained by the inventors after practice and careful research. Therefore, the discovery process of the above problems and the solutions proposed by the present disclosure for the above problems in the following text should be the contributions made by the inventors to the present disclosure during the process of the present disclosure.
[0040] To facilitate the understanding of the technical solutions of the present disclosure, the technical terms in the embodiments of the present disclosure are first explained as follows: Reinforcement Learning (RL) is an important branch of machine learning. Its core idea is to enable an agent to learn the optimal behavior strategy in a "trial and error" manner by interacting with the environment, so as to maximize the long-term cumulative reward. Reinforcement learning is widely applied in fields such as game AI (e.g., AlphaGo), robot control, autonomous driving, and recommendation systems.
[0041] An agent, that is, the entity that executes decisions, interacts with the environment by perceiving the environmental state and choosing actions.
[0042] The environment refers to everything outside the agent, including the state space, action space, and dynamic rules. After the agent executes an action, the environment will transfer from one state to another new state and return a reward. For example, in the field of autonomous driving, the environment includes the road state near the autonomous vehicle (such as road grade, whether it is an intersection, etc.), obstacle state (such as the moving direction, moving speed, moving acceleration of pedestrians or other vehicles), road driving rules (such as whether lane changing or overtaking is allowed), etc.
[0043] The state (State, S), that is, the set of features describing the environmental situation. The agent determines actions by observing the environmental state. For example, in the field of autonomous driving, the agent uses radar, image sensors, etc. to identify various types of obstacles in the environment, the moving direction, moving speed, moving acceleration of obstacles, or relevant information about the current road section determined by GPS, etc.
[0044] The action (Action, A), that is, the set of operations that the agent can execute in a certain state, representing the legal action space in environmental state s. For example, in the field of autonomous driving, the action can be the movable area, the movable speed range, the movable acceleration range, etc.
[0045] The reward (Reward, R), that is, the immediate feedback signal from the environment to the agent's action, is the core goal driving learning. The reward function needs to be reasonably defined to guide the agent to learn the expected behavior. For example, in Go, "winning the game" is +1, "losing the game" is -1, and in autonomous driving control, "reaching the destination" is +1, "not reaching the destination" is -1.
[0046] Policy (Policy, π): The policy model of the agent, describing the mapping from state to action, is divided into: Deterministic policy: , which represents choosing action a in environmental state s.
[0047] Stochastic policy: , which represents the probability of choosing action a in environmental state s.
[0048] Value Function: Evaluates the long-term reward expectation of a state-action pair and is used to guide policy optimization. It includes: Action Value Function (Q-function):
[0049] Represents the cumulative reward expectation of following policy π after executing action a from state s.
[0050] To facilitate the understanding of this embodiment, first, a detailed introduction to the training of a policy model disclosed in this embodiment of the present disclosure is provided. The execution subject of the training method of the policy model provided in this embodiment of the present disclosure is generally a computer device with certain computing capabilities. Such a computer device includes, for example: a terminal device or a server or other processing devices. The terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a vehicle-mounted device, a wearable device, etc. In some possible implementation manners, the training method of the policy model can be implemented by a processor calling computer-readable instructions stored in a memory.
[0051] It should be noted that the training method of the policy model provided in this embodiment of the present disclosure can be used in various scenarios, such as game and game theory scenarios, robot control, autonomous driving, recommendation systems, financial transactions, etc. Specifically, this embodiment of the present disclosure does not make any limitations.
[0052] The following explains the training method of the policy model provided in this embodiment of the present disclosure.
[0053] See Figure 1 As shown, it is a flowchart of the training method of the policy model provided in this embodiment of the present disclosure. The method includes steps S101 to S104, where: S101: Obtain a first environmental state and input the first environmental state into the policy model to be trained to obtain a first action corresponding to the first environmental state; S102: Use a pre-trained large language model to process the first environmental state and the first action to obtain a safety indication signal corresponding to the first action; the safety indication signal is used to indicate whether it is safe to execute the first action in the first environmental state; and S103: Interact with the environment based on the first action to obtain a second environmental state and a reward; S104: Construct interaction data based on the first environmental state, the first action, the safety indication signal, the second environmental state, and the reward, and train the policy model to be trained based on the interaction data to obtain a target policy model.
[0054] Among them, there is no logical order of execution between the above S102 and S103. S102 can be executed first and then S103, or S103 can be executed first and then S102; S102 and S103 can also be executed in parallel.
[0055] The above S101 to S104 will be described in detail below.
[0056] Regarding the above S101: In a specific implementation, when obtaining the first environmental state, for example, it is to obtain the first environmental state of the working environment of the agent that deploys the target policy model. This working environment is related to the application scenario of the agent, for example; when the application scenario of the agent is a game scenario, the working environment of the agent is the game scenario; this game scenario can be a virtual game scenario, a real game scenario, an augmented reality (AR) game scenario, or a virtual reality (VR) game scenario, which will vary according to actual application needs. When the application scenario of the agent is a robot control scenario, the working environment of the agent is the specific working environment of the robot. For example, the working scenario of an automatic sorting robot is the sorting site; the working scenario of an automatic search and rescue robot is the search and rescue site, etc. When the application scenario of the agent is in the field of autonomous driving, the working environment of the agent is the working environment corresponding to the road of autonomous driving, such as the road where the autonomous driving vehicle that deploys the agent is currently traveling, etc.
[0057] When the agent is used in different scenarios, the method of obtaining the first environmental state is different, and the obtained first environmental state is also different.
[0058] For example, when an agent is used in a game, the first environmental state is, for example, the game scene state. In the case of a board game (such as Go), this first environmental state is, for example, the layout state of the chess pieces on the board. In the case of a real person and an agent playing a game in reality, the way to obtain the first environmental state is, for example, to obtain an image of the current chessboard through an image sensor, and based on the image of the current chessboard, determine the state of the current chessboard, such as the positions of the chess pieces and the players to whom the chess pieces belong; in the case of a real person and an agent playing a game using a computer device, the way to obtain the first environmental state is, for example, to determine the state of the chessboard by reading the information recording the game progress currently on the computer device. Or when the agent is used for robot control, the first environmental state is, for example, the scene state in which the robot works; for example, if the robot is used to sort items on a conveyor belt, the first environmental state is the relevant distribution state of the items on the conveyor belt; the way to obtain the first environmental state is, for example, to obtain an image of the conveyor belt through an image sensor, and based on the obtained image of the conveyor belt, determine the distribution of the items on the conveyor belt, the conveying state, etc. When the agent is used in the field of autonomous driving, the first environmental state is, for example, the current road state in which the autonomous driving vehicle is traveling, such as the distribution of obstacles, the movement of obstacles, the state of the road, etc. For example, the state of obstacles and the road conditions within a certain range of the vehicle can be obtained through on-vehicle radar and on-vehicle image sensors.
[0059] After obtaining the first environmental state it can be input into the policy model to be trained, and the policy model to be trained is used to process the first environmental state to output a first action corresponding to the first environmental state .
[0060] Specifically, for different application scenarios, the first action corresponding to the first environmental state also varies. For example, when the agent is used in a game scene, the first action corresponding to the first environmental state is the next game operation; in a board game, the first action is, for example, placing a chess piece at a specific landing point. When the agent is used for robot control, the first action corresponding to the first environmental state is a specific control action for the robot; when the agent is used in the field of autonomous driving, the first action corresponding to the first environmental state is, for example, the next driving action, such as moving forward, backward, accelerating, decelerating, changing lanes, controlling the high beam to flash, etc. Specifically, according to different application scenarios, the first action also varies.
[0061] Regarding the above S102, a pre-trained large language model is used, for example, to evaluate the safety of environment state-action pairs, to evaluate whether it is safe to perform a certain action in a specific environment state, and to output a corresponding safety indication signal. In the embodiments of the present disclosure, the safety indication signal is used to indicate in the first environment state to perform the first action is safe or not.
[0062] In a possible implementation manner, the safety indication signal is, for example, a binary signal; when the value of the safety indication signal is the first numerical value, it indicates that it is safe to perform the first action in the first environment state to perform the first action is safe; when the value of the safety indication signal is the second numerical value, it indicates that it is unsafe to perform the first action in the first environment state to perform the first action is unsafe.
[0063] Exemplarily, it is assumed that when the value of the safety indication signal is 0, it indicates safety, and when the value of the safety indication signal is 1, it indicates unsafety.
[0064] In different scenarios, whether it is safe can be specifically set according to the specific application scenario; for example, in a game scenario, safety can be, for example, being attacked, a chess piece being eaten, losing a game, etc. Specifically, according to the differences of different games, the content corresponding to safety will be different. Another example is in the scenario of a robot performing item sorting control, safety can be, for example, not missing the items to be sorted; being unsafe means that the probability of missing the items to be sorted reaches a certain probability threshold. In the field of autonomous driving, safety is, for example, not having a collision; being unsafe means that the probability of colliding with nearby obstacles or driving out of the current lane reaches a certain probability threshold, etc. It can be specifically set according to actual needs, and the embodiments of the present disclosure do not make limitations.
[0065] In an alternative embodiment, there is also provided a specific method for using a pre-trained large language model to process the first environment state and the action to obtain a safety indication signal corresponding to the action, including: Converting the first environment state and the first action into a natural language description; Inputting the first environment state and the first action described in natural language into the large language model to obtain the safety indication signal.
[0066] In a specific implementation, for example, a state conversion module can be used to convert the first environmental state and the first action into a natural language description. The state conversion module can include, for example, two parts: state translation and action translation. State translation converts the numerical or symbolic environmental state encoding into a natural language with clear semantics. Action translation can discretize continuous actions into different options and convert them into a natural language description.
[0067] After both the first environmental state and the first action are converted into natural language descriptions, the natural language descriptions of the first environmental state and the first action are input into a large language model to obtain a safety indication signal.
[0068] The safety indication signal is, for example, expressed as: .
[0069] Regarding the above S103: When interacting based on the first action and the environment, for example, in the first environmental state, an agent with a policy model to be trained is controlled to execute the first action to obtain a second environmental state and a reward.
[0070] Specifically, the second environmental state, for example, is at a certain moment after the first action is executed in the first environmental state . The environment is affected by the first action and changes from the first environmental state to another new state, which is the second environmental state. It can be expressed, for example, as: . .
[0071] Reward is, for example, an immediate feedback signal from the environment to the first action .
[0072] Taking autonomous driving as an example, the first environmental state at a certain moment is , and this first environmental state is represented as a vector, which represents features such as the position, speed, and surrounding environment (such as the distribution, movement, and type of obstacles) of the autonomous driving vehicle. The policy model outputs the distribution of different actions based on this vector, and a driving action is obtained by sampling from this distribution , and this action is assumed to be going forward. The agent controls the vehicle to execute this action. At this time, the environment, position, speed, etc. of the vehicle have all changed, and a new environmental state is obtained. If the task is completed, such as moving to the desired position, a positive reward value (scalar) is returned. On the contrary, if a collision occurs after the first action is executed, a negative reward value is returned; if neither moving to the desired position nor a collision occurs, 0 is returned, that is, the reward value is 0.
[0073] For the above S104: After obtaining the first environmental state, the corresponding first action, the safety indication signal, the second environmental state, and the reward based on the above S101 - S103 the first environmental state, the corresponding action, the safety indication signal, the second environmental state, and the reward can be formed into a set of interaction data and the policy model to be trained can be trained using the interaction data to obtain the target policy model. the second environmental state the reward After that, the above - mentioned first environmental state the corresponding action the safety indication signal the second environmental state the reward constitute a set of interaction data, and the interaction data can be used to train the policy model to be trained to obtain the target policy model.
[0074] In another embodiment of the present disclosure, for example, at least one of the following methods a1 and a2 can be adopted to implement training the policy model to be trained using the interaction data: a1: Use the interaction data generated in the current training cycle to train the policy model to be trained for the current training cycle.
[0075] In a specific implementation, the policy model to be trained can be trained for multiple training cycles. In each training cycle, at least one execution of the steps described in the above S101 - S103 can be performed. That is, the interaction data corresponding to at least one moment can be obtained using the agent at at least one moment. That is, assuming that in one training cycle, the agent obtains the interaction data corresponding to n moments, there will be n sets of interaction data. n is an integer greater than 0.
[0076] In this case, the n sets of interaction data can be used to train the decision model to be trained for the current training cycle.
[0077] a2: Store the interaction data determined in the current training cycle in the experience replay pool; sample the interaction data stored in the experience replay pool to obtain the target interaction data corresponding to the current training cycle; based on the target interaction data, train the policy model to be trained for the current training cycle.
[0078] In a specific implementation, similar to the above a1, the policy model to be trained can also be trained for multiple training cycles. And in each training cycle, n sets of interaction data are obtained in the manner described in the above a1. After that, the n sets of interaction data can be stored in the experience replay pool. Then, at a certain sampling rate and / or sampling method, at least one set of target interaction data corresponding to the current training cycle is sampled from the experience replay pool, and based on the target interaction data, the policy model to be trained is trained for the current training cycle.
[0079] In this case, if the current training cycle is the first training cycle, the experience replay pool only includes the interaction data generated in one training cycle. Therefore, usually, when training the policy model to be trained in the first training cycle, the target interaction data includes at least part of the interaction data generated in the first training cycle.
[0080] If the current training cycle is the kth training cycle and it is other than the first training cycle, the experience replay pool includes not only the interaction data generated in the current training cycle but also the interaction data stored in the experience replay pool in the 1st to k - 1th training cycles. Therefore, when training the policy model in the kth training cycle, the target interaction data can include not only the interaction data generated in the kth training cycle but also at least part of the interaction data generated in the 1st to k - 1th training cycles.
[0081] In this way, the method of storing the interaction data of each training cycle in the experience replay pool and sampling the experience replay pool in each training cycle, so as to train the policy model to be trained with the sampled target interaction data, adopts the Off-policy training method. By storing historical interaction data in the experience replay pool, the utilization efficiency of interaction data can be improved, so that each training cycle can obtain sufficient target interaction data and improve the training efficiency of the decision network.
[0082] See Figure 2 As shown, the embodiments of the present disclosure also provide a specific method for training the policy model to be trained for the current training cycle based on the target interaction data, including: S201: Determine the policy model to be trained, the action value network, and the target network in the current training cycle; wherein, the network architectures of the target network and the action value network are the same; S202: Based on the target interaction data, update the action value network by minimizing the Bellman residual, update the policy model by maximizing the objective function, and update the target network through a slow update mechanism.
[0083] In a specific implementation, when the current training cycle is the first training cycle, the policy model to be trained, the action-value network, and the target network in the current training cycle are respectively an initialized policy model, an initialized action-value network, and an initialized target network; and the network architectures and internal network parameters of the initialized action-value network and the initialized target network are the same; when the current training cycle is other training cycles except the first training cycle, the policy model to be trained, the action-value network, and the target network in the current training cycle are the policy model to be trained updated in the previous training cycle of the current training cycle, the action-value network updated in the previous training cycle, and the target network updated in the previous training cycle.
[0084] In a specific implementation, there can be two action-value networks, including a first action-value network and a second action-value network respectively; there can also be two target networks, including a first target network and a second target network respectively.
[0085] Among them, the network architectures of the first action-value network and the second action-value network are the same, but the internal network parameters are different; the network architectures and internal network parameters of the initialized first action-value network and the initialized first target network are the same; the internal network parameters of the initialized second action-value network and the initialized second target network are the same.
[0086] In a specific implementation, when initializing the action-value network, the internal network parameters of the first action-value network and the second action-value network can be randomly initialized respectively to obtain the first action-value network and the second action-value network. Moreover, when initializing the first target network, the first target network is initialized according to the internal network parameters of the first action-value network; when initializing the second target network, the second target network is initialized according to the internal network parameters of the second action-value network.
[0087] In the subsequent training stage, since the action-value network and the target network are updated in each training cycle, as a result, in the subsequent training cycles, the internal network parameters of the first action-value network and the first target network will also be different, and the internal network parameters of the second action-value network and the second target network will also gradually be different. The first target network and the second target network can automatically regulate the degree of optimization during the optimization process of each round of the action-value network to solve the problem of possible overestimation of a certain action by a single action-value network, so as to make the training process more stable.
[0088] In another embodiment of the present disclosure, there is also provided a specific method for updating the action-value network of the current training cycle by minimizing the Bellman residual based on the target interaction data, updating the policy model of the current training cycle by maximizing the objective function, and updating the target network of the current training cycle by a slow update mechanism, including: (1): Based on preset weights, perform weighted summation processing on the network parameters of the target network in the current training cycle and the network parameters of the action-value network in the current training cycle to obtain the target network parameters of the target network in the current training cycle; update the target network in the current training cycle based on the target network parameters; (2): Use the action-value network of the current training cycle to process the first environmental state and the first action in the target interaction data to obtain the first cumulative reward expectation; and Use the target network of the current training cycle or the updated target network to process the second environmental state and the second action corresponding to the second environmental state in the target interaction data to obtain the second cumulative reward expectation; Substitute the first cumulative reward expectation, the second cumulative reward expectation, the safety indication information in the target interaction data, and the reward into a pre-constructed Bellman residual function, and update the action-value network of the current training cycle with the goal of minimizing the Bellman residual; (3): Substitute the first cumulative reward expectation into a pre-constructed objective function, and update the policy model of the current training cycle with the goal of maximizing the objective function.
[0089] In a specific implementation, the action-value network includes a first action-value network and a second action-value network , and the target network includes a first target network and a second target network as an example. Among them, the policy model is expressed as: . Among them, represents a policy, which consists of an action sequence composed of multiple actions.
[0090] Then the loss function is expressed as: (1) (2) (3) Among them, the above formula (3) is used to update the parameters of the target network, represents a weighting coefficient; represents the network parameters of the action-value network; Represent the network parameters of the target network. Usually, the update speed of the network parameters of the target network can be controlled by controlling the weighting coefficient.
[0091] The above formula (2) is the Bellman residual function, where, represents the security information; represents the second cumulative reward expectation; represents the first cumulative reward expectation; represents taking the minimum value of the second cumulative reward expectations obtained from the first target network and the second target network respectively; represents the reward; represents the weight parameter, which is a preset constant.
[0092] The above formula (1) is the objective function. Where, represents the first cumulative reward expectation; when there are two action value networks, it can be the mean, or the maximum value or the minimum value of the first cumulative reward expectations corresponding to the first action value network and the second action value network respectively. represents under the policy when the environmental state is execute the action probability.
[0093] In another embodiment of the present disclosure, in order to be able to flexibly control the influence of the security information on the training process, in the embodiments of the present disclosure, it may further include: determining the safety entropy weight parameter of the current training cycle; the safety entropy weight parameter is used to adjust the weight of the security information in the target interaction information when updating the action value network; when the current training cycle is the first training cycle, the safety entropy weight parameter of the current training cycle is a preset parameter; when the current training cycle is other training cycles except the first training cycle, the safety entropy weight parameter of the current training cycle is determined based on the security information in the target interaction data determined in the previous training cycle; In this case, during the process of training the policy model, for example, based on the target interaction data and the safety entropy weight parameter, the action value network of the current training cycle is updated by minimizing the Bellman residual, the policy model of the current training cycle is updated by maximizing the objective function, and the target network of the current training cycle is updated by a slow update mechanism.
[0094] In the embodiments of the present disclosure, the adaptive safety entropy represents the safety expectation value S of the policy performing different actions in the environmental state s, and is used to measure the safety of the policy. For example, it is defined as:
[0095] When determining the safety entropy temperature coefficient of the current training cycle, the optimization objective is the following formula (4): (4) represents the target safety entropy; it is a pre-set constant. represents the safety entropy weight parameter.
[0096] Taking this formula as the loss function, update it in the way of gradient descent That is, in the initial stage of training, the adaptive safety entropy reduces the access frequency to dangerous states through the action probability of uniform distribution; as the training progresses, the adaptive safety entropy dynamically adjusts according to the learning of external safety knowledge by the policy, increasing the access frequency to safe states; by automatically adjusting the safety entropy weight parameter, dynamically balance exploration and security.
[0097] In this case, when updating the action value network, the above formula (2) can be expressed as, for example:[[]]END]]
[0098] Thus, by adaptively adjusting the above safety entropy weight parameter based on the above formula (4) in different training cycles, dynamically control the weight of safety information when updating the action value network, so as to dynamically balance exploration and security.
[0099] In another embodiment of the present disclosure, in the interaction data, it may further include: a completion indication signal .
[0100] The completion indication signal is used to indicate the task execution status corresponding to the policy network after executing the first action in the first environmental state; wherein, the task execution status includes: there may be other actions to execute after executing the action; or, there is no other action to execute after executing the action. If the task execution status indicates that there may be subsequent actions to execute, the completion indication signal is "false", that is, not completed; if the task execution status indicates that no other actions can be executed subsequently, such as in the Go game scenario, "failure" or "victory" are regarded as not being able to execute other actions; in the autonomous driving scenario, "a collision occurred", "the destination of autonomous driving was reached", "the destination was not reached within the specified time", all are regarded as not being able to execute subsequent other actions. At this time, the completion indication signal can be set to: "true".
[0101] In the case where the task execution status indicates that no subsequent other actions can be performed, that is in the case of, the updating of the action value network of the current training cycle by minimizing the Bellman residual may include, for example: Using the action value network of the current training cycle, process the first environmental state and the first action in the target interaction data to obtain the first expected cumulative reward; substitute the first expected cumulative reward and the safety indication information in the target interaction data into the pre-constructed Bellman residual function, and update the action value network of the current training cycle with the goal of minimizing the Bellman residual.
[0102] Exemplarily, in this case, the formula (2) can be expressed as follows:
[0103] That is, since there are no other subsequent actions that can be executed, there will be no corresponding second environmental state at the next moment, and there will also be no reward corresponding to the current moment.
[0104] See Figure 3 As shown, the embodiments of the present disclosure also provide a specific example of training a decision-making model, including: (1): Preparation stage in advance: Pre-train a large language model, which is used to generate specific binary safety feedback signals, that is, safety signals, according to the environmental state and the corresponding action.
[0105] (2) Training stage: S301: Initialize the policy network to be trained and two action value networks and and two target networks and and the experience replay pool B.
[0106] Among them, and have exactly the same parameters, and have the same parameters, and have the same model structure but different internal parameters.
[0107] S302: Execute the k-th round of training, and determine the first environmental state in the k-th round of training , represents the environmental state at a certain moment. Input the first environmental state into the policy model to be trained to obtain the first action .
[0108] S303: Use the first environmental state and the corresponding first action , and interact with the environment to obtain the second environmental state at the next moment , Reward , and completion indication signal .
[0109] S304: Convert the first environmental state and the first action into a natural language description, and input the current environmental state and action of the natural language description into a pre-trained large language model to obtain a safety indication signal .
[0110] S305: Based on the first environmental state , the first action , the second environmental state , the reward , the completion indication signal , and the safety signal to form interaction data , store the interaction data in the experience replay pool B.
[0111] S306: Randomly sample from the experience replay pool B to obtain target interaction data b corresponding to the first round of training. The target interaction data b includes at least one set of interaction data.
[0112] S307: Use the target interaction data and the safety entropy temperature coefficient to update two action value networks by minimizing the Bellman residual and , update the policy network by maximizing the objective function , and update two target networks through a slow update mechanism and .
[0113] Where the loss function is expressed as:
[0114]
[0115] S308: Update the safety entropy temperature coefficient according to the adaptive safety entropy mechanism using the target interaction data b , and return to S302 for the next round of training.
[0116] The training method of the policy model provided by the embodiments of the present disclosure has the following beneficial effects: Large language model feedback mechanism: Through the large language model (LLM), the safety of the state-action pair is evaluated, and a densified safety constraint signal is generated. Using its rich prior knowledge, the large language model can quickly adapt to different task scenarios and provide accurate safety feedback.
[0117] Adaptive Safety Entropy Mechanism: Introduce Adaptive Safety Entropy as a security measurement index to dynamically adjust the security of the policy in different states. In the initial stage of training, Adaptive Safety Entropy reduces the access frequency to dangerous states through uniformly distributed action probabilities, and dynamically adjusts according to the learning of the policy in the later stage of training to increase the access frequency to safe states.
[0118] Off-policy Training Method: Adopt the Off-policy training method, store historical interaction data through an experience replay pool to improve data utilization efficiency. Combine the dual-action value network and the target network architecture to reduce fluctuations during training.
[0119] Optimization Objective and Convergence: Transform the security constraint problem into an unconstrained optimization problem through the Lagrange multiplier method, and use the policy gradient method to optimize the policy to accelerate the optimization and convergence of the objective.
[0120] Corresponding to the embodiment of the training method of the foregoing policy model, the present application also provides an embodiment of a training device for the policy model.
[0121] The embodiment of the training device for the policy model of the present application can be applied to a computer device. The device embodiment can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of the computer device where it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. From the hardware level, as Figure 4 shown, it is a hardware structure diagram of the computer device where the training device for the policy model of the present application is located. In addition to Figure 4 the processor, memory, network interface, and non-volatile memory shown, the computer device where the device is located in the embodiment usually also includes other hardware according to the actual functions of the training device for the policy model, which will not be elaborated here.
[0122] Please refer to Figure 5 . The training device for the policy model provided by the embodiment of the present disclosure includes: An acquisition module 51, configured to acquire a first environmental state, and input the first environmental state into a policy model to be trained, so as to obtain a first action corresponding to the first environmental state; A processing module 52, configured to process the first environmental state and the first action by using a pre-trained large language model to obtain a safety indication signal corresponding to the first action; the safety indication signal is used to indicate whether it is safe to execute the first action in the first environmental state; and Interact with the environment based on the first action to obtain a second environmental state and a reward; A training module 53, configured to construct interaction data based on the first environmental state, the first action, the safety indication signal, the second environmental state, and the reward, and train the to-be-trained policy model based on the interaction data to obtain a target policy model.
[0123] Optionally, when the processing module 52 processes the first environmental state and the first action using a pre-trained large language model to obtain a safety indication signal corresponding to the first action, it is configured to: Convert the first environmental state and the first action into a natural language description; Input the first environmental state and the first action described in natural language into the large language model to obtain the safety indication signal.
[0124] Optionally, when the training module 53 trains the to-be-trained policy model based on the interaction data, it is configured to: Use the interaction data generated in the current training cycle as target interaction data, and based on the target interaction data, train the to-be-trained policy model for the current training cycle; and / or, Store the interaction data determined in the current training cycle into an experience replay pool; Sample the interaction data stored in the experience replay pool to obtain target interaction data corresponding to the current training cycle; based on the target interaction data, train the to-be-trained policy model for the current training cycle.
[0125] Optionally, when the training module 53 trains the to-be-trained policy model for the current training cycle based on the target interaction data, it is configured to: Determine the to-be-trained policy model, action value network, and target network for the current training cycle; wherein, the to-be-trained policy model, action value network, and target network for the current training cycle are respectively an initialized policy model, an initialized action value network, and an initialized target network; and the network architectures and internal network parameters of the initialized action value network and the initialized target network are the same; or, they are the to-be-trained policy model updated in the previous training cycle of the current training cycle, the action value network updated in the previous training cycle, and the target network updated in the previous training cycle; Based on the target interaction data, update the action value network for the current training cycle by minimizing the Bellman residual, update the policy model for the current training cycle by maximizing the objective function, and update the target network for the current training cycle through a slow update mechanism.
[0126] Optionally, the action value network includes: a first action value network and a second action value network; the target network includes: a first target network and a second target network; the first action value network and the second action value network have the same network architecture but different network parameters; the initialized first action value network and the initialized first target network have the same network architecture and network parameters; the initialized second action value network and the initialized second target network have the same network parameters.
[0127] Optionally, when the training module 53 updates the action value network of the current training cycle by minimizing the Bellman residual, updates the policy model of the current training cycle by maximizing the objective function, and updates the target network of the current training cycle through a slow update mechanism, based on the target interaction data, it is used for: Performing a weighted sum processing on the network parameters of the target network of the current training cycle and the network parameters of the action value network of the current training cycle based on a preset weight to obtain the target network parameters of the target network in the current training cycle; and updating the target network of the current training cycle based on the target network parameters; Using the action value network of the current training cycle to process the first environmental state and the first action in the target interaction data to obtain a first cumulative reward expectation; and using the target network of the current training cycle or the updated target network to process the second environmental state and the second action corresponding to the second environmental state in the target interaction data to obtain a second cumulative reward expectation; substituting the first cumulative reward expectation, the second cumulative reward expectation, the safety indication information in the target interaction data, and the reward into a pre-constructed Bellman residual function, and updating the action value network of the current training cycle with the goal of minimizing the Bellman residual; Substituting the first cumulative reward expectation into a pre-constructed objective function, and updating the policy model of the current training cycle with the goal of maximizing the objective function.
[0128] Optionally, the interaction data further includes a completion indication signal; the completion indication signal is used to indicate the task execution status corresponding to the policy network after performing the first action in the first environmental state; When the task execution status indicates that no other actions can be performed, when the training module 53 updates the action value network of the current training cycle by minimizing the Bellman residual, it is used for: Using the action value network of the current training cycle, process the first environmental state and the first action in the target interaction data to obtain the first expected cumulative reward; substitute the first expected cumulative reward and the safety indication information in the target interaction data into the pre-constructed Bellman residual function, and update the action value network of the current training cycle with the goal of minimizing the Bellman residual.
[0129] Optionally, the determining module 51 is further configured to: determine the safety entropy weight parameter of the current training cycle; the safety entropy weight parameter is used to adjust the weight of the safety information in the target interaction information when updating the action value network; the safety entropy weight parameter of the current training cycle is a preset parameter, or is determined based on the safety information in the target interaction data of the previous training cycle; the training module, when updating the action value network of the current training cycle by minimizing the Bellman residual, is configured to: Based on the safety entropy weight parameter, update the action value network of the current training cycle by minimizing the Bellman residual.
[0130] The implementation processes of the functions and roles of each unit in the above device are specifically described in detail in the implementation processes of the corresponding steps in the above method, and will not be repeated here.
[0131] The embodiments of the present disclosure further provide a computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, it executes the steps of the training method of the policy model described in the above method embodiments. Among them, the storage medium may be a volatile or non-volatile computer-readable storage medium.
[0132] The embodiments of the present disclosure further provide a computer program product, which carries program codes, and the instructions included in the program codes can be used to execute the steps of the training method of the policy model described in the above method embodiments. For details, refer to the above method embodiments, and will not be repeated here.
[0133] Among them, the above computer program product can be specifically implemented in a manner of hardware, software, or a combination thereof. In an optional embodiment, the computer program product is specifically embodied as a computer storage medium. In another optional embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0134] The computer program or instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer program or instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium may be any available medium accessible by a computer or a data storage device such as a server or data center integrating one or more available media. The available medium may be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it may also be an optical medium, such as a digital video disc; or it may be a semiconductor medium, such as a solid-state drive. The computer-readable storage medium may be a volatile or non-volatile storage medium, or may include both volatile and non-volatile types of storage media.
[0135] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial descriptions of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this application. A person of ordinary skill in the art can understand and implement it without creative efforts.
[0136] The above are only the preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included within the scope of protection of this application.
Claims
1. A training method for a strategy model, characterized in that, The method includes: Obtaining a first environmental state and inputting the first environmental state into a policy model to be trained, to obtain a first action corresponding to the first environmental state; Processing the first environmental state and the first action by using a pre-trained large language model, to obtain a safety indication signal corresponding to the first action; the safety indication signal is used to indicate whether it is safe to execute the first action in the first environmental state; and Interacting with the environment based on the first action, to obtain a second environmental state and a reward; Constructing interaction data based on the first environmental state, the first action, the safety indication signal, the second environmental state, and the reward, and training the policy model to be trained based on the interaction data, to obtain a target policy model.
2. The training method according to claim 1, wherein The processing the first environmental state and the first action by using a pre-trained large language model, to obtain a safety indication signal corresponding to the first action, includes: Converting the first environmental state and the first action into natural language descriptions; Inputting the natural language descriptions of the first environmental state and the first action into the large language model, to obtain the safety indication signal.
3. The method according to claim 1 or 2, characterized in that, The training the policy model to be trained based on the interaction data includes: Taking the interaction data generated in the current training cycle as target interaction data, and training the policy model to be trained in the current training cycle based on the target interaction data; and / or, Storing the interaction data determined in the current training cycle into an experience replay pool; Sampling the interaction data stored in the experience replay pool, to obtain target interaction data corresponding to the current training cycle; and training the policy model to be trained in the current training cycle based on the target interaction data.
4. The method according to claim 3, wherein The training the policy model to be trained in the current training cycle based on the target interaction data includes: Determining the policy model to be trained, an action value network, and a target network in the current training cycle; wherein, the policy model to be trained, the action value network, and the target network in the current training cycle are respectively an initialized policy model, an initialized action value network, and an initialized target network; and the network architectures and network internal parameters of the initialized action value network and the initialized target network are the same; or, the policy model to be trained, the action value network, and the target network in the current training cycle are respectively the policy model to be trained updated in the previous training cycle of the current training cycle, the action value network updated in the previous training cycle, and the target network updated in the previous training cycle; Updating the action value network in the current training cycle by minimizing the Bellman residual based on the target interaction data, updating the policy model in the current training cycle by maximizing the objective function, and updating the target network in the current training cycle by a slow update mechanism.
5. The method according to claim 4, wherein The action value network includes: a first action value network and a second action value network; the target network includes: a first target network and a second target network; the first action value network and the second action value network have the same network architecture but different network internal parameters; the initialized first action value network and the initialized first target network have the same network architecture and network internal parameters; the initialized second action value network and the initialized second target network have the same network internal parameters.
6. The method according to claim 4 or 5, characterized in that Based on the target interaction data, update the action value network of the current training cycle by minimizing the Bellman residual, update the policy model of the current training cycle by maximizing the objective function, and update the target network of the current training cycle through a slow update mechanism, including: Based on a preset weight, perform a weighted sum processing on the network parameters of the target network in the current training cycle and the network parameters of the action value network in the current training cycle to obtain the target network parameters of the target network in the current training cycle; update the target network in the current training cycle based on the target network parameters; Use the action value network of the current training cycle to process the first environmental state and the first action in the target interaction data to obtain a first cumulative reward expectation; and use the target network or the updated target network in the current training cycle to process the second environmental state and the second action corresponding to the second environmental state in the target interaction data to obtain a second cumulative reward expectation; substitute the first cumulative reward expectation, the second cumulative reward expectation, the safety indication information in the target interaction data, and the reward into a pre-constructed Bellman residual function, and update the action value network of the current training cycle with the goal of minimizing the Bellman residual; Substitute the first cumulative reward expectation into a pre-constructed objective function, and update the policy model of the current training cycle with the goal of maximizing the objective function.
7. The method according to claim 6, characterized in that, In the interaction data, there is also a completion indication signal; the completion indication signal is used to indicate the task execution status corresponding to the policy network after performing the first action in the first environmental state; In the case where the task execution status indicates that no other actions can be performed, the updating of the action value network of the current training cycle by minimizing the Bellman residual includes: Use the action value network of the current training cycle to process the first environmental state and the first action in the target interaction data to obtain a first cumulative reward expectation; Substitute the first cumulative reward expectation and the safety indication information in the target interaction data into a pre-constructed Bellman residual function, and update the action value network of the current training cycle with the goal of minimizing the Bellman residual.
8. The method according to claim 6, wherein The method further includes: Determine the safety entropy weight parameter of the current training cycle; the safety entropy weight parameter is used to adjust the weight of the safety information in the target interaction information when updating the action value network; the safety entropy weight parameter of the current training cycle is a preset parameter, or is determined based on the safety information in the target interaction data of the previous training cycle; Updating the action value network of the current training cycle by minimizing the Bellman residual, including: Based on the safety entropy weight parameter, updating the action value network of the current training cycle by minimizing the Bellman residual.
9. A training device for a strategy model, characterized in that, The device includes: An acquisition module, configured to acquire a first environmental state and input the first environmental state into a policy model to be trained, so as to obtain a first action corresponding to the first environmental state; A processing module, configured to process the first environmental state and the first action by using a pre-trained large language model to obtain a safety indication signal corresponding to the first action; the safety indication signal is used to indicate whether it is safe to execute the first action in the first environmental state; and Interact with the environment based on the first action to obtain a second environmental state and a reward; A training module, configured to construct interaction data based on the first environmental state, the first action, the safety indication signal, the second environmental state, and the reward, and train the policy model to be trained based on the interaction data to obtain a target policy model.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, the steps of the method according to any one of claims 1-8 are implemented.
11. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the following steps are implemented: acquiring a first environmental state and inputting the first environmental state into a policy model to be trained to obtain a first action corresponding to the first environmental state; Processing the first environmental state and the first action by using a pre-trained large language model to obtain a safety indication signal corresponding to the first action; The safety indication signal is used to indicate whether it is safe to execute the first action in the first environmental state; and Interact with the environment based on the first action to obtain a second environmental state and a reward; Construct interaction data based on the first environmental state, the first action, the safety indication signal, the second environmental state, and the reward, and train the policy model to be trained based on the interaction data to obtain a target policy model.
Citation Information
Patent Citations
Safety reinforcement learning and safety control method and device, intelligent agent and storage medium
CN116415651A
Multi-level human intelligence enhanced automatic driving vehicle decision control method and system
CN117227758A
Large language model interaction method and device
CN117236416A
Unmanned system safety decision-making method and system based on action constraint and medium
CN118550194A
Edge server task cache model training and caching method and device
CN118567807A