Single-agent decision-making method and device based on deep reinforcement learning
Through the deep reinforcement learning method, the decision model of the agent is trained, the problem of optimal decision-making of the agent in complex environments is solved, and efficient action decision-making in multi-agent systems is realized.
Patent Information
- Application Number
- CN202411925548.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-27
AI Technical Summary
In complex environments, how to provide optimal action decision strategies for different types of agents, considering their own attributes and functions, especially when taking on different tasks in multiagent systems.
A single agent decision-making method based on deep reinforcement learning is adopted. By obtaining the agent sample data in the experience replay cache pool, the estimated network and the target network are trained, and the policy gradient and loss function are calculated. If the preset conditions are met, the corresponding network model is determined as the agent's decision model.
It realizes the optimal decision-making for single agents in complex environments, solves the problem of the correlation of neural network parameters updates in the actor-criticist structure, and is suitable for decision-making of continuous actions.
Smart Images

Figure CN120046643A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning, and particularly relates to a single-agent decision-making method and device based on deep reinforcement learning. Background Art
[0002] Based on the relevant fundamentals of deep learning and reinforcement learning, design and analysis are carried out for single-agent decision-making. Combining with the actual situation of the underwater environment, an optimal decision-making scheme for a single agent in a complex environment is constructed starting from the properties of the agent itself. Different types of agents have different detection or strike functions and undertake different tasks in a multi-agent system. For each agent, considering its own state information such as position and speed, and attribute information such as detection range, it is necessary to find an optimal action decision-making strategy that can achieve the expected goal. In view of different types of agents and the complexity of tasks, starting from the properties and functions of each agent itself, how to provide a method for generating an optimal decision for a single agent in complex tasks has become a technical problem that urgently needs to be solved in this field. Summary of the Invention
[0003] The purpose of the present invention is to provide a single-agent decision-making method and device based on deep reinforcement learning.
[0004] According to a first aspect of the present invention, there is provided a single-agent decision-making based on deep reinforcement learning, including:
[0005] Obtaining agent sample data in an experience replay buffer pool, where the sample data at least includes the observation state of the agent in the environment, the next observation state, the decision-making action, and the reward value;
[0006] Inputting the agent sample data into an estimation network and a target network respectively for training; the estimation network includes a policy network and a value network, and the target network includes a target policy network and a target value network;
[0007] Calculating the policy gradient of the estimation network and the loss function of the target network respectively;
[0008] If the policy gradient and the loss function meet a preset condition, the network models corresponding to the policy gradient and the loss function are determined as the decision-making model of the agent.
[0009] Optionally, the policy network is used for action decision-making, and the value network is used to evaluate the value of the action;
[0010] The target policy network has the same structure as the policy network, and the target value network has the same structure as the value network.
[0011] Optionally, the experience replay buffer pool is obtained by the following method:
[0012] Output a decision-making action corresponding to the current observation state according to the current observation state of the agent in the environment and the policy network;
[0013] Generate an execution action of the agent according to the decision-making action and Gaussian noise;
[0014] Generate a corresponding reward value according to the execution action of the agent;
[0015] Store the observation state, next observation state, decision-making action and reward value of the agent in the environment in the experience replay cache pool.
[0016] Optionally, the method further includes:
[0017] Update the evaluation network according to the sampled experience data in the experience replay cache pool;
[0018] Specifically include:
[0019] Train the evaluation network through forward propagation, with the input being the agent state and action pair, and the output being the evaluation value of the action;
[0020] Calculate the evaluation values of the next state and next action using the Bellman equation, and update the evaluation network using the TD mean square error.
[0021] Optionally, the method further includes:
[0022] Update the policy network according to the sampled experience data in the experience replay cache pool;
[0023] Specifically include:
[0024] Train the policy network through forward propagation, with the input being the agent state and the output being the action executed by the agent;
[0025] Maximize the evaluation value of the action using the gradient ascent method to update the policy network.
[0026] Optionally, the method further includes:
[0027] The target evaluation network and the target policy network are updated by copying the parameters of the evaluation network and the policy network respectively at preset time intervals.
[0028] According to the second aspect of the present invention, a single-agent decision-making device based on deep reinforcement learning is provided, including:
[0029] An acquisition module, configured to acquire agent sample data in an experience replay cache pool, where the sample data at least includes the observation state, next observation state, decision-making action and reward value of the agent in the environment;
[0030] A training module for training by respectively inputting the agent sample data into an estimation network and a target network; the estimation network includes a policy network and a value network, and the target network includes a target policy network and a target value network;
[0031] A calculation module for respectively calculating the policy gradient of the estimation network and the loss function of the target network;
[0032] A generation module for, if the policy gradient and the loss function meet a preset condition, determining the network models corresponding to the policy gradient and the loss function as the decision-making model of the agent.
[0033] Optionally, the policy network is used for making action decisions, and the value network is used for evaluating the value of actions;
[0034] The target policy network has the same structure as the policy network, and the target value network has the same structure as the value network.
[0035] Optionally, the experience replay buffer pool is obtained in the following manner:
[0036] According to the current observation state of the agent in the environment and the policy network, outputting a decision-making action corresponding to the current observation state;
[0037] Generating an execution action of the agent according to the decision-making action and Gaussian noise;
[0038] Generating a corresponding reward value according to the execution action of the agent;
[0039] Storing the observation state, the next observation state, the decision-making action, and the reward value of the agent in the environment in the experience replay buffer pool.
[0040] Optionally, the generation module is used for:
[0041] Updating the value network according to the sampled experience data in the experience replay buffer pool;
[0042] Specifically including:
[0043] Training the value network through forward propagation, with the input being the agent state and action pair, and the output being the evaluation value of the action;
[0044] Calculating the evaluation values of the next state and the next action using the Bellman equation, and updating the value network using the TD mean squared error.
[0045] Optionally, the generation module is used for:
[0046] Update the policy network according to the sampled experience data in the experience replay buffer pool;
[0047] Specifically, it includes:
[0048] Train the policy network through forward propagation, with the input being the agent state and the output being the action executed by the agent;
[0049] Use the gradient ascent method to maximize the evaluation value of the action and update the policy network.
[0050] Optionally, the generation module is used for:
[0051] The target evaluation network and the target policy network are updated by copying the parameters of the evaluation network and the policy network respectively at preset time intervals.
[0052] In a third aspect, the present application discloses an electronic device, which includes: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the method described in any of the above aspects.
[0053] In a fourth aspect, the present application discloses a non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by the processor of an electronic device, enabling the electronic device to execute the method described in any of the above aspects.
[0054] In a fifth aspect, the present application discloses a computer program product, when the instructions in the computer program product are executed by the processor of an electronic device, enabling the electronic device to execute the method described in any of the above aspects.
[0055] The beneficial effects brought by the present invention are as follows:
[0056] As can be seen from the above solution, the embodiments of the present invention provide a single-agent decision-making method and device based on deep reinforcement learning, including: obtaining agent sample data in an experience replay buffer pool, where the sample data at least includes the observation state, next observation state, decision-making action, and reward value of the agent in the environment; inputting the agent sample data into an estimation network and a target network respectively for training; the estimation network includes a policy network and an evaluation network, and the target network includes a target policy network and a target evaluation network; calculating the policy gradient of the estimation network and the loss function of the target network respectively; if the policy gradient and the loss function meet preset conditions, the network models corresponding to the policy gradient and the loss function are determined as the decision-making models of the agent, adopting an actor-critic reinforcement learning architecture, which solves the problem of correlation between neural networks before and after each parameter update in the actor-critic structure, and also solves the drawback that deep reinforcement learning cannot be used for continuous actions. Description of the Drawings
[0057] Figure 1 It is a schematic flow chart of a single-agent decision-making method based on deep reinforcement learning provided according to an embodiment;
[0058] Figure 2 It is a structure diagram of a policy network provided according to an embodiment;
[0059] Figure 3 It is a structure diagram of an evaluation network provided according to an embodiment;
[0060] Figure 4 It is a framework of a deep reinforcement learning method based on actor-critic provided according to an embodiment;
[0061] Figure 5 It is a process of generating empirical data provided according to an embodiment;
[0062] Figure 6 It is a training and updating process of an evaluation network provided according to an embodiment;
[0063] Figure 7 It is a training and updating process of a policy network provided according to an embodiment;
[0064] Figure 8 It is a block diagram of a single-agent decision-making device based on deep reinforcement learning of the present application.
[0065] Figure 9 It is a block diagram of an electronic device of the present application.
[0066] Figure 10 It is a block diagram of a computer-readable storage medium of the present application. Detailed Embodiments
[0067] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0068] Agents of different types have different detection or strike functions and undertake different tasks in a multi-agent system. For each agent, considering its own state information such as position and speed, and attribute information such as detection range, it searches for an optimal action decision-making strategy that can achieve the expected goal. In view of different types of agents and the complexity of tasks, starting from the attributes and functions of each agent itself, a method for generating the optimal decision of a single agent in complex tasks is studied.
[0069] Reference Figure 1 , a step flowchart of a single-agent decision-making method based on deep reinforcement learning according to the present application is shown. This method can be applied to an electronic device. Specifically, the method may include the following steps:
[0070] S101. Obtain agent sample data from the experience replay buffer pool. The sample data at least includes the observation state of the agent in the environment, the next observation state, the decision-making action, and the reward value;
[0071] S102. Input the agent sample data into the estimation network and the target network for training respectively. The estimation network includes a policy network and a value network, and the target network includes a target policy network and a target value network;
[0072] S103. Calculate the policy gradient of the estimation network and the loss function of the target network respectively;
[0073] S104. If the policy gradient and the loss function meet the preset conditions, the network models corresponding to the policy gradient and the loss function are determined as the decision-making models of the agent.
[0074] In the process of multi-agent game confrontation, each agent has independent decision-making ability. The deep deterministic policy gradient method is used to implement the action decision of a single agent, and the actor-critic reinforcement learning architecture is adopted. It solves the problem that there is a correlation between the neural network before and after each parameter update in the actor-critic structure, and also solves the disadvantage that deep reinforcement learning cannot be used for continuous actions.
[0075] Another embodiment of the present application further supplements and explains the single-agent decision-making method based on deep reinforcement learning provided in the above embodiment.
[0076] Based on the above agent action decision problem model, the detailed designs of the state space, action space, and establishment function of both platforms are as follows:
[0077] Platform W includes a state space, an action space, and a reward function. In the game confrontation system, the state space of Platform W is assumed to be represented as:
[0078] s i ={pos i , vel i , objective_pos i , detect_pos i , detect_time i}
[0079] Among them, pos i , vel i, objective_pos i , detect_pos i and detect_time i respectively represent the position, velocity, target point position of agent i, the position information of detecting platform D, and the number of consecutive detections of platform D. The number of consecutive detections of platform D in the state space is used to determine whether the position of platform D can be accurately located, thereby determining whether SN is in the active mode or the passive mode, and cooperating with the detection of platform W and DJ.
[0080] The action space is the set of all executable actions of the agent. Suppose the action space of platform W is represented as:
[0081] a i = {delta_vel i , sonar i , deep i}
[0082] Among them, delta_vel i , sonar i , deep i respectively represent the speed variable value of agent i, the working mode of SN, and the working depth of the platform. The speed variable value is the change amount of speed in the two-dimensional plane, including the heading and the speed. In this study, only the action interfaces for the working mode of SN and the working depth of the platform are reserved for future function expansion.
[0083] The reward function plays a guiding role in the action decision-making process, helping the agent obtain the optimal strategy. The reward function usually consists of two parts: state reward and action reward. Among them, the state reward is used to measure the value of the agent's current state, reflecting the degree of goal achievement. The action reward is used to measure the value of the agent taking a certain action, reflecting whether the action is beneficial to achieving the goal.
[0084] The reward function of platform W guides its detection, avoidance of platform D, and movement towards the target point, and is specifically expressed as:
[0085]
[0086] Among them, r 到达 and r 探测 respectively represent the rewards obtained for reaching and detecting platform D, and should be positive values; r 被探测 and r 被DJ respectively represent the penalties obtained for being detected by platform D and DJ, and should be negative values
[0087] Compared with platform W, platform D does not have the position information of the target point, and its state space is represented as:
[0088] si = {pos i , vel i , detect_pos i , notdetect_time i}
[0089] Among them, detect_pos i represents the position information of the platform W detected by the agent i, and notdetect_time i represents the number of consecutive times the platform W has not been detected, thereby determining the SN working mode.
[0090] (2) Action space:
[0091] The action space of the platform D is the same as that of the platform W, which is expressed as:
[0092]
[0093] Among them, delta_vel i , sonar i , deep i respectively represent the speed variable value of the agent i, the SN working mode, and the working depth of the platform.
[0094] (3) Reward function:
[0095] The reward function of the platform D guides it to avoid being detected by the platform W, while detecting and docking with the platform W. Specifically, it is expressed as:
[0096]
[0097] Among them, r 探测 and r DJ respectively represent the rewards obtained by the platform D for discovering and successfully docking with the platform W, and should be positive values; r 被探测 and r 被DJ respectively represent the penalties obtained by the platform D for being detected and docked by the platform W, and should be negative values.
[0098] For each agent, there are two network structures: a policy network and a value network. The policy network is used to learn the action decision with the highest evaluation value, while the value network is used to learn the true evaluation of the environment for the agent's actions. In addition, to avoid overestimation of the network, a target value network and a target policy network are introduced. The policy network and the value network are updated based on the Bellman equation, and the target policy network and the target value network are updated using the soft-copy method.
[0099] Both the policy network and the value network use the following components:
[0100] (1) Fully connected layer
[0101] In a fully connected layer (FC), each neuron is fully connected to all neurons in the previous layer, which can map the distributed feature representations learned by the network to the sample label space. The fully connected layer can be regarded as the simplest type of Convolutional Neural Network (CNN). The output of the hidden layer usually connects to the Relu activation function, and the output of the last fully connected layer generally outputs the result through the sigmoid function or the tanh function. All neural networks in this project adopt the fully connected method.
[0102] (2) Relu activation function
[0103] The Rectified Linear Unit (ReLU), also known as the rectified linear unit, is a non-saturating activation function that is not prone to gradient vanishing. Most hidden layers of forward neural networks use this activation function. The mathematical representation of the ReLU activation function is as follows:
[0104] ReLU(x) = max(0, x)
[0105] (3) Tanh activation function
[0106] The Tanh activation function, also known as the hyperbolic tangent function, is a saturating activation function with a value range of [-1, 1]. Its mathematical representation is:
[0107]
[0108] Optionally, the policy network is used for action decision-making, and the evaluation network is used to evaluate the value of the action;
[0109] The target policy network has the same structure as the policy network, and the target evaluation network has the same structure as the evaluation network.
[0110] The policy network is used to train the agent's action decision-making in a certain state. The input is the agent's state, and the output is the action. The evaluation network is used to train the evaluation function to evaluate the state and action, so that the evaluation is close to the true evaluation of the environment. The input is the combination of the state and action, and the output is the evaluation value of this combination. Moreover, the target policy network and the policy network have exactly the same structure, and the target evaluation network and the rating network have exactly the same structure. The specific structures of the two networks are as Figure 2 and Figure 3 shown.
[0111] The actor-critic based reinforcement learning method combines the policy gradient method and the value function method, and has the following advantages: 1) It has both a policy network that can generate action decisions and a critic network that evaluates the value of actions. 2) It uses an experience replay memory mechanism. On the one hand, it breaks the temporal correlation of training data, which is more conducive to the training of neural networks. On the other hand, through sampling, it improves the utilization rate of training data. 3) It introduces a target network to avoid overestimation of the neural network. The framework of the actor-critic based deep reinforcement learning method is as Figure 4 shown:
[0112] Optionally, the experience replay buffer pool is obtained in the following way:
[0113] According to the current observation state of the agent in the environment and the policy network, output the decision action corresponding to the current observation state;
[0114] According to the decision action and Gaussian noise, generate the execution action of the agent;
[0115] According to the execution action of the agent, generate the corresponding reward value;
[0116] Store the observation state, the next observation state, the decision action and the reward value of the agent in the environment in the experience replay buffer pool.
[0117] Specifically, each agent has four neural networks, namely: a policy network, a critic network, and the corresponding target policy network and target critic network. Among them, the policy network is used for action decision-making, and the critic network is used to evaluate the value of actions. The target policy network and the target critic network have the same structure as the policy network and the critic network respectively, which can improve the stability of network training.
[0118] (1) Experience Replay Mechanism
[0119] The agent uses the policy network for action decision-making, interacts with the environment continuously to generate a large amount of experience data, and stores each group of experience data in the buffer respectively. The generation process of the experience data is as Figure 5 shown. The policy network is represented as μ(s|θ μ ), takes the observation state s of the agent in the environment as input, passes through the policy network, and outputs the corresponding decision action a. In order to improve the agent's exploration ability of the environment and its search ability for the global optimal solution, Gaussian noise is added to the action a to obtain the execution action a of the agent. Then, the environment will return the reward value r according to the agent's action, complete an interaction process between the agent and the environment, and generate a group of experience data (s, a, s’, r), where s’ is the next observation state. Store the generated experience data in the buffer, and realize the training of the neural network by sampling the experience data.
[0120] Optionally, the method further includes:
[0121] Updating the evaluation network according to the sampled experience data in the experience replay buffer pool;
[0122] Specifically, it includes:
[0123] Training the evaluation network through forward propagation, with the input being the agent state and action pair, and the output being the evaluation value of the action;
[0124] Calculating the evaluation values of the next state and the next action using the Bellman equation, and updating the evaluation network using the TD mean squared error.
[0125] Specifically, updating the evaluation network
[0126] The evaluation network is represented as Q(s,a|θ Q ), and a small batch (batch size) of experience data is sampled from the buffer to train and update the network. Taking a set of experience data (s,a,s’,r) as an example, the training and updating process of the evaluation network is as Figure 6 shown. The evaluation network is trained through forward propagation, with the input being the state and action pair (s,a), and the output being the evaluation value Q(s,a) of the action. The true evaluation values of the next state and the next action (s’,a’) are calculated using the Bellman equation, and the evaluation network is updated using the TD mean squared error.
[0127] Loss Critic (θ Q ) = (r + Q(s,a) - Q - (s ′ ,a′)) 2
[0128]
[0129] where Q - (s ′ ,a′) represents the estimated evaluation value of the target evaluation network for the next state s’ and the next action a’, and α Q represents the update step size of the evaluation network.
[0130] Optionally, the method further includes:
[0131] Updating the policy network according to the sampled experience data in the experience replay buffer pool;
[0132] Specifically, it includes:
[0133] Training the policy network through forward propagation, with the input being the agent state, and the output being the action executed by the agent;
[0134] The gradient ascent method is used to maximize the evaluation value of the action, and the policy network is updated.
[0135] Specifically, in the embodiments of the present application, the policy network is updated as follows:
[0136] The policy network is represented as μ(s|θ μ ), which is updated in the same way as the evaluation network, and sampled experience data is used for network training and update. Taking a set of experience data (s, a, s’, r) as an example, the training and update process of the policy network is shown in the following figure. The policy network is trained through forward propagation, with the input being the state s and the output being the action apredict executed by the agent. The gradient ascent method is used to maximize the evaluation value of the action to update the policy network.
[0137] Loss Actor (θ μ )=-Q(s,a predict )
[0138]
[0139] where α μ represents the update step size of the policy network.
[0140] Optionally, the method further includes:
[0141] The target evaluation network and the target policy network are updated by copying the parameters of the evaluation network and the policy network respectively at preset time intervals.
[0142] Specifically, updating the target evaluation network and the target policy network
[0143] The target evaluation network and the target policy network do not update the network parameters through backpropagation, but are updated by copying the parameters of the evaluation network and the policy network respectively at regular intervals.
[0144]
[0145] First, four neural networks are designed for the agent: a policy network, an evaluation network, a target policy network, and a target evaluation network. Second, the policy network is called to generate the execution action for the agent, interact with the environment, and then extract the experience data from the data generated during the interaction and store it in the experience pool. When the storage capacity of the experience pool reaches a certain threshold, the sampled experience data is used to update the evaluation network and the policy network. At the same time, the target evaluation network and the policy network are updated by copying, and the specific algorithm process is as follows:
[0146] Initialize the state of the agent, the policy network, and the evaluation network, and set the number of training cycles M, the length of the training cycle T, and the experience replay buffer D;
[0147] FOR cycle = 1:M:
[0148] Initialize random noise, initialize the environment, the policy network and the evaluation network of the agent;
[0149] FOR time step = 1:T:
[0150] For each agent: According to the current policy and exploration noise, select and execute action a; Each agent executes the corresponding action, obtains the current reward value r, and transitions to the new state s'; Store the experience data (s, a, s', r) in the experience pool D;
[0151] IF the amount of experience data > the preset start training data amount, randomly sample batch size groups of training samples from D.
[0152] Calculate the evaluation network loss value according to the formula, update the evaluation network according to the formula, calculate the policy network loss value according to the formula, update the policy network according to the formula, update the target evaluation network and the target policy network according to the formula, and end the training of this cycle.
[0153] The embodiment of the present invention provides a single-agent decision-making method based on deep reinforcement learning, including: obtaining agent sample data in the experience replay buffer pool, where the sample data at least includes the observation state of the agent in the environment, the next observation state, the decision-making action and the reward value; inputting the agent sample data into the estimation network and the target network respectively for training; the estimation network includes a policy network and an evaluation network, and the target network includes a target policy network and a target evaluation network; calculate the policy gradient of the estimation network and the loss function of the target network respectively; if the policy gradient and the loss function meet the preset conditions, the network models corresponding to the policy gradient and the loss function are determined as the decision-making models of the agent, adopting the actor-critic reinforcement learning architecture, which solves the problem that there is a correlation between the neural network parameters before and after each parameter update in the actor-critic structure, and also solves the shortcoming that deep reinforcement learning cannot be used for continuous actions.
[0154] It should be noted that each implementable manner in this embodiment can be implemented alone, or can be combined in any combination manner without conflict. This application does not make a limitation.
[0155] Another embodiment of the present application provides a single-agent decision-making device based on deep reinforcement learning, which is used to execute the single-agent decision-making method based on deep reinforcement learning provided in the above embodiment.
[0156] Such as Figure 8As shown in the figure, it is a schematic structural diagram of a single-agent decision-making device based on deep reinforcement learning provided by an embodiment of the present application. The single-agent decision-making device based on deep reinforcement learning includes an acquisition module 801, a training module 802, a calculation module 803, and a generation module 804, where:
[0157] The acquisition module 801 is used to acquire the agent sample data in the experience replay buffer pool. Among them, the sample data at least includes the observation state of the agent in the environment, the next observation state, the decision-making action, and the reward value.
[0158] The training module 802 is used to input the agent sample data into the estimation network and the target network for training respectively; the estimation network includes a policy network and an evaluation network, and the target network includes a target policy network and a target evaluation network.
[0159] The calculation module 803 is used to calculate the policy gradient of the estimation network and the loss function of the target network respectively.
[0160] The generation module 804 is used to determine the network model corresponding to the policy gradient and the loss function as the decision-making model of the agent if the policy gradient and the loss function meet the preset conditions.
[0161] Regarding the device in this embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment related to the method, and will not be elaborated here.
[0162] Another embodiment of the present application further supplements the single-agent decision-making device based on deep reinforcement learning provided in the above embodiment.
[0163] Optionally, the policy network is used for action decision-making, and the evaluation network is used to evaluate the value of the action.
[0164] The target policy network has the same structure as the policy network, and the target evaluation network has the same structure as the evaluation network.
[0165] Optionally, the experience replay buffer pool is obtained in the following manner:
[0166] According to the current observation state of the agent in the environment and the policy network, output the decision-making action corresponding to the current observation state.
[0167] According to the decision-making action and Gaussian noise, generate the execution action of the agent.
[0168] According to the execution action of the agent, generate the corresponding reward value.
[0169] Store the observation state of the agent in the environment, the next observation state, the decision-making action, and the reward value in the experience replay buffer pool.
[0170] Optionally, a generation module is configured to:
[0171] Update the evaluation network according to the sampled experience data in the experience replay buffer pool;
[0172] Specifically, it includes:
[0173] Train the evaluation network through forward propagation, with the input being the agent state and action pair, and the output being the evaluation value of the action;
[0174] Calculate the evaluation values of the next state and the next action using the Bellman equation, and update the evaluation network using the TD mean square error.
[0175] Optionally, a generation module is configured to:
[0176] Update the policy network according to the sampled experience data in the experience replay buffer pool;
[0177] Specifically, it includes:
[0178] Train the policy network through forward propagation, with the input being the agent state and the output being the action executed by the agent;
[0179] Maximize the evaluation value of the action using the gradient ascent method to update the policy network.
[0180] Optionally, a generation module is configured to:
[0181] The target evaluation network and the target policy network are updated by copying the parameters of the evaluation network and the policy network respectively at preset time intervals.
[0182] For the apparatus embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For related parts, please refer to the corresponding descriptions in the method embodiments.
[0183] The embodiments of the present invention provide a single-agent decision-making method and apparatus based on deep reinforcement learning, including: obtaining agent sample data in the experience replay buffer pool, where the sample data at least includes the observation state of the agent in the environment, the next observation state, the decision-making action, and the reward value; inputting the agent sample data into the estimation network and the target network respectively for training; the estimation network includes a policy network and an evaluation network, and the target network includes a target policy network and a target evaluation network; calculating the policy gradient of the estimation network and the loss function of the target network respectively; if the policy gradient and the loss function meet the preset conditions, the network models corresponding to the policy gradient and the loss function are determined as the decision-making models of the agent. The actor-critic reinforcement learning architecture is adopted, which solves the problem of the correlation between the neural network parameters before and after each parameter update in the actor-critic structure, and also solves the shortcoming that deep reinforcement learning cannot be used for continuous actions.
[0184] Optionally, an embodiment of the present application further provides an electronic device, including: a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements each process of the foregoing method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0185] An embodiment of the present application further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, it implements each process of the foregoing method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium, such as a read-only memory (ROM for short), a random access memory (RAM for short), a magnetic disk, or an optical disc, etc.
[0186] Figure 9 FIG. 800 is a block diagram of an electronic device 800 shown in the present application. For example, the electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0187] Refer to Figure 9 , the electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0188] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the foregoing method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0189] The memory 804 is configured to store various types of data to support the operation of the device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, images, videos, and the like. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0190] The power supply component 806 provides power to various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.
[0191] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0192] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0193] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power-on button, and a lock button.
[0194] The sensor assembly 814 includes one or more sensors for providing an assessment of various aspects of the electronic device 800. For example, the sensor assembly 814 can detect the on / off state of the device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor assembly 814 can also detect a change in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and a change in the temperature of the electronic device 800. The sensor assembly 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0195] The communication component 816 is configured to facilitate communication between the electronic device 800 and other devices in a wired or wireless manner. The electronic device 800 can access a wireless network based on communication standards, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 5G), or a combination thereof. In an exemplary embodiment, the communication component 816 receives broadcast signals or broadcast operation information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0196] In an exemplary embodiment, the electronic device 800 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0197] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions, such as a memory 804 including instructions, is also provided. The above instructions can be executed by the processor 820 of the electronic device 800 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0198] Figure 10FIG. 0 is a block diagram of a computer-readable storage medium 1900 shown in the present application. For example, the computer-readable storage medium 1900 may be provided as a server.
[0199] Referring to Figure 10 , the computer-readable storage medium 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.
[0200] The computer-readable storage medium 1900 may further include a power component 1926 configured to perform power management of the computer-readable storage medium 1900, a wired or wireless network interface 1950 configured to connect the computer-readable storage medium 1900 to a network, and an input / output (I / O) interface 1958. The computer-readable storage medium 1900 may operate based on an operating system stored in the memory 1932, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM or the like.
[0201] It should be noted that, in this article, the terms "comprising", "including" or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the presence of additional identical elements in the process, method, article or device including the element.
[0202] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which may be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in various embodiments of the present application.
[0203] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.
[0204] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the embodiments of the present application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0205] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0206] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical or other form.
[0207] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0208] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0209] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0210] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art in the technical field disclosed by this application can easily think of changes or substitutions, and all of them should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0211] The above is the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A single-agent decision-making method based on deep reinforcement learning, characterized in that: include: Obtaining agent sample data in the experience replay buffer pool, wherein the sample data at least includes the agent's observation state, next observation state, decision action, and reward value in the environment; The agent sample data is respectively input into the estimation network and the target network for training; the estimation network includes a strategy network and an evaluation network, and the target network includes a target strategy network and a target evaluation network; Calculating the policy gradient of the estimation network and the loss function of the target network respectively; If the policy gradient and the loss function meet preset conditions, the network model corresponding to the policy gradient and the loss function is determined as the decision model of the agent.
2. The single-agent decision-making method based on deep reinforcement learning according to claim 1, characterized in that: The policy network is used to make action decisions, and the evaluation network is used to evaluate the value of actions; The target policy network has the same structure as the policy network, and the target evaluation network has the same structure as the evaluation network.
3. The single-agent decision-making method based on deep reinforcement learning according to claim 1, characterized in that: The experience replay buffer pool is obtained in the following way: According to the current observation state of the agent in the environment and the policy network, output a decision action corresponding to the current observation state; Generate an execution action of the agent according to the decision action and Gaussian noise; Generate a corresponding reward value according to the execution action of the agent; The observed state, next observed state, decision action and reward value of the agent in the environment are stored in the experience replay buffer pool.
4. The single-agent decision-making method based on deep reinforcement learning according to claim 1, characterized in that: The method further comprises: Updating the evaluation network according to the sampled experience data in the experience replay buffer pool; Specifically include: The evaluation network is trained according to the forward propagation, with the input being the agent state and action pair, and the output being the evaluation value of the action; The Bellman equation is used to calculate the evaluation values of the next state and the next action, and the TD mean square error is used to update the evaluation network.
5. The single-agent decision-making method based on deep reinforcement learning according to claim 1, characterized in that: The method further comprises: Updating the policy network according to the sampled experience data in the experience replay buffer pool; Specifically include: The policy network is trained through forward propagation, with the input being the agent state and the output being the action performed by the agent; The gradient ascent method is used to maximize the evaluation value of the action and update the policy network.
6. The single-agent decision-making method based on deep reinforcement learning according to claim 1, characterized in that: The method further comprises: The target evaluation network and the target strategy network are copied in a copying manner, and the parameters of the evaluation network and the strategy network are copied and updated at preset time intervals.
7. A single-agent decision-making device based on deep reinforcement learning, characterized in that: include: An acquisition module is used to acquire agent sample data in the experience replay buffer pool, wherein the sample data at least includes the agent's observation state, next observation state, decision action and reward value in the environment; A training module, used to input the agent sample data into the estimation network and the target network for training respectively; the estimation network includes a strategy network and an evaluation network, and the target network includes a target strategy network and a target evaluation network; A calculation module, used to calculate the policy gradient of the estimation network and the loss function of the target network respectively; A generation module is used to determine the network model corresponding to the policy gradient and the loss function as the decision model of the agent if the policy gradient and the loss function meet preset conditions.
8. The single-agent decision-making device based on deep reinforcement learning according to claim 7, characterized in that: The policy network is used to make action decisions, and the evaluation network is used to evaluate the value of actions; The target policy network has the same structure as the policy network, and the target evaluation network has the same structure as the evaluation network.
9. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the method according to any one of claims 1 to 6 when executed by the processor.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Brain-like reinforcement learning method and system based on hierarchical experience playback
CN121503573A
A brain-like reinforcement learning method and system based on hierarchical experience replay
CN121503573B
Storage medium configuration and processing methods and electronic devices
CN122672720A