Reinforcement Learning Method and Device for Control System Based on Maximum Entropy Projection

By using maximum entropy projection in reinforcement learning to identify the most valuable states and optimize the exploration strategy, the low learning efficiency problem caused by huge state space is solved, and faster training and task completion is achieved.

CN115526335BActive Publication Date: 2025-07-29709TH RESEARCH INSTITUTE CHINA STATE SHIPBUILDING CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211119789.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-15
Publication Date
2025-07-29
Estimated Expiration
2042-09-15

AI Technical Summary

Technical Problem

The existing reinforcement learning methods are difficult to effectively explore states with exploration value when the state space is huge, resulting in low learning efficiency.

Method used

By projecting the state vector of the control system into the maximum entropy space, identifying the most valuable states, combining a linear combination of exploration strategies and action strategies, the training process is optimized to improve exploration efficiency.

Benefits of technology

Speed up training speed, reduce the time of the agent to learn, and enable the control system to complete a given task faster.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115526335B_ABST
    Figure CN115526335B_ABST
Patent Text Reader

Abstract

The present invention discloses a reinforcement learning method for a control system based on maximum entropy projection: initialize model parameters; reset the reinforcement learning environment; at each moment, the agent generates an action according to a linear combination of an exploration policy and an action policy; execute the action in the environment, obtain a reward and a new environmental state, and add them to the training dataset; train and update the action policy; sample a subset for training the exploration policy from the training dataset, calculate the maximum entropy projection matrix, and train and update the exploration policy; after the reinforcement learning environment is executed, if the learning process converges, end the learning, otherwise return to continue learning. The method of the present invention can, by identifying the most exploratory valuable states, encourage the agent to explore these states, improve the exploration efficiency, accelerate the training speed, reduce the learning time of the agent, and enable the control system to start executing and complete the given task faster. The present invention also provides a corresponding reinforcement learning device for a control system based on maximum entropy projection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of reinforcement learning, and more specifically, relates to a reinforcement learning method and device for a control system based on maximum entropy projection. Background Art

[0002] Reinforcement learning (RL), also known as re-inforcement learning, evaluation learning or enhanced learning, is one of the paradigms and methodologies of machine learning, used to describe and solve the problem that an agent maximizes rewards or achieves specific goals by learning strategies during the interaction with the environment. Reinforcement learning is used to control systems facing the problem of incomplete exploration caused by a huge state space. Existing methods are difficult to find states with exploration value in the state space for exploration, which affects the learning efficiency. Summary of the Invention

[0003] In view of the above deficiencies or improvement requirements of the prior art, the present invention provides a reinforcement learning scheme based on maximum entropy projection. By projecting the state vector of the control system into the maximum entropy space, the most valuable state is found, the exploration efficiency of the state space is improved, and the training speed is further increased.

[0004] To achieve the above object, according to one aspect of the present invention, there is provided a reinforcement learning method for a control system based on maximum entropy projection, including the following steps:

[0005] S1 Initialize the model parameters, where the model parameters include an exploration policy parameters and an action policy parameters where the exploration policy represents the probability that, in order to fully explore, when the state is the agent selects the action , and the action policy represents the probability that, in order to obtain the maximum reward, when the state is the agent selects the action ; where the state is a vector composed of the readings of sensors deployed at multiple key positions of the control system, is a vector composed of the control quantities of multiple control units in the control system;

[0006] S2 Reset the reinforcement learning environment;

[0007] S3 At each moment , the agent generates an action according to the linear combination of the exploration policy and the action policy , where the weight indicates whether the current agent is more inclined to explore or to obtain the maximum return;

[0008] At time , the agent executes an action in the control system according to the current state of the control system , obtains a reward , the state of the control system becomes , and is added to the training data set , and the reflects the degree to which the control system correctly executes a given task;

[0009] S5 uses the training data set , and trains and updates the parameters of the action policy ;

[0010] S6 samples a subset from the training data set for training the exploration policy , calculates the maximum entropy projection matrix according to the training sample subset , and trains and updates the exploration policy ;

[0011] S7 checks the convergence condition. If it does not converge, it returns to S3. Otherwise, it ends.

[0012] In one embodiment of the present invention, the step S6 specifically includes:

[0013] S6-1 randomly samples training samples from the training data set to form a subset for training the exploration policy , where is a training sample in

[0014] S6-2 combines the current state vectors contained in all training samples into a state vector matrix , where shows that the th column is the state vector , the number of rows of is the same as the dimension of the state vector

[0015] S6-3 Through the formula Obtain the maximum entropy projection matrix , where is the state matrix after projection, is the projection matrix, Each column in is the state vector after projection by the projection matrix , is the matrix obtained by taking the mean value of row by row;

[0016] After the maximum entropy projection matrix projects the state matrix , each column in it is the state vector after projection. Count the number of times these column vectors appear in the state matrix after projection, find the column vector with the least number of appearances, and its subscript is . Take the th column in the state vector matrix as the exploration target state ;

[0017] S6-5 Construct new training data for training the exploration strategy , where the current state , action , next moment state in the training sample are the same as those in . The method for setting the reward in the training sample is as follows:

[0018]

[0019] S6-6 Use the new training data and adopt the reinforcement learning method to train and update the exploration strategy .

[0020] In an embodiment of the present invention, the method for initializing the parameters of the exploration strategy and the parameters of the action strategy is to generate random numbers with a uniform distribution in the interval.

[0021] In an embodiment of the present invention, step S2 is specifically: initialize the control system to return it to the initial state.

[0022] In one embodiment of the present invention, the reinforcement learning method in step S5 includes SAC (Soft Actor-Critic) or PPO (Proximal Policy Optimization).

[0023] In one embodiment of the present invention, the convergence condition in step S7 is that the number of iterations reaches the set maximum value.

[0024] In one embodiment of the present invention, the convergence condition in step S7 is that the mean square error of the obtained average cumulative return is less than a preset value.

[0025] In one embodiment of the present invention, the control unit in the control system is a motor or a relay.

[0026] In one embodiment of the present invention, the tasks to be completed by the control system are to maintain a certain motion state or move along a given trajectory.

[0027] According to another aspect of the present invention, there is also provided a reinforcement learning device for a control system based on maximum entropy projection, including at least one processor and a memory. The at least one processor and the memory are connected through a data bus. The memory stores instructions executable by the at least one processor. After being executed by the processor, the instructions are used to complete the above-mentioned reinforcement learning method for a control system based on maximum entropy projection.

[0028] Generally speaking, compared with the prior art by the above technical solution conceived by the present invention, the following beneficial effects are achieved:

[0029] The method of the present invention, by identifying the most valuable state for exploration, encourages the agent to explore this state, improves the exploration efficiency, speeds up the training speed, and for the control system, can reduce the learning time of the agent and enable the control system to start executing and complete the given task faster. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 is a flowchart showing the reinforcement learning method for a control system based on maximum entropy projection of the present invention;

[0031] Figure 2 is a flowchart showing a training exploration strategy in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] To make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0033] For a control system, its state is a vector composed of the states (readings) of sensors deployed at multiple key positions, and the input of this control system is a vector composed of the control quantities of multiple control units (such as motors, relays, etc.) in the system. This control system needs to complete a given task, such as maintaining a certain motion state, moving along a given trajectory, etc. The reward obtained after controlling this system is whether the system correctly executes the given task. This solution trains an agent through a reinforcement learning method based on maximum entropy projection, enabling it to autonomously control this system to complete the given task.

[0034] As Figure 1 shown, the present invention provides a reinforcement learning method for a control system based on maximum entropy projection, including:

[0035] S1 Initialize the model parameters, where the model parameters include the exploration policy parameters and the action policy parameters Among them, the exploration policy represents the probability that, in order to fully explore, when the state is the agent selects the action ; the action policy represents the probability that, in order to obtain the maximum reward, when the state is the agent selects the action ; the initialization method is to generate random numbers uniformly distributed in the interval; among them, the state is a vector composed of the readings of sensors deployed at multiple key positions of the control system, is a vector composed of the control quantities of multiple control units in the control system;

[0036] S2 Reset the reinforcement learning environment. In this embodiment, that is, initialize the control system to return it to the initial state;

[0037] S3 At each moment , an agent (hereinafter referred to as the agent) controlled by a reinforcement learning method based on maximum entropy projection provided by the present invention, according to the exploration policy and the action policy Linear combination Generate actions , where the weight indicates whether the current agent is more inclined to explore or to obtain the maximum reward, and the typical value is 0.1, which usually gradually decreases as learning progresses;

[0038] At time S4 , the agent executes an action in the environment (i.e., the control system) according to the current environmental state and obtains a reward . The environmental state (i.e., the state of the control system) changes to . Add to the training data set ; The reflects the degree to which the control system correctly executes a given task;

[0039] That is, if the task is completed well or one step closer to the given goal, a relatively large reward r will be obtained; otherwise, the value of r is small; if the task is not completed well or far from the given goal, a negative r will be obtained;

[0040] In S5, use the training data set and adopt common reinforcement learning methods, such as SAC (Soft Actor-Critic) or PPO (Proximal Policy Optimization), etc., to train and update the parameters of the action policy ; ;

[0041] In S6, sample a subset from the training data set for training the exploration policy . According to the training sample subset , calculate the maximum entropy projection matrix, and train and update the exploration policy ;

[0042] As shown in Figure 2 , the specific steps of S6 include:

[0043] In S6-1, randomly sample training samples from the training data set to form a subset for training the exploration policy s , where is a training sample in;

[0044] S6-2 Combine all the current state vectors contained in the training samples to form a state vector matrix , where the th column is the state vector , the number of rows of is the same as the dimension of the state vector ;

[0045] S6-3 Obtain the maximum entropy projection matrix through the formula , where is the projected state matrix, is the projection matrix, each column in is the state vector projected by the projection matrix , is the matrix obtained by taking the mean value of by row;

[0046] S6-4 For each column in the state matrix projected by the maximum entropy projection matrix , which is the projected state vector, count the number of times these column vectors appear in the projected state matrix . Find the column vector with the least number of occurrences, and its subscript is . Take the th column in the state vector matrix as the exploration target state ;

[0047] S6-5 Construct new training data for training the exploration strategy , where the current state , action , and next moment state in the training samples are the same as those in . The method for setting the reward in the training samples is as follows:

[0048]

[0049] S6-6 Use the new training data to train and update the exploration strategy using the reinforcement learning method;

[0050] S7 Check the convergence condition. If it does not converge, return to S3; otherwise, end. The convergence condition is usually that the number of iterations reaches the set maximum value, or the mean square error of the obtained average cumulative reward is less than the preset value.

[0051] Furthermore, the present invention also provides a control system reinforcement learning device based on maximum entropy projection, including at least one processor and a memory. The at least one processor and the memory are connected through a data bus. The memory stores instructions executable by the at least one processor. After being executed by the processor, the instructions are used to complete the above-mentioned control system reinforcement learning method based on maximum entropy projection.

[0052] Those skilled in the art can easily understand that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A reinforcement learning method for a control system based on maximum entropy projection, characterized in that, It includes the following steps: S1 initializes model parameters, including exploration strategy parameter and action strategies parameter , where the exploration strategy In order to fully explore, when the state is Agent chooses action The probability of action strategy In order to obtain the maximum reward, when the state is Agent chooses action The probability of state It is a vector composed of the readings of sensors deployed at multiple key locations in the control system. It is a vector composed of control quantities of multiple control units in the control system; S2 Reset the reinforcement learning environment; At each moment, the agent generates an action based on a linear combination of the exploration policy and the action policy , where the weight indicates whether the current agent is more inclined to explore or more inclined to obtain the maximum reward; At S4 the agent, according to the state of the current control system executes an action in the control system to obtain a reward and the state of the control system changes to and is added to the training data set wherein the reflects the degree to which the control system correctly executes a given task; Training dataset for S5 , and use the reinforcement learning method to train and update the parameters of the action policy ; ; S6 Sample a subset from the training dataset for training the exploration policy ; calculate the maximum entropy projection matrix according to the subset of training samples and train and update the exploration policy ; ; S7 Check the convergence condition. If it does not converge, return to S3; otherwise, end.

2. The reinforcement learning method for a control system based on maximum entropy projection according to claim 1, wherein, The specific content of step S6 includes: S6-1 in the training dataset randomly sample training samples to form a subset for training the exploration strategy , where is one training sample in ; S6-2 will compose the current state vectors included in all training samples into a state vector matrix , where the th column is the state vector , the number of rows of is the same as the dimension of the state vector ; S6-3 obtains the maximum entropy projection matrix through the formula where is the state matrix after projection, is the projection matrix, and each column in is the state vector after projection by the projection matrix ; , is the matrix obtained by taking the mean value of each row of . S6-4 passes through the maximum entropy projection matrix The state matrix after projection Each column in is the state vector after projection. Count the number of times these column vectors appear in the state matrix after projection Find the column vector with the least number of occurrences, and its subscript is marked as , and take the th column in the state vector matrix as the exploration target state ; ; S6-5 Constructing New Training Data for Training Exploration Strategies , where the current state in the training sample , action , next moment state are the same as those in , and the reward in the training sample is set as follows: S6-6 Use new training data , and train and update the exploration strategy using the reinforcement learning method .

3. The reinforcement learning method for a control system based on maximum entropy projection according to claim 1 or 2, characterized in that In step S1, for the exploration strategy parameters and the action strategy parameters The initialization method is to generate random numbers with a uniform distribution in the interval.

4. The reinforcement learning method for a control system based on maximum entropy projection according to claim 1 or 2, characterized in that The specific content of step S2 is: Initialize the control system to return it to the initial state.

5. The reinforcement learning method for a control system based on maximum entropy projection according to claim 1 or 2, characterized in that The reinforcement learning method in step S5 includes SAC (Soft Actor-Critic) or PPO (Proximal Policy Optimization).

6. The reinforcement learning method for a control system based on maximum entropy projection according to claim 1 or 2, characterized in that, The convergence condition in step S7 is: The number of iterations reaches the set maximum value.

7. The reinforcement learning method for a control system based on maximum entropy projection according to claim 1 or 2, characterized in that The convergence condition in step S7 is: The mean square error of the obtained average cumulative return is less than the preset value.

8. The reinforcement learning method for a control system based on maximum entropy projection according to claim 1 or 2, characterized in that The control unit in the control system is a motor or a relay.

9. The reinforcement learning method for a control system based on maximum entropy projection according to claim 1 or 2, characterized in that The task that the control system needs to complete is: Maintain a certain motion state or move along a given trajectory.

10. A control system reinforcement learning device based on maximum entropy projection, characterized in that: It includes at least one processor and a memory. The at least one processor and the memory are connected through a data bus. The memory stores instructions executable by the at least one processor. After being executed by the processor, the instructions are used to complete the control system reinforcement learning method based on maximum entropy projection according to any one of claims 1-9.