Unmanned ship path planning reinforcement learning method based on simplified environment and dynamics
By designing a multi-layer perceptron network and PPO algorithm in a simplified environment, the problem of training unmanned boats in difficult environments is solved, and efficient path planning and obstacle avoidance are achieved.
Patent Information
- Application Number
- CN202510711379.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-29
Smart Images

Figure CN120235212A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of control and information technology, and particularly relates to a reinforcement learning method for unmanned boat path planning based on a simplified environment and dynamics. Background Art
[0002] In the past decade, the path planning method of the reinforcement learning-based agent system has been rapidly developed in various application scenarios, especially in the field of obstacle avoidance and path optimization in complex environments. For example, path planning and obstacle avoidance problems, dynamic obstacle avoidance problems, and multi-objective path planning problems. As a special type of agent, the unmanned boat system has received much attention due to its wide application in the marine environment. In the path planning problem, the goal of the unmanned boat is to reach the target position from the initial position while avoiding collisions with obstacles, and the obstacles and complex dynamic characteristics in the environment make the path planning task highly challenging.
[0003] It is worth noting that most of the existing methods are trained in the environment set by the problem and for the corresponding unmanned boat model. However, this framework will lead to difficulties in training, slow convergence speed, or even non-convergence.
[0004] Under the above background, we propose a reinforcement learning method for unmanned boat path planning based on a simplified environment and dynamics. This method first considers a simple dynamic model in the simplest simulation environment and uses MLP+PPO to train a single unmanned boat to achieve path planning; secondly, an additional policy network is designed in the real environment and dynamic model to convert the control information in the simple environment into the control information in the real environment.
[0005] The difficulties in designing the reinforcement learning method for unmanned boat path planning based on a simplified environment and dynamics are as follows: First: How to design a simple environment so that the core of the basic path planning problem remains. Second: How to design a simple dynamic model so that its state evolution can be synchronized with the state evolution of the high-order dynamic model in the real environment through an additional policy network. Third: How to train in a simple environment, grasp the main core problems, and then expand to the real environment for supplementary training. Summary of the Invention
[0006] The purpose of the present invention is to propose a reinforcement learning method for unmanned boat path planning based on a simplified environment and dynamics, which is used to train the unmanned boat system to plan a path in a complex environment, and on the premise of avoiding collisions with obstacles, plan a path and run from the initial position to reach the target point.
[0007] To achieve the above object, the technical solution of the present invention is: a reinforcement learning method for unmanned boat path planning based on simplified environment and dynamics, specifically including the following steps: Step 1: Build a real environment E , set the state s of the unmanned boat to include the position of the unmanned boat, the reward function r to be the negative value of the length of the planned path of the unmanned boat, the unmanned boat has a dynamic model with high-order uncertain parameters, and the input control signal is u; Step 2: Build a simple environment E s , where the state s of the unmanned boat s , the reward function r s , the setting of environmental obstacles is the same as that of the real environment E , and the unmanned boat has a second-order integral series dynamic model, and the input control signal is u s ; Step 3: Design a first policy network based on a multi-layer perceptron; the first policy network interacts with the simple environment E s , and the input of the first policy network is the state s of the unmanned boat in the simple environment s , and the output is the control signal u of the simple environment s ; Step 4: Design a second policy network based on a multi-layer perceptron; the second policy network interacts with the first policy network and the real environment E , and the input of the second policy network is the control signal u of the simple environment s , and the output is the control signal u of the real environment; Step 5: Train the first policy network based on the PPO algorithm, and the training goal is to maximize the cumulative reward corresponding to the reward function r of the simple environment s ; Step 6: Train the second policy network based on the PPO algorithm, and the training goal is to minimize the error between the states and the error between the rewards in the two environments.
[0008] Preferably, the specific content of step 1 includes: The real environment includes changing sea conditions and obstacles; The state s of the unmanned boat includes the position of the unmanned boat. When the state s of the unmanned boat adopts a two-dimensional unmanned boat position, it is expressed as , and when the state s of the unmanned boat adopts a three-dimensional unmanned boat position, it is expressed as , where x , y , z represent the coordinates of the unmanned boat; From M intermediate path points divide the planned path between the initial position s 0 and the end target position s f intoM +1 segment, and the position of the path point in the planned path is denoted as s i , ; among them, when i takes 0 s i represents the initial position s 0, when i takes M +1 s i represents the end target position s f , when at this time s i represents the i th intermediate path point position; Set the reward function r as the negative path length, and the path length is obtained by summing the Euclidean distances between all adjacent path point positions. The reward function r is expressed as: Among them, is the i +1 and the i th path point positions, and the specific calculation formula is: Among them s ij and s i+1,j are respectively the i and the i +1th path point positions of the j th dimension coordinates. When using the two-dimensional unmanned boat position, the maximum dimension J is 2, and when using the three-dimensional unmanned boat position, the maximum dimension J is 3; The dynamic model of the unmanned boat with high-order and uncertain parameters is set as: Among them, s is the position of the unmanned boat; the dot above the parameter represents the first derivative; is s the function expression of the derivative; u is the input control signal; θ is the uncertain parameter of the dynamic model, including environmental factors and internal system parameters.
[0009] Preferably, the specific steps of step 2 include: The simple environment only includes simplified obstacles; The second-order integral series dynamic model is set as: Among them, s s is the state of the unmanned boat in a simple environment, including the position and speed of the unmanned boat; for a two-dimensional environment, the state of the unmanned boat ; for a three-dimensional environment, the state of the unmanned boat ; m s represents the intermediate state, , where represents the state vector of the unmanned boat in the n th intermediate state; u s represents the control signal input in the simple environment, , where represents the control amount of the control signal input by the unmanned boat in the k th direction; represents the proportionality factor in the dynamic model; The said reward function r s is: Among them, represents the path length in the simple environment; represents that the path from the starting point to the ending point in the simple environment is divided into M +1 segments, and the state of the unmanned boat at the i th path point obtained;
[0010] Preferably, the first policy network is used to implement path planning in a simple environment and a simple model; the multi-layer perceptron includes an input layer, L hidden layers and an output layer; between the input layer and the output layer, the network processes data through a fully connected layer; The input layer receives the state of the unmanned boat in the simple environment s s ; the output of each hidden layer is expressed as: Among them, l When taking 1, is the input layer , representing the state of the unmanned boat s s ; W l and b l are respectively the weight matrix and the bias term of the l th layer, is the activation function; The calculation of the output layer is: Among them, is the output of the last hidden layer L; and are the weight matrix and bias term of the output layer respectively; the output of the output layer is the control signal for the simple environment u s .
[0011] Preferably, the second policy network is used to convert the path planning control signal obtained from the simple environment and the simple model into the control signal in the real environment and the real model; the second policy network includes an input layer, L hidden layers and an output layer; the network input layer receives the control signal of the simple environment u s , and the output layer outputs the control signal u of the real environment.
[0012] Preferably, the loss function of the first policy network is specifically: wherein, represents the expectation calculation of the empirical data at time step t; is the probability ratio of the new and old policies for the control signal under the simple environment state ; represents the policy with parameter ; represents the probability of taking the control signal at state under the current policy; represents the probability of taking the control signal at state under the old policy; is the advantage function, indicating the superiority or inferiority of the current action relative to the average action; is the policy clipping parameter, used to limit the policy update amplitude; The control signal is simultaneously used as the output action of the policy network , that is , and the expression of the advantage function is as follows: wherein, represents the expected cumulative reward obtained by taking the action at state , represents the average policy value without considering the current action at state .
[0013] Preferably, the advantage function is calculated using Generalized Advantage Estimation, and the expression is: where d is an integer from 0 to T-t indicating the current time step t and a series of delayed time steps after it, T represents the maximum number of time steps of the current trajectory; is the temporal difference error; represents the immediate reward; is the discount factor, used to control the importance of future rewards; is a hyperparameter that balances the estimation variance and bias.
[0014] Preferably, the training objective of the second policy network is to minimize the total loss function to optimize the parameters of the second policy network, so as to achieve an effective mapping between the simple environment and the real environment through the optimized second policy network; the total loss function of the second policy network is specifically: where represents the error between the states in the two environments; represents the error between the rewards in the two environments; is the policy optimization loss based on the PPO algorithm; specifically as follows: where s I is the state of the unmanned boat in the real environment E for the I th training sample, is the state of the unmanned boat in the simple environment I for the th training sample, N represents the number of training samples; where r I is the reward of the I th training sample in the real environment E, is the reward of the I th training sample in the simple environment ; where is the policy ratio, representing the probability ratio between the current policy and the old policy.
[0015] Compared with the prior art, the present invention has the following beneficial effects: (1) By simplifying the environment setting and the dynamic model, the training efficiency of the path planning strategy is greatly improved, the convergence difficulty of the first policy network is reduced, and the number of layers required by the network is also less than that of the normal path planning strategy network; (2) By designing the second policy network, the control signal in the simplified environment can be converted into the control signal in the real environment, and the state and reward errors corresponding to the two environments will converge; (3) The first policy network and the second policy network are trained in a decoupled manner, so that the framework has strong generalization ability. For example, for different high-order dynamic models, the same first policy network can be used, and only the second policy network needs to be retrained, reducing the training time and training difficulty in the generalization process. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0017] The following combines the attached Figure 1 to specifically describe the technical solution of the present invention.
[0018] The present invention proposes a reinforcement learning method for unmanned boat path planning based on a simplified environment and dynamics, which specifically includes the following steps: Step 1: Build a real environment E Set the state s of the unmanned boat to include the position of the unmanned boat, the reward function r to be the negative value of the length of the planned path of the unmanned boat, the unmanned boat has a high-order dynamic model with uncertain parameters, and the input control signal is u; Step 2: Build a simple environment E s where the state s of the unmanned boat s , the reward function r s , the setting of environmental obstacles is the same as that of the real environment E , and the unmanned boat has a second-order integral series dynamic model, and the input control signal is u s ; Step 3: Design a first policy network based on a multi-layer perceptron; the first policy network interacts with the simple environment E s , and the input of the first policy network is the state s of the unmanned boat in the simple environment s , and the output is the control signal u in the simple environment s ; Step 4: Design a second policy network based on a multi-layer perceptron; the second policy network interacts with the first policy network and the real environment E , and the input of the second policy network is the control signal u in the simple environment s, the output is the control signal u in the real environment; Step 5: Train the first policy network based on the PPO algorithm, with the training objective of maximizing the simple environment reward function r s The corresponding cumulative reward to optimize the parameters of the first policy network. The learning rate of the PPO algorithm can be set to 0.01; Step 6: Train the second policy network based on the PPO algorithm, with the training objective of minimizing the error between the states and the error between the rewards in the two environments, that is, minimizing the mean square error of s and s s and r and r s to ensure that the control strategies between the real environment E and the simple environment Es are consistent.
[0019] In this embodiment, step 1 specifically includes: The real environment includes changing sea conditions and obstacles, such as various winds, waves, currents, tides, and reefs; The state s of the unmanned boat includes the position of the unmanned boat. When the state s of the unmanned boat adopts a two-dimensional unmanned boat position, it is expressed as When the state s of the unmanned boat adopts a three-dimensional unmanned boat position, it is expressed as , where x , y , z represent the coordinates of the unmanned boat; By M intermediate waypoints divide the planned path between the initial position s 0 and the end target position s f into M +1 segments. The position of the waypoint in the planned path is denoted as s i , ; among them, when i takes 0 s i represents the initial position s 0, when i takes M +1 s i represents the end target position s f , when s i represents the i th intermediate waypoint position; the intermediate waypoint position is expressed as: where represents the proportion of the position of the i th intermediate point in the total path; Set the reward function r as the negative path length, where the path length is obtained by summing the Euclidean distances between all adjacent path point positions, and the reward function r is expressed as: where is the i -th i and the -th path point positions, and the specific calculation formula is: s ij and s i+1,j are respectively the i -th i and the j -th J -th path point position of the J -th dimension coordinates. When using the two-dimensional unmanned boat position, the maximum dimension is 2, and when using the three-dimensional unmanned boat position, the maximum dimension where s is the unmanned boat position; the dot above the parameter represents the first derivative; is s the function expression of the derivative; u is the input control signal; θ is the uncertain parameter of the dynamic model, including environmental factors and internal system parameters.
[0020] where, the above-mentioned high-order dynamic model of the unmanned boat with uncertain parameters selects different models and parameters according to the actual unmanned boat used. For example, if the size of the unmanned boat is different, the size parameters in the model are also different.
[0021] In this embodiment, step 2 specifically includes: The simple environment only includes simplified obstacles, such as reefs, etc.; The second-order integral series dynamic model is set as: where s s is the state of the unmanned boat in the simple environment, including the position and speed of the unmanned boat; for a two-dimensional environment, the unmanned boat state ; for a three-dimensional environment, the unmanned boat state ; m s represents the intermediate state, where represents the state vector of the unmanned boat at the n -th intermediate state; us The control signal representing the simple environment input , where represents the control amount of the control signal input by the unmanned boat in the k th direction; represents the scale factor in the dynamic model; The reward function r s is: where represents the path length in the simple environment; represents that the path from the starting point to the ending point in the simple environment is divided into M +1 segments, and the state of the unmanned boat at the i th path point is obtained;
[0022] wherein, the dynamic model in step 2 is uniformly set as a second-order integral series type, that is, the selection of the simple model is not related to the actual unmanned boat model and parameters used.
[0023] In this embodiment, the first policy network is used to implement path planning in the simple environment and the simple model. Therefore, only a simple multi-layer perceptron architecture is needed to design the network; the multi-layer perceptron includes an input layer, L layers (for example, 24 layers) of hidden layers and an output layer; between the input layer and the output layer, the network processes data through a fully connected layer; The input layer receives the state of the unmanned boat in the simple environment s s ; the output of each hidden layer is expressed as: where l when taking 1, is the input layer , representing the state of the unmanned boat s s ; W l and b l are respectively the weight matrix and the bias term of the l th layer, is the activation function; The calculation of the output layer is: where is the output of the last hidden layer L; and are respectively the weight matrix and the bias term of the output layer; the output of the output layer is the control signal of the simple environmentu s 。
[0024] In this embodiment, the second policy network is used to convert the path planning control signal obtained from the simple environment and the simple model into the control signal in the real environment and the real model; the second policy network includes an input layer, L a hidden layer with a certain number of layers (for example, 24 layers) and an output layer; the network input layer receives the control signal of the simple environment u s , and the output layer outputs the control signal u of the real environment.
[0025] Although the second policy network cannot fully and truly implement the conversion of the control signal, after training, it can ensure that the error amplitude of the two control signals is limited, making the error amplitude of the final path planning relatively small. In many actual unmanned boat application scenarios, a small path planning error is acceptable.
[0026] In this embodiment, the cumulative reward is expressed as: where, represents the complete trajectory generated according to the policy with parameter , represents the path, represents the path, represents the policy with parameter , T represents the maximum number of time steps of each trajectory, t represents the current time step, is the immediate reward obtained when the unmanned boat is in state and takes the control signal in the simple environment, represents the state of the simple environment at the t th time step, represents the control action of the simple environment at the t th time step, is the discount factor.
[0027] In PPO, the maximization of the reward indirectly optimizes the policy through the advantage function. Therefore, the loss function of the first policy network is specifically: where, represents the expectation calculation of the empirical data at time step t; is the probability ratio of the old and new policies for the control signal in the simple environment state ; represents the parameter as The strategy represents the probability of taking the control signal at state under the current strategy; represents the probability of taking the control signal at state under the old strategy; is the advantage function, indicating the superiority or inferiority of the current action relative to the average action; is the policy clipping parameter, used to limit the magnitude of policy update; The said control signal is simultaneously used as the output action of the policy network , that is , and the said advantage function has the following expression: where represents the expected cumulative reward obtained by taking action at state , represents the average policy value without considering the current action at state .
[0028] In this embodiment, the said advantage function is calculated using generalized advantage estimation, and the expression is: where d is an integer from 0 to T-t , representing a series of delayed time steps after the current time step t , T represents the maximum number of time steps of the current trajectory; is the temporal difference error; represents the immediate reward; is the discount factor, used to control the importance of future rewards; is a hyperparameter that balances the estimation variance and bias.
[0029] In this embodiment, the total loss function of the second policy network is specifically: where represents the error between the states in the two environments; represents the error between the rewards in the two environments; is the policy optimization loss based on the PPO algorithm; specifically as follows: wheres I is the state of the unmanned boat of the I th training sample in the real environment E, is the state of the unmanned boat of the I th training sample in the simplified environment ; N represents the number of training samples; Among them, r I is the reward of the I th training sample in the real environment E, is the reward of the I th training sample in the simplified environment ; Among them, is the policy ratio, representing the probability ratio between the current policy and the old policy.
[0030] The training objective is to minimize , by the parameters of the second optimized policy network, Among them, represents the optimal parameters of the second optimized policy network obtained by optimization.
[0031] So far, all steps are completed.
[0032] The present invention studies how to design a reinforcement learning method for unmanned boat path planning based on a simplified environment and dynamics. This method can be used to train an unmanned boat system to perform path planning in a complex environment, and on the premise of avoiding collisions with obstacles, plan a trajectory and run from the initial position to reach the target point. For the path planning problem of unmanned boats, the existing algorithms are mainly based on optimization methods or learning methods, and are solved in the real environment. And this method is improved on the basis of the existing theory. By simplifying the real environment and the dynamic model, the training process becomes simple and efficient. Through the linked second policy network, the control signal in the simplified environment can be converted into the control signal in the real environment, and the state and reward errors corresponding to the two environments will converge. This characteristic shows that this method has strong generalization ability for different dynamic models.
[0033] The above are the preferred embodiments of the present invention. All changes made according to the technical solution of the present invention, when the functions and effects produced do not exceed the scope of the technical solution of the present invention, shall fall within the protection scope of the present invention.
Claims
1. An enhanced learning method for the path planning of an unmanned boat based on a simplified environment and dynamics, characterized in that, Specifically, it includes the following steps: Step 1: Build a real environment E , set the state s of the unmanned boat to include the position of the unmanned boat, the reward function r is the negative value of the length of the planned path of the unmanned boat, the unmanned boat has a high-order dynamic model with uncertain parameters, and the input control signal is u; Step 2: Build a simple environment E s , where the state s of the unmanned boat s , the reward function r s , the setting of environmental obstacles is the same as the real environment E , and the unmanned boat has a second-order integral series dynamic model, and the input control signal is u s ; Step 3: Design a first policy network based on a multi-layer perceptron; the first policy network interacts with the simple environment E s and the input of the first policy network is the state s of the unmanned boat in the simple environment s , and the output is the control signal u of the simple environment s ; Step 4: Design a second policy network based on a multi-layer perceptron; the second policy network interacts with the first policy network and the real environment E and the input of the second policy network is the control signal u of the simplified environment s and the output is the control signal u of the real environment; Step 5: Train the first policy network based on the PPO algorithm, with the training objective of maximizing the cumulative reward corresponding to the simple environment reward function r s The corresponding cumulative reward; Step 6: Train the second policy network based on the PPO algorithm, and the training objective is to minimize the error between the states and the error between the rewards in the two environments.
2. The reinforcement learning method for unmanned boat path planning based on simplified environment and dynamics according to claim 1, wherein The specific content of Step 1 includes: The real environment includes changing sea conditions and obstacles; The state s of the unmanned boat includes the position of the unmanned boat. When the state s of the unmanned boat adopts a two-dimensional position of the unmanned boat, it is expressed as , when the state s of the unmanned boat adopts a three-dimensional position of the unmanned boat, it is expressed as , where x , y , z represent the coordinates of the unmanned boat; From M intermediate path points divide the planned path between the initial position s 0 and the end target position s f into M +1 segments. The positions of the path points in the planned path are denoted as s i , ; among them, when i takes 0 s i represents the initial position s 0. When i takes M +1 s i represents the end target position s f . When is s i represents the position of the i th intermediate path point; Set the reward function r to the negative path length, where the path length is obtained by summing the Euclidean distances between all adjacent path point positions, and the reward function r is expressed as: Among them, is the Euclidean distance between the i +1-th and the i -th path point positions, and the specific calculation formula is: in s ij and s i+1,j They are i and i +1 waypoint location j dimensional coordinates, the maximum dimension when using the two-dimensional position of the unmanned boat J is 2, when the maximum dimension is used for the three-dimensional position of the unmanned boat J is 3; The dynamic model of the unmanned surface vehicle with high order and uncertain parameters is set as: Among them, s is the position of the unmanned boat; the · above the parameter represents the first derivative; is s the functional expression of the derivative; u is the input control signal; θ is the uncertain parameter of the dynamic model, including environmental factors and internal system parameters.
3. The method for reinforcement learning of the path planning of an unmanned boat based on a simplified environment and dynamics according to claim 2, wherein The specific content of Step 2 includes: The simple environment only includes simplified obstacles; The second-order integral series dynamic model is set as: Among them, s s is the state of the simple environment unmanned boat, including the position and speed of the unmanned boat; for a two-dimensional environment, the state of the unmanned boat ; for a three-dimensional environment, the state of the unmanned boat ; m s represents the intermediate state, , where represents the state vector of the unmanned boat in the n th intermediate state; u s represents the control signal input in the simple environment, , where represents the control amount of the control signal input by the unmanned boat in the k th direction; represents the proportionality factor in the dynamic model; The reward function r s is as follows: Among them, represents the path length in the simple environment; represents that the path from the starting point to the ending point in the simple environment is divided into M +1 segments, and the i th unmanned boat state at the path point.
4. The reinforcement learning method for unmanned surface vehicle path planning based on a simplified environment and dynamics according to claim 3, wherein The first policy network is used to implement path planning in a simple environment and a simple model; the multi-layer perceptron includes an input layer, L a hidden layer, and an output layer; between the input layer and the output layer, the network processes data through a fully connected layer; The input layer receives the state of the unmanned boat in a simple environment s s ; The output of each hidden layer is expressed as: Among them, l when taking 1, is the input layer , representing the state of the unmanned boat s s ; W l and b l are respectively the weight matrix and bias term of the l th layer, is the activation function; The calculation of the output layer is: Among them, is the output of the last hidden layer L; and are the weight matrix and bias term of the output layer respectively; the output of the output layer is the control signal of the simple environment u s .
5. The reinforcement learning method for unmanned boat path planning based on a simplified environment and dynamics according to claim 4, characterized in that The second policy network is used to convert the path planning control signal obtained from the simple environment and the simple model into the control signal in the real environment and the real model; the second policy network includes an input layer, L a hidden layer, and an output layer; the network input layer receives the control signal of the simple environment u s , and the output layer outputs the control signal u of the real environment.
6. The reinforcement learning method for unmanned boat path planning based on a simplified environment and dynamics according to claim 5, characterized in that The loss function of the first policy network Specifically: Among them, represents the expected value calculation of the empirical data at time step t; is the probability ratio of the control signal under the old and new policies in the simple environment state ; represents the policy with parameter ; represents the probability of taking the control signal in the current state under the current policy; ; represents the probability of taking the control signal in the state under the old policy; is the advantage function, indicating the superiority or inferiority of the current action relative to the average action; is the policy clipping parameter, used to limit the policy update amplitude; The control signal is simultaneously used as the output action of the policy network , that is , the expression of the advantage function is as follows: Among them, represents the expected cumulative reward obtained by taking action in state , and represents the average policy value in state without considering the current action.
7. The reinforcement learning method for unmanned boat path planning based on simplified environment and dynamics according to claim 6, wherein The said advantage function is calculated using Generalized Advantage Estimation, and the expression is: wherein, d is an integer from 0 to T - t representing the current time step t and a series of delayed time steps after T representing the maximum number of time steps of the current trajectory; is the time difference error; represents the immediate reward; is the discount factor used to control the importance of future rewards; is a hyperparameter that balances the estimation variance and bias.
8. The reinforcement learning method for unmanned boat path planning based on a simplified environment and dynamics according to claim 6, characterized in that The training objective of the second policy network is to minimize the total loss function , so as to optimize the parameters of the second policy network , so as to achieve an effective mapping between the simple environment and the real environment through the optimized second policy network; the total loss function of the second policy network is specifically as follows: Among them, represents the error between the states in two environments; represents the error between the rewards in two environments; is the policy optimization loss based on the PPO algorithm; specifically as follows: Among them, s I is the state of the unmanned boat of the I th training sample in the real environment E, is the state of the unmanned boat of the I th training sample in the simple environment , and N represents the number of training samples. Among them, r I is the I reward of the -th training sample in the real environment E, I and is the reward of the Among them, is the strategy ratio, representing the probability ratio between the current strategy and the old strategy.
Citation Information
Patent Citations
Unmanned ship hybrid sensing autonomous obstacle avoidance method and system based on reinforcement learning
CN111880535A
Unmanned ship dynamic target capturing method based on deep reinforcement learning, electronic equipment and storage medium
CN118276583A
Park logistics trolley path planning method based on map-free navigation
CN118730145A
Deep reinforcement learning quadruped robot motion control method and system based on constraint reward
CN119512184A
AUV action plan and operation control method based on reinforcement learning
JP2021034050A