Manta Ray-Type Bionic Fish Control Method, Device, and Storage Medium Based on Reinforcement Learning
Through the reinforcement learning algorithm and INS/GPS navigation system based on deep deterministic strategy gradient network, the problem of poor adaptability and robustness of the manta ray-style bionic fish control method is solved, and efficient control of autonomous swimming to the target point is achieved.
Patent Information
- Application Number
- CN202211009423.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-22
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-08-22
AI Technical Summary
The existing manta ray-style bionic fish control methods have problems such as poor adaptability, poor robustness, low training efficiency and poor practicality.
Using a reinforcement learning algorithm based on a deep deterministic strategy gradient network, combined with an INS and GPS combined navigation system, the bionic fish's swimming direction and velocity coefficient are controlled by observing the current motion state output direction and velocity coefficient of the bionic fish, and the servo control curve of the sine wave curve is constructed to achieve independent swimming.
It improves the control efficiency and robustness of bionic fish, enhances training efficiency, and improves practicality in complex environments.
Smart Images

Figure CN115390573B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a reinforcement learning algorithm in the fields of artificial intelligence and robotics, and particularly to a manta ray-like bionic fish control method, device, and storage medium based on reinforcement learning. Background Art
[0002] The ocean contains a large amount of resources. In recent years, the development and exploration of ocean resources have been increasingly emphasized. As an important exploration tool, bionic fish have also been accelerating their development pace in recent years. To achieve the motion control of bionic fish, scholars at home and abroad have also conducted a large amount of research on the control algorithms of bionic fish, mostly limited to non-intelligent control algorithms. Wang Ming, Yu Junzhi, Tan Min. CPG Control and Implementation of Pectoral Fin Propelled Robotic Fish. Robotics Journal, 20100315 discloses that the bionic robotic fish based on CPG realizes the smooth switching and stability of modalities such as forward, turning, and backward of the bionic fish, but it cannot achieve autonomous swimming and has disadvantages such as low control efficiency and poor robustness.
[0003] In recent years, there have been more and more cases of applying deep reinforcement learning algorithms to the control of bionic robots, and some scholars have also applied them to the control of bionic fish. Chinese Patent CN 110909859 discloses a bionic robotic fish motion control method based on adversarial structured control. This method uses deep reinforcement learning as an optimization algorithm, takes the accuracy and speed of moving to the target point as the reward item, and takes the sum of the servo powers as the loss item to construct an optimization objective function. However, this method controls the swimming of the bionic fish by directly controlling the servo, and the training efficiency is low. Tianhao Zhang, Runyu Tian, Chen Wang, Guangming Xie. Path-following Control of Fish-like Robots: A Deep Reinforcement Learning Approach[J]. IFAC PapersOnLine, 2020, 53-2(8166-8167) discloses that the bionic fish based on reinforcement learning control can achieve autonomous swimming through path tracking, but it is difficult to apply the method of shooting the pose state of the bionic fish by a camera to real environments such as rivers and lakes, and the practicability is poor.
[0004] The existing control methods applied to manta ray-like bionic fish have problems such as poor adaptability and poor robustness, while the bionic fish based on reinforcement learning control has problems such as low training efficiency and poor practicability. Summary of the Invention
[0005] Objective of the Invention: The present invention aims to provide a control method for a manta ray - like biomimetic fish based on reinforcement learning, which can improve training efficiency, practicality, control efficiency, and robustness. The present invention also provides a control device for the control method of the manta ray - like biomimetic fish based on reinforcement learning.
[0006] Technical Solution: The control method for the manta ray - like biomimetic fish based on reinforcement learning according to the present invention includes the following steps:
[0007] (1) Establish a world coordinate system: Use the initial point as the coordinate origin, the east direction as the positive X - direction, and the north direction as the positive Y - direction;
[0008] (2) Construct an INS and GPS integrated navigation system, output the position information of the biomimetic fish at time t, and deduce the error angle err_yaw and the distance err_dist between the current position and the target point at time t;
[0009] (3) Construct a reinforcement learning model, use DDPG as the reinforcement learning model framework, input the state at time t, state = [err_yaw, v, err_dist, v_yaw], where v is the real - time moving speed of the manta ray - like biomimetic fish, and v_yaw is the turning speed, and output the action action = [Kv, Kt], where Kv is the speed coefficient and Kt is the direction coefficient;
[0010] (4) Train the reinforcement learning model through a simulation system;
[0011] (5) After the training is completed, the reinforcement learning model outputs the action action = [Kv, Kt] as the input value of the control curve of the servo of the manta ray - like biomimetic fish to control the speed and swimming direction of the fish body.
[0012] Further, in step (2), the formula for the deviation err_yaw between the real - time heading angle and the target heading angle is as follows:
[0013] err_yaw = new_yaw - tar_yaw
[0014]
[0015]
[0016] where new_yaw is the current yaw angle, taw_yaw is the target heading angle, p x and p y respectively represent the abscissa and ordinate of the target point, n x and n y respectively represent the X - coordinate and Y - coordinate of the current position output by the INS and GPS integrated navigation system, e x and e yThey respectively represent the difference in the X coordinate and the difference in the Y coordinate between the target point and the current position.
[0017] The formula for the distance err_dist between the current position and the target point is as follows:
[0018]
[0019] Furthermore, the formula for the reward function r of the reinforcement learning model in step (3) is as follows:
[0020] r = r s + r c
[0021] r s = r yaw + r dist + r v + r v_yaw
[0022] r c = r a + r d
[0023] Among them, r s is the daily reward, r c is the settlement reward at the end of the round, r yaw is the direction reward, r dist is the distance reward, r v is the speed reward, r v_yaw is the rotational speed reward, r a is the task completion reward, r d is the device damage reward.
[0024] Direction reward err_yaw represents the deviation between the real-time heading angle and the target heading angle; Distance reward D represents the distance between the manta ray bionic fish and the target point at the initial moment; Speed reward v represents the real-time moving speed of the manta ray bionic fish, with the unit of m / s; Rotational speed reward v_yaw represents the turning speed, with the unit of ° / s.
[0025] Task completion reward r a The formula is as follows:
[0026]
[0027] Among them, represents the horizontal distance between the current position of the manta ray bionic fish and the target position point, with the unit of m; Δh represents the depth deviation between the manta ray bionic fish and the target point, with the unit of m; s is the number of training steps in a round, s max The maximum number of training steps;
[0028] Device damage reward r d The formula is as follows:
[0029]
[0030] Where x i , y i , z i are the three-dimensional coordinates of the current manta ray bionic fish, and x max , y max , z max are the three-dimensional coordinates of the maximum movement position of the manta ray bionic fish.
[0031] Furthermore, in step (5), the control curve of the servo motor of the manta ray bionic fish is a sine wave curve, and the angle output control function of the servo motor is:
[0032]
[0033]
[0034]
[0035] Where α l is the rotation angle of the left flapping servo motor, α0 is the maximum pitch angle of the flapping servo motor, ω is the angular frequency, θ l is the rotation angle of the left rotating servo motor, K tl represents the left direction coefficient, θ0 is the maximum rotation angle of the rotating servo motor, φ is the phase difference between the two servo motors on the same side, α r is the rotation angle of the right flapping servo motor, θ r is the rotation angle of the right rotating servo motor, K tr represents the right direction coefficient; the output range of the speed coefficient Kv is [0, 1], and the output range of the direction coefficient Kt is [-1, 1].
[0036] A manta ray bionic fish control device, comprising:
[0037] INS and GPS combined navigation module to obtain the combined navigation results of the speed v, position, and rotation speed of the bionic fish;
[0038] Reinforcement learning module, using DDPG as the reinforcement learning model framework to train the speed coefficient and direction coefficient of the bionic fish based on the deep deterministic policy gradient network;
[0039] Control module, the speed coefficient and direction coefficient are used as the input values of the control curve of the servo motor of the bionic fish to control the movement of the bionic fish by the servo motor;
[0040] The deep deterministic policy gradient network includes a policy network and an evaluation network;
[0041] The policy network includes a first action state estimation module, a first action state reality module, and a policy gradient module;
[0042] The evaluation network includes a second action state estimation module, a second action state reality module, and a loss function module;
[0043] The first action state estimation module is connected to the second action state estimation module and the policy gradient module, and the first action state estimation module outputs the state parameters and receives the reward value;
[0044] The first action state reality module is connected to the second action state reality module, and the first action state reality module receives the reward value;
[0045] The policy gradient module is connected to the first action state estimation module and the second action state estimation module;
[0046] The second action state estimation module is connected to the first action state estimation module, the policy gradient module, and the loss function module, and receives the reward value;
[0047] The second action state reality module is connected to the first action state reality module and the loss function module, and receives the reward value;
[0048] The loss function module is connected to the second action state reality module and the second action state estimation module.
[0049] An electronic device includes: a memory and a processor; the memory is used to store a computer program; wherein, the processor executes the computer program in the memory to implement the above method.
[0050] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it is used to implement the above method.
[0051] Advantageous effects: Compared with the prior art, the significant advantages of the present invention are: 1. The present invention controls the manta ray bionic fish through a reinforcement learning algorithm to achieve the task of autonomous swimming of the bionic fish to the target point, improving the control efficiency and robustness; 2. The present invention observes the current motion state of the bionic fish as the input of the reinforcement learning algorithm, and outputs the direction and speed coefficient to control the swimming direction and speed of the bionic fish, improving the training efficiency compared with the method of controlling the joint motion angle; 3. The present invention uses GPS / INS as the positioning and navigation system, which is more practical than the positioning method of photographing the pose of the bionic fish. Description of the Drawings
[0052] Figure 1It is the flowchart of the control method of the manta ray - like bionic fish in the present invention;
[0053] Figure 2 It is the schematic diagram of the loosely - coupled GPS / INS integrated structure;
[0054] Figure 3 It is the schematic diagram of the framework structure of the reinforcement learning algorithm;
[0055] Figure 4 It is the schematic diagram of the joint simulation environment of pycharm and webots;
[0056] Figure 5 It is the schematic diagram of the structure of the pectoral fin and the servo of the bionic fish;
[0057] Figure 6 It is the control curve graph of the servo. Specific implementation mode
[0058] The present invention will be further described below with reference to the accompanying drawings.
[0059] As Figure 1 shown, the manta - ray - like bionic fish control method based on reinforcement learning of the present invention includes the following steps:
[0060] (1) Obtain the target point position sent by the upper computer, establish a northeast - sky coordinate system, with the initial point as the coordinate origin, the east direction as the positive X - direction, and the north direction as the positive Y - direction.
[0061] (2) Construct an INS and GPS integrated navigation system. Obtain the initial point longitude and latitude data from GPS, and obtain the position longitude and latitude data of the bionic fish at time t. Obtain JY901 data, including the current yaw angle, three - axial accelerations, angular velocities, and the current geomagnetic field and other data. Through the calculation of the three - axial accelerations, angular velocities, and geomagnetic field data, inertial navigation positioning can be realized to obtain information such as the current speed (v), position, etc.
[0062] To realize the complementary characteristics of GPS and INS, a loosely - coupled GPS / INS integrated architecture is adopted, as Figure 2 shown. Both GPS and INS work independently and each provides the results of navigation parameters. To improve the navigation accuracy, usually the position and speed of GPS are input into the filter. At the same time, the position, speed, and attitude of INS are also used as inputs to the filter. The filter compares the differences between the two, establishes an error model to estimate the error of INS. Using these errors to correct the inertial navigation results, the combined navigation results of speed (v), position, and rotation speed (v_yaw) are obtained.
[0063] Among them, the INS inertial navigation system needs to use an error model to analyze and estimate various error sources related to the INS, including errors along the geoid (latitude error, precision error, altitude error), velocity errors along the Earth system (eastward velocity error Ve, northward velocity error Vn, upward velocity error Vu), and errors of three attitude angles (pitch, roll, yaw). It also includes the bias of the accelerometer and the drift of the gyroscope.
[0064] The position information obtained by the GPS / INS inertial navigation combination can be calculated to obtain the error angle (err_yaw) at time t and the distance (err_dist) at time t.
[0065] The deviation err_yaw formula between the real-time heading angle and the target heading angle is as follows:
[0066] err_yaw = new_yaw - tar_yaw
[0067]
[0068]
[0069] Among them, new_yaw is the current yaw angle obtained from the sensor, and taw_yaw is the target heading angle calculated through the arctangent function. p x and p y respectively represent the abscissa and ordinate of the target point, n x and n y respectively represent the X coordinate and Y coordinate of the current position output by the 1NS and GPS integrated navigation system. e x and e y respectively represent the difference between the X coordinates and the difference between the Y coordinates of the target point and the current position.
[0070] The formula for the distance err_dist between the current position and the target point is as follows:
[0071]
[0072] The state value at time t required as input for the reinforcement learning algorithm is state = [err_yaw, v, err_dist, v_yaw]. So far, the speed (v) is obtained through the inertial navigation system, the rotational speed (v_yaw) is obtained through the sensor, and the error angle (err_yaw) and the distance (err_dist) are obtained through calculation.
[0073] (3) Build a reinforcement learning model. Obtain the state value state at time t, and output the action action = [Kv, Kt] as the direction coefficient and action coefficient. After the bionic fish takes an action and interacts with the environment again, let t := t + 1, and jump to step (2) to obtain the state value at time t + 1.
[0074] The basic reinforcement learning algorithm of the present invention is the Deep Deterministic Policy Gradient algorithm (DDPG). As Figure 3 shown in the schematic diagram of the reinforcement learning algorithm architecture, it is mainly divided into a main network and a target network. The actor network in the main network outputs actions, and then after the bionic fish interacts with the environment, the state value is obtained. The critic network makes an evaluation. After continuous interactions, the bionic fish can continuously perform autonomous learning, update the strategy, and finally can make optimal actions according to different environmental states. To achieve the fixed-point swimming control task, it is necessary to construct an observation space, an action space, and a reward function. Then, simulation training is carried out in the simulation environment we built.
[0075] The described observation space is the set of all states in the world where the manta ray bionic fish is located, and it consists of two parts: the current pose state of the manta ray bionic fish and the target position point state. Therefore, the observation space of the entire system consists of three parts. The first part, the current pose state, includes the three-dimensional coordinate of the current position information of the manta ray bionic fish, the yaw angle, the sailing speed, and the turning speed. The second part includes the three-dimensional coordinate information of the target position of the manta ray bionic fish.
[0076] The described action space refers to the set of all actions that the manta ray bionic fish can take. We use the action output of the reinforcement learning as the input of the sine wave curve, so the action output should be the direction coefficient Kt and the speed coefficient Kv.
[0077] The described reward function is to enable the reinforcement learning to complete the set fixed-point swimming control task. We divide the reward value into daily rewards and episode termination rewards. The daily rewards include direction rewards, speed rewards, distance rewards, and rotation speed rewards. The settlement rewards include task completion and device damage rewards.
[0078] For the described direction reward, the greater the deviation between the real-time heading angle and the target heading angle, the greater the negative reward given. When the deviation between the real-time heading angle and the target heading angle is 0, the reward value is 0. The direction reward formula is as follows:
[0079]
[0080] where err_yaw represents the deviation between the real-time heading angle and the target heading angle.
[0081] Regarding the distance reward, we hope that the manta ray - like biomimetic fish is as close as possible to the target position on the horizontal plane. Therefore, a negative reward for the horizontal - plane distance is set. The closer the manta ray - like biomimetic fish is to the target position, the smaller the negative reward. By setting the negative reward for distance, it can better promote the manta ray - like biomimetic fish to approach the target point. The distance - reward formula is as follows:
[0082]
[0083] where err_dist represents the distance between the manta ray - like biomimetic fish and the target point, and D represents the distance between the manta ray - like biomimetic fish and the target point at the initial moment.
[0084] Regarding the rotational - speed reward, too high a turning speed is likely to cause the situation that the manta ray - like biomimetic fish cannot stop in time when adjusting to the correct heading angle, which is not conducive to the control of the manta ray - like biomimetic fish. Therefore, a negative reward for limiting the turning speed needs to be added. When the turning speed exceeds the acceptable maximum turning speed, a negative reward based on overspeed is given. The rotational - speed reward formula is as follows:
[0085]
[0086] where v yaw represents the turning speed, with the unit of ° / s. When the turning speed exceeds 2° / s, a fixed negative reward is given.
[0087] Regarding the speed reward, when the manta ray - like biomimetic fish stops in place, a negative reward will be given to encourage the manta ray - like biomimetic fish to move. The speed - reward formula is as follows:
[0088]
[0089] where v represents the real - time moving speed of the manta ray - like biomimetic fish, with the unit of m / s. When the moving speed is lower than 0.05 m / s, a fixed negative reward is given.
[0090] The final result of the daily reward is
[0091] r s = r yaw + r dist + r v + r v_yaw
[0092] Regarding the task - completion reward, when the position of the manta ray - like biomimetic fish in the world coordinate system falls within the cylindrical inner region with the target position as the center, the allowable error as the radius, and the allowable depth error, the task is considered completed. The task - completion reward formula is as follows:
[0093]
[0094] In the above formula represents the horizontal distance between the position of the current manta ray - like bionic fish and the target position point, and Δh represents the depth deviation between the manta ray - like bionic fish and the target point. When the number of training steps in a round is less than or equal to the maximum number of training steps and both the horizontal distance and the depth deviation are less than 0.5 meters, the task is determined to be completed, the round ends, and the settlement reward is 0. When the manta ray - like bionic fish has consumed the maximum number of round steps set in this round of training but has not reached the target position, the round terminates and the task is determined to fail, and the settlement reward is - 1.
[0095] Regarding the device damage reward, during the training process, when the manta ray device swims to the maximum active position in the world, if the manta ray - like bionic fish still chooses to continue performing actions beyond the boundary, it will cause the manta ray - like bionic fish to be damaged. At this time, the round terminates and the task is determined to fail. The device damage reward function is specifically as follows:
[0096]
[0097] The settlement reward function is specifically:
[0098] r c =r a +r d
[0099] After determining the daily reward function during the training process and the settlement reward function at the end of the round, the reward function for the entire deep reinforcement learning task is determined as follows:
[0100] r = r s +r c
[0101] (4) Train the reinforcement learning model through the simulation system. The simulation environment is as Figure 4 shown. First, create the basic model of the manta ray - like bionic fish and the simulation experiment environment in Webots. Second, include the library files of Webots into the project in Pycharm; finally, set the controller as an external controller in Webots, and then the robot and the simulation environment in Webots can be controlled through Pycharm.
[0102] (5) After completing the training, the reinforcement learning model outputs the action action = [Kv, Kt] as the input value of the control curve of the servo of the manta ray - like bionic fish to control the speed and swimming direction of the fish body.
[0103] To make the manta ray - like bionic fish produce turning movements, the present invention makes an analogy between the swimming control task of the manta ray - like bionic fish and the autonomous driving of a car. First, the driving direction problem needs to be considered, which is the heading angle of the manta ray - like bionic fish. We must ensure the correct driving direction to ensure that it can swim to the target point. Secondly, the driving speed problem needs to be considered, which is the navigation speed problem of the manta ray - like bionic fish. On the premise of the correct direction, we need to swim to the target point as quickly as possible. Then the angle output control functions of the four servos of the manta ray - like bionic fish are as follows:
[0104]
[0105] As Figure 5 shown in the schematic diagram of the left pectoral fin structure, the No. 1 servo is responsible for the flapping movement, and the rotation angle of the servo is α l . The No. 2 servo is responsible for the rotational movement, and the rotation angle of the servo is θ l . In the above formula, Kv represents the speed coefficient, Ktl represents the left - hand direction coefficient, and Ktr represents the right - hand direction coefficient. Among them, the output range of the speed coefficient Kv is [0, 1], and the output range of the direction coefficient Kt is [-1, 1]. Make corresponding processing on the direction coefficient Kt:
[0106]
[0107]
[0108] As Figure 6 shown in the control curves of the four servos, Curve I is the left - hand flapping servo, Curve II is the left - hand rotational servo, Curve III is the right - hand flapping servo, and Curve IV is the right - hand rotational servo. When the speed and direction coefficient change, the corresponding servo control curves change, thus changing the swimming state of the bionic fish.
Claims
1. A control method for a manta ray - like bionic fish based on reinforcement learning, characterized in that, It includes the following steps: (1) Establish a world coordinate system: Use the initial point as the coordinate origin, the east direction as the positive X direction, and the north direction as the positive Y direction; (2) Construct an INS and GPS integrated navigation system to output the position information of the biomimetic fish at time t, and deduce the error angle err_yaw and the distance err_dist between the current position and the target point at time t; (3) Construct a reinforcement learning model, use DDPG as the reinforcement learning model framework, input the state at time t state = [err_yaw, v, err_dist, v_yaw], where v is the real-time moving speed of the manta ray biomimetic fish, and v_yaw is the turning speed, and output the action action = [Kv, Kt], where Kv is the speed coefficient and Kt is the direction coefficient; (4) Train the reinforcement learning model through a simulation system; (5) After training is completed, the reinforcement learning model outputs the action action = [Kv, Kt] as the input value of the control curve of the servo of the manta ray biomimetic fish to control the speed and swimming direction of the fish body; In step (2), the formula for the deviation err_yaw between the real-time heading angle and the target heading angle is as follows: err_yaw = new_yaw - tar_yaw Among them, new_yaw is the current yaw angle, tww_yaw is the target heading angle, p x and p y respectively represent the abscissa and ordinate of the target point, n x and n y respectively represent the X coordinate and Y coordinate of the current position output by the INS and GPS integrated navigation system, e x and e y respectively represent the differences between the X coordinates and Y coordinates of the target point and the current position; The formula for the distance err_dist between the current position and the target point is as follows:
2. The manta ray - type bionic fish control method based on reinforcement learning according to claim 1, wherein In step (3), the reward function r of the reinforcement learning model is as follows: r=r s +r c r s =r yaw +r dist +r v +r v_yaw r c =r a +r d Among them, r s is the daily reward, r c is the settlement reward at the end of the round, r yaw is the direction reward, r dist is the distance reward, r v is the speed reward, r v_yaw is the rotational speed reward, r a is the task completion reward, r d is the device damage reward.
3. The manta ray - type bionic fish control method based on reinforcement learning according to claim 2, wherein Direction Reward err_yaw represents the deviation between the real-time heading angle and the target heading angle; Distance Reward D represents the distance between the manta ray bionic fish and the target point at the initial moment; Speed Reward v represents the real-time moving speed of the manta ray bionic fish, with the unit of m / s; Rotation Speed Reward v_yaw represents the turning speed, with the unit of ° / s.
4. The manta ray - type bionic fish control method based on reinforcement learning according to claim 2, wherein Reward r for completing the task a The formula is as follows: Among them, represents the horizontal distance between the position of the current manta ray bionic fish and the target position point, with the unit of m; Δh represents the depth deviation between the manta ray bionic fish and the target point, with the unit of m; s is the number of training steps in a round, s max the maximum number of training steps; Device damage reward r d The formula is as follows: where x i , y i , z i are the three-dimensional coordinates of the current manta ray bionic fish, and x max , y max , z max are the three-dimensional coordinates of the maximum movement position of the manta ray bionic fish.
5. The manta ray - type bionic fish control method based on reinforcement learning according to claim 1, characterized in that In step (5), the control curve of the servo of the manta ray biomimetic fish is a sine wave curve, and the angle output control function of the servo is: Among them, α l is the rotation angle of the left flapping servo, α0 is the maximum pitch angle of the flapping servo, ω is the angular frequency, θ l is the rotation angle of the left rotating servo, K tl represents the left direction coefficient, θ0 is the maximum rotation angle of the rotating servo, φ is the phase difference between the two servos on the same side, α r is the rotation angle of the right flapping servo, θ r is the rotation angle of the right rotating servo, K tr represents the right direction coefficient; the output range of the speed coefficient Kv is [0, 1], and the output range of the direction coefficient Kt is [-1, 1].
6. The manta ray - type bionic fish control method based on reinforcement learning according to any one of claims 1 - 5, characterized in that, The method is implemented through a manta ray biomimetic fish control device, and the manta ray biomimetic fish control device includes an INS and GPS integrated navigation module to obtain the integrated navigation results of the speed v, position, and rotation speed of the biomimetic fish; a reinforcement learning module that uses DDPG as the reinforcement learning model framework to train the speed coefficient and direction coefficient of the biomimetic fish based on the deep deterministic policy gradient network; a control module that uses the speed coefficient and direction coefficient as the input value of the control curve of the servo of the biomimetic fish to control the movement of the biomimetic fish by the servo; The deep deterministic policy gradient network includes a policy network and a value network; The policy network includes a first action state estimation module, a first action state reality module, and a policy gradient module; The value network includes a second action state estimation module, a second action state reality module, and a loss function module; The first action state estimation module is connected to the second action state estimation module and the policy gradient module, and the first action state estimation module outputs the state parameters and receives the reward value; The first action state reality module is connected to the second action state reality module, and the first action state reality module receives the reward value; The policy gradient module is connected to the first action state estimation module and the second action state estimation module; The second action state estimation module is connected to the first action state estimation module, the policy gradient module, and the loss function module, and receives the reward value; The second action state reality module is connected to the first action state reality module and the loss function module, and receives the reward value; The loss function module is connected to the second action state reality module and the second action state estimation module.
7. An electronic device, comprising: A memory and a processor; The memory is used to store a computer program; wherein, the processor executes the computer program in the memory to implement the method according to any one of claims 1-6.
8. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, it is used to implement the method according to any one of claims 1-6.
Citation Information
Patent Citations
Robot fish control method and device, equipment and storage medium
CN111191399A