Control system of mobile engineering equipment operation mechanism based on reinforcement learning PID controller and use method

Through the PID controller based on reinforcement learning, the PID parameters are automatically adjusted to adapt to different working conditions, which solves the shortcomings of traditional PID controllers in complex environments and realizes more efficient control of mobile engineering equipment operation mechanisms.

CN120065687AActive Publication Date: 2025-05-30ZHEJIANG UNIV OF TECH +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510126446.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-30
Estimated Expiration
2045-01-27

AI Technical Summary

Technical Problem

When traditional PID controllers face complex and changeable working environments and nonlinear system characteristics, they are difficult to meet the requirements of high-precision control, especially in terms of dynamic response speed, anti-interference ability and stability.

Method used

Using a PID controller based on reinforcement learning, the training agent automatically adjusts the parameters of the PID controller, so that it can adaptively optimize the control effect under different working conditions.

Benefits of technology

It realizes more accurate and efficient control of the unmanned mobile engineering equipment operation mechanism, can be adjusted according to the operation mode switching and control mode, and improves the system's adaptability and control performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120065687A_ABST
    Figure CN120065687A_ABST
Patent Text Reader

Abstract

The invention discloses a control system of a mobile engineering equipment operation mechanism based on a reinforcement learning PID controller and a use method, and aims to realize switching management of two operation modes of remote driving and automatic driving, control of instruction stream transmission and rapid switching and control of two joint control modes. The intelligent agent is trained through a reinforcement learning algorithm to automatically adjust parameters of the PID controller, so that the controller can adaptively optimize the control effect under different working conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of unmanned mobile engineering equipment, and particularly relates to a control system and a using method for a working mechanism of a mobile engineering equipment based on a reinforcement learning PID controller, so as to realize more accurate and efficient control of the working mechanisms of mobile engineering equipment such as push rakes, excavators, loaders, etc. Background Art

[0002] For mobile engineering equipment such as push rakes, excavators, and loaders, their working mechanisms usually include working mechanisms such as boom arms and buckets, and these working mechanisms are usually driven by hydraulic rods. For unmanned mobile engineering equipment, it is usually necessary to meet multiple working modes such as remote driving and autonomous driving, and during actual operation, the operator needs to perform switching control of the working modes according to the actual production conditions. Among the operation methods for the working mechanism, joint speed control and joint position control have their respective advantages. Speed control is suitable for situations where there are high requirements for action speed control and low requirements for position accuracy, such as excavation and loading; position control is suitable for situations where there are high requirements for position accuracy, such as leveling operations. If the operation method can be switched according to the operation needs during actual operation, the operation complexity can be reduced and the operation efficiency can be improved.

[0003] On the other hand, the control performance of these devices is directly related to work efficiency, safety, and energy consumption. In the control system, the PID controller is widely used in various industrial automation scenarios because of its simple structure, easy understanding, and implementation. However, in the face of complex and changeable working environments and non-linear system characteristics, the fixed parameters of the traditional PID controller are often difficult to meet the requirements of high-precision control, especially in terms of dynamic response speed, anti-interference ability, and stability. Summary of the Invention

[0004] The invention proposes a control system and a using method for a working mechanism of a mobile engineering equipment based on a reinforcement learning PID controller to solve the above problems, aiming to realize the switching management of two working modes of remote driving and autonomous driving, the control of instruction flow transmission, and the switching and control of two joint control modes; the intelligent agent is trained by a reinforcement learning algorithm to automatically adjust the parameters of the PID controller, so that the controller can adaptively optimize the control effect under different working conditions.

[0005] A control system for a working mechanism of a mobile engineering equipment based on a reinforcement learning PID controller is composed of a remote control and management platform, an autonomous driving control unit, an on-vehicle computer, a virtual commissioning system, a TCP-CAN protocol conversion unit, an electro-hydraulic proportional valve, a joint drive hydraulic rod, and a stroke sensor; the stroke sensor receives the actual position and actual speed of the joint drive hydraulic rod and transmits them to the TCP-CAN protocol conversion unit;

[0006] The remote control platform and the autonomous driving control unit send control instructions of the working mechanism to the vehicle-mounted computer via Ethernet. The content of the control instructions includes the working mode, control mode of the working mechanism, and the joint target speed or joint target position.

[0007] The vehicle-mounted computer includes a data communication server. The data communication server receives the control instructions from the remote control platform and the autonomous driving control unit, and transmits the control instructions to the working mode switching unit respectively. At the same time, the data communication server receives the information of the actual position and actual speed from the TCP-CAN protocol conversion unit, transmits the actual position to the position PID controller in the position controller, and transmits the actual speed to the speed PID controller in the speed controller.

[0008] The working mode switching unit selects control instructions from the data communication server according to the working mode and sends them to the joint control mode switching unit. The joint control mode switching unit judges the selected joint control mode. If it is the speed control mode, it sends the target speed data to the speed PID controller. If it is the position mode, it sends the target position to the position PID controller.

[0009] The position PID controller obtains the actual position of the joint driving hydraulic rod from the data communication server, calculates the position error through the difference between the actual position and the target position, and transmits the position error to the position control agent. The position control agent outputs a set of PID parameter increments to the position PID controller according to the current position error. The position PID controller calculates the target speed for controlling the joint driving hydraulic rod according to the new PID parameters and the current position error, and sends the target speed data to the speed PID controller. The speed PID controller calculates the speed error according to the received target speed and simultaneously obtains the actual speed data of the joint hydraulic rod in the data communication server, and transmits the speed error to the speed control agent. The speed control agent outputs a set of PID parameter increments to the speed PID controller according to the current speed error. The speed PID controller calculates the target valve opening of the electro-hydraulic proportional valve for controlling the joint driving hydraulic rod according to the new PID parameters and the current speed error, and sends it to the electro-hydraulic proportional valve through the TCP-CAN protocol conversion unit to realize the control of the joint driving hydraulic rod.

[0010] More specifically, the data communication server includes a "control section of the management and control platform", a "control section of autonomous driving", and a "vehicle status feedback section". Among them, the "control section of the management and control platform" is used to store the control data from the remote control platform; the "control section of autonomous driving" is used to store the control instructions from the intelligent control module; the "vehicle status feedback section" is used to store the vehicle status and sensor data fed back by the vehicle controller.

[0011] More specifically, the operation mode switching unit polls the "operation mode register" in the "system control area" of the "control platform control segment" at a frequency of 10 Hz, defining "1" as the remote driving mode and "2" as the automatic operation mode;

[0012] If the operation mode is "1", it is determined that the operation mode is the remote driving mode, and the data in the "control platform control segment" is sent as control data to the joint control mode switching unit;

[0013] If the operation mode is "2", it is determined that the operation mode is the automatic operation mode, and the data in the "automatic driving control segment" is sent as control data to the joint control mode switching unit;

[0014] The joint control mode switching unit determines the control mode selected by the remote control platform or the automatic driving control unit according to the data of the "joint control mode register" in the data of the "control platform segment" or the "automatic driving control segment" output by the "operation mode switching unit";

[0015] If the control mode is the "speed mode", the data of the "target speed register" of each joint in the control data is used as the reference input of the speed PID controller of each joint, and this reference input is sent to the speed controller;

[0016] If the control mode is "position control", the data of the "target position register" of each joint in the corresponding control segment data is used as the reference input of the position PID controller of each joint, and this reference input is sent to the position controller. The position PID controller calculates the desired speed of the controlled joint hydraulic rod, and then this desired speed is used as the reference input of the joint speed PID controller and sent to the speed controller.

[0017] More specifically, the usage method of the control system of the working mechanism of the mobile engineering equipment based on the reinforcement learning PID controller is as follows:

[0018] A. Speed mode;

[0019] A1. Define the state space s of the speed control agent as the current speed error e v (t) of the joint drive hydraulic rod,

[0020] e v (t) = S set - S actual (t)(1)

[0021] where S set is the target speed, and S actual is the actual speed;

[0022] A2. Define the action space, which is defined as the increments of the three gain values of the PID controller, i.e., ΔK p , ΔK i , ΔK d ;

[0023] A3. Design the reward function:

[0024] R t = -ω 1 |e t | + ω 2 I(|e t | < ε) + ω 3 ( - |I t |) + ω 4 ( - |D t |)

[0025] + ω 5 ( - (|ΔK p,t | + |ΔK i,t | + |ΔK d,t |)) (2)

[0026] where e t is the control error at the t-th step. For the speed controller, it is the speed error e v (t), and for the position controller, it is the position error e p (t); I(|e t | < ε) is an indicator function that is 1 when |e t | < ε and 0 otherwise; ε represents the acceptable error range; I t is the integral term at the t-th step; D t is the derivative term at the t-th step; ΔK p,t , ΔK i,t , ΔK d,t are the incremental adjustments to the PID gains at the t-th step respectively; ω 1 , ω 2 , ω 3 , ω 4 , ω 5 are non - negative weight parameters that are adjusted according to specific application requirements;

[0027] A4. Build a reinforcement learning agent:

[0028] Initialize the action network. Use the ReLU function as the activation function. At each time step, sample a specific action a according to the Gaussian distribution output by the action network t :

[0029] a t = μ t + σ t ·∈ (3)

[0030] where a t is the sampled action; μ t and σ t are the mean and standard deviation of the action distribution output by the action network respectively; ∈ is a random number sample drawn from the standard normal distribution;

[0031] Initialize the value network, use the ReLU function as the activation function, and the last layer outputs a scalar representing the value V(s t );

[0032] Initialize the optimizer and hyperparameters: Use the Adam optimizer and set the learning rate α. Usually, α is set to a small value to ensure the stability of the training process; Set the number of steps (n_steps) collected before each update, the discount factor γ, the maximum number of steps per episode, the batch size (batch_size), and the truncation constant ε;

[0033] A5. Train the constructed agent:

[0034] A5.1 The remote control platform and the autonomous driving control unit send the target speed and speed mode commands, which are transmitted to the joint control mode switching unit through the data communication server and the operation mode switching unit in sequence; The joint control mode switching unit switches to the speed mode and transmits the target speed to the speed PID controller; At the same time, the speed PID controller obtains the actual speed of the joint drive hydraulic rod from the data communication server, and calculates the current speed error e v (t) to obtain the state s t of the speed control agent. Input the state into the action network of the speed control agent. According to the action distribution output by the action network of the speed control agent, sample a specific action a t ; The sampled action a t is the increment of the speed PID controller parameter, and is accumulated with the speed PID controller parameter of the previous moment to obtain a new PID parameter; The speed PID controller calculates the control quantity u v (t), and u v (t) is the opening percentage of the electro-hydraulic proportional valve in this system;

[0035]

[0036] A5.2 Interactive training of the speed control agent;

[0037] A5.2.1 If the joint drive hydraulic rod is selected to train the speed control agent, the speed PID controller will use u v(t) is input into the electro-hydraulic proportional valve through the TCP-CAN protocol conversion unit, and then the joint driving hydraulic rod is controlled; the joint driving hydraulic rod outputs new speed data, and the speed data is transmitted back to the speed PID controller through the stroke sensor, the TCP-CAN protocol conversion unit and the data communication server; the speed PID controller calculates a new state s by combining the currently given target control speed t+1 , and calculates the reward r t , calculates the advantage function, and stores the experience (s t , a t , r t , s t+1 ) in the experience pool of the speed control agent, completing an interaction between the speed control agent and the environment;

[0038] A5.2.2 If the virtual commissioning system is selected to train the speed control agent, then the speed PID controller will input u v (t) into the virtual commissioning system; the virtual commissioning system outputs new speed data, and the data communication server transmits it back to the speed PID controller; the speed PID controller calculates a new state s by combining the currently given target control speed t+1 , and calculates the reward r t , calculates the advantage function, and stores the experience (s t , a t , r t , s t+1 ) in the experience pool of the speed control agent, completing an interaction between the speed control agent and the environment;

[0039] A5.3 Repeat the selected speed control agent interaction training in A5.2 for n_steps times, and use generalized advantage estimation to calculate the advantage function

[0040]

[0041] δ t+l = r t+l + γV(s t+l+1 ) - V(s t+l ) (6)

[0042] where t is the current time step. l is the offset starting from the current time step t. γ is the discount factor, indicating the importance of future rewards; λ is the smoothing parameter of GAE, controlling the time range of advantage estimation; δ t+l is the TD error at time step t + l, representing the difference between the immediate reward and the value estimate; r t+l is the reward at time step t + l, calculated by Equation (2);

[0043] At each time step, calculate the generalized advantage estimate and stored in the experience pool of the speed control agent together with the experience for guiding the update of the action network;

[0044] A5.4 The speed control agent updates the action network according to the objective function of the action network: The update objective of the action network is to maximize the expected cumulative reward while ensuring that the new policy does not deviate too far from the old policy. The objective function of the action network is:

[0045]

[0046] where, is the policy ratio, ε is the truncation ratio, is the mathematical expectation;

[0047] Perform gradient ascent on the action network according to the above objective function for multiple mini-batch updates to maximize the objective function; As the parameters of the action network of the speed control agent are updated, the PID parameters generated by the action network can make the control performance of the speed PID controller tend to the optimal response speed and control error;

[0048] A5.5 The speed control agent updates the value network according to the loss function of the value network: The value network performs mini-batch sampling based on the data in the experience pool; Calculate the loss function L VF (θ v ) of the value network, and adjust the parameters of the value network by gradient descent to reduce the estimation error of the value network for the reward value that can be obtained by the action network adopting the current policy; The loss function L VF (θ v ) of the value network is the mean square error, which is used to minimize the gap between the predicted value and the target value:

[0049]

[0050] where, V(s t ) is the value estimate output by the value network, is the generalized advantage estimate, θ v is the parameter of the value network; V θv (s t ) is the predicted value output by the value network in the state s t ; represents the expected value of the variance of all predicted values and target values in the time step t; V t target is the target value, which is calculated using the generalized advantage estimate:

[0051]

[0052] V(s t ) is the value estimate output by the value network, is the generalized advantage estimation;

[0053] Finally, the value network estimates the expected cumulative reward that can be obtained by following the policy of the current action network starting from the given state s, calculates the advantage function, and uses it to calculate the objective function of the action network to update the action network;

[0054] A5.6 Repeat the training steps of A5.4 - A5.5 for the speed control agent, continuously update the action network and the value network until the average reward tends to be stable or reaches the predetermined maximum number of iterations; after the update is completed, stop the training and save the parameters of the action network and the value network. At this time, the PID control parameters output by the action network can make the speed tracking performance of the speed PID controller reach the best response speed and the smallest control error as much as possible;

[0055] A5.7 If in A5.2, the speed control agent is trained using the virtual debugging system in A5.2.2, then additionally execute A5.2.1 to train the speed control agent using the joint - driven hydraulic rod, and repeat the training process of A5.2 - A5.6 to further adaptively train the speed control agent to optimize the performance loss of the agent caused by the mathematical model error between the virtual debugging system and the real system;

[0056] A6. Put the speed mode into actual application. The remote management and control platform and the autonomous driving control unit send the target speed and speed mode instructions, which are transmitted to the joint control mode switching unit through the data communication server and the operation mode switching unit in sequence; the joint control mode switching unit switches to the speed mode and transmits the target speed to the speed PID controller;

[0057] The action network of the speed control agent generates a set of PID parameter increments according to the current speed error, accumulates this parameter increment to the original PID controller, and then the speed PID controller calculates the target opening of the electro - hydraulic proportional valve to control the joint - driven hydraulic rod according to the control error between the current target speed and the actual speed. The action network outputs different PID parameter increments according to different input states, so that the speed controller can implement the best PID parameters to control the joint - driven hydraulic rod according to the current input state to achieve the best speed tracking control effect;

[0058] B. Position mode;

[0059] B1. Define the state space s as the current speed error e p (t) of the joint - driven hydraulic rod, e p (t) = L set - L actual (t) (9) where L set is the target position, L actualis the actual position;

[0060] Execute A2 - A4 on the position PID controller, construct the position control agent, and initialize the parameter settings; B2. Train the constructed position control agent:

[0061] B2.1 The position controller interacts with the controlled object to obtain experience: The remote management and control platform and the autonomous driving control unit send the target position and the position mode instruction, which are transmitted to the joint control mode switching unit through the data communication server and the operation mode switching unit in sequence; The joint control mode switching unit switches to the position mode and transmits the target position to the position PID controller; At the same time, the position PID controller obtains the actual position of the joint drive hydraulic rod from the data communication server, and calculates the current position error e p (t) as the state s of the position control agent t , input the state into the action network of the position control agent, and sample a specific action a according to the action distribution output by the action network of the position control agent t ; The sampled action a t is the parameter increment of the position PID controller, and is accumulated with the parameters of the position PID controller at the previous moment to obtain the new parameters of the position PID controller, and calculate the control quantity u p (t), take u p (t) as the reference input of the speed controller and input it into the speed PID controller, and the speed PID controller calculates the opening of the electro - hydraulic proportional valve;

[0062] B2.2.1 If it is selected to use the joint drive hydraulic rod to train the speed position control agent, then the speed PID controller will input u v (t) into the electro - hydraulic proportional valve through the TCP - CAN protocol conversion unit, and then control the joint drive hydraulic rod; The joint drive hydraulic rod outputs new position data, and transmits it back to the position PID controller through the stroke sensor, the TCP - CAN protocol conversion unit and the data communication server; The speed PID controller calculates the new state s t+1 , and calculates the reward r t , calculate the advantage function, and store the experience (s t , a t , r t , s t+1 ) in the experience pool of the position control agent, and complete an interaction between the position control agent and the environment;

[0063] B2.2.2 If it is selected to use the virtual commissioning system to train the position control agent, then the speed PID controller will input u v(t) Input virtual debugging system; the virtual debugging system outputs new position data, which is sent back to the position PID controller through the data communication server; the position PID controller calculates a new state s by combining the current given target position t+1 , and calculates the reward r t , calculates the advantage function, and stores the experience (s t , a t , r t , s t+1 ) in the experience pool of the position control agent, completing an interaction between the position control agent and the environment;

[0064] B2.3 After the position control agent interacts with the environment for n_steps times, calculate the generalized advantage estimate according to Equation (5) For each time step t, calculate the generalized advantage estimate and store it in the experience pool of the position controller together with the experience, which is used to guide the update of the action network of the position control agent;

[0065] B2.4 Calculate the objective function of the action network of the position control agent according to Equation (7), and perform gradient ascent on the parameters of the action network of the position control agent according to the objective function to maximize its objective function;

[0066] B2.5 Calculate the loss function of the value network of the position control agent according to Equation (8), and the loss function of the value network of the position control agent performs gradient descent to adjust the parameters of the value network of the position control agent, so as to reduce the estimation error of the reward value that the value network can obtain by the action network adopting the current policy under the given state, so as to more accurately evaluate the quality of the current action network of the position control agent, and further more accurately guide the update of the action network of the position control agent;

[0067] B2.6 Execute the training steps of B2.1~B2.5 on the position controller. The position control agent continuously adjusts the action and value networks until the average reward tends to be stable or reaches the predetermined maximum number of iterations; after the update is completed, the PID control parameters output by the action network can make the position tracking performance of the position PID controller reach the best response speed and as small a control error as possible; when the performance of the position control agent reaches the expected goal, stop training and save the parameters of the action network and the value network;

[0068] B2.7 If the position control agent is trained using a virtual debugging system in B2.2.2, then additionally execute B2.2.1 to train the speed control agent using a joint-driven hydraulic rod, and repeat the training process of B2.2 - B2.5 to further adaptively train the position control agent to optimize the performance loss of the agent caused by the mathematical model error between the virtual debugging system and the real system;

[0069] B3. Put the position mode into practical application. The remote management and control platform and the autonomous driving control unit send the target position and the position mode instruction, which are transmitted to the joint control mode switching unit through the data communication server and the operation mode switching unit in sequence; the joint control mode switching unit switches to the position mode and transmits the target position to the position PID controller;

[0070] The action network of the position control agent generates a set of PID parameter increments based on the error between the actual position and the target position, and accumulates this parameter increment to the position PID controller. The position PID controller then calculates the target speed of the joint-driven hydraulic rod using the new PID parameters according to the control error between the target position and the actual position, and this target speed is tracked and controlled by the speed PID controller; the action network of the position PID controller outputs different PID parameter increments according to different input states, so that the position PID controller can implement control on the joint-driven hydraulic rod with the best PID parameters according to the current input state.

[0071] The beneficial effects produced by the present invention are as follows:

[0072] 1. By designing the data communication server, the operation mode switching unit, the joint control mode switching unit, the position controller, and the speed controller, the operation mode switching control and the operation method switching control of the working mechanism of the unmanned mobile engineering equipment are realized. The advantage is that the operator can quickly and conveniently switch the operation mode and the operation method at any time according to the actual operation conditions.

[0073] 2. A joint hydraulic servo control method based on a PID controller of reinforcement learning is proposed. By designing and training the speed control agent and the position control agent, the speed controller and the position controller can adaptively adjust the PID controller parameters in different control error states, exert the best control effect of the controller, without human participation, and improve the adaptability of the hydraulic servo system based on the PID controller to complex environmental changes. Brief Description of the Drawings

[0074] Figure 1 is the structural diagram of the control system of the working mechanism of the mobile engineering equipment based on the reinforcement learning PID controller of the present invention.

[0075] Figure 2 It is a schematic diagram of the register partition of the data communication server of the present invention.

[0076] Figure 3 It is a control flow chart for switching between the operation mode and the control mode of the present invention.

[0077] Figure 4 It is a schematic diagram of the learning principle of the speed control agent and the position control agent of the present invention. Detailed implementation manners

[0078] The present invention will be further described in detail below with reference to the accompanying drawings.

[0079] In this embodiment, a Komatsu D65E push rake is adopted. The working mechanism of the push rake has two degrees of freedom, which are respectively driven by hydraulic rods, the opening of the hydraulic circuit is controlled by an electro-hydraulic proportional valve, and the control is carried out through the CAN bus. A high-performance vehicle-mounted computer is equipped at the vehicle end, and communicates with the remote management and control platform and the autonomous driving control unit through Ethernet. The vehicle-mounted computer controls the electro-hydraulic proportional valve through a TCP-CAN protocol converter. A stroke sensor is installed on the driven hydraulic rod to measure the current length of the hydraulic rod, and the data is fed back to the vehicle-mounted computer through the TCP-CAN protocol converter at a frequency of 100 Hz.

[0080] As Figure 1 shown, the control system involved in a control method for the working mechanism of a mobile engineering equipment based on a reinforcement learning PID controller proposed by the present invention is composed of a remote management and control platform, an autonomous driving control unit, a vehicle-mounted computer, a virtual commissioning system, a TCP-CAN protocol conversion unit, an electro-hydraulic proportional valve, a joint-driven hydraulic rod and a stroke sensor.

[0081] Among them, the remote control platform and the autonomous driving control unit send control instructions of the working mechanism to the vehicle-mounted computer through Ethernet, including the working mode, control mode of the working mechanism, and the joint target speed or joint target position. The vehicle-mounted computer is a computing carrier for realizing the control of the working mechanism of mobile engineering equipment. It receives control instructions from the remote control platform and the autonomous driving unit, obtains data from the travel sensor, calculates the opening degree of the electro-hydraulic proportional valve of the joint drive hydraulic rod according to the current state and instructions, and sends the electro-hydraulic proportional valve opening degree data to the virtual commissioning system or sends it to the electro-hydraulic proportional valve through the TCP-CAN protocol conversion unit. The virtual commissioning system is used to replace the real controlled object, receive the electro-hydraulic proportional valve opening degree instruction from the vehicle-mounted computer, and output the current joint drive hydraulic rod position through internal simulation calculation, which is used to train the constructed PID controller based on reinforcement learning. The role of the TCP-CAN protocol conversion unit is to perform protocol conversion and data forwarding between the TCP / IP communication protocol used by the vehicle-mounted computer and the CAN communication protocol used by the travel sensor and the electro-hydraulic proportional valve. The electro-hydraulic proportional valve is used for the opening degree control of the hydraulic circuit of the joint drive hydraulic rod, and adjusts the valve opening degree according to the received valve target opening degree value. The joint drive hydraulic rod is used to drive the working mechanism of mobile engineering equipment to act. The travel sensor is installed in parallel with the joint drive hydraulic rod, and is used to measure the actual length of the joint drive hydraulic rod in real time, and send the measured length data of the joint hydraulic rod to the vehicle-mounted computer through the TCP-CAN protocol conversion unit.

[0082] The vehicle-mounted computer includes a data communication server, a working mode switching unit, a joint control mode switching unit, and a position controller and a speed controller based on reinforcement learning.

[0083] The data communication server receives the operation mode data, control mode data, and joint target speed or joint target position command data of the working mechanism from the remote control platform and the autonomous driving control unit, and at the same time receives the travel sensor data forwarded by the TCP-CAN protocol conversion unit; the operation mode switching unit selects control instructions from the data communication server according to the operation mode and sends them to the joint control mode switching unit; the joint control mode switching unit judges the selected joint control mode. If it is the speed control mode, it sends the target speed data to the speed controller. If it is the position mode, it sends the target position to the position controller; the position controller obtains the actual position of the joint drive hydraulic rod from the data communication server, calculates the target speed for controlling the joint drive hydraulic rod in combination with the target position data, and sends the target speed data to the speed controller; the speed controller calculates the target valve opening of the electro-hydraulic proportional valve for controlling the joint drive hydraulic rod according to the received target speed and at the same time obtains the actual speed data of the joint hydraulic rod in the data communication server, and sends it to the electro-hydraulic proportional valve through the TCP-CAN protocol conversion unit to realize the control of the joint drive hydraulic rod.

[0084] The construction steps of the data communication server are as follows: First, allocate a section of memory in the vehicle-mounted computer as a data register, and divide the register into a "control platform control section", an "autonomous driving control section", and a "vehicle status feedback section" according to the source of the control instructions. Among them, the "control platform control section" is used to store the control data from the remote control platform; the "autonomous driving control section" is used to store the control instructions from the intelligent control module; the "vehicle status feedback section" is used to store the vehicle status and sensor data feedback by the vehicle controller. The data communication server conducts data communication with the remote control platform, the autonomous driving control unit, and the TCP-CAN protocol conversion unit through the TCP / IP protocol respectively.

[0085] Furthermore, divide the "control platform control section" into a "motion control area" and a "system control area". Allocate and define the "target position register" and "target speed register" of each joint of the working mechanism in the "motion control area"; allocate and define the "operation mode register" and "joint control mode register" in the "system control area".

[0086] Divide the "autonomous driving control section" into a "motion control area" and a "control mode area". Allocate and define the "target position register" and "target speed register" of each joint of the working mechanism in the "motion control area"; allocate and define the "joint control mode register" in the "control mode area".

[0087] The operation mode switching unit reads the operation mode data in the data communication server in real time, and selects to obtain the control mode and motion control data from either the "control platform driving section" or the "automatic driving control section" according to the operation mode, and sends them to the joint control mode switching unit. The specific judgment process is as follows:

[0088] The operation mode switching unit polls the "operation mode register" in the "system control area" of the "control platform control section" at a frequency of 10 Hz, and defines "1" as the remote driving mode and "2" as the automatic operation mode.

[0089] If the operation mode is "1", it is determined that the operation mode is the remote driving mode, and the data in the "control platform control section" is sent as control data to the joint control mode switching unit;

[0090] If the operation mode is "2", it is determined that the operation mode is the automatic operation mode, and the data in the "automatic driving control section" is sent as control data to the joint control mode switching unit.

[0091] The joint control mode switching unit determines the control mode selected by the remote control platform or the automatic driving control unit according to the data of the "joint control mode register" in the data of the "control platform section" or the "automatic driving control section" output by the "operation mode switching unit".

[0092] If the control mode is the "speed mode", the data of the "target speed register" of each joint in the control data is used as the reference input of the speed PID controller of each joint, and this reference input is sent to the speed controller.

[0093] If the control mode is "position control", the data of the "target position register" of each joint in the corresponding control section data is used as the reference input of the position PID controller of each joint, and this reference input is sent to the position controller. The position PID controller calculates the expected speed of the controlled joint hydraulic rod, and then this expected speed is used as the reference input of the joint speed PID controller and sent to the speed controller.

[0094] The actual position of the driving hydraulic rod is obtained from the stroke sensor through the TCP-CAN protocol conversion unit. The actual speed of the driving hydraulic rod is obtained by differentiating the actual position, and is stored in the "vehicle state feedback section" to provide state feedback to the PID controller.

[0095] The usage method of the control system of the above Komatsu D65E push rake is as follows:

[0096] A. Speed mode;

[0097] A1. Define the state space s of the speed control agent as the current speed error e of the joint driving hydraulic rod v(t),

[0098] e v (t) = S set -S actual (t) (1)

[0099] where S set is the target speed and S actual is the actual speed;

[0100] A2. Define the action space, which is defined as the increments of the three gain values of the PID controller, i.e., ΔK p , ΔK i , ΔK d ;

[0101] A3. Design the reward function:

[0102] R t = -ω 1 |e t | + ω 2 I(|e t | < ε) + ω 3 (-|I t |) + ω 4 (-|D t |)

[0103] + ω 5 (-(|ΔK p,t | + |ΔK i,t | + |ΔK d,t |)) (2

[0104] where e t is the control error at the t-th step. For the speed controller, it is the speed error e v (t), and for the position controller, it is the position error e p (t); I(|e t | < ε) is the indicator function, which is 1 when |e t | < ε and 0 otherwise; ε represents the acceptable error range; I t is the integral term at the t-th step; D t is the derivative term at the t-th step; ΔK p,t , ΔK i,t , ΔK d,t are the incremental adjustments to the PID gains at the t-th step respectively; ω 1 , ω 2 , ω 3 , ω 4 , ω 5 are non - negative weight parameters, which are adjusted according to specific application requirements;

[0105] A4. Constructing a Reinforcement Learning Agent:

[0106] Initialize the action network with the ReLU function as the activation function. At each time step, sample a specific action a according to the Gaussian distribution output by the action network t :

[0107] a t = μ t + σ t ·∈(3)

[0108] where a t is the sampled action; μ t and σ t are the mean and standard deviation of the action distribution output by the action network respectively; ∈ is a random number sample drawn from the standard normal distribution;

[0109] Initialize the value network with the ReLU function as the activation function. The last layer outputs a scalar representing the value V(s t );

[0110] Initialize the optimizer and hyperparameters: Use the Adam optimizer and set the learning rate α. Usually, α is set to a small value to ensure the stability of the training process; Set the number of steps (n_steps) collected before each update, the discount factor γ, the maximum number of steps per episode, the batch size (batch_size), and the truncation constant ε;

[0111] A5. Training the Constructed Agent:

[0112] A5.1 The remote control platform and the autonomous driving control unit send the target speed and speed mode commands, which are transmitted to the joint control mode switching unit through the data communication server and the operation mode switching unit in sequence; The joint control mode switching unit switches to the speed mode and transmits the target speed to the speed PID controller; At the same time, the speed PID controller obtains the actual speed of the joint drive hydraulic rod from the data communication server and calculates the current speed error e v (t) to obtain the state s of the speed control agent t , and input the state into the action network of the speed control agent. According to the action distribution output by the action network of the speed control agent, sample a specific action a t ; The sampled action a t is the increment of the speed PID controller parameter and is accumulated with the speed PID controller parameter at the previous moment to obtain a new PID parameter; The speed PID controller calculates the control quantity u v (t), and u v (t) is the opening percentage of the electro-hydraulic proportional valve in this system;

[0113]

[0114] A5.2 Speed control intelligent agent interactive training;

[0115] A5.2.1 If the joint-driven hydraulic rod is selected to train the speed control intelligent agent, the speed PID controller will input u v (t) into the electro-hydraulic proportional valve through the TCP-CAN protocol conversion unit, and then control the joint-driven hydraulic rod; the joint-driven hydraulic rod outputs new speed data, and transmits it back to the speed PID controller through the stroke sensor, TCP-CAN protocol conversion unit and data communication server; the speed PID controller calculates a new state s t+1 and calculates the reward r t , calculates the advantage function, and stores the experience (s t , a t , r t , s t+1 ) in the experience pool of the speed control intelligent agent to complete an interaction between the speed control intelligent agent and the environment;

[0116] A5.2.2 If the virtual commissioning system is selected to train the speed control intelligent agent, the speed PID controller will input u v (t) into the virtual commissioning system; the virtual commissioning system outputs new speed data, and the data communication server transmits it back to the speed PID controller; the speed PID controller calculates a new state s t+1 and calculates the reward r t , calculates the advantage function, and stores the experience (s t , a t , r t , s t+1 ) in the experience pool of the speed control intelligent agent to complete an interaction between the speed control intelligent agent and the environment;

[0117] A5.3 Repeat the selected speed control intelligent agent interactive training in A5.2 for n_steps times, and use generalized advantage estimation to calculate the advantage function

[0118]

[0119] δ t+l = r t+l + γV(s t+l+1 ) - V(s t+l ) (6)

[0120] where \(t\) is the current time step, \(l\) is the offset starting from the current time step \(t\), \(\gamma\) is the discount factor representing the importance of future rewards, \(\lambda\) is the smoothing parameter of GAE that controls the time range of advantage estimation, and \(\delta\) t+l is the TD error at time step \(t + l\), representing the difference between the immediate reward and the value estimate; \(r\) t+l is the reward at time step \(t + l\), calculated by Equation (2);

[0121] At each time step, the generalized advantage estimate is calculated and stored in the experience pool of the speed control agent together with the experience to guide the update of the action network;

[0122] A5.4 The speed control agent updates the action network according to the objective function of the action network: The update objective of the action network is to maximize the expected cumulative reward while ensuring that the new policy does not deviate too far from the old policy. The objective function of the action network is:

[0123]

[0124] where is the policy ratio and \(\epsilon\) is the truncation ratio, is the mathematical expectation;

[0125] Perform gradient ascent according to the above objective function to perform multiple mini - batch updates on the action network to maximize the objective function; as the parameters of the action network of the speed control agent are updated, the PID parameters generated by the action network can make the control performance of the speed PID controller tend to the optimal response speed and control error;

[0126] A5.5 The speed control agent updates the value network according to the loss function of the value network: The value network performs mini - batch sampling based on the data in the experience pool; calculate the loss function \(L\) of the value network VF (\(\theta\) v ), and adjust the parameters of the value network by gradient descent to reduce the estimation error of the value network for the reward value that the action network can obtain by taking the current policy; the loss function \(L\) of the value network VF (\(\theta\) v ) is the mean square error, used to minimize the gap between the predicted value and the target value:

[0127]

[0128] where \(V(s\) t ) is the value estimate output by the value network, is the generalized advantage estimate, \(\theta\) v are the parameters of the value network; \(V\) θv (s t ) is the value network at state \(s\) tThe predictive value of the following output; Denotes taking the expected value of the variance between all predicted values and the target value at time step t; V t target is the target value, calculated using Generalized Advantage Estimation:

[0129]

[0130] V(s t ) is the value estimate output by the value network, is the Generalized Advantage Estimation;

[0131] Finally, the value network estimates the expected cumulative reward that can be obtained by following the policy of the current action network starting from the given state s, calculates the advantage function, and uses it to calculate the objective function of the action network to update the action network;

[0132] A5.6 Repeat the training steps of A5.4 - A5.5 for the speed control agent, continuously updating the action network and the value network until the average reward stabilizes or reaches the predetermined maximum number of iterations; after the update is completed, stop the training and save the parameters of the action network and the value network. At this time, the PID control parameters output by the action network can make the speed tracking performance of the speed PID controller reach the best response speed and the smallest control error;

[0133] A5.7 If in A5.2, A5.2.2 is executed to train the speed control agent using the virtual debugging system, then additionally execute A5.2.1 to train the speed control agent using the joint-driven hydraulic rod, and repeat the training process of A5.2 - A5.6 to further adaptively train the speed control agent to optimize the performance loss of the agent caused by the mathematical model error between the virtual debugging system and the real system;

[0134] A6. Put the speed mode into actual application. The remote management and control platform and the autonomous driving control unit send the target speed and speed mode commands, which are transmitted to the joint control mode switching unit through the data communication server and the operation mode switching unit in sequence; the joint control mode switching unit switches to the speed mode and transmits the target speed to the speed PID controller;

[0135] The action network of the speed control agent generates a set of PID parameter increments based on the current speed error, accumulates the parameter increments to the original PID controller, and the speed PID controller then calculates the target opening of the electro-hydraulic proportional valve to control the joint drive hydraulic rod according to the control error between the current target speed and the actual speed; the action network outputs different PID parameter increments according to different input states, enabling the speed controller to implement control on the joint drive hydraulic rod with the optimal PID parameters according to the current input state to achieve the best speed tracking control effect;

[0136] B. Position mode;

[0137] B1. Define the state space s as the current speed error e p (t) of the joint drive hydraulic rod, e p (t) = L set -L actual (t) (9) where L set is the target position, L actual is the actual position;

[0138] Execute A2 - A4 on the position PID controller, construct the position control agent, and initialize the parameter settings; B2. Train the constructed position control agent:

[0139] B2.1 The position controller interacts with the controlled object to obtain experience: The remote management and control platform and the autonomous driving control unit send the target position and the position mode instruction, which are transmitted to the joint control mode switching unit through the data communication server and the operation mode switching unit in sequence; the joint control mode switching unit switches to the position mode and transmits the target position to the position PID controller; meanwhile, the position PID controller obtains the actual position of the joint drive hydraulic rod from the data communication server, calculates the current position error e p (t) as the state s t of the position control agent, inputs the state into the action network of the position control agent, samples a specific action a t according to the action distribution output by the action network of the position control agent; the sampled action a t is the parameter increment of the position PID controller, and is accumulated with the parameters of the position PID controller at the previous moment to obtain the new parameters of the position PID controller, calculates the control quantity u p (t), takes u p (t) as the reference input of the speed controller, inputs it into the speed PID controller, and the speed PID controller calculates the opening of the electro-hydraulic proportional valve;

[0140] B2.2.1 If the joint-driven hydraulic rod is selected to train the velocity-position control agent, the velocity PID controller will input u v (t) into the electro-hydraulic proportional valve through the TCP-CAN protocol conversion unit, and then control the joint-driven hydraulic rod; the joint-driven hydraulic rod outputs new position data, and transmits it back to the position PID controller through the stroke sensor, TCP-CAN protocol conversion unit and data communication server; the velocity PID controller calculates a new state s t+1 by combining the currently given target position, and calculates the reward r t , calculates the advantage function, and stores the experience (s t , a t , r t , s t+1 ) in the experience pool of the position control agent, completing one interaction between the position control agent and the environment;

[0141] B2.2.2 If the virtual commissioning system is selected to train the position control agent, the velocity PID controller will input u v (t) into the virtual commissioning system; the virtual commissioning system outputs new position data and transmits it back to the position PID controller through the data communication server; the position PID controller calculates a new state s t+1 by combining the currently given target position, and calculates the reward r t , calculates the advantage function, and stores the experience (s t , a t , r t , s t+1 ) in the experience pool of the position control agent, completing one interaction between the position control agent and the environment;

[0142] After the position control agent interacts with the environment for n_steps times, calculate the generalized advantage estimation according to Equation (5) For each time step t, calculate the generalized advantage estimation and store it in the experience pool of the position controller together with the experience, which is used to guide the update of the action network of the position control agent;

[0143] B2.4 Calculate the objective function of the action network of the position control agent according to Equation (7), and perform gradient ascent on the parameters of the action network of the position control agent according to the objective function to maximize its objective function;

[0144] B2.5 Calculate the loss function of the position control agent value network according to Equation (8). The loss function of the position control agent value network performs gradient descent to adjust the parameters of the position control agent value network, so as to reduce the estimation error of the reward value that its value network can obtain by the action network adopting the current policy under the given state, thereby more accurately evaluating the quality of the current position control agent action network, and further more accurately guiding the update of the position control agent action network;

[0145] B2.6 Execute the training steps of B2.1 - B2.5 for the position controller. The position control agent continuously adjusts the action and value networks until the average reward tends to be stable or reaches the predetermined maximum number of iterations; after the update is completed, the PID control parameters output by the action network can make the position tracking performance of the position PID controller reach the best response speed and as small a control error as possible; stop training when the performance of the position control agent reaches the expected goal, and save the parameters of the action network and the value network;

[0146] B2.7 If in B2.2, B2.2.2 is executed to train the position control agent using the virtual commissioning system, then additionally execute B2.2.1 to train the speed control agent using the joint - driven hydraulic rod, and repeat the training process of B2.2 - B2.5 to perform further adaptive training on the position control agent to optimize the performance loss of the agent caused by the mathematical model error between the virtual commissioning system and the real system;

[0147] B3. Put the position mode into practical application. The remote management and control platform and the autonomous driving control unit send the target position and the position mode command, which are transmitted to the joint control mode switching unit through the data communication server and the operation mode switching unit in sequence; the joint control mode switching unit switches to the position mode and transmits the target position to the position PID controller;

[0148] The action network of the position control agent generates a set of PID parameter increments according to the error between the actual position and the target position, and accumulates this parameter increment to the position PID controller. The position PID controller then calculates the target speed of the joint - driven hydraulic rod using the new PID parameters according to the control error between the target position and the actual position. This target speed is tracked and controlled by the speed PID controller; the action network of the position PID controller outputs different PID parameter increments according to different input states, so that the position PID controller can implement control on the joint - driven hydraulic rod with the best PID parameters according to the current input state.

[0149] It should be emphasized that the implementation cases described in the present invention are illustrative rather than restrictive. Therefore, the present invention includes but is not limited to the embodiments described in the specific implementation plans. Any other similar implementation manners obtained by those skilled in the art based on the technical solution of the present invention also fall within the protection scope of the present invention.

Claims

1. A control system for a mobile engineering equipment operating mechanism based on a reinforcement learning PID controller, characterized in that: It consists of a remote control platform, an automatic driving control unit, an on-board computer, a virtual debugging system, a TCP-CAN protocol conversion unit, an electro-hydraulic proportional valve, a joint drive hydraulic rod and a stroke sensor; the stroke sensor receives the actual position and actual speed of the joint drive hydraulic rod and transmits it to the TCP-CAN protocol conversion unit; The remote control platform and the automatic driving control unit send control instructions of the operating mechanism to the on-board computer via Ethernet. The content of the control instructions includes the operating mode, control mode and joint target speed or joint target position of the operating mechanism; The on-board computer includes a data communication server, which receives control instructions from the remote control platform and the automatic driving control unit, and transmits the control instructions to the operation mode switching unit respectively; at the same time, the data communication server receives information on the actual position and actual speed from the TCP-CAN protocol conversion unit, transmits the actual position to the position PID controller in the position controller, and transmits the actual speed to the speed PID controller in the speed controller; The operation mode switching unit selects a control instruction from the data communication server according to the operation mode and sends it to the joint control mode switching unit; The joint control mode switching unit determines the selected joint control mode, and if it is a speed control mode, sends the target speed data to the speed PID controller; If it is position mode, the target position is sent to the position PID controller; The position PID controller obtains the actual position of the joint-driven hydraulic rod from the data communication server, calculates the position error through the difference between the actual position and the target position, and transmits the position error to the position control agent. The position control agent outputs a set of PID parameter increments to the position PID controller according to the current position error. The position PID controller calculates the target speed of controlling the joint-driven hydraulic rod according to the new PID parameters and the current position error, and sends the target speed data to the speed PID controller; the speed PID controller obtains the actual speed data of the joint hydraulic rod in the data communication server according to the received target speed, calculates the speed error, and transmits the speed error to the speed control agent; the speed control agent outputs a set of PID parameter increments to the speed PID controller according to the current speed error. The speed PID controller calculates the target valve opening of the electro-hydraulic proportional valve that controls the joint-driven hydraulic rod according to the new PID parameters and the current speed error, and sends it to the electro-hydraulic proportional valve through the TCP-CAN protocol conversion unit to realize the control of the joint-driven hydraulic rod.

2. According to claim 1, a control system for a mobile engineering equipment operating mechanism based on a reinforcement learning PID controller is characterized in that: The data communication server contains the "control platform control segment", "autonomous driving control segment" and "vehicle status feedback segment"; the "control platform control segment" is used to store control data from the remote control platform; the "autonomous driving control segment" is used to store control instructions from the intelligent control module; the "vehicle status feedback segment" is used to store the vehicle status and sensor data fed back by the vehicle controller.

3. According to claim 2, a control system for a mobile engineering equipment operating mechanism based on a reinforcement learning PID controller is characterized in that: The operation mode switching unit polls the "operation mode register" of the "system control area" in the "control platform control section" at a frequency of 10 Hz, defining "1" as the remote driving mode and "2" as the automatic operation mode; If the operation mode is "1", the operation mode is judged to be the remote driving mode, and the data in the "control platform control segment" is sent to the joint control mode switching unit as control data; If the operation mode is "2", the operation mode is determined to be the automatic operation mode, and the data in the "automatic driving control section" is sent to the joint control mode switching unit as control data; The joint control mode switching unit determines the control mode selected by the remote control platform or the automatic driving control unit according to the data of the "joint control mode register" in the "control platform segment" or "automatic driving control segment" data output by the "operation mode switching unit"; If the control mode is "speed mode", the "target speed register" data of each joint in the control data is used as the reference input of the speed PID controller of each joint, and this reference input is sent to the speed controller; If the control mode is "position control", the data of the "target position register" of each joint in the corresponding control segment data will be used as the reference input of each joint position PID controller, and this reference input will be sent to the position controller. The position PID controller will calculate the expected speed of the hydraulic rod of the controlled joint, and then send this expected speed as the reference input of the joint speed PID controller to the speed controller.

4. According to claim 1, a method for using a control system of a mobile engineering equipment operating mechanism based on a reinforcement learning PID controller comprises the following steps: A.Speed ​​mode; A1. Define the state space s of the speed control agent as the current speed error e of the joint-driven hydraulic rod v (t), e v (t)=S set -S actual (t) (1) Where S set is the target speed, S actual is the actual speed; A2. Define the action space, which is defined as the increment of the three gain values ​​of the PID controller, namely ΔK p , ΔK i , ΔK d ; A3. Design reward function: R t =-ω1|e t |+ω2I(e t |<ε)+ω3(-|I t |)+ω4(-|D t |)+ω5(-(|ΔK p,t |+|ΔK i,t |+|ΔK d,t |)) (2) Among them, e t is the control error of step t, and for the speed controller it is the speed error e v (t), for the position controller it is the position error e p (t);I(|e t |<ε) is the indicator function, when |e t |<ε is 1, otherwise it is 0; ε represents the acceptable error range; I t is the integral term of the tth step; D t is the differential term at step t; ΔK p,t , ΔK i,t , ΔK d,t are the incremental adjustments to the PID gain in step t; ω1, ω2, ω3, ω4, ω5 are non-negative weight parameters, which are adjusted according to specific application requirements; A4. Building a reinforcement learning agent: Initialize the action network, use the ReLU function as the activation function, and sample the specific action a according to the Gaussian distribution output by the action network at each time step t : a t =μ t +s t ·∈ (3) Among them, a t is the sampled action; μ t and σ t are the mean and standard deviation of the action distribution output by the action network respectively; ∈ is a random number sample drawn from a standard normal distribution; Initialize the value network, use the ReLU function as the activation function, and the last layer outputs a scalar representing the value of the state V(s t ); Initialize the optimizer and hyperparameters: Use the Adam optimizer and set the learning rate α. Usually α is set to a small value to ensure the stability of the training process; set the number of steps collected before each update (n_steps), the discount factor γ, the maximum number of steps per round, the batch size (batch_size), and the truncation constant ε; A5. Train the constructed agent: A5.1 The remote control platform and the automatic driving control unit send target speed and speed mode instructions, which are transmitted to the joint control mode switching unit through the data communication server and the operation mode switching unit in turn; the joint control mode switching unit switches to the speed mode and transmits the target speed to the speed PID controller; at the same time, the speed PID controller obtains the actual speed of the joint drive hydraulic rod from the data communication server, and calculates the current speed error e in combination with the given target control speed. v (t) Get the state s of the speed control agent t , input the state into the action network of the speed control agent, and sample a specific action a according to the action distribution output by the action network of the speed control agent t ; Sampling action a t That is, the speed PID controller parameter increment, and the new PID parameters are accumulated with the speed PID controller parameters at the previous moment; the speed PID controller calculates the control quantity u according to formula (4): v (t),u v (t) In this system, it is the opening percentage of the electro-hydraulic proportional valve; A5.2 Speed ​​control agent interactive training; A5.2.1 If you choose to use the joint-driven hydraulic rod to train the speed control agent, the speed PID controller will u v (t) is input to the electro-hydraulic proportional valve through the TCP-CAN protocol conversion unit, and then controls the joint drive hydraulic rod; the joint drive hydraulic rod outputs new speed data and transmits it back to the speed PID controller through the stroke sensor, TCP-CAN protocol conversion unit and data communication server; the speed PID controller calculates the new state s based on the current given target control speed t+1 , and calculate the reward r t , calculate the advantage function, and convert the experience (s t ,a t ,r t ,s t+1 ) is stored in the experience pool of the speed control agent, completing an interaction between the speed control agent and the environment; A5.2.2 If you choose to use the virtual commissioning system to train the speed control agent, the speed PID controller will u v (t) Input the virtual debugging system; the virtual debugging system outputs new speed data, and the data communication server transmits it back to the speed PID controller; the speed PID controller calculates the new state s based on the current given target control speed t+1 , and calculate the reward r t , calculate the advantage function, and convert the experience (s t ,a t ,r t ,s t+1 ) is stored in the experience pool of the speed control agent, completing an interaction between the speed control agent and the environment; A5.3 Repeat n_steps times the speed control agent interaction training selected in A5.2, using the generalized advantage estimate to calculate the advantage function δ t+l =r t+l +γV(s t+l+1 )-V(s t+l ) (6) Where t is the current time step. l is the offset from the current time step t. γ is the discount factor, which indicates the importance of future rewards. λ is the smoothing parameter of GAE, which controls the time range of advantage estimation. δ t+l is the TD error at time step t+l, representing the difference between the immediate reward and the value estimate; r t+l is the reward at time step t+l, calculated by equation (2); At each time step, compute the generalized advantage estimate And stored together with the experience in the experience pool of the speed control agent to guide the update of the action network; A5.4 Speed ​​Control The agent updates the action network according to the objective function of the action network: The update goal of the action network is to maximize the expected cumulative reward while ensuring that the new strategy does not deviate too far from the old strategy. The objective function of the action network is: in, is the strategy ratio, ε is the cutoff ratio, is the mathematical expectation; According to the above objective function, the action network is updated multiple times in small batches by performing gradient ascent to maximize the objective function; as the action network parameters of the speed control agent are updated, the PID parameters generated by the action network can make the control performance of the speed PID controller tend to the optimal response speed and control error; A5.5 Speed ​​Control The agent updates the value network according to the loss function of the value network: the value network performs small batch sampling based on the data in the experience pool; calculates the loss function L of the value network VF (θ ν ), gradient descent adjusts the parameters of the value network to reduce the value network's estimation error of the reward value that the action network can obtain by taking the current strategy; the loss function L of the value network VF (θ v ) is the mean squared error, which is used to minimize the difference between the predicted value and the target value: Among them, V(s t ) is the value estimate of the output of the value network, is the generalized advantage estimate, θ ν is the parameter of the value network; V θν (s t ) is the value network in state s t The predicted value of the next output; V represents the expected value of the variance between all predicted values ​​and target values ​​in time step t; t target is the target value, calculated using the generalized advantage estimate: V(s t ) is the value estimate of the output of the value network, is the generalized advantage estimate; Finally, the value network estimates the expected cumulative reward that can be obtained by following the current action network's strategy starting from a given state s, and calculates the advantage function, which is used to calculate the objective function of the action network to update the action network; A5.6 Repeat the training steps of A5.4-A5.5 for the speed control agent, and continuously update the action network and the value network until the average reward becomes stable or the predetermined maximum number of iterations is reached; after the update is completed, stop the training and save the parameters of the action network and the value network. At this time, the PID control parameters output by the action network can make the speed tracking performance of the speed PID controller achieve the best response speed and the smallest possible control error; A5.7 If A5.2.2 is executed in A5.2 to train the speed control agent using the virtual debugging system, then A5.2.1 is additionally executed to train the speed control agent using the joint-driven hydraulic rod, and the training process of A5.2 to A5.6 is repeated to further adaptively train the speed control agent to optimize the agent performance loss caused by the mathematical model error between the virtual debugging system and the real system; A6. Put the speed mode into practical use. The remote control platform and the automatic driving control unit send the target speed and speed mode instructions, which are transmitted to the joint control mode switching unit through the data communication server and the operation mode switching unit in sequence; the joint control mode switching unit switches to the speed mode and transmits the target speed to the speed PID controller; The action network of the speed control agent generates a set of PID parameter increments according to the current speed error, and adds the parameter increments to the original PID controller. The speed PID controller then uses the new PID parameters to calculate the target opening of the electro-hydraulic proportional valve to control the joint drive hydraulic rod according to the control error between the current target speed and the actual speed; the action network outputs different PID parameter increments according to different input states, so that the speed controller adopts the best PID parameters to control the joint drive hydraulic rod according to the current input state, so as to achieve the best speed tracking control effect; B. Position mode; B1. Define the state space s as the current velocity error e of the joint drive hydraulic rod p (t), e p (t) = L set -L actual (t)(9) where L set is the target position, L actual is the actual location; Execute A2 to A4 for the position PID controller, build the position control agent, and initialize the parameter settings; B2. Train the constructed position control agent: B2.1 The position controller interacts with the controlled object to gain experience: The remote control platform and the automatic driving control unit send target position and position mode instructions, which are transmitted to the joint control mode switching unit through the data communication server and the operation mode switching unit in turn; the joint control mode switching unit switches to the position mode and transmits the target position to the position PID controller; at the same time, the position PID controller obtains the actual position of the joint drive hydraulic rod from the data communication server, and calculates the current position error e in combination with the given target position. p (t) is the state s of the position control agent t , input the state into the action network of the position control agent, and sample a specific action a according to the action distribution output by the action network of the position control agent t ; Sampling action a t That is, the parameter increment of the position PID controller, and the new position PID controller parameters are accumulated with the position PID controller parameters at the previous moment. The control quantity u is calculated according to formula (4): p (t), u p (t) is used as the reference input of the speed controller and input into the speed PID controller, which calculates the opening of the electro-hydraulic proportional valve; B2.2.1 If you choose to use the joint-driven hydraulic rod to train the speed-position control agent, the speed PID controller will u v (t) is input to the electro-hydraulic proportional valve through the TCP-CAN protocol conversion unit, and then controls the joint drive hydraulic rod; the joint drive hydraulic rod outputs new position data and transmits it back to the position PID controller through the stroke sensor, TCP-CAN protocol conversion unit and data communication server; the speed PID controller calculates the new state s based on the current given target position t+1 , and calculate the reward r t , calculate the advantage function, and convert the experience (s t ,a t ,r t ,s t+1 ) is stored in the experience pool of the position control agent, completing an interaction between the position control agent and the environment; B2.2.2 If you choose to use the virtual debugging system to train the position control agent, the speed PID controller will u v (t) Input the virtual debugging system; the virtual debugging system outputs the new position data and transmits it back to the position PID controller through the data communication server; the position PID controller calculates the new state s based on the current given target position t+1 , and calculate the reward r t , calculate the advantage function, and convert the experience (s t ,a t ,r t ,s t+1 ) is stored in the experience pool of the position control agent, completing an interaction between the position control agent and the environment; B2.3 After the position control agent interacts with the environment n_steps times, the generalized advantage estimate is calculated according to formula (5) For each time step t, compute the generalized advantage estimate And it is stored together with the experience in the experience pool of the position controller to guide the update of the action network of the position control agent; B2.4 Calculate the action network objective function of the position control agent according to formula (7), and update the action network parameters of the position control agent according to the objective function by performing the gradient ascent method so that the objective function tends to be maximized; B2.5 Calculate the loss function of the position control agent value network according to formula (8). The position control agent value network loss function performs the gradient descent method to adjust the parameters of the position control agent value network to reduce the estimation error of the value network for the reward value that can be obtained by the action network taking the current strategy under a given state, thereby more accurately evaluating the quality of the current position control agent action network and further more accurately guiding the update of the position control agent action network; B2.6 performs the training steps of B2.1 to B2.5 on the position controller, and the position control agent continuously adjusts the action and value networks until the average reward tends to be stable or reaches the predetermined maximum number of iterations; after the update is completed, the PID control parameters output by the action network can achieve the best response speed and the smallest possible control error for the position tracking performance of the position PID controller; when the performance of the position control agent reaches the expected goal, the training is stopped, and the parameters of the action network and the value network are saved; B2.7 If B2.2.2 is executed in B2.2 to train the position control agent using the virtual debugging system, then B2.2.1 is additionally executed to train the speed control agent using the joint-driven hydraulic rod, and the training process of B2.2 to B2.5 is repeated to further adaptively train the position control agent to optimize the agent performance loss caused by the mathematical model error between the virtual debugging system and the real system; B3. Put the position mode into practical application. The remote control platform and the automatic driving control unit send the target position and position mode instructions, which are transmitted to the joint control mode switching unit through the data communication server and the operation mode switching unit in turn; the joint control mode switching unit switches to the position mode and transmits the target position to the position PID controller; The action network of the position control agent generates a set of PID parameter increments according to the error between the actual position and the target position, and accumulates the parameter increments to the position PID controller. The position PID controller then uses the new PID parameters to calculate the target speed of the joint-driven hydraulic rod according to the control error between the target position and the actual position. The target speed is tracked and controlled by the speed PID controller. The action network of the position PID controller outputs different position PID parameter increments according to different input states, so that the position PID controller adopts the optimal PID parameters to control the joint-driven hydraulic rod according to the current input state.

Citation Information

Patent Citations

  • Robot pose control method for nut feeding suite

    CN118605135A

  • Quadruped robot motion control method, system and equipment and storage medium

    CN119200658A

  • Device parameter setting support system

    JP4681082B1