Reinforcement learning control method and system based on physical space feedback

Through domain randomized training in the simulation environment and reinforcement learning methods combining multimodal sensor data and security constraints in real environments, the problems of high policy training cost and high security risks in robot control are solved, and efficient and safe physical system control is achieved.

CN120195985APending Publication Date: 2025-06-24CHANGZHOU UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510341435.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In existing robot control, strategy training costs are high and security risks are high, especially in complex environments, which are difficult to achieve efficient and safe control.

Method used

A reinforcement learning control method based on physical spatial feedback is adopted, and policy optimization and security correction in the real environment is achieved through field random training in the simulation environment, combining multimodal sensor data and security constraints.

Benefits of technology

It reduces the training cost and security risks in real environments, improves the adaptability and accuracy of robots in complex environments, and achieves high sample efficiency and high safety physical system control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120195985A_ABST
    Figure CN120195985A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and robot control, in particular to a reinforcement learning control method and system based on physical space feedback, and the method comprises the steps: training an initial strategy through domain randomization in a simulation environment; fusing multi-modal sensor data in a real environment, and constructing an environment state; deploying the strategy network after the strategy network parameters are optimized into a real environment, combining a model base and model-free reinforcement learning, utilizing simulation data and real data to jointly optimize the strategy network, and compensating a dynamical model error through online fine adjustment; and correcting the action instruction in real time based on the security constraint. The system comprises a sensor module, a strategy network module, a security constraint module and a simulation-real migration module. According to the method, the dynamical model error is compensated through online fine adjustment; through a mixed model base and a model-free reinforcement learning architecture, and in combination with multi-modal sensor data and a security constraint mechanism, high-sample-efficiency and high-security physical system control is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence and robot control, and particularly relates to a reinforcement learning control method and system based on physical space feedback. Background Art

[0002] Robot technology has achieved rapid development in the past few decades and is widely used in many fields such as industrial manufacturing, medical care, logistics distribution, home service, and space exploration. With the continuous progress of technology, the requirements for the adaptability, flexibility, and accuracy of robots in complex environments are increasing day by day, which promotes the continuous innovation and breakthrough of robot strategy training and motion control technologies.

[0003] The application of traditional reinforcement learning (RL) in physical environments faces challenges such as high sample costs (requiring a large number of real interactions), insufficient safety (policies may generate dangerous actions), and sensitivity to noise and delay (unstable sensor data).

[0004] When the strategies of existing models based on simulation training are transferred to the real environment, the performance degrades due to dynamic differences (such as mismatches in friction coefficients and inertial parameters). Summary of the Invention

[0005] The technical problem to be solved by the present invention is the problems of high strategy training cost and high safety risk in existing robot control.

[0006] To this end, the present invention provides a reinforcement learning control method and system based on physical space feedback, which solves the problems of high strategy training cost and high safety risk in the prior art through a physical feedback closed-loop and a simulation-real transfer mechanism.

[0007] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0008] A reinforcement learning control method based on physical space feedback includes the following steps:

[0009] Step 1, training an initial strategy through domain randomization in a simulation environment;

[0010] Step 2, fusing multi-modal sensor data in the real environment to construct an environmental state;

[0011] Step 3, deploying the policy network after optimizing the policy network parameters to the real environment, combining model-based and model-free reinforcement learning, jointly optimizing the policy network using simulation data and real data, and compensating for dynamic model errors through online fine-tuning;

[0012] Step 4, real-time correcting action instructions based on safety constraints.

[0013] Further, in Step 2, Kalman filtering is used for state estimation: where the vector z t is the observed value, the matrix K t is the Kalman gain, and the vector is the prior state estimate, that is, the predicted value based on the previous state and the dynamics model, is the current state estimate value, the optimal state estimate obtained through sensor data fusion.

[0014] Further, in Step 3, the parameters are optimized by combining model-based and model-free learning: where α is the learning rate, λ is the weight of the simulation data, and k is the number of iterations; R(τ) is the trajectory cumulative reward, and p real is the probability distribution, representing the real environment trajectory distribution and reflecting the dynamics characteristics of the real physical system; p model is the probability distribution, representing the simulation environment trajectory distribution, which is the virtual data distribution generated based on the dynamics model; is the model-free part, is the model-based part.

[0015] Further, the range of the learning rate α is 1×10 -4 ~1×10 -3 . When the real environment is a high-dynamic environment (such as UAV obstacle avoidance), set α to 1×10 -4 . When the real environment is a static environment (such as industrial assembly), set α to 1×10 -3 .

[0016] Further, the range of the simulation data weight λ is 0.3~0.7. When high simulation accuracy is required, set λ to 0.7. When the difference between the simulation environment and the real environment is large, set λ to 0.3.

[0017] Further, the learning rate is dynamically adjusted by an adaptive adjustment method. The adaptive adjustment method is linear attenuation or gradient-based adjustment. When linear attenuation is used for adjustment, set α = 3e -4 at the initial stage of training, and it is reduced to le -4 at the later stage; when the gradient-based adjustment method is used for adjustment, then when the reward fluctuation in 10 consecutive iterations > 20%, set α to be less than half of α in the previous iteration.

[0018] Further, in Step 4, the safety constraints include barrier functions or convex optimization constraints. The barrier function is a continuously differentiable function, and its value is positive in the safe area and negative in the dangerous area.

[0019] Further, in step four, for the original action a t perform safety monitoring, and optimize and correct the original action through the safety constraint by means of the barrier function B(s) to output the safe action a t ': '

[0020] where B(s t+1 ) ≥ 0 is the safety constraint condition of the barrier function, and ∥a t - a t ∥ 2 represents the minimum adjustment amount of the action.

[0021] Further, in step four, the barrier function is combined with the design of the reward function of reinforcement learning, and the reward function is used to guide the policy to learn safe actions.

[0022] A reinforcement learning control system based on physical space feedback includes: a sensor module, a policy network module, a safety constraint module, and a simulation-real transfer module. The sensor module is connected to the physical environment and the data processing module. The data processing module is connected to the policy network module. The simulation-real transfer module acts on the simulation environment. The policy network module is connected to the simulation environment and the safety constraint module. The safety constraint module outputs action instructions to the physical environment.

[0023] The beneficial effect of the present invention is that the present invention discloses a reinforcement learning control method and system based on physical space feedback. In this application, parameter randomization is introduced in simulation training to improve policy generalization, and online fine-tuning is used to compensate for dynamic model errors; through a hybrid model-based and model-free reinforcement learning architecture, combined with multi-modal sensor data and a safety constraint mechanism, high-sample-efficiency and high-safety physical system control are achieved. The present invention is applicable to fields such as robots and industrial automation, and can effectively reduce the training cost and safety risk in the real environment.

[0024] Specifically, this application obtains real-time parameters from the physical environment through the sensor module, and based on the real-time parameters, optimizes and trains the policy network in the simulation environment and the real environment, uses safety constraints to correct the actions output by the policy in real time, and outputs the corrected safe actions to the physical environment, thereby forming a closed loop to solve the problem of dangerous actions in physical system interaction. Brief Description of the Drawings

[0025] The present invention will be further described below with reference to the drawings and embodiments.

[0026] Figure 1 is the structural framework diagram of the reinforcement learning control system based on physical space feedback in the present invention.

[0027] Figure 2It is a flowchart of the reinforcement learning control method based on physical space feedback in the present invention.

[0028] Figure 3 It is a logical schematic diagram of the safety constraint module in the present invention.

[0029] Figure 4 It is a flowchart of continuously monitoring safety indicators and triggering a recovery mechanism in the present invention. Detailed implementation manners

[0030] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.

[0031] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", "axial", "radial", "circumferential", etc. are based on the orientation or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, so it cannot be understood as a limitation of the present invention. In addition, the features defined as "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.

[0032] In the description of the present invention, it should be noted that, unless otherwise clearly defined and limited, the terms "installed", "connected", "connected" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected, or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0033] A reinforcement learning control system based on physical space feedback includes a sensor module, a data processing module, a policy network module, a safety constraint module, and a simulation-real migration module.

[0034] Among them, the sensor module collects multi-modal data such as force, position, and vision and transmits the multi-modal data to the data processing module; the data processing module filters, fuses, and estimates the state of the raw data and sends the results to the policy network module; the simulation-real transfer module realizes policy transfer through domain randomization and online adaptation. The simulation-real transfer module sends the randomization parameters to the simulation environment and sends the initially trained policy in the simulation environment to the policy network module; the policy network module generates control instructions based on a deep neural network (such as the PPO or SAC algorithm) and sends them to the safety constraint module; the safety constraint module ensures action safety based on mathematical constraints (such as Convex Optimization), feeds back the safety state to the policy network module, and sends the corrected safe action to the device physical switching to adjust the action.

[0035] A reinforcement learning control method based on physical space feedback includes the following steps:

[0036] Step 1, initialize the reinforcement learning agent, select the simulation or real environment mode, and load the pre-trained dynamics model (if any).

[0037] Step 2, in the simulation environment, generate diverse physical parameters (such as mass, friction coefficient) through domain randomization and train the initial policy.

[0038] Build a dynamics model in the simulation environment, generate diverse parameters through domain randomization, and optimize the policy network parameter θ:

[0039]

[0040] Among them, p sim is the dynamics of the simulation environment; γ is the discount factor. Since the actions in this application require multi-objective navigation, γ = 0.99 is set. If the action is a single grasp, then γ = 0.95 is set; r(s t ,a t ) is the reward function. The reward function includes task objectives and safety constraints. Among them, s t is the task objective, and a t is the action.

[0041] Step 3, in the real environment, collect physical feedback data in real time through sensors (force sense, vision, IMU), perform Kalman filtering and state estimation, construct the environmental state, and solve the problem of dangerous actions in physical system interaction through sensor fusion and real-time safety correction.

[0042] The physical feedback data includes, but is not limited to, data such as the collected force F, position q, and vision I. After the sensor module obtains these data, it transmits them to the data processing module, and the data processing module uses Kalman filtering for state estimation: Among them, the vector z t is the observed value, the matrix K t is the Kalman gain, and the vector is the prior state estimate, that is, the predicted value based on the previous state and the dynamics model, is the current state estimate value, the optimal state estimate obtained through sensor data fusion.

[0043] Step 4: Deploy the policy network after optimizing the policy network parameters into the real environment. Combine model-based and model-free reinforcement learning, use simulation data and real data to jointly optimize the policy network, introduce an adaptive transfer mechanism, introduce parameter randomization in simulation training, and compensate for the dynamics model error through online fine-tuning to narrow the difference between the simulation and the real environment.

[0044] Update the policy through real environment interaction data, and optimize the parameters by combining model-based and model-free learning:

[0045]

[0046] Among them, α is the learning rate, and its range is 1×10 -4 ~1×10 -3 . When the real environment is a high-dynamic environment (such as UAV obstacle avoidance), set α to 1×10 -4 . When the real environment is a static environment (such as industrial assembly), set α to 1×10 -3 ; λ is the simulation data weight, and its range is 0.3~0.7. When high simulation accuracy is required, set λ to 0.7. When the difference between the simulation environment and the real environment is large, set λ to 0.3; k is the number of iterations; R(τ) is the trajectory cumulative reward, the total reward of a single trajectory, which measures the performance of the policy in the environment; p real is the probability distribution, representing the real environment trajectory distribution, reflecting the dynamics characteristics of the real physical system; p model is the probability distribution, representing the simulation environment trajectory distribution, which is the virtual data distribution generated based on the dynamics model; is the model-free part, is the model-based part.

[0047] In the simulation stage, model-based pre-training is adopted. Build a dynamics model in the simulation environment, generate diverse parameters through domain randomization, and optimize the objective to maximize the simulation cumulative reward Model-based pre-training in the simulation stage has high sample efficiency and avoids real hardware losses.

[0048] In the real stage, model-free fine-tuning is adopted to adapt to real environment noise and uncertainty. Collect data in the real environment, update the network parameters through policy gradients (such as PPO), and optimize the objective to maximize the real cumulative reward

[0049] Model-based data provides a large number of low-cost samples during the simulation phase, covering a wide exploration of the state space; model-free data corrects the simulation bias during the real phase to ensure the effectiveness of the policy in the real environment, thus overall synergistically optimizing the parameters and improving the effect of parameter optimization.

[0050] It should be noted that according to the learning rate setting principle, the setting method and adaptive adjustment method of the learning rate and data weight are as follows:

[0051] Table 1 Learning Rate Setting Principle

[0052]

[0053] Table 2 Data Weight Setting Principle

[0054]

[0055] Among them, corresponding to the hybrid training phase, the learning rate can adopt an adaptive adjustment method, and the adaptive adjustment method can adopt linear decay or gradient-based adjustment. Linear decay: Set α = 3e -4 , and decrease it to le -4 in the later stage; Gradient-based adjustment: If the reward fluctuation > 20% for 10 consecutive iterations, then α < 0.5α.

[0056] Through refined parameter configuration, the safety, efficiency and generalization ability can be balanced to achieve the reliable deployment of the physical reinforcement learning system.

[0057] Step 5, based on safety constraints (such as barrier functions or convex optimization constraints, action amplitude limits), perform real-time correction on the actions output by the policy to avoid dangerous states.

[0058] Perform safety monitoring on the original action a t , and optimize and correct the original action through safety constraints to output the safe action a t ':

[0059]

[0060] In the real environment, the action a output by the policy t is constrained by the barrier function B(s):

[0061]

[0062] Among them, the action a t is represented as a vector, such as [-10, 10] N·m; B(s) is defined as a safety region indicator function (such as the limit of the robot joint angle), and the design of the barrier function is a continuously differentiable function, and its value is positive in the safety region and negative in the dangerous region.

[0063] In the reinforcement learning control method based on physical space feedback, the Barrier Function is the core mechanism to ensure system safety. Its role is to force the actions of the agent (such as a robot) to always be within a preset safe range through mathematical constraints, thereby avoiding dangerous states (such as joint limit violations, collisions, or overloads). Compared with the existing hard-coded rule-based obstacle avoidance method, this method can optimize and adjust actions in real time, maintain task continuity, and dynamically adapt to complex constraints (such as multi-joint collaborative limit).

[0064] Furthermore, the barrier function can be combined with the design of the reinforcement learning reward function to guide policy learning of safe actions and optimize performance in the long term. For example:

[0065] r safe =r task +w·log(B(s))

[0066] where r safe is the task-based reward (such as successful execution of grasping); w is the safety weight coefficient used to control the balance between safety and task performance. By combining safety constraints with the reward function, in the initial stage of training, the agent is punished for exploring dangerous areas and quickly converges to a safe policy; in the later stage of training, the task reward is maximized within the safety constraints to achieve efficient and safe action output.

[0067] Step six, deploy the policy to the physical system, continuously monitor the safety conditions and trigger the recovery mechanism (such as emergency stop or policy rollback).

[0068] Continuously monitor the safety conditions: where ∧: logical "and" operation, indicating that all safety constraints need to be satisfied simultaneously. Output: If true, the system is safe; if false, trigger the recovery mechanism

[0069] Conditions for triggering the recovery mechanism: 1. B i (s t )<0, the i-th safety barrier function (such as grasping force, joint angle constraint) is less than 0, that is, the action is in a dangerous state; 2. The current i-th state variable (such as joint angle, torque) s t,i , is lower than the i-th safety threshold s i,limit , ∥s t,i -s i,limit ∥>c i , c i is the safety tolerance (such as c i =0.1rad). If any of the above conditions are met, the recovery mechanism is implemented.

[0070] The recovery mechanism includes emergency stop and policy rollback. Emergency stop means that when the safety constraint is violated, zero action is forced to be output (such as shutting down the motor), a t = 0; Policy rollback means that if the safety alarm is continuously triggered more than N times (such as N = 5), the safety policy parameter θ of the previous version safe will replace the current policy network parameter θ.

[0071] Example 1

[0072] In this example, the action is for the robotic arm to grasp, and the target task is to grasp a random mass object of 0.2 - 0.5 kg. The reward function is designed as follows:

[0073] r = w1 · grasping success - w2 · ∥q - q desired ∥ 2 - w3 · ∥τ∥ 2

[0074] where w1 is the task completion weight, which encourages successful grasping of the target object, w1 = 10; w2 is the trajectory tracking weight: it penalizes the deviation of the joint angle from the expected value to ensure motion accuracy, w2 = 0.1; w3 is the energy consumption weight: it penalizes high torque output to reduce energy consumption and mechanical losses, w3 = 0.01; q is the actual joint angle, which is the real-time angle value of each joint of the robotic arm obtained through sensors; q desired is the expected joint angle, which is the theoretical joint angle calculated according to the task target through inverse kinematics; τ is the joint output torque, which is the actual torque value driven by the motor obtained through sensors.

[0075] S1. Simulation pre-training

[0076] Build a dynamic model:

[0077]

[0078] where M is the inertia matrix and τ is the joint torque.

[0079] The policy network is updated using the PPO algorithm:

[0080]

[0081] S2. Safety constraint:

[0082] Set the barrier function B(s) to constrain the grasping force:

[0083]

[0084] When the grasping force F grip < F max , B(F) > 0, which is a safe state and no correction is needed; when F grip ≥ Fmax When B(F) ≤ 0, it is a dangerous state and action correction is triggered. Through the correction of the barrier function, the safest action a t closest to the original action a t ’ can be found to ensure the safety of the next state S t+1 and avoid hardware damage or personal injury at the minimum adjustment cost (such as torque change).

[0085] S3 Simulation Pretraining: Train the robotic arm grasping strategy in MuJoCo, randomizing the load mass and the desktop friction coefficient.

[0086] S4 Real Environment Deployment

[0087] The grasping force is detected in real time through a force sensor. If it exceeds the threshold, the strategy correction is triggered: In the embodiment, the grasping force is monitored in real time. If it exceeds 50N, the torque is set to zero.

[0088] Use an online fine-tuning algorithm (such as Meta-RL) to adjust the weights of the policy network to compensate for the joint damping differences of the real robotic arm: In the embodiment, the Meta-RL algorithm is used to adapt to the joint friction differences of the real robotic arm (measured friction coefficient 0.4 vs simulation 0.3).

[0089] Embodiment 2

[0090] The action in this embodiment is the walking of a quadruped robot. The safety constraint design of this action is to set the safe ranges of joint angles and contact forces, and the action is optimized in real time through quadratic programming (QP).

[0091] In this embodiment, the extended Kalman filter (EKF) fuses the IMU and joint encoder data, and the position estimation error < 2 cm.

[0092] In this embodiment, when the center of mass laterally deviates: When the offset > 5 cm, trigger quadratic programming (QP) to adjust the gait, and the solution time < 1 ms.

[0093] When the robot imbalance is detected, switch to a predefined safety strategy (such as a crawling posture).

[0094] Inspired by the ideal embodiments of the present invention described above, through the above description, relevant workers can make various changes and modifications completely within the scope of the technical idea of this invention. The technical scope of this invention is not limited to the content in the specification, and its technical scope must be determined according to the scope of the claims.

Claims

1. A reinforcement learning control method based on physical space feedback, characterized in that: The following steps are involved: Step 1: Train the initial strategy through domain randomization in a simulation environment; Step 2: Fuse multimodal sensor data in the real environment to construct the environment state; Step 3: Deploy the policy network after optimizing the policy network parameters into the real environment, combine the model-based and model-free reinforcement learning, use the simulation data and real data to jointly optimize the policy network, and compensate for the dynamic model error through online fine-tuning; Step 4: Correct the action instructions in real time based on safety constraints.

2. The reinforcement learning control method based on physical space feedback according to claim 1, characterized in that: In step 2, the Kalman filter is used for state estimation: Among them, the vector z t is the observed value, the matrix K t is the Kalman gain, the vector is the prior state estimate, that is, the predicted value based on the state at the previous moment and the dynamic model, is the current state estimate, the optimal state estimate obtained by sensor data fusion.

3. The reinforcement learning control method based on physical space feedback according to claim 1, characterized in that: In step 3, the model-based and model-free learning optimization parameters are combined: Among them, α is the learning rate, λ is the simulation data weight, k is the number of iterations; R(τ) is the cumulative reward of the trajectory, p real is a probability distribution, which represents the distribution of real environment trajectories and reflects the dynamic characteristics of the real physical system; model is a probability distribution, which represents the trajectory distribution of the simulation environment, and is the virtual data distribution generated based on the dynamics model; For the model-free part, The base part of the model.

4. The reinforcement learning control method based on physical space feedback according to claim 3 is characterized in that: The range of the learning rate α is 1×10 -4 ~1×10 -3 When the real environment is a highly dynamic environment (such as drone obstacle avoidance), set α to 1×10 -4 When the real environment is a static environment (such as industrial assembly), set α to 1×10 -3 .

5. The reinforcement learning control method based on physical space feedback according to claim 4 is characterized in that: The range of the simulation data weight λ is 0.3 to 0.

7. When high simulation accuracy is required, λ is set to 0.

7. When the simulation environment is greatly different from the real environment, λ is set to 0.

3.

6. The reinforcement learning control method based on physical space feedback according to claim 5, characterized in that: The learning rate is dynamically adjusted using an adaptive adjustment method, which is linear attenuation or gradient-based adjustment. When linear attenuation is used for adjustment, α=3e is set at the beginning of training. -4 , and later dropped to le -4 ; When using the gradient-based adjustment method, if the reward fluctuation is > 20% for 10 consecutive iterations, set α to less than half of α in the previous iteration.

7. The reinforcement learning control method based on physical space feedback according to claim 4, characterized in that: In step 4, the safety constraint includes a barrier function or a convex optimization constraint, where the barrier function is a continuous differentiable function whose value is positive in the safe area and negative in the dangerous area.

8. The reinforcement learning control method based on physical space feedback according to claim 7, characterized in that: In step 4, for the original action a t Perform safety monitoring, optimize and modify the original action through the barrier function B(s) through safety constraints, and output the safety action a t ′: Among them, B(s t+1 )≥0 is the safety constraint condition of the barrier function, ∥a ′ t -a t ∥ 2 Indicates the minimum adjustment amount of the action.

9. The reinforcement learning control method based on physical space feedback according to claim 1, characterized in that: In step 4, the barrier function is combined with the reward function design of reinforcement learning to guide the strategy to learn safe actions through the reward function.

10. A reinforcement learning control system based on physical space feedback for implementing the reinforcement learning control method based on physical space feedback as claimed in any one of claims 1 to 9, characterized in that: include: A sensor module, a policy network module, a safety constraint module and a simulation-to-reality migration module, wherein the sensor module is connected to a physical environment and a data processing module, the data processing module is connected to a policy network module, the simulation-to-reality migration module acts on a simulation environment, the policy network module is connected to a simulation environment and a safety constraint module, and the safety constraint module outputs action instructions to the physical environment.

Citation Information

Cited By

  • Virtual-real migration method based on general multi-agent parallel reinforcement learning framework

    CN121168566A

  • Robot hybrid control method and system based on simulation-reality alignment

    CN122194686A