A robot arm valve screwing training method based on pose gating and multi-stage rewards
Patent Information
- Application Number
- CN202610944142.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-25
AI Technical Summary
由于大角度旋拧具有强非线性与连续接触的物理特性,若缺乏严密的逻辑门控约束与分阶段的引导,机械臂极易在旋拧途中破坏原有的正交包络姿态,导致夹爪滑脱、脱扣或发生严重的刚性物理干涉
[0031]经由上述的技术方案可知,与现有技术相比,本发明公开提供了一种基于姿态门控与多阶段奖励的机械臂阀门旋拧训练方法,能够通过空间门控约束与有效抓取判定,校验机械臂操作中的动态正交态势与物理合法性,并利用全局复合奖励函数与最大熵软动作评价算法进行强化学习训练,引导机械臂自发维持合规位姿,使其在长序列连续旋拧任务中克服奖励稀疏与灾难性遗忘难题,降低夹爪滑脱与物理刚性干涉的隐患。本发明具备较好的自适应泛化性,能够在无需依赖昂贵力传感器的情况下,广泛应用于阀门的大角度连续旋拧操作中,为复杂高危工业场景下的自动化管路维护与机械臂交互作业提供端到端的底层技术方案。
Smart Images

Figure CN122807873A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and more specifically to a training method for valve turning of a robotic arm based on spatial gating constraints and multi-stage rewards. Background Technology
[0002] In recent years, with the improvement of industrial automation, the maintenance needs in fields such as intelligent manufacturing, unmanned factories, and chemical pipeline maintenance have been increasing. Using robotic arms to replace manual labor in high-risk valve operations has become an industry trend. Relying on deep reinforcement learning and high-fidelity physical simulation technology, robotic arms can break away from the dependence on traditional complex inverse kinematic trajectory planning and achieve end-to-end autonomous control.
[0003] In continuous valve turning operations, a stable and compliant pre-grasping posture is a prerequisite for ensuring effective convergence of reinforcement learning training. After establishing the pre-grasping training starting point, the transition from the static pre-grasping stage to the dynamic large-angle continuous turning process still faces extremely high operational risks. Due to the strong nonlinearity and continuous contact physical characteristics of large-angle turning, without rigorous logical gating constraints and phased guidance, the robotic arm is highly susceptible to disrupting its original orthogonal envelope posture during turning, leading to gripper slippage, disengagement, or severe rigid physical interference. Existing reinforcement learning algorithms, when handling such long-sequence continuous control tasks, often employ simple distance penalties or sparse success rewards, resulting in a large agent exploration space, numerous invalid collisions, difficulty in model convergence, or high failure rates in practical deployment. Therefore, a training method is urgently needed that can verify posture legitimacy in real time during the grasping and turning stages and provide joint incentives based on the valve rotation process to ensure the kinematic stability and anti-disengagement capability of the robotic arm throughout the entire dynamic turning cycle.
[0004] Therefore, this invention proposes a method for continuous valve turning control of a robotic arm based on attitude gating and multi-stage reward. Summary of the Invention
[0005] In view of the above problems, this invention is proposed to provide a robotic arm valve turning training method based on spatial gating constraints and multi-stage rewards to overcome or at least partially solve the above problems. This method uses deep reinforcement learning as its core, taking the robotic arm's pre-grasping posture of the valve as a priori starting point, and extends the definition of the interaction space for valve turning tasks. Through serial spatial envelope verification, effective grasping judgment, and multi-stage dynamic incentives, a globally composite reward mechanism with strict constraints is constructed to dynamically evaluate and iteratively train the reinforcement learning policy network. This method can guide the robotic arm to spontaneously maintain a compliant orthogonal envelope posture and lock the gripper in a timely manner during continuous valve turning tasks, overcoming the risks of posture distortion, physical rigidity interference, and gripper slippage that are easily induced in long and complex operations, thus realizing valve turning operations.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] In a first aspect, embodiments of the present invention provide a robotic arm valve turning training method based on posture gating and multi-stage rewards, comprising: Step 1: Use the pre-grabbed posture as the initial pose constraint to perform Markov process modeling for the valve turning task; Step 2: Construct a spatial envelope-based gating condition based on S1; Step 3: Based on the gating conditions of the spatial envelope, construct a gripper closure reward based on the spatial envelope and distance threshold; Step 4: Based on the gating conditions of the spatial envelope and the gripper closure reward, construct a valve turning joint reward based on effective gripping logic gating; Step 5: Synthesize a global composite reward function based on the gripper closure reward and the valve turning reward; Step 6: Using a soft motion evaluation algorithm based on maximum entropy, the policy network is iteratively trained through the global composite reward function to output a control strategy model for continuous valve turning of the robotic arm.
[0008] Preferably, step 1 specifically includes: Step S101: Construct the basic environment for physical simulation A high-fidelity physical simulation engine is established at the bottom layer, and a simulation model of a robotic arm and valve containing fine geometric collision bodies, dynamic friction coefficients and joint physical extreme parameters is loaded. Based on the simulation model, multiple parallel and independent physical sub-environments are constructed to form a set of valve continuous turning interactive environments. Step S102: Define the state and action space of the Markov decision process. Based on the physical simulation environment, reinforcement learning core parameters are constructed for the continuous valve turning task, including state space, action space, round initial condition and round termination condition; the state space is configured as a multi-dimensional vector that integrates the joint state of the robotic arm, the gripper state, the action state of the previous time step, and the real-time rotation angle of the valve; the action space is configured as a composite action vector that includes the six-degree-of-freedom pose increment command of the robotic arm end effector and the one-dimensional opening and closing control command. Step S103: Apply round-initial constraints based on attitude alignment At the start of each training round, the pre-grabbing posture output by the posture alignment strategy model is obtained as the initial pose constraint in the round initial condition; based on the initial pose constraint, the state of the robotic arm and the valve is reset in the physical sub-environment; the state reset includes: on the valve base coordinates with introduced random position perturbation, positioning the robotic arm end effector above the valve handle gripping point, and controlling the principal axis of the local coordinate system of the end effector to remain spatially perpendicular to the valve handle, while randomly initializing the rotation angle and closed state of the robotic arm end effector, setting the joint velocity of the robotic arm to zero, and randomly setting the initial rotation angle of the valve to a value between 0 and 90 degrees; Step S104: Perform a round termination determination based on physical interaction. During the continuous physical interaction between the robotic arm and the valve according to the instructions output from the motion space, the interaction status is monitored in real time to determine whether the round termination condition is met: when the physical interaction time step of the current round reaches the preset maximum interaction time threshold, or when the valve rotation joint angle exceeds the preset upper and lower bounds of the safety angle, a termination signal is triggered and the current training round ends.
[0009] Preferably, the gating condition for the spatial envelope is:
[0010]
[0011]
[0012]
[0013] in, For the spatial envelope, For distance-related logic gating conditions, For attitude-related logic gating conditions, , Let represent the cosine similarity between the z-axis of the robotic arm end effector and the y-axis of the valve handle, and the cosine similarity between the y-axis of the robotic arm end effector and the x-axis of the valve handle, respectively. A pre-defined gating threshold for vertical attitude. For location-dependent logic gating conditions, and These represent the vector differences between the left and right fingertips of the gripper and the valve, respectively. For angle-dependent logic gating conditions, The vector is the line connecting the left and right fingertips of the gripper. Let x be the x-axis of the local coordinate system of the valve handle, perpendicular to the direction of the handle rod. This is a preset angle threshold.
[0014] Preferably, step 3 specifically includes: The first-stage reward is constructed by utilizing the gating conditions of the spatial envelope and, based on the relative positions of the left and right fingertips of the end effector and the valve handle, stimulating the robotic arm end effector to close to the estimated grasping closing distance from a distance angle. A second-stage reward is constructed, using the grasping closure distance as a closure gating condition to incentivize the end effector of the robotic arm to close further from the joint angle dimension.
[0015] Preferably, the reward for the first stage is:
[0016] in, The estimated gripping and closing distance based on the width of the valve handle. This is the first phase of rewards. The gating condition for the spatial envelope; The second-stage reward is:
[0017]
[0018] in, Indicates the position of the end effector drive joint. For closed gating conditions, For the second phase of rewards, and These represent the vector differences between the left and right fingertips of the gripper and the valve, respectively.
[0019] Preferably, step 4 specifically includes: The gating conditions of the spatial envelope are combined with the closure gating conditions to generate effective grasping logic gating conditions; Based on the aforementioned effective capture logic gating conditions, a valve turning joint reward system comprising continuous and phased rewards is constructed. :
[0020]
[0021]
[0022] in, For continuous rewards, This is a stage reward, where k is a proportionality coefficient. For the current valve angle, The valve angle at the start of the round. For each stage, the angle threshold, Reward value at each stage This represents a logical judgment function.
[0023] Preferably, the global composite reward function is:
[0024] in, To penalize the smoothness of motion for action commands in adjacent time steps, to Let R be the weight of each reward function, and let R be the global composite reward function.
[0025] Preferably, step 6 specifically includes: Build parameters are Policy network The parameters are respectively and Twin value network and The output layer of the policy network is used to construct a Gaussian distribution of the continuous action space, and the output layer of the value network is used to output the Q value. During training, the immediate reward is calculated using the global composite reward function, the twin value network parameters are optimized by minimizing the mean squared error loss function, and the expected cumulative return and action entropy are maximized by minimizing the KL divergence, thereby optimizing the policy network parameters. When the loss functions of the twin value network and policy network tend to stabilize and the cumulative reward reaches the convergence limit, the trained control policy model is output.
[0026] Preferably, the twin value network loss function is:
[0027] in, This represents the loss function of the twin value network. Let be the spatial state observation value at time t. Let t be the action performed by the agent. The instantaneous reward at time t. The spatial state observation value at time t+1 As a reward discount factor, An experience replay pool for storing historical interaction trajectory data; Indicates from the experience replay pool The expected value obtained from the sampled state transition data set. Indicates that the parameter is The value network for the current state Next action Action value estimation; The soft state value at the next time step is calculated using the following formula:
[0028] in, Indicates action Follow the current policy network Mathematical expectation under given distribution conditions; Indicates the first A value network for the next time-state-value pair Action value assessment value, This means taking the smaller of the evaluation values from the outputs of the two twin value networks, thereby achieving the overestimation of Q-values in consistent reinforcement learning. The parameter is The policy network in a given state Down Output Action The probability distribution; It is a learnable temperature parameter used to adjust the randomness of the strategy.
[0029] Preferably, the policy network loss function is:
[0030] in, This represents the loss function of the policy network, where α is a learnable temperature parameter. Representing state Sampled from the empirical playback pool ,action Sampling from policy network The expected value obtained at that time; The parameter is The policy network in a given state Down Output Action The probability distribution; This means taking the smaller of the two value networks' evaluation values for the current state-action pair to avoid overestimating the value.
[0031] As can be seen from the above technical solution, compared with the prior art, this invention discloses a robotic arm valve turning training method based on posture gating and multi-stage rewards. It can verify the dynamic orthogonal posture and physical legitimacy of the robotic arm operation through spatial gating constraints and effective grasping judgment. Furthermore, it utilizes a global composite reward function and a maximum entropy soft motion evaluation algorithm for reinforcement learning training, guiding the robotic arm to spontaneously maintain compliant postures. This overcomes the problems of reward sparsity and catastrophic forgetting in long-sequence continuous turning tasks, reducing the risks of gripper slippage and physical rigidity interference. This invention possesses good adaptive generalization capabilities and can be widely applied to large-angle continuous valve turning operations without relying on expensive force sensors, providing an end-to-end underlying technical solution for automated pipeline maintenance and robotic arm interactive operations in complex and high-risk industrial scenarios. Attached Figure Description
[0032] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0033] Figure 1 This is a flowchart of a robotic arm valve turning training method based on posture gating and multi-stage reward provided in an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the implementation process of the precondition attitude alignment strategy in an embodiment of the present invention; Figure 3 This is a schematic diagram of the network structure of the soft action evaluation algorithm based on maximum entropy provided in an embodiment of the present invention. Detailed Implementation
[0034] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0035] See Figure 1 The diagram shown is a flowchart of a robotic arm valve turning training method based on posture gating and multi-stage rewards according to the present invention.
[0036] The method in this embodiment includes the following steps: Step 1, as follows Figure 1As shown in Figure S1, the pre-grabbed posture output by the posture alignment strategy model is used as the initial pose constraint to construct a reinforcement learning framework for valve turning tasks.
[0037] The algorithm implementation process of the pose alignment strategy is as follows: Figure 2 As shown, the initial "random posture" of the robotic arm in the workspace is first obtained; then the algorithm outputs control commands to guide the robotic arm to perform the "end-effector posture constraint" process first, that is, to adjust the spatial angle of the end effector so that it establishes a relative alignment relationship with the local coordinate system of the target valve; on the basis of satisfying the posture constraint, the algorithm further drives the "robotic arm end-effector approaching the valve" process, dynamically reducing the spatial distance between the end gripper and the valve target grasping point; after the above continuous motion control, the robotic arm finally reaches and hovers in the "pre-grasping posture".
[0038] Will Figure 2 The "pre-grasping posture" of the robotic arm relative to the valve serves as the initial pose constraint for the robotic arm valve turning model training in this stage. At the beginning of each training round, the robotic arm and the target valve have achieved posture alignment, that is, the end effector of the robotic arm reaches above the grasping point of the valve handle, and the principal axis of the local coordinate system of the end effector is relatively perpendicular to the valve handle in spatial posture, but the rotation and closing of the gripper are random.
[0039] The continuous dynamic process of valve turning by the robotic arm is uniformly modeled as a discrete-time Markov decision process, and a low-level physical simulation framework is constructed: a low-level high-fidelity physical simulation engine is used to construct a set of valve continuous turning interaction environments consisting of multiple parallel and independent physical sub-environments; each parallel environment fully inherits and loads the robotic arm and valve simulation model containing fine geometric collision bodies, dynamic friction coefficients and joint physical extreme parameters.
[0040] Based on the underlying physical framework, a reinforcement learning core parameter structure is constructed for the continuous valve turning task, including the state space. Action space Initial conditions of the round Termination conditions .
[0041] state space robotic arm state space gripper state space and the action state of the previous time step Based on the valve rotation angle Expanding on this, the expression is: Action space Six-DOF pose increment commands from the end effector of the robotic arm and one-dimensional opening and closing control commands Composition, that is .
[0042] Initial conditions of the round The expression is as follows:
[0043] Among them, initial conditions It refers to the valve position. Define it. The valve position at the start of each round, under the initial pose constraints, i.e., the valve position in the pre-grab state, is defined in the base coordinate system. Add random small perturbations to Domain randomization; initial conditions The object is the joint pose of the robotic arm. At the beginning of each round, the initial pose of the robotic arm is set. Reconstructed to the initial pose constraints in step 1 , set the initial velocity of the joint Set to 0, that is, for The robotic arm automatically aligns to the corresponding pre-grabbing posture and remains stationary at the randomly selected valve position; initial conditions For valve angle Randomization of the field randomizes the initial valve angle to any value between 0 and 90 degrees.
[0044] End of round condition Based on the maximum interaction time per round, the expression for the valve turning task is extended as follows:
[0045] in, The termination condition under the constraint of maximum interaction time in a round, based on the maximum interaction time threshold. Perform an end determination; For valve angle The termination condition is when the robotic arm fails to perform a turning operation or an interference collision occurs, and the valve rotation joint angle exceeds the set upper and lower limits of the safety angle. and This caused an abnormal termination.
[0046] Step 2, as follows Figure 1 As shown in Figure S2, a gating condition based on spatial envelope is constructed.
[0047] Based on the physical simulation environment and reinforcement learning parameters in step 1, the spatial envelope gating conditions for the robot arm to rotate the valve are defined. The gate condition is formed by the intersection of four sub-conditions, and its expression is as follows:
[0048] For distance-related logic gating conditions, when the distance difference between the robotic arm and the standard pre-grasp point... Exceeding the predetermined threshold When the logic is zero, the end effector is restricted to the pre-grab position.
[0049] As an attitude-related logical gating condition, the spatial attitude of the end effector and valve handle relative to each other in the valve turning task is measured by calculating the axial cosine similarity in three-dimensional space. The expression is as follows:
[0050] in , Let represent the cosine similarity between the z-axis of the robotic arm end effector and the y-axis of the valve handle, and the cosine similarity between the y-axis of the robotic arm end effector and the x-axis of the valve handle, respectively. This is a preset gating threshold for vertical orientation.
[0051] This is a location-dependent logic gating condition. The gating is in... end gripper Under the premise of constraining the planar attitude, the center position of the end effector in the horizontal plane is determined. Align the position with the valve handle. Take measurements at the left and right fingertips of the grippers. and By subtracting from the center of the valve handle, we obtain the vector difference between the left and right fingertips of the gripper and the valve. and at the threshold Under gating, obtain The condition, expressed as follows:
[0052] The logic gate condition is angle-dependent. The vector connecting the left and right fingertips of the gripper is taken. Measured by cosine similarity in the horizontal direction For the case perpendicular to the valve handle, the expression is as follows:
[0053] in Let x be the x-axis of the local coordinate system of the valve handle, perpendicular to the direction of the handle rod. This is a preset angle threshold.
[0054] Step 3, as follows Figure 1As shown in Figure S3, a gripper closure reward based on spatial envelope and distance threshold is constructed.
[0055] The gripper closure reward includes two phases of rewards.
[0056] Phase 1 Rewards Gating conditions based on the spatial envelope in step 2 Based on this, and according to the relative positions of the left and right fingertips of the end effector and the valve handle, the robot arm end effector is excited to close from a distance angle, as expressed below:
[0057] in, The estimated gripping and closing distance for the width of the valve handle.
[0058] Because the estimated closure distance may have errors, the closure distance threshold from the previous stage is set as a gating condition. Thus, you will receive the second stage reward. The reward, derived from the joint angle, motivates the robotic arm's end effector to close further, as expressed below:
[0059]
[0060] in, This shows the position of the end effector's drive joint.
[0061] Step 4, as follows Figure 1 As shown in Figure S4, a valve turning joint reward system based on effective capture logic gating is constructed.
[0062] The spatial envelope gating conditions obtained in steps 2 and 3 are used as follows: With closed gate conditions Combining these conditions yields effective capture logic gating conditions. ,Right now .
[0063] Valve turning joint reward based on effective capture logic gating , by continuous rewards Stage rewards It consists of two parts, and its reward function The expression is as follows:
[0064] Continuous rewards Where k is a proportionality coefficient, For the current valve angle, The valve angle is the initial angle of the round.
[0065] Stage Rewards The reward is divided into three stages, and its expression is:
[0066] in For each stage, the angle threshold, Reward value at each stage This represents a logical judgment function.
[0067] Step 5, as follows Figure 1 As shown in Figure S5, the global composite reward function is synthesized.
[0068] Based on the reward functions inherited and extended in steps 1 to 4, the approximation, closure, and turning reward terms are linearly weighted with the action change rate penalty term to obtain a global composite reward function for valve turning tasks. The expression is as follows:
[0069] in To penalize the smoothness of motion for action commands in adjacent time steps, to These represent the weights of each reward function.
[0070] Step 6, as follows Figure 1 As shown in Figure S6, the soft action evaluation algorithm based on maximum entropy is used for iterative training.
[0071] Based on the physical simulation environment, reinforcement learning parameters, and reward function for valve turning obtained in steps 1 and 5, the overall training is based on a Markov decision process (at time t, the state is...). Action selection is Instant rewards are The next state is ), use as Figure 3 The maximum entropy soft action evaluation algorithm is shown. The specific training process is as follows: Construct parameters respectively Policy network The parameters are respectively and Twin value network and The hidden layers are all configured as three fully connected layers with [256, 128, 128] neurons, and use the exponential linear unit ELU(x) as the activation function. Specifically, the output layer of the policy network is designed with two parts: mean and standard deviation, to construct a Gaussian distribution in the continuous action space; the output layer of the value network is a single node with linear activation, used to output the Q-value.
[0072] In terms of algorithm configuration optimization, the Adam optimizer is uniformly used for gradient descent updates. The capacity of the experience replay pool is set to [value missing]. Batch size is 256; initial learning rates for the policy network and value network. Reward Discount Factor ; Soft update coefficients of the target network during training .
[0073] The model is iteratively trained using a multi-round alternating cycle of "environment interaction - experience playback - network update": (1) In each iteration of training, each parallel sub-environment Independently execute exploration tasks. The intelligent agent interacts with the physical simulation environment in real time to generate interaction trajectory experience in the current environment. All interaction trajectory experiences are stored in a high-capacity experience replay pool. In the process of inputting data into the network, the system maps the data of different dimensions in the observation state space S to a standard distribution space with zero mean using an empirical normalization method.
[0074] (2) During the network update phase, from the experience replay pool Gradient calculation is performed using randomly sampled small batches of data. A learnable temperature parameter α is introduced, and an adaptive temperature loss function is designed for α to achieve automatic adjustment of the dynamic balance between exploration and utilization.
[0075] a) For the value network, the mean squared error (MSE) loss function is used to minimize the soft Bellman residual, and the parameters of the twin value network are calculated. Gradient optimization. Its loss function. The expression is as follows:
[0076] Among them, soft state value
[0077] b) For the policy network, the expected cumulative return and action entropy are maximized by minimizing the KL divergence, and the policy network parameters are then calculated. Gradient optimization. Its loss function. The expression is as follows:
[0078] Based on a multi-round alternating cycle of "environmental interaction - experience playback - network update", when the loss functions of the twin value network and policy network tend to stabilize, and the cumulative reward... Joint reward with valve turning When the convergence limit is reached, the strategy network that has completed global parameter optimization is frozen and extracted, and the control strategy model for continuous valve turning of the robotic arm is output. The strategy model is used to control the robotic arm to complete the valve turning task.
[0079] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0080] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A training method for valve turning in a robotic arm based on posture gating and multi-stage rewards, characterized in that, include: Step 1: Use the pre-grabbed posture as the initial pose constraint to perform Markov process modeling for the valve turning task; Step 2: Construct a spatial envelope-based gating condition based on S1; Step 3: Based on the gating conditions of the spatial envelope, construct a gripper closure reward based on the spatial envelope and distance threshold; Step 4: Based on the gating conditions of the spatial envelope and the gripper closure reward, construct a valve turning joint reward based on effective gripping logic gating; Step 5: Synthesize a global composite reward function based on the gripper closure reward and the valve turning reward; Step 6: Using a soft motion evaluation algorithm based on maximum entropy, the policy network is iteratively trained through the global composite reward function to output a control strategy model for continuous valve turning of the robotic arm.
2. The method as described in claim 1, characterized in that, Step 1 specifically includes: Step S101: Construct the basic environment for physical simulation A high-fidelity physical simulation engine is established at the bottom layer, and a simulation model of a robotic arm and valve containing fine geometric collision bodies, dynamic friction coefficients and joint physical extreme parameters is loaded. Based on the simulation model, multiple parallel and independent physical sub-environments are constructed to form a set of valve continuous turning interactive environments. Step S102: Define the state and action space of the Markov decision process. Based on the physical simulation environment, reinforcement learning core parameters are constructed for the continuous valve turning task, including state space, action space, round initial condition and round termination condition; the state space is configured as a multi-dimensional vector that integrates the joint state of the robotic arm, the gripper state, the action state of the previous time step, and the real-time rotation angle of the valve; the action space is configured as a composite action vector that includes the six-degree-of-freedom pose increment command of the robotic arm end effector and the one-dimensional opening and closing control command. Step S103: Apply round-initial constraints based on attitude alignment At the start of each training round, the pre-grabbing posture output by the posture alignment strategy model is obtained as the initial pose constraint in the round initial condition; based on the initial pose constraint, the state of the robotic arm and the valve is reset in the physical sub-environment; the state reset includes: on the valve base coordinates with introduced random position perturbation, positioning the robotic arm end effector above the valve handle gripping point, and controlling the principal axis of the local coordinate system of the end effector to remain spatially perpendicular to the valve handle, while randomly initializing the rotation angle and closed state of the robotic arm end effector, setting the joint velocity of the robotic arm to zero, and randomly setting the initial rotation angle of the valve to a value between 0 and 90 degrees; Step S104: Perform a round termination determination based on physical interaction. During the continuous physical interaction between the robotic arm and the valve according to the instructions output from the motion space, the interaction status is monitored in real time to determine whether the round termination condition is met: when the physical interaction time step of the current round reaches the preset maximum interaction time threshold, or when the valve rotation joint angle exceeds the preset upper and lower bounds of the safety angle, a termination signal is triggered and the current training round ends.
3. The method as described in claim 1, characterized in that, The gating condition for the spatial envelope is: in, For the spatial envelope, For distance-related logic gating conditions; For attitude-related logic gating conditions, , Let represent the cosine similarity between the z-axis of the robotic arm end effector and the y-axis of the valve handle, and the cosine similarity between the y-axis of the robotic arm end effector and the x-axis of the valve handle, respectively. A pre-defined gating threshold for vertical attitude; For location-dependent logic gating conditions, and These represent the vector differences between the left and right fingertips of the gripper and the valve, respectively. To constrain the threshold of the vector sum, the positions of the left and right fingertips are located near the valve; For angle-dependent logic gating conditions, The vector is the line connecting the left and right fingertips of the gripper. Let x be the x-axis of the local coordinate system of the valve handle, perpendicular to the direction of the handle rod. This is a preset angle threshold.
4. The method as described in claim 3, characterized in that, Step 3 specifically includes: The first-stage reward is constructed by utilizing the gating conditions of the spatial envelope and, based on the relative positions of the left and right fingertips of the end effector and the valve handle, stimulating the robotic arm end effector to close to the estimated grasping closing distance from a distance angle. A second-stage reward is constructed, using the grasping closure distance as a closure gating condition to incentivize the end effector of the robotic arm to close further from the joint angle dimension.
5. The method as described in claim 4, characterized in that, The first stage reward is: in, The estimated gripping and closing distance based on the width of the valve handle. This is the first phase of rewards. The gating condition for the spatial envelope; The second-stage reward is: in, Indicates the position of the end effector drive joint. For closed gating conditions, For the second phase of rewards, and These represent the vector differences between the left and right fingertips of the gripper and the valve, respectively.
6. The method as described in claim 5, characterized in that, Step 4 specifically includes: The gating conditions of the spatial envelope are combined with the closure gating conditions to generate effective grasping logic gating conditions; Based on the aforementioned effective capture logic gating conditions, a valve turning joint reward system comprising continuous and phased rewards is constructed. : in, For continuous rewards, This is a stage reward, where k is a proportionality coefficient. For the current valve angle, The valve angle at the start of the round. For each stage, the angle threshold, Reward value at each stage This represents a logical judgment function.
7. The method as described in claim 1, characterized in that, The global composite reward function is: in, To penalize the smoothness of motion for action commands in adjacent time steps, to Let R be the weight of each reward function, and let R be the global composite reward function.
8. The method as described in claim 1, characterized in that, Step 6 specifically includes: Build parameters are Policy network The parameters are respectively and Twin value network and The output layer of the policy network is used to construct a Gaussian distribution of the continuous action space, and the output layer of the value network is used to output the Q value. During training, the immediate reward is calculated using the global composite reward function, the twin value network parameters are optimized by minimizing the mean squared error loss function, and the expected cumulative return and action entropy are maximized by minimizing the KL divergence, thereby optimizing the policy network parameters. When the loss functions of the twin value network and policy network tend to stabilize and the cumulative reward reaches the convergence limit, the trained control policy model is output.
9. The method as described in claim 8, characterized in that, The loss function of the twin value network is: in, Represents the loss function of the twin value network; Let be the spatial state observation value at time t. Let t be the action performed by the agent. The instantaneous reward at time t. The spatial state observation value at time t+1 As a reward discount factor, An experience replay pool for storing historical interaction trajectory data; Indicates from the experience replay pool The expected value obtained from the sampled state transition data set. Indicates that the parameter is The value network for the current state Next action Action value estimation; The soft state value at the next time step is calculated using the following formula: in, Indicates action Follow the current policy network Mathematical expectation under given distribution conditions; Indicates the first A value network for the next time-state-value pair Action value assessment value, This means taking the smaller of the evaluation values from the outputs of the two twin value networks, thereby achieving the overestimation of Q-values in consistent reinforcement learning. The parameter is The policy network in a given state Down Output Action The probability distribution; It is a learnable temperature parameter used to adjust the randomness of the strategy.
10. The method as described in claim 8, characterized in that, The policy network loss function is: in, The policy network loss function is represented as follows. Representing state Sampled from the empirical playback pool ,action Sampling from policy network The expected value obtained at that time; The parameter is The policy network in a given state Down Output Action The probability distribution; This indicates taking the smaller of the two value network evaluation values for the current state-action pair.