Mechanical arm grabbing intelligent optimization method and system based on reinforcement learning
By constructing dynamic deformation state vectors and generating grab velocity control instructions in reinforcement learning networks, the problem of strategy convergence of robotic arms when grabbing dynamic deformation objects is solved, and the autonomy and robustness of robotic arms in complex scenarios is improved.
Patent Information
- Application Number
- CN202510408948.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-25
AI Technical Summary
When existing robotic arms grasp dynamically deformed objects, it is difficult to stably perceive and model the position and contact force state of flexible objects, which leads to difficulty in converging strategies and the inability to generate stable and efficient grasping trajectories, limiting their autonomy and robustness in complex scenarios.
By obtaining the dynamic deformation characteristic data of the target object and the real-time position information of the end effector of the robot arm, a dynamic deformation state vector is constructed, and the grab velocity control instructions are generated in combination with the reinforcement learning strategy network, the Jacobian matrix is used for velocity limiting processing, and the reward function weight is dynamically adjusted, and the motion control instruction set is iteratively optimized to meet the convergence conditions.
It significantly improves the independent decision-making ability and stability of the robotic arm in the dynamic deformation object grasping task, ensures that the grab process is efficiently executed within the safety threshold, reduces the uncertainty of physical response, and improves task generalization.
Smart Images

Figure CN120363181A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mechanical control, and more specifically, to an intelligent optimization method and system for robotic arm grasping based on reinforcement learning. Background Art
[0002] The development of industrial automation and intelligent robot technology has gradually extended the robotic arm grasping task from structured scenarios to complex unstructured scenarios; in practical applications such as logistics sorting and medical assistance, robotic arms often need to operate flexible objects with dynamic deformation characteristics (such as soft packaging materials, biological tissues, etc.). Flexible objects are prone to non-rigid deformation due to material characteristics during the grasping process, making it difficult to stably sense and model key state parameters such as their pose and contact force. Traditional grasping control methods usually rely on preset rules or static environment assumptions and are difficult to adapt to the non-linear response characteristics of dynamically deformed objects, restricting the autonomy and robustness of robotic arms in complex scenarios.
[0003] When existing robotic arm grasping methods face dynamically deformed objects, due to the dramatic increase in the state space dimension caused by object deformation and the uncertainty of physical responses, the problem of sparse reward signals becomes prominent during the policy training process. Specifically, the reinforcement learning model is difficult to effectively associate the complex coupling relationship between grasping actions and the state changes of deformed objects through the trial-and-error mechanism, resulting in the policy converging to a local optimal solution and being unable to generate stable and efficient grasping trajectories, restricting the actual application effect of robotic arms in flexible object operation tasks. Summary of the Invention
[0004] To overcome the above-mentioned defects of the prior art, embodiments of the present invention provide an intelligent optimization method and system for robotic arm grasping based on reinforcement learning to solve the problems raised in the above background art.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] An intelligent optimization method for robotic arm grasping based on reinforcement learning, comprising the following steps:
[0007] S1. Obtain the dynamic deformation characteristic data of the target object and the real-time pose information of the end effector of the robotic arm;
[0008] S2. Generate deformation energy parameters based on the dynamic deformation characteristic data and the real-time pose information, perform joint encoding in combination with the robotic arm joint motion parameters, and construct a dynamic deformation state vector;
[0009] S3. Determine a correction coefficient according to the correlation relationship between historical deformation recovery data and the real-time deformation rate distribution to correct the dynamic deformation state vector;
[0010] S4. Input the corrected dynamic deformation state vector into the reinforcement learning policy network to generate a motion control instruction set including the grasping force component;
[0011] S5. Based on the Jacobian matrix of the dynamic deformation state vector and the motion control instruction set, perform amplitude limiting processing on the grasping force component;
[0012] S6. According to the change trend of the deformation energy parameter and the grasping force component after amplitude limiting processing, dynamically adjust the reward function weight distribution coefficient, and iteratively optimize the motion control instruction set until the preset convergence condition is met.
[0013] In a preferred embodiment, S1 includes:
[0014] S1-1. Obtain the deformation displacement distribution of the surface contact area of the target object through a vision sensor, and obtain the corresponding real-time deformation rate distribution of the deformation displacement distribution through a tactile sensor;
[0015] S1-2. Obtain the real-time position coordinates of the end effector through the robotic arm joint encoder, and obtain the real-time attitude angle of the end effector through the inertial measurement unit;
[0016] S1-3. Synchronize and align the deformation displacement distribution, real-time deformation rate distribution, real-time position coordinates, and real-time attitude angle according to a preset time window to generate dynamic deformation characteristic data and real-time pose information.
[0017] In a preferred embodiment, S2 includes:
[0018] S2-1. According to the displacement amount of each data point in the deformation displacement distribution, generate the corresponding local deformation elastic potential energy parameter through the elastic potential energy calculation formula;
[0019] S2-2. According to the rate value of each data point in the real-time deformation rate distribution and the displacement amount of the same data point in the corresponding deformation displacement distribution, generate the local deformation kinetic energy parameter through the nonlinear damping correction formula;
[0020] S2-3. Dynamically adjust the superposition weight coefficient of the local deformation elastic potential energy parameter and the local deformation kinetic energy parameter based on the statistical variance value of the real-time deformation rate distribution to obtain the deformation energy parameter;
[0021] S2-4. Map the deformation energy parameter, the real-time position coordinates and joint torques in the robotic arm joint motion parameters to the same vector space according to a preset coding rule, and construct a dynamic deformation state vector through normalization processing.
[0022] In a preferred embodiment, S3 includes:
[0023] S3-1. Calculate the historical deformation recovery rate distribution based on the difference in the deformation displacement distribution between adjacent time windows in the historical deformation recovery data, where the historical deformation recovery data is the aligned deformation displacement distribution sequence within the preset time window in step S1.
[0024] S3-2. Calculate the point-by-point ratio of the real-time deformation rate distribution and the historical deformation recovery rate distribution at the same data point positions to generate the deformation rate recovery ratio distribution.
[0025] S3-3. Generate the correction coefficient of the dynamic deformation state vector based on the relationship between the ratio of each data point in the deformation rate recovery ratio distribution and the preset recovery ratio threshold, where the preset recovery ratio threshold is calibrated based on the stress relaxation characteristics of the target object material.
[0026] S3-4. Perform a multiplication operation on the correction coefficient and the deformation energy parameter of the corresponding data point in the dynamic deformation state vector to obtain the corrected dynamic deformation state vector.
[0027] In a preferred embodiment, S4 includes:
[0028] S4-1. Expand the dimension of the corrected dynamic deformation state vector, convert the real-time position coordinates and real-time attitude angles into polar coordinate system parameters, and splice them with the deformation energy parameter to form an extended state vector.
[0029] S4-2. Input the extended state vector into the multi-head attention layer of the reinforcement learning policy network to generate the initial weight distribution of the grasping force components based on the preset action space decoupling rule.
[0030] S4-3. Perform non-linear activation processing on the initial weight distribution of the grasping force components through an independent fully connected layer to generate the independent action probability distribution of each component.
[0031] S4-4. Perform anti-normalization processing on the independent action probability distribution according to the maximum driving threshold of the robotic arm joint motion parameters to generate a motion control instruction set including the grasping force components.
[0032] In a preferred embodiment, the initial weight distributions of the translational displacement components and the rotational angle components are also generated based on the preset action space decoupling rule; non-linear activation processing is also performed on the initial weight distributions of the translational displacement components and the rotational angle components; the motion control instruction set also includes the translational displacement components and the rotational angle components.
[0033] In a preferred embodiment, S5 includes:
[0034] S5-1. Calculate the Jacobian matrix of the end effector pose with respect to the joint angles based on the real-time position coordinates and joint torques in the robotic arm joint motion parameters.
[0035] S5-2. Identify the deformation-sensitive regions where the deformation energy parameters exceed the preset energy threshold based on the distribution of deformation energy parameters in the dynamic deformation state vector;
[0036] S5-3. Perform a dot product operation on the Jacobian matrix and the gradient of the deformation energy parameters in the deformation-sensitive regions to generate the spatial gradient influence value of the grasping force component on the deformation energy parameters;
[0037] S5-4. Calculate the amplitude limit threshold of the grasping force component according to the ratio relationship between the spatial gradient influence value and the preset safety factor, where the preset safety factor is calibrated based on the maximum deformation tolerance of the target object material;
[0038] S5-5. Truncate the values in the grasping force component that exceed the amplitude limit threshold, and update the truncated grasping force component to the motion control instruction set.
[0039] In a preferred embodiment, S6 includes:
[0040] S6-1. Calculate the exponentially weighted moving average of the deformation energy accumulation rate according to the temporal change rate of the deformation energy parameters in the dynamic deformation state vector;
[0041] S6-2. Generate a stability decay coefficient characterizing the grasping stability fluctuation based on the variance value of the grasping force component after amplitude limiting processing;
[0042] S6-3. Input the exponentially weighted moving average of the deformation energy accumulation rate and the stability decay coefficient into a preset weight allocation function to dynamically adjust the weight allocation coefficients of the deformation energy accumulation term and the grasping stability term in the reward function;
[0043] S6-4. Recalculate the reward function value based on the updated weight allocation coefficients, and iteratively optimize the motion control instruction set through the policy gradient algorithm until the joint convergence index of the deformation energy parameters and the grasping force component meets the preset convergence conditions.
[0044] On the other hand, the present invention provides a robotic arm grasping intelligent optimization system based on reinforcement learning, including:
[0045] Deformation pose acquisition module: Obtain the dynamic deformation characteristic data of the target object and the real-time pose information of the end effector of the robotic arm;
[0046] Energy state encoding module: Generate deformation energy parameters based on the dynamic deformation characteristic data and the real-time pose information, and perform joint encoding in combination with the robotic arm joint motion parameters to construct a dynamic deformation state vector;
[0047] State Dynamic Correction Module: Determine the correction coefficient according to the correlation between historical deformation recovery data and real-time deformation rate distribution to correct the dynamic deformation state vector;
[0048] Strategy Instruction Generation Module: Input the corrected dynamic deformation state vector into the reinforcement learning strategy network to generate a set of motion control instructions including the grasping force component;
[0049] Force Limiting Control Module: Based on the Jacobian matrix of the dynamic deformation state vector and the set of motion control instructions, perform amplitude limiting processing on the grasping force component;
[0050] Weight Dynamic Optimization Module: Dynamically adjust the reward function weight distribution coefficient according to the change trend of the deformation energy parameter and the grasping force component after amplitude limiting processing, and iteratively optimize the set of motion control instructions until the preset convergence condition is met.
[0051] Compared with the prior art, the present invention has the following beneficial effects:
[0052] 1. Through multi-dimensional state representation and closed-loop feedback mechanism, effectively improve the autonomous decision-making ability and stability of the robotic arm in the task of grasping dynamically deformed objects. In view of the non-linear characteristics of the deformation process of flexible objects, introduce the deformation energy parameter as the core state index, and combine the joint coding mechanism of the robotic arm joint motion parameters to map the complex deformation state and the robotic arm action space into a high-dimensional dynamic vector, significantly enhancing the understanding ability of the reinforcement learning model for the coupling relationship between deformation and action; By real-time correcting the deformation state vector and dynamically adjusting the reward function weight, the system can adaptively balance the deformation energy control and the grasping stability goal, avoid the problem of local convergence of the strategy caused by sparse reward signals in traditional methods, and at the same time suppress the excessive interference of the robotic arm action on the object deformation, ensuring that the grasping process is efficiently executed within the safety threshold;
[0053] 2. Through mechanical gradient analysis and hierarchical optimization strategy, construct a complete closed-loop logic from perception to control; Based on the force limiting mechanism of the Jacobian matrix, accurately quantify the gradient impact of the grasping action on the deformation energy, and realize real-time safety truncation of the force component; The multi-objective dynamic weight distribution and iterative optimization mechanism endow the system with the ability of adaptive adjustment to complex dynamic scenarios, not only reducing the uncertainty of physical response in the operation of dynamically deformed objects, but also improving the task generalization ability of the robotic arm in unstructured environments. Brief Description of the Drawings
[0054] Figure 1 It is a flowchart of an intelligent optimization method for robotic arm grasping based on reinforcement learning according to the present invention;
[0055] Figure 2 It is a structural schematic diagram of an intelligent optimization system for robotic arm grasping based on reinforcement learning according to the present invention. Specific Embodiments
[0056] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0057] Embodiment 1: Figure 1 A method for intelligent optimization of robotic arm grasping based on reinforcement learning of the present invention is given, which includes the following steps:
[0058] S1. Obtain the dynamic deformation characteristic data of the target object and the real-time pose information of the end effector of the robotic arm;
[0059] S2. Generate deformation energy parameters based on the dynamic deformation characteristic data and the real-time pose information, perform joint encoding in combination with the robotic arm joint motion parameters, and construct a dynamic deformation state vector;
[0060] S3. Determine the correction coefficient according to the correlation relationship between the historical deformation recovery data and the real-time deformation rate distribution to correct the dynamic deformation state vector;
[0061] S4. Input the corrected dynamic deformation state vector into the reinforcement learning policy network to generate a motion control instruction set including the grasping force component;
[0062] S5. Perform amplitude limiting processing on the grasping force component based on the Jacobian matrix of the dynamic deformation state vector and the motion control instruction set;
[0063] S6. Dynamically adjust the reward function weight distribution coefficient according to the change trend of the deformation energy parameter and the grasping force component after amplitude limiting processing, and iteratively optimize the motion control instruction set until the preset convergence condition is met.
[0064] S1. Obtain the dynamic deformation characteristic data of the target object and the real-time pose information of the end effector of the robotic arm, including:
[0065] S1-1. Obtain the deformation displacement distribution of the surface contact area of the target object through a vision sensor, and obtain the real-time deformation rate distribution corresponding to the deformation displacement distribution through a tactile sensor;
[0066] S1-2. Obtain the real-time position coordinates of the end effector through a robotic arm joint encoder, and obtain the real-time attitude angle of the end effector through an inertial measurement unit;
[0067] S1-3. Synchronize and align the deformation displacement distribution, real-time deformation rate distribution, real-time position coordinates, and real-time attitude angles according to a preset time window to generate dynamic deformation characteristic data and real-time pose information.
[0068] In step S1-1, the deformation displacement distribution of the contact area on the surface of the target object is obtained through a vision sensor. The vision sensor uses an RGB-D camera, and its depth information is used to capture the three-dimensional deformation displacement data on the object surface. Specifically, during implementation, the feature points preset on the surface of the target object are tracked, and the deformation displacement distribution is generated based on the change amount of the three-dimensional coordinates of the feature points in adjacent frames. Each data point in the deformation displacement distribution represents the displacement amount at the corresponding position in the contact area on the object surface, with the unit of millimeter. At the same time, the real-time deformation rate distribution corresponding to the deformation displacement distribution is obtained through a tactile sensor. The tactile sensor is a distributed pressure sensor array attached to the grasping surface of the end effector of the robotic arm, which measures the change in the dynamic pressure distribution in the object contact area in real time, and calculates the real-time deformation rate distribution according to the rate of change of the pressure distribution over time.
[0069] Each data point in the real-time deformation rate distribution represents the pressure change amount at the corresponding position per unit time, with the unit of Pascal per second. For example, when the end effector contacts an inflated packaging bag, the tactile sensor captures the process of the pressure in the contact area rising from the initial value to the stable value step by step, and its change rate is a component of the real-time deformation rate distribution.
[0070] In step S1-2, the real-time position coordinates of the end effector are obtained through the robotic arm joint encoders. The robotic arm joint encoders are installed at each rotating joint, and record the rotation angles of each joint in real time. Based on the forward solution of the robotic arm kinematic model, the three-dimensional coordinates of the end effector in the base coordinate system are calculated, that is, the real-time position coordinates, with the unit of meter. The real-time position coordinates include the components in the X-axis, Y-axis, and Z-axis directions. At the same time, the real-time attitude angles of the end effector are obtained through an inertial measurement unit. The inertial measurement unit is fixed at the end of the end effector and measures the rotation angles of the end effector around the X-axis, Y-axis, and Z-axis in real time, that is, the roll angle, pitch angle, and yaw angle, with the unit of degree.
[0071] The real-time attitude angles are used to characterize the spatial orientation of the end effector during the grasping process. For example, when the end effector approaches the target object, the roll angle output by the inertial measurement unit is 5 degrees, the pitch angle is -3 degrees, and the yaw angle is 0 degrees, indicating that the end effector is in a slightly tilted state.
[0072] In step S1-3, the deformation displacement distribution, the real-time deformation rate distribution, the real-time position coordinates, and the real-time attitude angles are synchronously aligned according to a preset time window. The length of the preset time window is determined according to the maximum sampling frequency of each sensor. For example, the sampling frequency of the vision sensor is 30 Hz, the sampling frequency of the tactile sensor is 100 Hz, the sampling frequency of the robotic arm joint encoder is 1 kHz, and the sampling frequency of the inertial measurement unit is 200 Hz. Then the preset time window is set to 33 ms to match the data period of the sensor with the lowest sampling frequency.
[0073] The specific method of synchronous alignment is as follows: taking the data acquisition moment of the vision sensor as the reference timestamp, interpolation processing is performed on the data of the tactile sensor, the robotic arm joint encoder, and the inertial measurement unit within the preset time window before and after the reference timestamp to generate a time-aligned deformation displacement distribution, real-time deformation rate distribution, real-time position coordinates, and real-time attitude angles.
[0074] The synchronously aligned data is integrated into dynamic deformation characteristic data and real-time pose information, where the dynamic deformation characteristic data includes the deformation displacement distribution and the real-time deformation rate distribution, and the real-time pose information includes the real-time position coordinates and the real-time attitude angles.
[0075] The above implementation process uses multi-sensor data fusion and time synchronization technology to ensure the spatio-temporal consistency of the dynamic deformation characteristic data and the real-time pose information, providing reliable input for the subsequent generation of deformation energy parameters and the construction of state vectors. The combined use of the vision sensor and the tactile sensor can simultaneously capture the geometric characteristics and dynamic response characteristics of deformation, while the collaborative work of the robotic arm joint encoder and the inertial measurement unit accurately quantifies the spatial pose state of the end effector. The time window synchronization mechanism solves the problem of time misalignment caused by the difference in the data acquisition frequencies of heterogeneous sensors. For example, the high-frequency pressure data of the tactile sensor is aligned with the low-frequency deformation data of the vision sensor through interpolation, avoiding control instruction oscillations caused by data asynchrony.
[0076] S2. Generate deformation energy parameters based on the dynamic deformation characteristic data and the real-time pose information, and perform joint coding in combination with the robotic arm joint motion parameters to construct a dynamic deformation state vector, including:
[0077] S2-1. According to the displacement amount of each data point in the deformation displacement distribution, generate the corresponding local deformation elastic potential energy parameter through the elastic potential energy calculation formula;
[0078] S2-2. According to the rate value of each data point in the real-time deformation rate distribution and the displacement amount of the same data point in the corresponding deformation displacement distribution, generate the local deformation kinetic energy parameter through the non-linear damping correction formula;
[0079] S2-3. Dynamically adjust the superposition weight coefficient of the local deformation elastic potential energy parameter and the local deformation kinetic energy parameter based on the statistical variance value of the real-time deformation rate distribution to obtain the deformation energy parameter;
[0080] S2-4. Map the deformation energy parameter, the real-time position coordinates and joint torques in the robotic arm joint motion parameters to the same vector space according to the preset coding rule, and construct a dynamic deformation state vector through normalization processing.
[0081] In step S2-1, generate the local deformation elastic potential energy parameter according to the displacement amount of each data point in the deformation displacement distribution. The deformation displacement distribution is derived from the three-dimensional displacement data of the object surface contact area obtained by the vision sensor in step S1-1. The elastic potential energy calculation formula is based on Hooke's law principle. Specifically, in implementation, multiply the displacement amount of each data point by its corresponding material stiffness coefficient, then square the result, and then multiply by the preset elastic coefficient constant to obtain the local deformation elastic potential energy parameter. The material stiffness coefficient is pre-calibrated according to the material type of the target object. For example, when the target object is made of rubber, the material stiffness coefficient is set to the first preset value; when the target object is made of silicone, the material stiffness coefficient is set to the second preset value. The elastic coefficient constant is determined by fitting experimental data. For example, in the test scenario of grasping a soft packaging bag, the elastic coefficient constant is set to 0.5.
[0082] In step S2-2, generate the local deformation kinetic energy parameter according to the rate value of each data point in the real-time deformation rate distribution and the displacement amount of the same data point in the corresponding deformation displacement distribution. The real-time deformation rate distribution is derived from the pressure change rate data obtained by the tactile sensor in step S1-1. The specific implementation method of the non-linear damping correction formula is: multiply the square of the rate value by the preset mass coefficient and then by the correction factor, where the correction factor is 1 minus the ratio of the displacement amount to the preset maximum displacement threshold.
[0083] The preset maximum displacement threshold is set according to the material characteristics of the target object. For example, when grasping an inflated packaging bag, the maximum displacement threshold is set to 10 mm. The preset mass coefficient is calibrated through experiments. For example, in the same test scenario, the mass coefficient is set to 0.3. By introducing a correction factor related to the displacement amount, the kinetic energy calculation deviation in the large displacement area is suppressed.
[0084] In step S2-3, the superposition weight coefficient of the local deformation elastic potential energy parameter and the local deformation kinetic energy parameter is dynamically adjusted based on the statistical variance value of the real-time deformation rate distribution. The calculation method of the statistical variance value is as follows: the variance of the rate values of all data points in the real-time deformation rate distribution is calculated to obtain the statistical variance value representing the degree of rate fluctuation. The dynamic adjustment rule of the superposition weight coefficient is as follows: when the statistical variance value is greater than the preset variance threshold, the weight coefficient of the local deformation elastic potential energy parameter is increased; when the statistical variance value is less than or equal to the preset variance threshold, the weight coefficient of the local deformation kinetic energy parameter is increased. The preset variance threshold is determined according to historical data statistics. For example, in multiple grasping tests, the variance threshold is set to 5 pascals per second squared. The adjustment range of the superposition weight coefficient is limited to 0.3 to 0.7 to ensure the balanced contribution of the two types of energy parameters.
[0085] In step S2-4, the deformation energy parameter, the real-time position coordinates and joint torques in the robotic arm joint motion parameters are mapped to the same vector space. The specific implementation method of the preset coding rule is as follows: the X-axis, Y-axis, and Z-axis components of the real-time position coordinates are respectively converted into the radial distance, azimuth angle, and altitude angle in the polar coordinate system, the joint torque is converted into a per-unit value, and the deformation energy parameter retains the original scalar value. The specific method of normalization is as follows: the maximum and minimum normalization is performed on the radial distance, azimuth angle, and altitude angle in the polar coordinate system respectively, and the per-unit value joint torque and the deformation energy parameter are Z-Score standardized to eliminate the dimension difference. All the normalized parameters are concatenated into a dynamic deformation state vector in the preset order. For example, they are arranged in the order of polar coordinate components, per-unit value joint torque, and deformation energy parameter. The dimension of the dynamic deformation state vector is the same as the number of parameters. For example, when the polar coordinate components include 3 dimensions, the per-unit value joint torque includes 6 dimensions, and the deformation energy parameter includes 1 dimension, the state vector is 10-dimensional.
[0086] The above implementation process constructs a high-dimensional vector that can simultaneously represent the deformation energy characteristics of the object and the motion state of the robotic arm through staged energy parameter calculation, dynamic weight adjustment, and multi-source data fusion coding. The separate calculation of the elastic potential energy parameter and the kinetic energy parameter avoids the linear deviation problem of energy indicators in traditional methods, and the dynamic weight adjustment mechanism adaptively balances the energy contribution according to the fluctuation characteristics of the deformation rate. The polar coordinate system conversion and normalization processing ensure the comparability of heterogeneous data in the vector space. For example, the position coordinates of the robotic arm and the deformation energy parameter are in the same numerical order of magnitude after normalization, avoiding the convergence difficulty caused by data scale differences in the subsequent training process of the reinforcement learning policy network.
[0087] S3. Determine a correction coefficient according to the correlation between the historical deformation recovery data and the real-time deformation rate distribution to correct the dynamic deformation state vector, including:
[0088] S3-1. Calculate the historical deformation recovery rate distribution based on the difference in the deformation displacement distribution between adjacent time windows in the historical deformation recovery data, where the historical deformation recovery data is the aligned deformation displacement distribution sequence within the preset time window in step S1.
[0089] S3-2. Calculate the point-by-point ratio of the real-time deformation rate distribution and the historical deformation recovery rate distribution at the same data point positions to generate the deformation rate recovery ratio distribution.
[0090] S3-3. Generate the correction coefficient of the dynamic deformation state vector based on the relationship between the ratio of each data point in the deformation rate recovery ratio distribution and the preset recovery ratio threshold, where the preset recovery ratio threshold is calibrated based on the stress relaxation characteristics of the target object material.
[0091] S3-4. Perform a multiplication operation on the correction coefficient and the deformation energy parameter of the corresponding data point in the dynamic deformation state vector to obtain the corrected dynamic deformation state vector.
[0092] In step S3-1, the historical deformation recovery rate distribution is calculated based on the difference in the deformation displacement distribution between adjacent time windows in the historical deformation recovery data. The historical deformation recovery data is from the aligned deformation displacement distribution sequence in step S1-3. Specifically, in implementation, take the deformation displacement distribution data of the current time window and the previous time window, calculate the difference in the displacement amount of each data point, and then divide by the time window length to obtain the historical deformation recovery rate distribution. For example, when the time window length is 33 milliseconds, the displacement amount of a certain data point in the previous window is 2 millimeters, and the displacement amount in the current window is 1.5 millimeters, then the historical deformation recovery rate value is (2 - 1.5) / 0.033 ≈ 15.15 millimeters per second. The historical deformation recovery rate distribution reflects the natural deformation recovery ability of the object without external intervention.
[0093] In step S3-2, the point-by-point ratio of the real-time deformation rate distribution and the historical deformation recovery rate distribution is calculated at the same data point positions. The real-time deformation rate distribution is from the real-time deformation rate distribution data obtained by the tactile sensor in step S1-1, and the historical deformation recovery rate distribution is from the calculation result of step S3-1. The specific implementation method of the point-by-point ratio calculation is: divide the real-time deformation rate value of each data point by the historical deformation recovery rate value at the corresponding position to generate the deformation rate recovery ratio distribution. For example, when the real-time deformation rate of a certain data point is 20 millimeters per second and the historical deformation recovery rate is 15 millimeters per second, the ratio is 20 / 15 ≈ 1.33. The deformation rate recovery ratio distribution characterizes the relative rate difference between the real-time deformation and the historical natural recovery.
[0094] In step S3-3, a correction coefficient is generated according to the relationship between the ratio of each data point in the strain rate recovery ratio distribution and a preset recovery ratio threshold. The preset recovery ratio threshold is calibrated based on the stress relaxation characteristics of the target object material. Specifically, in implementation, the stress decay curve of the object under a fixed strain is measured through a material tensile experiment, and the reciprocal of the time required for the stress to decay to 50% of the initial value is used as the recovery ratio threshold. For example, if it takes 5 seconds for the stress of a rubber material to decay to 50% under a 10% tensile strain, the recovery ratio threshold is set to 1 / 5 = 0.2 seconds. -1 For each data point in the strain rate recovery ratio distribution, if the ratio is greater than the preset recovery ratio threshold, the correction coefficient is set to 0.8; if the ratio is less than or equal to the preset recovery ratio threshold, the correction coefficient is set to 1.2. The role of the correction coefficient is to inhibit the accumulation of deformation energy beyond the natural recovery ability.
[0095] In step S3-4, the correction coefficient is multiplied by the deformation energy parameter of the corresponding data point in the dynamic deformation state vector. The dynamic deformation state vector is derived from the vector data constructed in step S2-4, and the deformation energy parameter corresponds to the scalar value in the vector that represents the accumulation of the object's deformation energy. Specifically, in implementation, each deformation energy parameter value of the dynamic deformation state vector is multiplied by the correction coefficient of the corresponding data point generated in step S3-3, and other parameters (such as real-time position coordinates, joint torques) remain unchanged. For example, if the deformation energy parameter of a certain data point is 50 joules and the correction coefficient is 0.8, the corrected deformation energy parameter is 50×0.8 = 40 joules. The corrected dynamic deformation state vector is input into the reinforcement learning policy network in the subsequent step S4 to ensure that the policy generation process takes into account the dynamic constraints of the deformation recovery ability.
[0096] The above implementation process quantifies the correlation between real-time deformation and historical recovery rates, dynamically adjusts the weight of the deformation energy parameter, and solves the problem of policy overshoot caused by traditional methods ignoring the natural recovery characteristics of materials. The calculation of the historical deformation recovery rate is based on the time window difference method to avoid the introduction of complex models; the calibration of the preset recovery ratio threshold combines material mechanics experimental data to ensure the physical rationality of the correction coefficient. The point-by-point multiplication operation of the correction coefficient and the deformation energy parameter accurately inhibits abnormal energy accumulation in local areas. For example, when grasping an inflated packaging bag, the local strain rate far exceeds the historical recovery ability, and the correction coefficient automatically reduces the corresponding energy parameter to avoid excessive squeezing actions in policy generation.
[0097] S4. Input the corrected dynamic deformation state vector into the reinforcement learning policy network to generate a motion control instruction set including a grasping force component, including:
[0098] S4-1. Expand the dimension of the corrected dynamic deformation state vector, convert the real-time position coordinates and real-time attitude angles into polar coordinate system parameters, and splice them with the deformation energy parameter to form an extended state vector;
[0099] S4-2. Input the extended state vector into the multi-head attention layer of the reinforcement learning policy network, and generate the initial weight distributions of the translational displacement component, the rotational angle component, and the grasping force component based on the preset action space decoupling rules;
[0100] S4-3. Perform non-linear activation processing on the initial weight distributions of the translational displacement component, the rotational angle component, and the grasping force component respectively through independent fully-connected layers to generate the independent action probability distributions of each component;
[0101] S4-4. According to the maximum driving threshold of the robotic arm joint motion parameters, perform inverse normalization processing on the independent action probability distributions to generate a motion control instruction set including the translational displacement component, the rotational angle component, and the grasping force component.
[0102] In step S4-1, the dimension of the corrected dynamic deformation state vector is extended. The corrected dynamic deformation state vector is derived from the correction result of step S3-4, and it includes the deformation energy parameter, the real-time position coordinates, and the joint torque. The specific implementation method of dimension extension is: convert the X-axis, Y-axis, and Z-axis components of the real-time position coordinates into the radial distance, azimuth angle, and elevation angle in the polar coordinate system. The conversion rule is calculated according to the mathematical relationship from the three-dimensional rectangular coordinate system to the polar coordinate system. The converted polar coordinate system parameters and the deformation energy parameter are concatenated in a preset order, for example, arranged in the order of radial distance, azimuth angle, elevation angle, and deformation energy parameter to form the extended state vector.
[0103] In step S4-2, input the extended state vector into the multi-head attention layer of the reinforcement learning policy network to generate the initial weight distribution. The specific implementation method of the multi-head attention layer is: split the extended state vector into multiple sub-vectors according to the preset number of heads, and perform dot product operations on each sub-vector with the query vectors of the preset translational displacement component, rotational angle component, and grasping force component respectively to generate the attention weights of each action component. The preset action space decoupling rule is that each action component corresponds to an independent query vector. For example, the dimension of the query vector of the translational displacement component is 3, the rotational angle component is 3, and the grasping force component is 1. The initial weight distribution is obtained by weighted summation of the attention weights of each sub-vector. For example, when the extended state vector is split into 4 sub-vectors, the initial weight of the translational displacement component is the average of the weights corresponding to the 4 sub-vectors.
[0104] In step S4-3, the initial weight distribution is non-linearly activated through an independent fully-connected layer. The specific implementation of the independent fully-connected layer is as follows: independent fully-connected neural network layers are respectively set for the initial weight distributions of the translational displacement component, the rotation angle component, and the grasping force component. Each layer contains a preset number of hidden neurons. For example, the fully-connected layer of the translational displacement component contains 64 hidden neurons, the fully-connected layer of the rotation angle component contains 64 hidden neurons, and the fully-connected layer of the grasping force component contains 32 hidden neurons. The non-linear activation processing uses the ReLU function to map the output value of the fully-connected layer to the non-negative interval, generating the independent action probability distribution of each component. For example, the action probability distribution of the translational displacement component is the normalized value within the interval [-1, 1], representing the displacement ratio of the robotic arm in the X, Y, and Z axis directions.
[0105] In step S4-4, the independent action probability distribution is inverse-normalized according to the maximum driving threshold of the robotic arm joint motion parameters. The maximum driving threshold comes from the physical driving limits of each joint of the robotic arm. For example, the maximum driving threshold of the translational displacement component is ±0.5 m / s, the maximum driving threshold of the rotation angle component is ±30 deg / s, and the maximum driving threshold of the grasping force component is 50 N. The specific method of inverse-normalization is: multiplying the normalized action probability distribution value by the maximum driving threshold of the corresponding component to obtain the actual motion control instruction. For example, if the action probability distribution value of the translational displacement component is 0.5, the actual translational displacement instruction is 0.5 × 0.5 = 0.25 m / s. The finally generated motion control instruction set includes the actual control values of the translational displacement component, the rotation angle component, and the grasping force component.
[0106] The above implementation process ensures that the generation process of the action components not only retains the characteristics of deformation energy but also conforms to the physical driving constraints of the robotic arm through polar coordinate transformation, multi-head attention decoupling, and independent fully-connected layer design. The polar coordinate system transformation enhances the spatial representation ability of the position parameters. For example, the azimuth angle and the elevation angle can more intuitively reflect the motion direction of the end effector; the multi-head attention mechanism improves the ability of the policy network to capture complex state features through parallel calculation of multiple sub-vectors; the independent fully-connected layer avoids parameter coupling of different action components. For example, the training of the grasping force component is not interfered by the gradient of the translational displacement component. The inverse-normalization processing maps the abstract probability distribution to the actual driving range. For example, the 50 N upper limit of the grasping force component can prevent the robotic arm from being damaged due to overload.
[0107] S5. Perform amplitude limiting processing on the grasping force component based on the Jacobian matrix of the dynamic deformation state vector and the motion control instruction set, including:
[0108] S5-1. Calculate the Jacobian matrix of the end effector pose with respect to the joint angles according to the real-time position coordinates and joint torques in the robotic arm joint motion parameters;
[0109] S5-2. Identify the deformation-sensitive regions where the deformation energy parameters exceed the preset energy threshold based on the distribution of deformation energy parameters in the dynamic deformation state vector.
[0110] S5-3. Perform a dot product operation on the Jacobian matrix and the gradient of the deformation energy parameters in the deformation-sensitive regions to generate the spatial gradient influence value of the grasping force component on the deformation energy parameters.
[0111] S5-4. Calculate the amplitude limit threshold of the grasping force component according to the ratio relationship between the spatial gradient influence value and the preset safety factor, where the preset safety factor is calibrated based on the maximum deformation tolerance of the target object material.
[0112] S5-5. Truncate the values in the grasping force component that exceed the amplitude limit threshold, and update the truncated grasping force component to the motion control instruction set.
[0113] In step S5-1, calculate the Jacobian matrix of the end effector pose with respect to the joint angles based on the real-time position coordinates and joint torques in the robotic arm joint motion parameters. The real-time position coordinates are from the three-dimensional coordinates of the end effector obtained by the robotic arm joint encoder in step S1-2, and the joint torques are from the real-time torque feedback data of the robotic arm drive system in step S2-4.
[0114] The specific calculation method of the Jacobian matrix is as follows: According to the forward kinematic formula of the robotic arm kinematic model, calculate the partial derivatives of the real-time position coordinates (X-axis, Y-axis, Z-axis) and attitude angles (roll angle, pitch angle, yaw angle) of the end effector with respect to each joint angle respectively, and form a Jacobian matrix with 6 rows and N columns, where N is the number of robotic arm joints. For example, when the robotic arm has 6 revolute joints, the Jacobian matrix is a 6x6 matrix, and each column corresponds to the influence degree of a joint angle change on the end effector pose.
[0115] In step S5-2, identify the deformation-sensitive regions based on the distribution of deformation energy parameters in the dynamic deformation state vector. The deformation energy parameters are from the scalar values generated in step S2-3, and the dynamic deformation state vector is from the correction result in step S3-4. The identification rule for the deformation-sensitive regions is: Traverse all the deformation energy parameter data points in the dynamic deformation state vector. If the deformation energy parameter value of a certain data point exceeds the preset energy threshold, then determine that the region where this data point is located is the deformation-sensitive region.
[0116] The preset energy threshold is calibrated according to the yield strength of the target object material. For example, when grasping an object made of rubber material, the preset energy threshold is set to 80% of the material yield strength value. If the yield strength of rubber measured through a material tensile experiment is 10 joules, then the preset energy threshold is 8 joules. The determination result of the deformation-sensitive region is stored in the form of a binary mask, where the position value marked as the sensitive region is 1 and the non-sensitive region is 0.
[0117] In step S5-3, the Jacobian matrix is dot-multiplied with the gradient of the deformation energy parameter in the deformation-sensitive region to generate a spatial gradient influence value. The gradient of the deformation energy parameter is derived from the distribution of the deformation energy parameter after superimposing weights in step S2-3, and the gradient value is obtained by calculating the difference in the deformation energy parameter between adjacent data points. The specific implementation of the dot product operation is as follows: for each data point in the deformation-sensitive region, the gradient vector of its deformation energy parameter (including components in the X, Y, and Z directions) is extracted, and this gradient vector is dot-multiplied with the column vector corresponding to the end-effector pose in the Jacobian matrix to obtain the spatial gradient influence value of this data point.
[0118] For example, when the gradient vector of a certain data point is (0.2, 0.3, 0.1), and the column vectors of the Jacobian matrix corresponding to the X, Y, and Z positions of the end-effector are (0.5, 0.6, 0.4), then the dot product result is 0.2×0.5 + 0.3×0.6 + 0.1×0.4 = 0.1 + 0.18 + 0.04 = 0.32. The spatial gradient influence values of all data points in the sensitive region form a gradient influence distribution map.
[0119] In step S5-4, the amplitude limit threshold is calculated based on the ratio relationship between the spatial gradient influence value and the preset safety factor. The preset safety factor is calibrated based on the maximum deformation tolerance of the target object material. The maximum deformation tolerance is measured by the maximum compression deformation of the object before fracture through a material compression experiment and converted into an energy threshold. For example, when the maximum compression deformation of silicone material is 20 mm, the corresponding energy threshold is 15 joules, and the safety factor is set to the reciprocal of the energy threshold, that is, 1 / 15 ≈ 0.067 joules -1 。
[0120] The calculation formula for the amplitude limit threshold is: the safety factor divided by the absolute value of the spatial gradient influence value. When the spatial gradient influence value is 0.32, the amplitude limit threshold is 0.067 / 0.32 ≈ 0.21 N -1 This threshold represents the upper limit of the contribution of unit grasping force to the accumulation of deformation energy.
[0121] In step S5-5, the values in the grasping force component that exceed the amplitude limit threshold are truncated. The grasping force component is derived from the force command value in the motion control instruction set generated in step S4-4. The specific method of truncation is as follows: compare the command value of each grasping force component with the amplitude limit threshold. If the command value is greater than the threshold, set the command value to the threshold; if the command value is less than or equal to the threshold, retain the original value. For example, when the force command value for a certain grasp is 25 N and the amplitude limit threshold is 20 N, the truncated force command value is 20 N. The truncated grasping force component is updated to the motion control instruction set to ensure that the subsequent iterative optimization process in step S6 is carried out within the safety boundary.
[0122] In the above implementation process, through mechanical gradient analysis and dynamic safety threshold setting, the influence intensity of the grasping force on the deformation energy is accurately controlled. The introduction of the Jacobian matrix takes into account the mechanical relationship between the robotic arm joint motion and the end effector pose. For example, when the joint torque of the robotic arm increases, the magnitude of the corresponding column vector of the Jacobian matrix rises, resulting in an increase in the spatial gradient influence value, thereby reducing the amplitude limit threshold and suppressing the grasping force command. The identification of the deformation-sensitive area is based on the calibration of the material yield strength to ensure that the force limit is automatically triggered when approaching the material deformation limit. The truncation process adopts a point-by-point comparison mechanism to avoid the decline in grasping stability caused by global force suppression. For example, only the force commands in high-risk areas are restricted, and the original command values are maintained in other areas.
[0123] S6. Dynamically adjust the weight distribution coefficient of the reward function according to the change trend of the deformation energy parameter and the grasping force component after amplitude limit processing, and iteratively optimize the motion control instruction set until the preset convergence condition is met, including:
[0124] S6-1. Calculate the exponentially weighted moving average of the deformation energy accumulation rate according to the time series change rate of the deformation energy parameter in the dynamic deformation state vector;
[0125] S6-2. Generate a stability decay coefficient representing the fluctuation of grasping stability based on the variance value of the grasping force component after amplitude limit processing;
[0126] S6-3. Input the exponentially weighted moving average of the deformation energy accumulation rate and the stability decay coefficient into the preset weight distribution function to dynamically adjust the weight distribution coefficients of the deformation energy accumulation term and the grasping stability term in the reward function;
[0127] S6-4. Recalculate the reward function value based on the updated weight distribution coefficient, and iteratively optimize the motion control instruction set through the policy gradient algorithm until the joint convergence index of the deformation energy parameter and the grasping force component meets the preset convergence condition.
[0128] In step S6-1, the exponentially weighted moving average of the deformation energy accumulation rate is calculated according to the temporal change rate of the deformation energy parameter in the dynamic deformation state vector. The dynamic deformation state vector is derived from the correction result of step S3-4 and includes the corrected deformation energy parameter, the real-time position coordinates, and the joint torque. The calculation method of the temporal change rate is as follows: take the difference between the deformation energy parameter at the current moment and the previous moment, and divide it by the time interval. For example, when the time interval is 0.1 second, the current deformation energy parameter is 50 joules, and the previous moment is 48 joules, then the temporal change rate is (50 - 48) / 0.1 = 20 joules / second. The exponentially weighted moving average is calculated by weighted summing the historical change rates using an exponential decay factor, and the decay factor is calibrated according to the material relaxation characteristics of the target object. For example, the decay factor for rubber material is set to 0.9, and for silicone material is set to 0.85. The exponentially weighted moving average reflects the long-term trend of deformation energy accumulation and suppresses the interference of instantaneous fluctuations.
[0129] In step S6-2, the stability decay coefficient is generated based on the variance value of the grasping force component after amplitude limiting processing. The grasping force component after amplitude limiting processing is derived from the truncation result of step S5-5. The calculation method of the variance value is as follows: take the numerical sequence of the grasping force component within a continuous preset time window and calculate its variance. For example, when the time window contains 10 sampling points, the numerical sequence of the grasping force component is [20 N, 21 N, 19 N, 20 N, 22 N], and its variance is calculated as the average of the sum of the squared deviations of each value from the mean. The generation rule of the stability decay coefficient is: divide the variance value by the preset variance reference value, and then take the natural logarithm to obtain the resulting stability decay coefficient. The stability decay coefficient characterizes the stability fluctuation degree of the grasping action, and the larger the coefficient, the worse the stability.
[0130] In step S6-3, the exponentially weighted moving average of the deformation energy accumulation rate and the stability decay coefficient are input into a preset weight allocation function to dynamically adjust the weight allocation coefficient. The specific implementation method of the preset weight allocation function is: linearly combine the exponentially weighted moving average and the stability decay coefficient according to a preset ratio to generate the weight allocation coefficients for the deformation energy accumulation term and the grasping stability term. For example, when the exponentially weighted moving average is 15 joules / second, the stability decay coefficient is 0.693, and the preset ratio is that the weight proportion of the deformation energy accumulation term is 60% and the grasping stability term is 40%, the weight allocation coefficient of the deformation energy accumulation term is 15×0.6 = 9, and the weight allocation coefficient of the grasping stability term is 0.693×0.4≈0.277. After normalizing the weight allocation coefficients, the weight of the deformation energy accumulation term is 9 / (9 + 0.277)≈97%, and the weight of the grasping stability term is 3%. The weight allocation coefficient is dynamically updated with each iteration.
[0131] In step S6-4, the reward function value is recalculated based on the updated weight distribution coefficient, and the motion control instruction set is iteratively optimized through the policy gradient algorithm. The reward function value is calculated as follows: the numerical values of the deformation energy accumulation term and the grasping stability term are multiplied by their corresponding weight distribution coefficients respectively and then added together. For example, when the numerical value of the deformation energy accumulation term is 50 joules, the weight distribution coefficient is 0.97, the numerical value of the grasping stability term is 80 (dimensionless stability score), and the weight distribution coefficient is 0.03, the reward function value is 50×0.97 + 80×0.03 = 48.5 + 2.4 = 50.9. The specific implementation method of the policy gradient algorithm is: calculate the gradient direction of the policy network according to the reward function value, and update the network parameters along the gradient direction to maximize the reward value. The termination condition for iterative optimization is that the following two preset convergence conditions are satisfied simultaneously: the change rate of the deformation energy parameter is less than 1 joule / second, and the variance value of the grasping force component is less than 1 N². When the joint convergence index reaches the above conditions, the final motion control instruction set is output.
[0132] The above implementation process ensures that the policy optimization process takes into account both deformation energy control and grasping stability through dynamic weight adjustment and multi-objective joint convergence verification. The exponentially weighted moving average suppresses the misjudgment of the instantaneous noise on the energy accumulation trend. For example, when grasping an inflated packaging bag, a brief deformation surge will not trigger a weight mutation; the stability decay coefficient quantifies the force fluctuation into an adjustable penalty term. For example, when the force variance is too large, the stability weight is automatically increased; the iterative optimization of the policy gradient algorithm is based on the normalized weight distribution coefficient, avoiding the convergence oscillation caused by target conflicts. The setting of the joint convergence index combines the deformation energy and the force stability. For example, when the change rate of the deformation energy meets the standard and the force fluctuation is controllable, it is determined to converge, ensuring the physical feasibility and reliability of the grasping action.
[0133] Embodiment 2: Figure 2 The structural schematic diagram of an intelligent optimization system for robotic arm grasping based on reinforcement learning according to the present invention is given. An intelligent optimization system for robotic arm grasping based on reinforcement learning includes:
[0134] Deformation pose acquisition module: Obtain the dynamic deformation characteristic data of the target object and the real-time pose information of the end effector of the robotic arm;
[0135] Energy state encoding module: Generate deformation energy parameters based on the dynamic deformation characteristic data and the real-time pose information, and perform joint encoding in combination with the robotic arm joint motion parameters to construct a dynamic deformation state vector;
[0136] State dynamic correction module: Determine the correction coefficient according to the correlation between the historical deformation recovery data and the real-time deformation rate distribution to correct the dynamic deformation state vector;
[0137] Strategy instruction generation module: Input the corrected dynamic deformation state vector into the reinforcement learning policy network to generate a set of motion control instructions including the grasping force component;
[0138] Force limit control module: Based on the Jacobian matrix of the dynamic deformation state vector and the motion control instruction set, perform amplitude limiting processing on the grasping force component;
[0139] Weight dynamic optimization module: According to the change trend of the deformation energy parameter and the grasping force component after amplitude limiting processing, dynamically adjust the reward function weight distribution coefficient, and iteratively optimize the motion control instruction set until the preset convergence condition is met.
[0140] The above formulas are all dimensionless and take their numerical calculations. The formulas are obtained by collecting a large amount of data for software simulation to obtain a formula closest to the actual situation. The preset parameters and threshold selection in the formulas are set by those skilled in the art according to the actual situation.
[0141] It should be noted that the present invention can be deployed on the device itself to achieve embedded applications, or can also run on a PC or other terminals with a user interface, so as to meet various hardware environments and usage requirements.
[0142] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on the computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wire (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0143] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described systems, devices, and modules can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0144] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or modules can be in electrical, mechanical, or other forms.
[0145] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules. They can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0146] In addition, in each embodiment of the present application, the functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.
[0147] If the function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or this part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0148] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0149] Finally, the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An intelligent optimization method for robotic arm grasping based on reinforcement learning, characterized in that, It includes the following steps: S1. Obtain the dynamic deformation characteristic data of the target object and the real-time pose information of the end effector of the robotic arm; S2. Generate deformation energy parameters based on the dynamic deformation characteristic data and the real-time pose information, perform joint encoding in combination with the robotic arm joint motion parameters, and construct a dynamic deformation state vector; S3. Determine a correction coefficient according to the correlation relationship between the historical deformation recovery data and the real-time deformation rate distribution to correct the dynamic deformation state vector; S4. Input the corrected dynamic deformation state vector into the reinforcement learning policy network to generate a motion control instruction set including the grasping force component; S5. Based on the Jacobian matrix of the dynamic deformation state vector and the motion control instruction set, perform amplitude limiting processing on the grasping force component; S6. Dynamically adjust the reward function weight distribution coefficient according to the change trend of the deformation energy parameter and the grasping force component after amplitude limiting processing, and iteratively optimize the motion control instruction set until the preset convergence condition is met.
2. The intelligent optimization method for robotic arm grasping based on reinforcement learning according to claim 1, wherein S1 includes: S1-1. Obtain the deformation displacement distribution of the surface contact area of the target object through a vision sensor, and obtain the corresponding real-time deformation rate distribution of the deformation displacement distribution through a tactile sensor; S1-2. Obtain the real-time position coordinates of the end effector through the robotic arm joint encoder, and obtain the real-time attitude angle of the end effector through an inertial measurement unit; S1-3. Synchronize and align the deformation displacement distribution, the real-time deformation rate distribution, the real-time position coordinates, and the real-time attitude angle according to a preset time window to generate the dynamic deformation characteristic data and the real-time pose information.
3. A method for intelligent optimization of robotic arm grasping based on reinforcement learning according to claim 1, characterized in that S2 It includes: S2-1. Generate the corresponding local deformation elastic potential energy parameter according to the displacement amount of each data point in the deformation displacement distribution through the elastic potential energy calculation formula; S2-2. Generate the local deformation kinetic energy parameter according to the rate value of each data point in the real-time deformation rate distribution and the displacement amount of the same data point in the corresponding deformation displacement distribution through the non-linear damping correction formula; S2-3. Dynamically adjust the superposition weight coefficient of the local deformation elastic potential energy parameter and the local deformation kinetic energy parameter based on the statistical variance value of the real-time deformation rate distribution to obtain the deformation energy parameter; S2-4. Map the deformation energy parameter, the real-time position coordinates and the joint torque in the robotic arm joint motion parameters to the same vector space according to a preset encoding rule, and construct a dynamic deformation state vector through normalization processing.
4. The intelligent optimization method for robotic arm grasping based on reinforcement learning according to claim 1, wherein S3 It includes: S3-1. Calculate the historical deformation recovery rate distribution according to the difference in the deformation displacement distribution of adjacent time windows in the historical deformation recovery data, where the historical deformation recovery data is the sequence of aligned deformation displacement distributions within the preset time window in step S1; S3-2. Perform point-by-point ratio calculation on the real-time deformation rate distribution and the historical deformation recovery rate distribution at the same data point positions to generate the deformation rate recovery ratio distribution; S3-3. Generate the correction coefficient of the dynamic deformation state vector according to the relationship between the ratio of each data point in the deformation rate recovery ratio distribution and a preset recovery ratio threshold, where the preset recovery ratio threshold is calibrated based on the stress relaxation characteristics of the target object material. S3-4. Multiply the correction coefficient by the deformation energy parameter of the corresponding data point in the dynamic deformation state vector to obtain the corrected dynamic deformation state vector.
5. The intelligent optimization method for robotic arm grasping based on reinforcement learning according to claim 1, wherein S4 Including: S4-1. Expand the dimension of the corrected dynamic deformation state vector, convert the real-time position coordinates and real-time attitude angles into polar coordinate system parameters, and splice them with the deformation energy parameters to form an extended state vector; S4-2. Input the extended state vector into the multi-head attention layer of the reinforcement learning policy network, and generate the initial weight distribution of the grasping force components based on the preset action space decoupling rule; S4-3. Perform non-linear activation processing on the initial weight distribution of the grasping force components through an independent fully connected layer to generate the independent action probability distribution of each component; S4-4. According to the maximum driving threshold of the robotic arm joint motion parameters, perform inverse normalization processing on the independent action probability distribution to generate a motion control instruction set including the grasping force components.
6. The intelligent optimization method for robotic arm grasping based on reinforcement learning according to claim 5, characterized in that Based on the preset action space decoupling rule, the initial weight distributions of the translational displacement components and the rotational angle components are also generated; it also includes performing non-linear activation processing on the initial weight distributions of the translational displacement components and the rotational angle components; the motion control instruction set also includes the translational displacement components and the rotational angle components.
7. The intelligent optimization method for robotic arm grasping based on reinforcement learning according to claim 1, wherein S5 Including: S5-1. Calculate the Jacobian matrix of the end effector pose with respect to the joint angles according to the real-time position coordinates and joint torques in the robotic arm joint motion parameters; S5-2. Based on the deformation energy parameter distribution in the dynamic deformation state vector, identify the deformation sensitive area where the deformation energy parameter exceeds the preset energy threshold; S5-3. Perform a dot product operation on the Jacobian matrix and the deformation energy parameter gradient of the deformation sensitive area to generate the spatial gradient influence value of the grasping force component on the deformation energy parameter; S5-4. Calculate the amplitude limit threshold of the grasping force component according to the ratio relationship between the spatial gradient influence value and the preset safety factor, where the preset safety factor is calibrated based on the maximum deformation tolerance of the target object material; S5-5. Truncate the values in the grasping force component that exceed the amplitude limit threshold, and update the truncated grasping force component to the motion control instruction set.
8. The intelligent optimization method for robotic arm grasping based on reinforcement learning according to claim 1, wherein S6 Including: S6-1. Calculate the exponentially weighted moving average of the deformation energy accumulation rate according to the time series change rate of the deformation energy parameter in the dynamic deformation state vector; S6-2. Generate a stability decay coefficient representing the grasping stability fluctuation based on the variance value of the grasping force component after amplitude limiting processing; S6-3. Input the exponentially weighted moving average of the deformation energy accumulation rate and the stability decay coefficient into the preset weight allocation function to dynamically adjust the weight allocation coefficients of the deformation energy accumulation term and the grasping stability term in the reward function; S6-4. Recalculate the reward function value based on the updated weight allocation coefficients, and iteratively optimize the motion control instruction set through the policy gradient algorithm until the joint convergence index of the deformation energy parameter and the grasping force component meets the preset convergence condition.
9. A robotic arm grasping intelligent optimization system based on reinforcement learning, which is used to implement the robotic arm grasping intelligent optimization method according to any one of claims 1-8, characterized in that, Including: Deformation pose acquisition module: Obtain the dynamic deformation characteristic data of the target object and the real-time pose information of the end effector of the robotic arm; Energy state encoding module: Generate deformation energy parameters based on dynamic deformation characteristic data and real-time pose information, perform joint encoding in combination with robotic arm joint motion parameters, and construct a dynamic deformation state vector; State dynamic correction module: Determine a correction coefficient according to the correlation between historical deformation recovery data and real-time deformation rate distribution to correct the dynamic deformation state vector; Strategy instruction generation module: Input the corrected dynamic deformation state vector into a reinforcement learning strategy network to generate a motion control instruction set including a grasping force component; Force amplitude limiting control module: Perform amplitude limiting processing on the grasping force component based on the Jacobian matrix of the dynamic deformation state vector and the motion control instruction set; Weight dynamic optimization module: Dynamically adjust the reward function weight distribution coefficient according to the change trend of the deformation energy parameter and the grasping force component after amplitude limiting processing, and iteratively optimize the motion control instruction set until the preset convergence condition is met.
Citation Information
Patent Citations
Autonomous unknown object pick and place
CA3130215A1
Mechanical arm autonomous moving grabbing method based on visual-tactile fusion under complex illumination condition
CN113696186A
Control method for preventing over-deformation grasping and converging at specified time
CN116604569A
Two-finger clamp deformable object fine grabbing system based on tactile feedback
CN117656039A
Cited By
Mechanical arm intelligent control method and system based on deep reinforcement learning
CN120572540A
Operation and maintenance manipulator intelligent control method and system based on visual identification
CN120680525A
An intelligent control method and system for an operation and maintenance robot based on visual recognition
CN120680525B
Method and system for constructing intelligent simulation data set of electric power scene
CN120822432A
Container loading space optimization method based on robot collaboration
CN121093406A