Robotic Arm Control Method and Device, Electronic Device, and Storage Medium
By combining the reward value reshaping and importance setting sampling technology of visually perceptual excitation in the robotic arm control model, the problems of low sampling efficiency and sparse rewards in the traditional method are solved, and more efficient robotic arm control performance is achieved.
Patent Information
- Application Number
- CN202510388879.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-31
AI Technical Summary
Traditional robot arm control methods have problems with low algorithm sampling efficiency and sparse rewards under deep learning methods, making it difficult to effectively improve the algorithm performance during robot arm control.
By constructing a robotic arm control model based on deep reinforcement learning algorithm, combining the reward value based on visual perception excitation, generating the reshaping reward value, and setting and sampling the sample data based on this, model parameters are optimized to improve control performance.
It improves the sampling efficiency of the robotic arm control algorithm, reduces reward sparsity, enhances the adaptability and intelligence level of the model in complex environments, and solves the problem of improving algorithm performance in traditional methods.
Smart Images

Figure CN119871469B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a robotic arm control method and device, an electronic device, and a storage medium. Background Technique
[0002] With the development of technology, robot research has always been an important area of technological development and is widely used in fields such as domestic service, manufacturing, national defense technology, aerospace, and medical assistance. With the development of artificial intelligence technology, combining artificial intelligence with robot technology to study robotic control systems with autonomous decision-making and learning capabilities has become an important branch and frontier focus of robot research. Among them, the robotic arm, as an important execution mechanism of the robot, is widely used in key areas of production and life. The research and development of the robotic arm control system is of great significance to the development of robots. The robotic arm control system is based on the robot's own perception, decision-making, planning, and control capabilities, and the operation target is to reach the target state from the initial state.
[0003] Traditional robotic arm control methods can enable robots to continuously interact with the environment and learn through trial and error through reinforcement learning to autonomously obtain optimized strategies and adapt to dynamic external environments. However, when traditional robotic arm control methods are used for robotic arm control through deep learning, the problems of low algorithm sampling efficiency and sparse rewards are not considered. Therefore, how to improve the sampling efficiency of the algorithm and solve the sparse rewards of the algorithm during the robotic arm control process are problems that need to be solved urgently at present. Summary of the Invention
[0004] This application provides a robotic arm control method and device, an electronic device, and a storage medium to at least solve the problem of how to improve the sampling efficiency of the algorithm and solve the sparse rewards of the algorithm during the robotic arm control process in related technologies.
[0005] This application provides a robotic arm control method, including:
[0006] Obtain the state information, action information, state identifier, and reshaped reward value of the robotic arm generated during the robotic arm control process based on the robotic arm control model, where the robotic arm control model is a model constructed based on a deep reinforcement learning algorithm, and the reshaped reward value is a reward value obtained by reshaping the reward value in combination with a visual perception incentive-based method;
[0007] Generate sample data according to the state information, action information, reshaped reward value, and state identifier, and store the sample data in the sample pool;
[0008] Perform importance setting processing on the sample data in the sample pool based on the reshaped reward value to obtain the importance corresponding to the sample data, and perform data sampling processing on the sample data in the sample pool based on the importance of the sample data to obtain sampling data;
[0009] Optimize the model parameters of the robotic arm control model according to the sampled data to obtain the optimized robotic arm control model;
[0010] Control the robotic arm according to the optimized robotic arm control model.
[0011] This application also provides a robotic arm control device, including:
[0012] An acquisition unit, configured to acquire the state information, action information, state identifier, and reshaped reward value of the robotic arm generated during the control process of the robotic arm based on the robotic arm control model, where the robotic arm control model is a model constructed based on a deep reinforcement learning algorithm, and the reshaped reward value is a reward value obtained by reshaping the reward value in combination with a visual perception incentive-based method;
[0013] A generation unit, configured to generate sample data according to the state information, action information, reshaped reward value, and state identifier, and store the sample data in a sample pool;
[0014] A setting unit, configured to perform importance setting processing on the sample data in the sample pool based on the reshaped reward value to obtain the importance corresponding to the sample data;
[0015] A sampling unit, configured to perform data sampling processing on the sample data in the sample pool based on the importance of the sample data to obtain sampled data;
[0016] An optimization unit, configured to optimize the model parameters of the robotic arm control model according to the sampled data to obtain the optimized robotic arm control model;
[0017] A control unit, configured to control the robotic arm according to the optimized robotic arm control model.
[0018] This application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above robotic arm control methods when executing the computer program.
[0019] This application also provides a computer-readable storage medium, in which a computer program is stored, where the computer program, when executed by a processor, implements the steps of any of the above robotic arm control methods.
[0020] This application also provides a computer program product, including a computer program, where the computer program, when executed by a processor, implements the steps of any of the above robotic arm control methods.
[0021] Through this application, a robotic arm control method and device, an electronic device, and a storage medium are provided. During the process of controlling the robotic arm by the robotic arm control model, a reward value reshaping process is carried out by combining a vision perception incentive method to determine the reshaped reward value, which can reduce reward sparsity. Various information generated during the robotic arm control process by the robotic arm control model is used to generate sample data. Then, based on the reshaped reward value, importance settings are made for the sample data, and sampling is carried out according to the importance of the sample data. Compared with uniform sampling in the prior art, the sampling efficiency is improved, and the technical problems of how to improve the sampling efficiency of the algorithm during the robotic arm control process and how to solve the reward sparsity of the algorithm can be solved, achieving the technical effects of improving the sampling efficiency of the algorithm and solving the reward sparsity of the algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] To more clearly illustrate the embodiments of the present application, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0023] Figure 1 It is a schematic flowchart of a robotic arm control method provided by an embodiment of the present application;
[0024] Figure 2 It is a schematic flowchart of information acquisition provided by an embodiment of the present application;
[0025] Figure 3 It is a schematic diagram of the principle of a method for importance setting provided by an embodiment of the present application;
[0026] Figure 4 It is a schematic diagram of a vision perception incentive strategy provided by an embodiment of the present application;
[0027] Figure 5 It is a schematic structural diagram of a robotic arm control device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0029] It should be noted that in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0030] In order to enable those skilled in the art of this technology to better understand the solution of this application, the following further describes this application in detail with reference to the drawings and specific embodiments.
[0031] Figure 1 The flowchart of a robotic arm control method provided for an embodiment of this application is shown. In combination with the execution process of the robotic arm control method, the method is described in detail.
[0032] As Figure 1 shown, the robotic arm control method includes:
[0033] Step 101, obtain the state information, action information, state identifier, and reshaped reward value of the robotic arm generated during the control process of the robotic arm based on a robotic arm control model, where the robotic arm control model is a model constructed based on a deep reinforcement learning algorithm, and the reshaped reward value is a reward value obtained by reshaping the reward value in combination with a visual perception incentive-based method.
[0034] In the embodiment of this application, the robotic arm control model is an intelligent decision-making model constructed based on a deep reinforcement learning algorithm, and its core adopts but is not limited to: the improved Deep Deterministic Policy Gradient (DDPG) algorithm. The algorithm collaborates four neural networks (Actor evaluation network, Actor target network, Critic evaluation network, Critic target network) through the Actor-Critic framework to handle the control problem of the robotic arm in a continuous action space.
[0035] The state information is a multi-dimensional data set describing the real-time motion state of the robotic arm during the control process, which includes first state information and second state information. The first state information includes the state of the robotic arm before performing an action, and the second state information includes the state of the robotic arm after performing an action. The first state information and the second state information both include but are not limited to: the angle characteristics of each joint of the robotic arm, the position information of the end effector, and the spatial relationship with the target position.
[0036] The action information refers to the control instructions output by the robotic arm control model, including the actions that the robotic arm needs to execute, which are manifested as the angular velocity or angle adjustment amount of each joint at the next time step. For example, for a robotic arm with N degrees of freedom, the action information can be represented as a vector , where is the angle change amount of the th joint.
[0037] The status flag is a key signal used to mark the progress of the robotic arm control task, usually represented in the form of a boolean variable (e.g., done) to indicate whether the task has terminated or completed. For example, when the end effector of the robotic arm reaches within the allowable error range of the target position, the status flag is marked as the termination state.
[0038] Among them, the reward value obtained by reshaping the reward value in combination with the method based on visual perception incentive means that during the process of the robotic arm control model controlling the robotic arm, first, the sparse final reward value is decomposed into continuous intermediate reward signals to obtain the first reward value, and then during the process of the robotic arm control model controlling the robotic arm, the second reward value is calculated through the method based on visual perception incentive. After that, the first reward value and the second reward value are combined to obtain the reshaped reward value. Further, regarding the method based on visual perception incentive, for example: the internal and external incentive control method based on visual perception, etc.
[0039] The reshaped reward value is a dynamic reward calculation mechanism designed to solve the problem of sparse rewards in traditional reinforcement learning. Through the reward value reshaping process, the sparse final reward value is decomposed into continuous intermediate reward signals. Specifically, the calculation of the reshaped reward value is based on the real-time distance change between the end effector of the robotic arm and the target position. For example, if an action information consists of 5 actions, the traditional method is to generate a reward value after the five actions are completed. The reshaped reward value is based on a preset number of actions (e.g., 1 action). At this time, the robotic arm generates a first reward value for each action it performs. After that, the first reward value is combined with the second reward value obtained through the method based on visual perception incentive to obtain the reshaped reward value, and the problem of the reward value coefficient can be solved through continuous intermediate reward signals.
[0040] Step 102: Generate sample data according to the state information, action information, reshaped reward value, and status flag, and store the sample data in the sample pool.
[0041] In the embodiments of the present application, the sample data is a set composed of multi-dimensional information generated during the interaction between the robotic arm and the environment. Each sample data includes, but is not limited to: the current state information in the state information, that is, the first state information (s), the executed action information ( ), the obtained reshaping reward value (r), the next state information, i.e., the second state information (s'), and the state flag (done). Specifically, after each action of the robotic arm is executed, the robotic arm control model records the state information at the current moment, such as the joint angles and the end - effector position, i.e., the second state information. Combining with the action instruction output by the robotic arm control model (e.g., the joint angle adjustment amount), it collects the reward value feedback by the environment after executing this action (including: the superposition result of the reshaping reward value and the external reward), and detects whether the task terminates (the state flag). Various information is combined into a complete sample data {s, , r, s', done}, which represents the transition process of the robotic arm from the current state to the next state.
[0042] The sample pool is a database for storing sample data, and its essence is an experience replay set (which can be denoted as D). During the deep reinforcement learning training process, the robotic arm continuously interacts with the environment to generate new sample data and continuously stores it in the sample pool. The sample pool adopts a first - in - first - out or dynamic capacity management strategy to ensure that the stored sample data covers the interaction experiences of different training stages. For example: when the sample pool reaches the preset capacity, the new sample data will replace the oldest sample data to maintain the timeliness and diversity of the data. Through the accumulation of the sample pool, the robotic arm control model can learn from a large number of historical experiences rather than relying only on the current interaction data, thereby improving the stability of training and the generalization ability of the policy.
[0043] The generation and storage process of the sample data includes, but is not limited to, the following steps: In each optimization of the robotic arm control model, first, according to the first state information (s), the Actor network generates action information ( ), and after superimposing exploration noise (e.g., Gaussian noise), it is sent to the robotic arm for execution. After the execution is completed, the environment returns the second state information (s'), the calculated reshaping reward value (r), and the state flag (done). Subsequently, the robotic arm control model stores the five - tuple {s, , r, s', done} as a complete sample data into the sample pool. This process is cycled in the continuous action selection and state transition of the robotic arm until the task terminates or reaches the preset training steps.
[0044] Step 103, perform importance setting processing on the sample data in the sample pool based on the reshaping reward value to obtain the importance corresponding to the sample data, and perform data sampling processing on the sample data in the sample pool based on the importance of the sample data to obtain the sampled data.
[0045] In the embodiments of the present application, the importance setting process is a process of assigning dynamic weights (i.e., importance) to each sample data in the sample pool, aiming to distinguish the value contributions of different sample data to the current training stage. The calculation of importance is based on the deviation degree between the reshaped reward value (r) of the sample and the average value of the reshaped reward values in the sample pool ( ) and combines with the dynamic adjustment strategy of the training stage.
[0046] The data sampling process is a process of screening training data from the sample pool based on the importance of the sample data. When performing the data sampling process, it is first necessary to convert the importance into a sampling probability so that samples with high importance are more likely to be selected for model update.
[0047] The sampled data is a subset selected from the sample pool based on the importance of the sample data, which reflects the emphasis on high-value experiences in the current training stage. Compared with uniform sampling, sampling based on the importance of the sample data enables the robotic arm control model to more efficiently utilize successful experiences (such as: action sequences approaching the target) and key failure cases (such as: collisions or deviations from the trajectory), thereby accelerating policy iteration.
[0048] Step 104, perform model parameter optimization processing on the robotic arm control model according to the sampled data to obtain an optimized robotic arm control model.
[0049] In the embodiments of the present application, the model parameter optimization process is a step for the robotic arm control model to learn from the sampled data. By iteratively adjusting the weight parameters of the internal neural network of the robotic arm control model, the action strategy output by the robotic arm control model maximizes the long-term cumulative reward. The optimization process includes but is not limited to: two parallel stages, namely, the Q-value function approximation of the Critic network and the policy gradient enhancement of the Actor network. Specifically, the sampled data {s, , r, s′, done} sampled from the sample pool is input into the Critic and Actor networks, and the errors between the predicted Q-values (i.e., the second expected values) and the actual target Q-values (i.e., the first expected values) are calculated respectively, as well as the policy gradient direction. Then, the network parameters are updated through the backpropagation algorithm.
[0050] The optimized robotic arm control model is a stable policy network obtained through multiple rounds of parameter iterative updates, which can generate accurate action instructions in real time according to the environmental state. For example, in a grasping task, the joint angle adjustment amount output by the model can make the end effector approach the target along the optimal trajectory while avoiding joint over-limit or collision. Compared with the robotic arm control model, the optimized robotic arm control model not only has higher action decision-making efficiency but also can adapt to complex scenarios such as dynamic changes in the target position and environmental noise interference.
[0051] Step 105: Control the robotic arm according to the optimized robotic arm control model.
[0052] In an embodiment of the present application, when the optimized robotic arm control model controls the robotic arm, the model optimization process will still be executed. Specifically, it includes but is not limited to: obtaining new state information, new action information, new state identifiers, and new reshaped reward values of the robotic arm generated during the robotic arm control process based on the optimized robotic arm control model;
[0053] If the new state identifier is the termination state, end the process of optimizing the model parameters of the optimized robotic arm control model; if the state identifier is a non-termination state, continue to optimize the model parameters of the optimized robotic arm control model. That is, generate new sample data according to the new state information, new action information, new reshaped reward value, and new state identifier, and store the new sample data in the sample pool;
[0054] Perform importance setting processing on the new sample data in the sample pool based on the new reshaped reward value to obtain the importance corresponding to the new sample data, perform data sampling processing on the new sample data in the sample pool based on the importance of the new sample data to obtain new sampled data; optimize the model parameters of the optimized robotic arm control model according to the new sampled data to obtain a secondarily optimized robotic arm control model; control the robotic arm according to the secondarily optimized robotic arm control model.
[0055] Repeat the above steps, continuously optimize the robotic arm control model during the robotic arm control process, and improve the adaptability and intelligence level of the robotic arm in an unstructured environment.
[0056] A robotic arm control method, device, electronic device, and storage medium of the present application, by combining the method of reshaping the reward value based on visual perception excitation during the process of the robotic arm control model controlling the robotic arm to determine the reshaped reward value, can reduce reward sparsity, and generate sample data from various information generated by the robotic arm control model during the robotic arm control process. Then, based on the reshaped reward value, importance setting is performed on the sample data, and sampling is performed according to the importance of the sample data. Compared with uniform sampling in the prior art, the sampling efficiency is improved, which can solve the technical problems of how to improve the sampling efficiency of the algorithm during the robotic arm control process and solve the reward sparsity of the algorithm, and achieve the technical effects of improving the sampling efficiency of the algorithm and solving the reward sparsity of the algorithm.
[0057] In an achievable embodiment of the present application, regarding the acquisition of the state information, action information, state identifier, and reshaped reward value of the robotic arm, the present application provides a flow schematic diagram for information acquisition, as Figure 2 shown, including:
[0058] Step 201, determine the target position information of the target position corresponding to the robotic arm and the first state information of the robotic arm, where the target position is the position expected to be reached after controlling the robotic arm, and the first state information includes at least the current state of the robotic arm before performing an action.
[0059] In the embodiments of the present application, the target position information is the final position that the robotic arm needs to reach during the control process, usually represented by three-dimensional coordinates (e.g., ), or polar coordinates (e.g., horizontal radius , angle and depth ). For example: In a grasping task, the target position information may include, but is not limited to: the spatial coordinates of the center point of the object to be operated. The target position information is input into the robotic arm control model as the ultimate goal of the control task, and is used to guide policy generation and deviation evaluation.
[0060] The first state information is a state description of the robotic arm before performing an action, including but not limited to: the real-time angles of the joints of the robotic arm ( ), the current position of the end effector of the robotic arm, the joint speed, and the initial relative distance from the target position. Specifically, if the first state information is represented in the form of a feature vector, for example:
[0061]
[0062]
[0063] where N is the number of degrees of freedom of the robotic arm. The first state information provides the initial environmental perception basis for the action decision of the control model, is the feature vector form of the first state information, that is, the state vector.
[0064] Step 202, according to the target position information, generate the action information of the robotic arm through the robotic arm control model, where the action information includes at least the actions performed by the robotic arm.
[0065] In the embodiments of the present application, the action information is a set of instructions generated by the robotic arm control model, representing the motion parameters of the joints of the robotic arm at the next moment (e.g., angle increment or angular velocity). For example: For a six-degree-of-freedom robotic arm, the action information can be expressed as , where each is normalized to adapt to the physical motion range of the robotic arm. The action information is calculated by the Actor network in the robotic arm control model based on the first state information and the target position, and its essence is a mapping strategy from the state space to the action space.
[0066] Step 203: The robotic arm control model controls the robotic arm to execute an action according to the action information, and obtains the second state information of the robotic arm. The second state information at least includes the current state of the robotic arm after executing the action. The state information includes the first state information and the second state information.
[0067] In an embodiment of the present application, the second state information is the state of the robotic arm after executing the action corresponding to the action information, and includes the joint angles after the action is executed ( ), the new position of the end effector, and the real-time relative distance from the target position. The second state information is collected in real time by robotic arm sensors (such as encoders, vision systems) and fed back into the robotic arm control model for evaluating the action effect and updating subsequent strategies. The continuity of the state information (the transition from the first state to the second state) reflects the dynamic process of the interaction between the robotic arm and the environment.
[0068] Step 204: Determine the position deviation between the robotic arm and the target position according to the second state information and the target position information, and determine the state identifier according to the position deviation, where the state identifier is the identifier of the second state information.
[0069] In an embodiment of the present application, the position deviation is the distance (such as the Euclidean distance) between the current position of the end effector of the robotic arm, that is, the current position of the robotic arm in the second state information ( ), and the target position ( ), and can be represented by, but not limited to, the following method: . The position deviation is used to quantify the proximity of the state of the robotic arm after the action execution to the task target, and is the basis for the calculation of the reward value and the determination of the state identifier. For example: when the position deviation is less than the preset threshold , it is determined that the robotic arm has successfully reached the target position, and the state identifier is the termination state.
[0070] The state identifier is a task progress marker based on the position deviation, and usually uses a boolean variable (such as: done) to indicate whether the task is terminated. The introduction of the state identifier enables the control model to distinguish the training stage and reset the environment at the end of the task to start a new control cycle.
[0071] Step 205: During the process of the robotic arm control model controlling the robotic arm to execute an action according to the action information, a reward value is generated every time a preset number of actions are executed until all the actions in the action information are executed, obtaining the first reward value, and combining the first reward value with the second reward value calculated by the method of visual perception incentive to obtain the reshaped reward value.
[0072] In an embodiment of the present application, the reshaped reward value is an incentive signal dynamically generated by the robotic arm control model during the execution of actions. The reshaped reward value is related to the first reward value obtained based on the position deviation and the second reward value obtained based on visual perception incentives. The calculation of the first reward value is closely related to the change in the position deviation. After every preset number of actions (e.g., every 5 steps, every 1 step) are executed, the robotic arm control model generates an intermediate reward based on the difference between the latest position deviation and the historical deviation. For example, if the robotic arm gradually approaches the target during consecutive actions, a positive reward is obtained; if it deviates from the target or stalls, a negative penalty is obtained.
[0073] Specifically, the calculation of the first reward value can be expressed by, but not limited to, formula (1):
[0074] Formula (1)
[0075] Wherein, is the first reward value, is the preset number of times, is the position of the robotic arm at the moment during the execution of the action by the robotic arm, is during the execution of the action by the robotic arm, the position of the robotic arm at the previous actions during the action, that is, the position is obtained through the position after actions, is the position deviation between the position and the target position, is the position deviation between the
[0076] Furthermore, the reshaped reward value , wherein, is the reshaped reward value, is the second reward value.
[0077] The full - process closed - loop optimization of the robotic arm control task is achieved through information acquisition. First, the fusion input of the target position information and multi - dimensional state information enables the robotic arm control model to accurately perceive the relative spatial relationship between the robotic arm and the target position, providing a reliable decision - making basis for motion generation. Second, the dynamic calculation of the position deviation and the real - time determination of the state identifier enhance the system's autonomous evaluation ability of the task progress, ensuring a rapid response to abnormal situations (such as target deviation or hardware failures) in complex environments. Third, the periodic reshaping reward value generation mechanism decomposes the sparse final target reward into progressive intermediate incentives, significantly improving the learning efficiency and policy stability of the model in long - sequence actions. In addition, through the continuous update of the first state and the second state, the robotic arm control model can adapt to the changes in robotic arm dynamics (such as load changes and joint wear) online, maintaining high - precision control performance.
[0078] In an implementable embodiment of the present application, for determining the position deviation and determining the state identifier according to the position deviation, the following methods can be used but are not limited to: determining the current position coordinates of the robotic arm according to the joint angle information of the robotic arm included in the second state information; performing coordinate comparison processing on the current position coordinates and the target position coordinates of the target position included in the target position information to obtain the position deviation between the robotic arm and the target position; when the position deviation is less than or equal to the preset deviation threshold, determining the state identifier of the second state information as the termination state; when the position deviation is greater than the preset deviation threshold, determining the state identifier of the second state information as the non - termination state.
[0079] In the embodiment of the present application, the joint angle information is the real - time angle data of each joint of the robotic arm after the action is executed, usually collected by a joint encoder or a potentiometer. The current position coordinates are the real - time positions of the end - effector of the robotic arm in three - dimensional space, obtained through forward kinematics calculation. For example: the current position of the robotic arm in the second state information ( ), and the target position coordinates are the three - dimensional space positions that the end - effector is expected to reach in the robotic arm control task, that is, the target position ( ).
[0080] The coordinate comparison processing is a process of calculating the spatial distance between the current position coordinates and the target position coordinates. The position deviation can be represented by .
[0081] The preset deviation threshold ( ) is the maximum error radius allowed between the position of the end - effector of the robotic arm and the target position, which is a user - defined threshold, usually set according to the task precision requirements. For example: in a precision assembly task, can be set to the millimeter level (such as: 1mm), while in a material handling task, It can be relaxed to the centimeter level (e.g., 5 cm). The selection of the threshold needs to balance control accuracy and execution efficiency. An overly small threshold may lead to excessive adjustment, while an overly large threshold may reduce the quality of task completion.
[0082] The termination state is an identifier for the second state information when the position deviation ≤ indicating that the robotic arm has successfully reached the target position and the current control task is completed. At this time, the robotic arm control model stops generating subsequent action instructions and triggers a task end signal (e.g., reset the robotic arm or start the next task process). The accurate determination of the termination state avoids meaningless extra actions of the robotic arm after reaching the target and improves control efficiency.
[0083] The non - termination state is an identifier for the second state information when the position deviation > indicating that the robotic arm has not met the task completion conditions and needs to continue generating adjustment actions. In this state, the robotic arm control model continuously optimizes the action strategy according to the latest state information to drive the end - effector of the robotic arm to further approach the target position. The dynamic maintenance of the non - termination state ensures the continuous adaptability of the robotic arm in complex environments, such as coping with dynamic offsets of the target position or external disturbances.
[0084] Through joint angle analysis and dynamic position deviation evaluation, an efficient closed - loop management of the robotic arm control task is achieved. Based on real - time coordinate calculation of forward kinematics, it can accurately reflect the spatial position of the end - effector, providing a reliable data basis for deviation analysis. The intelligent comparison between the position deviation and the preset deviation threshold makes the task termination determination both strict and flexible, avoiding positioning deviations caused by premature termination and preventing resource waste caused by excessive iteration.
[0085] In an implementable embodiment of the present application, for the generation of sample data, it can also be implemented by, but not limited to, the following method: perform five - tuple data generation processing according to the first state information, the second state information, action information, reshaped reward value, and state identifier in the state information to obtain sample data.
[0086] In the embodiment of the present application, the five - tuple data generation processing is a step of encoding multi - source information generated during the interaction between the robotic arm and the environment into sample data. The five - tuple consists of the first state information (s), action information ( ), reshaped reward value (r), second state information (s′), and state identifier (done).
[0087] Through the structured generation and storage of five-tuple data, the training efficiency and policy generalization ability of the robotic arm control model are significantly improved, enabling the robotic arm control model to learn the long-term value of actions from historical experience and avoid over-reliance on real-time interaction data. The standardized format of the five-tuple data facilitates cooperation with mechanisms such as priority sampling and experience replay.
[0088] In an implementable embodiment of the present application, for the importance setting process, the following methods can be adopted but are not limited to: setting an initial importance for the sample data in the sample pool, and averaging the reshaped reward values of all sample data in the sample pool to obtain an average reward value; in the sample pool, for the sample data whose reshaped reward value is greater than or equal to the average reward value, parameter calculation processing is performed through a first preset algorithm to obtain a first importance parameter, and for the sample data whose reshaped reward value is less than the average reward value, parameter calculation processing is performed through a second preset algorithm to obtain a second importance parameter; obtaining the number of optimization times for the robotic arm control model to perform model parameter optimization processing, and calculating the importance corresponding to the sample data based on the number of optimization times, the first importance parameter, and the second importance parameter.
[0089] In the embodiment of the present application, regarding the setting of the importance of the sample data in the sample pool, the present application provides a schematic diagram of the principle of a method for setting importance, as Figure 3 shown, where the agent is the robotic arm control model, the experience replay pool is the sample pool, the initialization of sample importance is to set the initial importance, the update of sample importance, and the acquisition of sample data are the processes of calculating the importance corresponding to the sample data based on the number of optimization times, the first importance parameter, and the second importance parameter.
[0090] The first preset algorithm and the second preset algorithm are algorithms set by custom in advance. Regarding the calculation of the first importance parameter and the second importance parameter, it can be determined by but not limited to formula (2):
[0091] Formula (2)
[0092] Among them, is the first importance parameter, is the second importance parameter, is the reshaped reward value greater than or equal to the average reward value, is the reshaped reward value less than the average reward value, is the average reward value.
[0093] The number of optimizations is the number of iterations in which the robotic arm control model has completed parameter updates, and is used to dynamically adjust the weight coefficient in importance calculation. The process and principle of calculating the importance corresponding to sample data based on the number of optimizations include, but are not limited to, the following methods: In the early stage of training, when the number of optimizations is small, in order to prevent successful samples from being submerged, the importance of samples with a reward value greater than the average reward value is increased. In the middle and late stages of training, when the number of optimizations is large, in order to prevent falling into a local optimal strategy and increase sample diversity, the importance of samples with a reward value less than the average reward value is increased.
[0094] Through the dynamic importance calculation mechanism, the learning efficiency and policy robustness of the robotic arm control model are significantly improved. Based on the hierarchical evaluation of the average reward value, it can intelligently distinguish high-value samples from potential value samples, providing a quantitative basis for priority sampling. The adjustment of the weight coefficient driven by the number of optimizations enables the model to adaptively balance exploration and exploitation at different training stages, quickly converging to a feasible strategy in the initial stage and finely optimizing the strategy details in the later stage. The uniform setting of the initial importance provides a fair exploration basis for the initial stage of training, avoiding early policy solidification caused by priority bias.
[0095] In an implementable embodiment of the present application, for calculating the importance corresponding to sample data based on the number of optimizations, the first importance parameter, and the second importance parameter, it can be implemented in, but not limited to, the following ways: When the number of optimizations is less than the preset number threshold, calculate the importance of sample data with a reshaped reward value greater than or equal to the average reward value according to the first preset parameter and the first importance parameter; calculate the importance of sample data with a reshaped reward value less than the average reward value according to the second preset parameter and the second importance parameter; when the number of optimizations is greater than or equal to the preset number threshold, calculate the importance of sample data with a reshaped reward value greater than or equal to the average reward value according to the third preset parameter and the first importance parameter; calculate the importance of sample data with a reshaped reward value less than the average reward value according to the fourth preset parameter and the second importance parameter, where the fourth preset parameter is greater than the second preset parameter.
[0096] In an embodiment of the present application, the preset number threshold is a threshold set by the user and is a parameter for dividing the training stage, usually set according to the task complexity and experience. For example: in a scenario where the total number of training times is 5000 times, the preset number threshold can be set to 2000 times, indicating that the first 2000 times are the early stage of training, and the subsequent is the middle and late stage of training. The setting of the threshold affects the calculation strategy of sample importance and needs to be verified through experiments to adapt to the specific task requirements.
[0097] The first preset parameter ( ), and the second preset parameter ( ) For calculating the importance of samples. At this stage, the model focuses on quickly accumulating high-value experiences, so higher weights are assigned to samples with high rewards. For example: set = 0.8, = 0.2,
[0098] Specifically, for calculating the importance of sample data with reshaped reward values greater than or equal to the average reward value, formula (3) can be used but is not limited to:
[0099] Formula (3)
[0100] For calculating the importance of sample data with reshaped reward values less than the average reward value, formula (4) can be used but is not limited to:
[0101] Formula (4)
[0102] Wherein, is the importance.
[0103] Furthermore, regarding the application of the number of optimization times, for example: in the first 2000 rounds of training (i.e., the number of optimization times is less than 2000, and 2000 is the preset number threshold), and are 0.8 and 0.2 respectively, the greater the reshaped reward value the greater the importance the greater, the more important the sample data with reshaped reward values greater than In the last 2000 rounds (i.e., the number of optimization times is greater than or equal to 2000), increase the value to 0.4, and improve the importance of sample data with reshaped reward values less than By adjusting the importance parameters in stages, the training efficiency and policy robustness of the robotic arm control model are improved simultaneously.
[0104] In an implementable embodiment of the present application, for data sampling processing of sample data, the following methods can be used but are not limited to: calculating the sampling probability corresponding to each sample data in the sample pool according to the importance; arranging the sample data in the sample pool for data sampling processing in descending order of the sampling probability to obtain the sampled data.
[0105] In the embodiment of the present application, for calculating the sampling probability, formula (5) can be used but is not limited to:
[0106]
[0107] Formula (5)
[0108] Wherein, is the sampling probability, is the number of sampled data.
[0109] The training efficiency and strategy generalization ability of the robotic arm control model are significantly improved through the dynamic probability allocation and intelligent sampling mechanism.
[0110] In an implementable embodiment of the present application, after obtaining the state information, action information, and state identifier, the following method can also be adopted: determine whether the state identifier is a termination state; if the state identifier is a termination state, end the process of optimizing the model parameters of the robotic arm control model; if the state identifier is a non-termination state, continue to optimize the model parameters of the robotic arm control model.
[0111] In an embodiment of the present application, in each round of training iteration, after the robotic arm control model executes an action and obtains environmental feedback, first parse the state identifier. If the state identifier is a termination state (done = 1), immediately terminate the current training cycle, save the optimized model parameters, and reset the robotic arm and the environment to the initial state to start a new round of training; if the state identifier is a non-termination state (done = 0), continue to sample data from the sample pool, perform parameter updates of the Critic and Actor networks, and enter the next control step.
[0112] The training efficiency and resource utilization rate of the robotic arm control model are significantly improved through the intelligent determination of the termination state and the dynamic control of the training process.
[0113] In an implementable embodiment of the present application, for optimizing the model parameters of the robotic arm control model, it can be implemented by, but not limited to, the following methods: determine the first evaluation parameter corresponding to the first evaluation network, the first target parameter corresponding to the first target network, the second evaluation parameter corresponding to the second evaluation network, and the second target parameter corresponding to the second target network in the robotic arm control model; the first target network performs action generation processing based on the second state information in the sampled data to obtain predicted action information, and the second evaluation network performs parameter generation processing based on the first state information and action information in the sampled data to obtain the first expected value; the second target network performs parameter generation processing based on the second state information and the predicted action information to obtain the second expected value, and performs parameter generation processing based on the second expected value and the reshaped reward value in the sampled data to obtain the target value; construct a first loss function based on the target value and the first expected value, construct a second loss function based on the first expected value, and optimize the model parameters of the robotic arm control model according to the first loss function and the second loss function to obtain the optimized robotic arm control model.
[0114] In an embodiment of the present application, the first evaluation network and the first target network are two instances of the Actor network, respectively assuming the roles of policy generation and target value calculation. Among them, the first evaluation network (Actor evaluation network) generates a real-time action policy based on the current state, and its parameters (the first evaluation parameter is denoted as θ) are updated through policy gradients; the first target network (Actor target network) is a delayed mirror of the evaluation network, and its parameters (the first target parameter, denoted as θ′) are synchronized with θ through a soft update mechanism, and are used to generate stable target actions to reduce value estimation bias.
[0115] The second evaluation network and the second target network are two instances of the Critic network, respectively responsible for action value evaluation and target value calculation. The second evaluation network (Critic evaluation network) predicts the Q value based on the current state and action, and its parameters (the second evaluation parameter, denoted as ) are updated through mean squared error loss; the second target network (Critic target network) is a delayed mirror of the evaluation network, and the parameters (the second target parameter, denoted as ) are synchronized with through soft update, and are used to calculate the target Q value to stabilize the training process.
[0116] Action generation processing is performed by the first target network. In the parameter optimization stage, the first target network receives the second state information (s′) in the sampled data and generates predicted action information ( ) through forward propagation. This predicted action is used to evaluate the expected action value of the next state and avoid the estimation fluctuations caused by directly using the evaluation network parameters.
[0117] The first expected value (Q) is calculated by the second evaluation network, and its inputs are the first state information (s) and action information ( ) in the sampled data, and the output is the predicted Q value (the first expected value) of the current state-action pair, that is: Q( ), and the first expected value represents the expected long-term cumulative reward for executing the action in the state , and is the core evaluation index of the Critic network.
[0118] The second expected value ( ) is calculated by the second target network, and its inputs are the second state information (s′) and the predicted action ( ) generated by the first target network, and the output is the predicted Q value (the second expected value) of the next state-action pair, that is: , and the second expected value represents the expected long-term cumulative reward for executing the action in the state , and is used to construct the target Q value.
[0119] The target value ( is the benchmark value for optimizing the Critic network, which is calculated by fusing the reshaped reward value (r) with the discounted second expected value. Regarding the calculation of the target value, it can be carried out through, but not limited to, formula (6):
[0120] Formula (6)
[0121] where is the discount factor (0 ≤ ≤ 1), which is used to balance the immediate reward and the long-term benefit. If the state flag is the termination state (done = 1), then the target value only retains the current reward ( = ). represents the reshaped reward value of the th data in the sampled data, represents the second state information of the th data in the sampled data, represents the predicted action information of the th data in the sampled data.
[0122] The first loss function (Critic loss) is the mean square error between the predicted Q value (the first expected value (Q)) and the target value ( ), which is used to optimize the parameters of the second evaluation network. Regarding the first loss function, it can be carried out through formula (7):
[0123] Formula (7)
[0124] where is the number of sampled data, represents the first state information of the th data in the sampled data, represents the action information of the th data in the sampled data. By minimizing this loss through the gradient descent algorithm, the Q value prediction of the Critic network gradually approaches the true action value distribution.
[0125] The second loss function (Actor loss) is the negative mean value of the Q value evaluated by the Critic network, which is used to optimize the parameters θ of the first evaluation network. Regarding the second loss function, it can be carried out through formula (8):
[0126] Formula (8)
[0127] where is the second loss function. By maximizing the Q value, the policy generated by the Actor network is guided towards the high-value action direction, thereby improving the task completion efficiency of the robotic arm.
[0128] Through the dual - network architecture and collaborative optimization mechanism, the policy stability and convergence efficiency of the robotic arm control model are significantly improved. The introduction of the target network effectively alleviates the problem of value estimation bias. For example, the delayed synchronization mechanism suppresses the drastic fluctuations of Q - values, making the training process smoother. The mean - square error optimization of the Critic network and the policy gradient improvement of the Actor network form a closed - loop feedback, enabling the action policy to gradually approach the optimal solution under the guidance of value evaluation. The soft - update mechanism avoids the risk of policy mutation caused by traditional hard updates through progressive parameter synchronization, ensuring the action coherence and safety of the robotic arm in the real environment. In addition, the collaborative effect of the dual - loss functions enables the model to optimize both the action value estimation and policy generation capabilities simultaneously.
[0129] In an implementable embodiment of the present application, for optimizing the model parameters of the robotic arm control model according to the first loss function and the second loss function, the following implementation methods can be adopted but are not limited to: optimizing the second evaluation parameter according to the first loss function to obtain the optimized second evaluation parameter, and optimizing the first evaluation parameter according to the second loss function to obtain the optimized first evaluation parameter; performing function transfer processing on the first target parameter according to the optimized first evaluation parameter to obtain the optimized first target parameter, and performing function transfer processing on the second target parameter according to the optimized second evaluation parameter to obtain the optimized second target parameter; based on the optimized first evaluation parameter, the optimized first target parameter, the optimized second evaluation parameter, and the optimized second target parameter, obtaining the optimized robotic arm control model.
[0130] In the embodiment of the present application, according to the first loss function, the error gradient is backpropagated through a gradient - descent algorithm (such as the Adam optimizer) to adjust the value, so that the Q - value prediction of the Critic network gradually approaches the true action value distribution. The optimized second evaluation parameter ( 1) can more accurately evaluate the long - term rewards of different actions. For example, in the robotic arm grasping task, it can accurately distinguish the cumulative reward differences between efficient paths and inefficient actions.
[0131] According to the second loss function, the policy gradient is backpropagated through a gradient - ascent algorithm to adjust the value of θ, so that the action policy generated by the Actor network can obtain a higher Q - value evaluation. The optimized first evaluation parameter ( ) can generate better joint control instructions. For example, in the robotic arm obstacle - avoidance task, it can output a smooth and safe trajectory adjustment amount.
[0132] The function transfer processing is a soft - update mechanism for the target network parameters, which gradually synchronizes the optimization results of the evaluation network to the target network through weighted averaging. Specifically, it includes but is not limited to:
[0133] Update of the first target parameter ( ): The optimized first evaluation parameter 1 and the current first target parameter are fused according to the weight coefficient (e.g., 0.001): , is the optimized first evaluation parameter, which makes the parameters of the Actor target network slowly approach the evaluation network and maintains the stability of policy generation.
[0134] Update of the second target parameter ( ): The optimized second evaluation parameter 1 and the current second target parameter are fused according to the same coefficient : , is the optimized second evaluation parameter, which makes the Q-value prediction of the Critic target network change with a delay and suppresses the fluctuation of value estimation.
[0135] Through the step-by-step parameter optimization and soft synchronization mechanism, the training stability and policy performance of the robotic arm control model are significantly improved.
[0136] In an implementable embodiment of the present application, when combining the first reward value with the second reward value obtained by calculating the reward value in the manner of visual perception incentive, the following methods can also be used but are not limited to: performing information dimensionality reduction processing on the image information obtained by the robotic arm control model during the robotic arm control process to obtain a first image vector, and converting the first state information into a state vector and performing data fusion processing with the first image vector to obtain a first target vector; inputting the first target vector into a preset prediction model for prediction processing to obtain a second target vector, and calculating a prediction error according to the second target vector and the first true vector converted from the second state information to obtain a first prediction error; inputting the second target vector into a preset prediction model for prediction processing to obtain a third target vector, and calculating a prediction error according to the second true vector converted from the third state information to obtain a second prediction error, where the third state information is the next state information of the second state information; until a preset number of prediction errors are obtained, performing an average value calculation process on the preset number of prediction errors to obtain an average prediction error, and combining the average prediction error with the external reward value of the robotic arm control model during the robotic arm control process to obtain a second reward value; combining the second reward value with the first reward value to obtain a reshaped reward value.
[0137] In the embodiments of the present application, in the actual application of the robotic arm, different from the simulation environment, simply inputting joint information into the algorithm cannot guarantee the effect. In the actual scenario, the state information input of the robotic arm is mostly based on visual images, which is also the main medium for the interaction between the robotic arm and the environment. However, using the original high-dimensional perception information as input and outputting the corresponding control information will have problems such as information redundancy and increased training complexity for the deep reinforcement learning control algorithm.
[0138] When applying deep reinforcement learning in the real robotic arm operation environment, in addition to the need for an efficient representation of the environmental state, there is also the problem of the reward mechanism. In the real environment, intermediate rewards cannot be manually set as in the simulation environment, or the reward setting is too complex, resulting in too high a degree of manual intervention, thereby reducing the intelligence level of the system. Therefore, when directly applying the deep reinforcement learning method to the real robotic arm control environment, problems such as the lack of external rewards or sparse rewards are often encountered.
[0139] To facilitate understanding of the implementation process of the embodiments of the present application, the embodiments of the present application provide a schematic diagram of an incentive strategy for visual perception, as Figure 4 shown, where the state prediction network is a preset prediction model, built into the intelligent agent (robotic arm control model). The preset prediction model is a custom convolutional neural network model, and the control module is the module in the intelligent agent (robotic arm control model) for controlling the robotic arm.
[0140] The original image information will be compressed into a low-dimensional and efficient space using a convolutional neural network, represented as the first image vector , and at the same time, the state vector corresponding to the fused state information to better capture the state information of the object and the robotic arm during the motion control process of the robotic arm.
[0141] A neural network structure is used as the state prediction model to and the action vector corresponding to the executed action information as input to predict the next state (the second target vector ), the real next state (the first real vector ) and calculate the first prediction error .
[0142] For the most recent consecutive prediction models, the average prediction error needs to be calculated for each model to characterize the quality of the predicted action. The average prediction error can be expressed by formula (9):
[0143] Formula (9)
[0144] where is the average prediction error, is the preset quantity, is the th prediction error among the prediction errors of the preset quantity.
[0145] The obtained reward can be used as an incentive signal for the agent (robotic arm control model), prompting the agent to select actions that can maximize the learning process, thereby learning the operation task more efficiently. To introduce this self-motivating reward into the robotic arm control model, it can be combined with an external reward and expressed by formula (10):
[0146] Formula (10)
[0147] where, is the external reward value, is the second reward value.
[0148] Through the reward design driven by visual information fusion and prediction error, the performance of the robotic arm control model in complex scenarios is significantly improved.
[0149] In summary, the present application can achieve the following technical effects:
[0150] A robotic arm control method, device, electronic device, and storage medium of the present application, by combining a method based on visual perception incentive to perform reward value reshaping processing during the process of the robotic arm control model controlling the robotic arm, determining the reshaped reward value, can reduce reward sparsity, and generate sample data from various information generated during the process of the robotic arm control model controlling the robotic arm. Then, based on the reshaped reward value, importance settings are made for the sample data, and sampling is performed according to the importance of the sample data. Compared with uniform sampling in the prior art, the sampling efficiency is improved, and the technical problems of how to improve the sampling efficiency of the algorithm during the process of robotic arm control and how to solve the reward sparsity of the algorithm can be solved, achieving the technical effects of improving the sampling efficiency of the algorithm and solving the reward sparsity of the algorithm.
[0151] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0152] An embodiment of the present application also provides a robotic arm control device, Figure 5 is a schematic structural diagram of a robotic arm control device provided by the application, as Figure 5 shown, including:
[0153] An acquisition unit 51, configured to acquire the state information, action information, state identifier, and reshaped reward value of the robotic arm generated during the control process of the robotic arm based on a robotic arm control model, where the robotic arm control model is a model constructed based on a deep reinforcement learning algorithm, and the reshaped reward value is a reward value obtained by reshaping the reward value in combination with a visual perception incentive-based method;
[0154] A generation unit 52, configured to generate sample data according to the state information, action information, reshaped reward value, and state identifier, and store the sample data in a sample pool;
[0155] A setting unit 53, configured to perform importance setting processing on the sample data in the sample pool based on the reshaped reward value to obtain the importance corresponding to the sample data;
[0156] A sampling unit 54, configured to perform data sampling processing on the sample data in the sample pool based on the importance of the sample data to obtain sampled data;
[0157] An optimization unit 55, configured to perform model parameter optimization processing on the robotic arm control model according to the sampled data to obtain an optimized robotic arm control model;
[0158] A control unit 56, configured to control the robotic arm according to the optimized robotic arm control model.
[0159] In an embodiment of the present application, the acquisition unit 51 is further configured to:
[0160] Determine the target position information of the target position corresponding to the robotic arm and the first state information of the robotic arm, where the target position is the position expected to be reached after controlling the robotic arm, and the first state information includes at least the current state of the robotic arm before performing the action;
[0161] According to the target position information, generate the action information of the robotic arm through the robotic arm control model, where the action information includes at least the action performed by the robotic arm;
[0162] The robotic arm control model controls the robotic arm to perform the action according to the action information to obtain the second state information of the robotic arm, where the second state information includes at least the current state of the robotic arm after performing the action, and the state information includes the first state information and the second state information;
[0163] Determine the position deviation between the robotic arm and the target position according to the second state information and the target position information, and determine the state identifier according to the position deviation, where the state identifier is the identifier of the second state information;
[0164] During the process that the robotic arm control model controls the robotic arm to execute actions according to the action information, a reward value is generated every time a preset number of actions are executed until all the actions in the action information are completed, obtaining a first reward value, and combining and processing the first reward value with a second reward value calculated by a reward value calculation method based on visual perception incentives to obtain a reshaped reward value.
[0165] In an embodiment of the present application, the obtaining unit 51 is further configured to:
[0166] Determine the current position coordinates corresponding to the robotic arm according to the joint angle information of the robotic arm included in the second state information;
[0167] Perform coordinate comparison processing on the current position coordinates and the target position coordinates of the target position included in the target position information to obtain the position deviation between the robotic arm and the target position;
[0168] In the case that the position deviation is less than or equal to a preset deviation threshold, determine the status flag of the second state information as the termination state;
[0169] In the case that the position deviation is greater than the preset deviation threshold, determine the status flag of the second state information as the non-termination state.
[0170] In an embodiment of the present application, the generating unit 52 is further configured to perform five-tuple data generation processing according to the first state information, the second state information, the action information, the reshaped reward value, and the status flag in the state information to obtain sample data.
[0171] In an embodiment of the present application, the setting unit 53 is further configured to:
[0172] Set an initial importance degree for the sample data in the sample pool, and perform an averaging process on the reshaped reward values of all the sample data in the sample pool to obtain an average reward value;
[0173] In the sample pool, perform parameter calculation processing on the sample data with a reshaped reward value greater than or equal to the average reward value through a first preset algorithm to obtain a first importance degree parameter, and perform parameter calculation processing on the sample data with a reshaped reward value less than the average reward value through a second preset algorithm to obtain a second importance degree parameter;
[0174] Obtain the number of optimization times for the robotic arm control model to perform model parameter optimization processing, and calculate the importance degree corresponding to the sample data based on the number of optimization times, the first importance degree parameter, and the second importance degree parameter.
[0175] In an embodiment of the present application, the setting unit 53 is further configured to:
[0176] When the number of optimization times is less than the preset number threshold, calculate the importance of the sample data whose reshaping reward value is greater than or equal to the average reward value according to the first preset parameter and the first importance parameter;
[0177] Calculate the importance of the sample data whose reshaping reward value is less than the average reward value according to the second preset parameter and the second importance parameter;
[0178] When the number of optimization times is greater than or equal to the preset number threshold, calculate the importance of the sample data whose reshaping reward value is greater than or equal to the average reward value according to the third preset parameter and the first importance parameter;
[0179] Calculate the importance of the sample data whose reshaping reward value is less than the average reward value according to the fourth preset parameter and the second importance parameter, where the fourth preset parameter is greater than the second preset parameter.
[0180] In an embodiment of the present application, the sampling unit 54 is further configured to:
[0181] Calculate the sampling probability corresponding to each sample data in the sample pool according to the importance;
[0182] Perform data sampling processing on the sample data in the sample pool in the order of the sampling probability from large to small to obtain sampling data.
[0183] In an embodiment of the present application, the optimization unit 55 is further configured to:
[0184] Determine whether the status flag is the termination status;
[0185] If the status flag is the termination status, end the process of optimizing the model parameters of the robotic arm control model;
[0186] If the status flag is not the termination status, continue to optimize the model parameters of the robotic arm control model.
[0187] In an embodiment of the present application, the optimization unit 55 is further configured to:
[0188] Determine the first evaluation parameter corresponding to the first evaluation network, the first target parameter corresponding to the first target network, the second evaluation parameter corresponding to the second evaluation network, and the second target parameter corresponding to the second target network in the robotic arm control model;
[0189] The first target network performs action generation processing according to the second state information in the sampling data to obtain predicted action information, and the second evaluation network performs parameter generation processing according to the first state information and the action information in the sampling data to obtain the first expected value;
[0190] The second target network performs parameter generation processing according to the second state information and the predicted action information to obtain a second expected value, and performs parameter generation processing according to the second expected value and the reshaped reward value in the sampled data to obtain a target value;
[0191] Construct a first loss function according to the target value and the first expected value, construct a second loss function according to the first expected value, and perform model parameter optimization processing on the robotic arm control model according to the first loss function and the second loss function to obtain an optimized robotic arm control model.
[0192] In an embodiment of the present application, the optimization unit 55 is further configured to:
[0193] Perform parameter optimization processing on the second evaluation parameter according to the first loss function to obtain an optimized second evaluation parameter, and perform parameter optimization processing on the first evaluation parameter according to the second loss function to obtain an optimized first evaluation parameter;
[0194] Perform function transfer processing on the first target parameter according to the optimized first evaluation parameter to obtain an optimized first target parameter, and perform function transfer processing on the second target parameter according to the optimized second evaluation parameter to obtain an optimized second target parameter;
[0195] Based on the optimized first evaluation parameter, the optimized first target parameter, the optimized second evaluation parameter, and the optimized second target parameter, obtain an optimized robotic arm control model.
[0196] In an embodiment of the present application, the acquisition unit 51 is further configured to:
[0197] Perform information dimensionality reduction processing on the image information obtained by the robotic arm control model during the robotic arm control process to obtain a first image vector, and convert the first state information into a state vector and perform data fusion processing with the first image vector to obtain a first target vector;
[0198] Input the first target vector into a preset prediction model for prediction processing to obtain a second target vector, and calculate a prediction error according to the second target vector and the first true vector converted from the second state information to obtain a first prediction error;
[0199] Input the second target vector into a preset prediction model for prediction processing to obtain a third target vector, and calculate a prediction error according to the second true vector converted from the third state information to obtain a second prediction error, where the third state information is the next state information of the second state information;
[0200] Until a preset number of prediction errors are obtained, perform an average calculation process based on the preset number of prediction errors to obtain an average prediction error, and combine the average prediction error with the external reward value in the process of controlling the robotic arm by the robotic arm control model to obtain a second reward value;
[0201] Combine the second reward value with the first reward value to obtain a reshaped reward value.
[0202] For the description of the features in the corresponding embodiments of the robotic arm control device, reference can be made to the relevant descriptions in the corresponding embodiments of the robotic arm control method, which will not be elaborated here one by one.
[0203] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above embodiments of the robotic arm control method.
[0204] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above embodiments of the robotic arm control method when running.
[0205] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drive, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disk, magnetic disk or optical disc and other various media that can store computer programs.
[0206] An embodiment of the present application further provides a computer program product. The above computer program product includes a computer program, and the computer program realizes the steps in any of the above embodiments of the robotic arm control method when executed by a processor.
[0207] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and the computer program realizes the steps in any of the above embodiments of the robotic arm control method when executed by a processor.
[0208] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered as exceeding the scope of this application.
[0209] The above has introduced in detail a robotic arm control method and device, an electronic device, and a storage medium provided by this application. Specific examples are used herein to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A robot arm control method, characterized in that: include: Acquire state information, action information, state identification, and remodeling reward value of the robotic arm generated in the process of controlling the robotic arm based on the robotic arm control model, wherein the robotic arm control model is a model constructed based on a deep reinforcement learning algorithm, the remodeling reward value is a reward value obtained by remodeling the reward value in combination with a visual perception incentive method, the state information includes first state information and second state information, the first state information includes at least the current state of the robotic arm before performing an action, and the second state information includes at least the current state of the robotic arm after performing an action; Generate sample data according to the state information, the action information, the reshaping reward value and the state identifier, and store the sample data in a sample pool; Based on the remodeling reward value, the sample data in the sample pool is processed for importance setting to obtain the importance corresponding to the sample data, and based on the importance of the sample data, data sampling is processed on the sample data in the sample pool to obtain sample data; Performing model parameter optimization processing on the robot arm control model according to the sampling data to obtain an optimized robot arm control model; Controlling the robotic arm according to the optimized robotic arm control model; Acquiring the state identifier includes: determining a position deviation between the robotic arm and the target position according to the second state information and target position information of the target position corresponding to the robotic arm, and determining the state identifier according to the position deviation, wherein the state identifier is an identifier of the second state information, and the target position is a position expected to be reached after the robotic arm is controlled; Obtaining the reshaping reward value includes: in the process of the robotic arm control model controlling the robotic arm to perform an action according to the action information, a reward value is generated each time a preset number of actions are performed until all actions in the action information are executed, thereby obtaining a first reward value, and combining the first reward value with a second reward value obtained by calculating the reward value based on visual perception incentives to obtain the reshaping reward value.
2. The robot arm control method according to claim 1, characterized in that: The acquisition of the state information and action information of the robot arm generated during the control of the robot arm based on the robot arm control model includes: Determine the target position information of the target position corresponding to the robotic arm and the first state information of the robotic arm; According to the target position information, generating the action information of the robotic arm through the robotic arm control model, wherein the action information at least includes the action performed by the robotic arm; The robotic arm control model controls the robotic arm to perform an action according to the action information to obtain the second state information of the robotic arm.
3. The robot arm control method according to claim 2, characterized in that: Determining the position deviation between the robot arm and the target position according to the second state information and the target position information, and determining the state identifier according to the position deviation includes: Determine the current position coordinates corresponding to the robotic arm according to the joint angle information of the robotic arm included in the second state information; Performing coordinate comparison processing according to the current position coordinates and the target position coordinates of the target position included in the target position information to obtain a position deviation between the mechanical arm and the target position; When the position deviation is less than or equal to a preset deviation threshold, determining the state identifier of the second state information as a termination state; When the position deviation is greater than a preset deviation threshold, the state identifier of the second state information is determined as a non-terminal state.
4. The robot arm control method according to claim 2, characterized in that: Generating sample data according to the state information, the action information, the remodeling reward value and the state identifier includes: A five-tuple data generation process is performed according to the first state information and the second state information in the state information, the action information, the reshaping reward value, and the state identifier to obtain the sample data.
5. The robot arm control method according to claim 1, characterized in that: The importance setting process of the sample data in the sample pool based on the remodeling reward value to obtain the importance corresponding to the sample data includes: Setting an initial importance for the sample data in the sample pool, and performing average processing on the reshaping reward values of all the sample data in the sample pool to obtain an average reward value; In the sample pool, the sample data whose remodeling reward value is greater than or equal to the average reward value is subjected to parameter calculation processing by a first preset algorithm to obtain a first importance parameter, and the sample data whose remodeling reward value is less than the average reward value is subjected to parameter calculation processing by a second preset algorithm to obtain a second importance parameter; The number of optimizations of the robot arm control model for model parameter optimization processing is obtained, and the importance corresponding to the sample data is calculated based on the number of optimizations, the first importance parameter, and the second importance parameter.
6. The robot arm control method according to claim 5, characterized in that: The calculating the importance corresponding to the sample data based on the number of optimizations, the first importance parameter, and the second importance parameter comprises: When the number of optimizations is less than a preset number threshold, calculating the importance of the sample data whose remodeling reward value is greater than or equal to the average reward value according to the first preset parameter and the first importance parameter; Calculating the importance of the sample data whose remodeling reward value is less than the average reward value according to the second preset parameter and the second importance parameter; When the number of optimizations is greater than or equal to the preset number threshold, calculating the importance of the sample data whose remodeling reward value is greater than or equal to the average reward value according to a third preset parameter and the first importance parameter; The importance of the sample data whose reshaping reward value is less than the average reward value is calculated according to a fourth preset parameter and the second importance parameter, wherein the fourth preset parameter is greater than the second preset parameter.
7. The robot arm control method according to claim 1, characterized in that: The data sampling process is performed on the sample data in the sample pool based on the importance of the sample data to obtain the sample data, which includes: Calculate the sampling probability corresponding to each sample data in the sample pool according to the importance; According to the order of the sampling probabilities from large to small, data sampling processing is performed on the sample data in the sample pool to obtain the sample data.
8. The robot arm control method according to claim 1, characterized in that: After obtaining the state information, action information, state identification, and reshaping reward value of the robotic arm generated during the robotic arm control process based on the robotic arm control model, the method further includes: Determine whether the status indicator is a terminated state; If the state is marked as a termination state, the process of performing model parameter optimization processing on the robot arm control model is terminated; If the state is identified as a non-terminal state, the model parameter optimization process of the robot arm control model continues.
9. The robot arm control method according to claim 2, characterized in that: The performing model parameter optimization processing on the robot arm control model according to the sampling data to obtain the optimized robot arm control model comprises: Determine a first evaluation parameter corresponding to the first evaluation network, a first target parameter corresponding to the first target network, a second evaluation parameter corresponding to the second evaluation network, and a second target parameter corresponding to the second target network in the manipulator control model; The first target network performs action generation processing according to the second state information in the sampled data to obtain predicted action information, and the second evaluation network performs parameter generation processing according to the first state information and the action information in the sampled data to obtain a first expected value; The second target network performs parameter generation processing according to the second state information and the predicted action information to obtain a second expected value, and performs parameter generation processing according to the second expected value and the reshaping reward value in the sampled data to obtain a target value; A first loss function is constructed according to the target value and the first expected value, a second loss function is constructed according to the first expected value, and model parameters of the robotic arm control model are optimized according to the first loss function and the second loss function to obtain the optimized robotic arm control model.
10. The robot arm control method according to claim 9, characterized in that: The performing model parameter optimization processing on the robotic arm control model according to the first loss function and the second loss function to obtain the optimized robotic arm control model comprises: Performing parameter optimization processing on the second evaluation parameter according to the first loss function to obtain an optimized second evaluation parameter, and performing parameter optimization processing on the first evaluation parameter according to the second loss function to obtain an optimized first evaluation parameter; Performing function transfer processing on the first target parameter according to the optimized first evaluation parameter to obtain the optimized first target parameter, and performing function transfer processing on the second target parameter according to the optimized second evaluation parameter to obtain the optimized second target parameter; The optimized robotic arm control model is obtained based on the optimized first evaluation parameter, the optimized first target parameter, the optimized second evaluation parameter and the optimized second target parameter.
11. The robot arm control method according to claim 2, characterized in that: The combining of the first reward value and the second reward value obtained by calculating the reward value based on visual perception incentive to obtain the reshaping reward value includes: Performing information dimension reduction processing on image information acquired by the robot control model during the robot control process to obtain a first image vector, and converting the first state information into a state vector and performing data fusion processing with the first image vector to obtain a first target vector; Inputting the first target vector into a preset prediction model for prediction processing to obtain a second target vector, and performing prediction error calculation based on the second target vector and the first true vector converted from the second state information to obtain a first prediction error; The second target vector is input into the preset prediction model for prediction processing to obtain a third target vector, and a prediction error is calculated according to the second real vector converted from the third state information to obtain a second prediction error, wherein the third state information is the next state information of the second state information; until a preset number of prediction errors are obtained, average value calculation processing is performed according to the preset number of prediction errors to obtain an average prediction error, and the average prediction error is combined with an external reward value of the robot arm control model during the robot arm control process to obtain the second reward value; The second reward value is combined with the first reward value to obtain the reshaping reward value.
12. A robot arm control device, characterized in that: include: an acquisition unit, for acquiring state information, action information, state identification, and remodeling reward value of the robotic arm generated in a process of controlling the robotic arm based on a robotic arm control model, wherein the robotic arm control model is a model constructed based on a deep reinforcement learning algorithm, the remodeling reward value is a reward value obtained by remodeling the reward value in combination with a visual perception incentive-based method, the state information includes first state information and second state information, the first state information includes at least a current state of the robotic arm before performing an action, and the second state information includes at least a current state of the robotic arm after performing an action; A generating unit, configured to generate sample data according to the state information, the action information, the reshaping reward value and the state identifier, and store the sample data in a sample pool; A setting unit, configured to perform importance setting processing on the sample data in the sample pool based on the remodeling reward value to obtain the importance corresponding to the sample data; A sampling unit, configured to perform data sampling processing on the sample data in the sample pool based on the importance of the sample data to obtain sample data; An optimization unit, used for performing model parameter optimization processing on the robotic arm control model according to the sampling data to obtain an optimized robotic arm control model; A control unit, used for controlling the robot arm according to the optimized robot arm control model; The acquisition unit is further used to determine the position deviation between the robotic arm and the target position according to the second state information and the target position information of the target position corresponding to the robotic arm, and determine the state identifier according to the position deviation, wherein the state identifier is the identifier of the second state information, and the target position is the position expected to be reached after the robotic arm is controlled; The acquisition unit is also used to generate a reward value every time a preset number of actions are performed during the process in which the robotic arm control model controls the robotic arm to perform an action according to the action information, until all actions in the action information are performed, thereby obtaining a first reward value, and combining the first reward value with a second reward value obtained by calculating the reward value based on visual perception incentives to obtain the reshaping reward value.
13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the robot arm control method as claimed in any one of claims 1 to 11 when executing the computer program.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the robot arm control method according to any one of claims 1 to 11.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the robot arm control method according to any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Mechanical arm intelligent impedance control method and system based on deep reinforcement learning
CN116587275A