Mechanical arm grabbing model training method and device, electronic equipment and storage medium
By using a piecewise reward function and network layer loss optimization, the problem of sparse rewards in the training of robotic arm grasping models was solved, thereby improving the grasping success rate.
Patent Information
- Application Number
- CN202310234468.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-03
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2043-03-03
AI Technical Summary
In existing technologies, robotic arm grasping models suffer from low grasping success rates due to sparse rewards during training, making it impossible to effectively improve grasping strategies.
A piecewise reward function is adopted. By acquiring the environmental state information of the object to be grasped by the robotic arm, the information is input into the pre-constructed piecewise reward function to obtain reward information. The model to be trained is then trained based on the environmental state information, action information, and reward information, including the combination of progressive reward function and grasping reward function, and the loss update of the commentator and actor network layers is optimized.
This effectively avoids sparse rewards, improves the training effect of the robotic arm grasping model, and increases the success rate of grasping objects.
Smart Images

Figure CN116197909B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of machine vision, and in particular to a mechanical arm grasping model training method and device, electronic equipment and a storage medium. BACKGROUND
[0002] With the development of artificial intelligence technology, relevant intelligent robots are widely used in various industries. They play a crucial role in improving industrial production efficiency, reducing production costs, and improving product quality.
[0003] In the prior art, reinforcement learning has been introduced into the control and planning of mechanical arms, so that the mechanical arm has certain recognition, judgment, comparison, identification, memory and self-adjustment capabilities in the interaction with the environment.
[0004] At present, reinforcement learning has the problem of sparse rewards in the environment reward, that is, when the mechanical arm fails to grasp the object, the reward obtained is always 0, and no positive reward can be obtained to improve the grasping strategy of the mechanical arm, resulting in poor training effect of the mechanical arm grasping model and low success rate of object grasping using the mechanical arm grasping model. SUMMARY
[0005] The present application provides a mechanical arm grasping model training method, device, electronic equipment and storage medium to improve the training accuracy of the mechanical arm grasping model and thus improve the success rate of object grasping using the mechanical arm grasping model.
[0006] According to an aspect of the present application, a mechanical arm grasping model training method is provided, comprising:
[0007] obtaining environment state information of an object to be grasped by a mechanical arm;
[0008] inputting the environment state information of the object to be grasped by the mechanical arm into a pre-constructed segmented reward function to obtain reward information;
[0009] training a to-be-trained model based on the environment state information of the object to be grasped by the mechanical arm, action information corresponding to the environment state information of the object to be grasped by the mechanical arm, and the reward information to obtain a mechanical arm grasping model.
[0010] According to another aspect of the present application, a mechanical arm grasping model training device is provided, comprising:
[0011] an environment state information acquisition module configured to obtain environment state information of an object to be grasped by a mechanical arm;
[0012] a reward information determination module configured to input the environment state information of the object to be grasped by the mechanical arm into a pre-constructed segmented reward function to obtain reward information;
[0013] The grasping model training module is configured to train the to-be-trained model based on the environment state information of the object to be grasped by the robot arm, the action information corresponding to the environment state information of the object to be grasped by the robot arm, and the reward information, and obtain a robot arm grasping model.
[0014] According to another aspect of the present application, an electronic device is provided, which comprises:
[0015] at least one processor;
[0016] and a memory connected in communication with the at least one processor;
[0017] wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the training method of the robot arm grasping model according to any one of the embodiments of the present application.
[0018] According to another aspect of the present application, a computer readable storage medium is provided, which stores computer instructions for enabling a processor to perform the training method of the robot arm grasping model according to any one of the embodiments of the present application when executed by the processor.
[0019] The technical solution of the embodiments of the present application obtains the environment state information of the object to be grasped by the robot arm, then inputs the environment state information of the object to be grasped by the robot arm into a pre-constructed segmented reward function to obtain reward information, and then trains a to-be-trained model based on the environment state information of the object to be grasped by the robot arm, the action information corresponding to the environment state information of the object to be grasped by the robot arm, and the reward information, to obtain a robot arm grasping model. The above technical solution determines the reward information through the segmented reward function, which can effectively avoid the problem of sparse reward, thereby improving the training effect of the robot arm grasping model, and further improving the success rate of using the robot arm grasping model to grasp the object.
[0020] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0022] Figure 1 is a flow chart of a training method of a mechanical arm grasping model according to an embodiment of the present application;
[0023] Figure 2 is a flow chart of a training method of a mechanical arm grasping model according to an embodiment of the present application;
[0024] Figure 3 is a flow chart of a training method of a mechanical arm grasping model according to an embodiment of the present application;
[0025] Figure 4 is a structural schematic diagram of a network to be trained according to an embodiment of the present application;
[0026] Figure 5 is a result schematic diagram of model simulation according to an embodiment of the present application;
[0027] Figure 6 is a structural schematic diagram of a training device of a mechanical arm grasping model according to an embodiment of the present application;
[0028] Figure 7 is a structural schematic diagram of an electronic device for implementing the training method of the mechanical arm grasping model according to an embodiment of the present application. DETAILED DESCRIPTION
[0029] In order to make the personnel in the art better understand the present application scheme, the technical solutions in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by the person skilled in the art without creative labor should belong to the scope of protection of the present application.
[0030] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0031] Embodiment one
[0032] Figure 1 A flowchart of a training method of a mechanical arm grasping model is provided for Embodiment One of the present application. The present embodiment can be applied to the training of a mechanical arm grasping model based on visual reinforcement learning. The method can be performed by a training device of a mechanical arm grasping model, which can be realized in the form of hardware and / or software, and can be configured in a computer terminal and / or a server. As shown in Figure 1 The method comprises the following steps.
[0033] In S110, the environmental state information of the object to be grasped by the mechanical arm is acquired.
[0034] In the present embodiment, the environmental state information (State) refers to the state of the agent in the current environment, which can include but is not limited to image length, image width and RGB (Red Green Blue) color.
[0035] Specifically, the environmental state information of the object to be grasped by the mechanical arm can be acquired from a preset storage location of the electronic device; or the environmental state information of the object to be grasped by the mechanical arm can be acquired from other electronic devices connected to the electronic device, which can be a camera or the like arranged on the mechanical arm.
[0036] In S120, the environmental state information of the object to be grasped by the mechanical arm is input into a pre-constructed segmented reward function to obtain reward information.
[0037] In the present embodiment, the segmented reward function is a segmented reward function designed according to the grasping distance, which can be used to determine the reward of the action taken. The reward information (Reward) is the reward of the action taken in the environmental state.
[0038] In some embodiments, the segmented reward function can be composed of two reward sub-functions; in some embodiments, the segmented reward function can also be composed of more than two reward sub-functions, which is not limited in the present embodiment.
[0039] It should be noted that, compared with the existing sparse reward, the segmented reward function can provide finely divided rewards to improve the grasping strategy of the mechanical arm, thereby improving the training effect of the mechanical arm grasping model.
[0040] In S130, the trained model is trained based on the environmental state information of the object to be grasped by the mechanical arm, the action information corresponding to the environmental state information of the object to be grasped by the mechanical arm, and the reward information, to obtain a mechanical arm grasping model.
[0041] In the present embodiment, the action information (Action) is the action that can be taken by the agent in the environmental state.
[0042] Specifically, the plurality of environment state information, the action information corresponding to the plurality of environment state information, and the reward information corresponding to the plurality of environment state information can be used as model training samples of the to-be-trained model, and then the to-be-trained model is trained according to the model training samples until a to-be-trained model stop training condition is met, and a robot arm grabbing model is obtained.
[0043] On the basis of the above embodiments, after the to-be-trained model is trained based on the environment state information of the object to be grabbed by the robot arm, the action information corresponding to the environment state information of the object to be grabbed by the robot arm, and the reward information, the method further comprises: obtaining real environment state information of the object to be grabbed by the robot arm; and inputting the real environment state information of the object to be grabbed by the robot arm into the robot arm grabbing model to complete the grabbing action of the object to be grabbed.
[0044] The real environment state information refers to the environment state information collected by the real robot arm.
[0045] For example, the trained robot arm grabbing model is saved and deployed to a real robot arm. The real robot arm can obtain real environment state information of an object to be grabbed by the robot arm, and then input the real environment state information of the object to be grabbed by the robot arm into the robot arm grabbing model to control the real robot arm to complete the action of grabbing the object.
[0046] The technical scheme of the embodiment of the application determines the reward information through a segmented reward function, which can effectively avoid the problem of sparse rewards, thereby improving the training effect of the robot arm grabbing model and further improving the success rate of using the robot arm grabbing model to grab the object.
[0047] Embodiment two
[0048] Figure 2 A flowchart of a training method of a robot arm grabbing model provided in the second embodiment of the application, the method of the present embodiment can be combined with each optional scheme in the training method of the robot arm grabbing model provided in the above embodiments. The training method of the robot arm grabbing model provided in the present embodiment is further optimized. Optionally, the segmented reward function comprises a progressive reward function and a grabbing reward function. Correspondingly, the inputting of the environment state information of the object to be grabbed by the robot arm into the pre-constructed segmented reward function to obtain reward information comprises: inputting the environment state information of the object to be grabbed by the robot arm into the progressive reward function to obtain progressive reward information; inputting the environment state information of the object to be grabbed by the robot arm into the grabbing reward function to obtain grabbing reward information; and determining the reward information based on the progressive reward information and the grabbing reward information.
[0049] For example, Figure 2As shown, the method comprises:
[0050] S210, obtaining environment state information of an object to be grabbed by the robot arm.
[0051] S220, inputting the environment state information of the object to be grabbed by the robot arm into the progressive reward function to obtain progressive reward information.
[0052] S230, inputting the environment state information of the object to be grabbed by the robot arm into the grabbing reward function to obtain grabbing reward information.
[0053] S240, determining reward information based on the progressive reward information and the grabbing reward information.
[0054] S250, training a to-be-trained model based on the environment state information of the object to be grabbed by the robot arm, action information corresponding to the environment state information of the object to be grabbed by the robot arm, and the reward information, to obtain a robot arm grabbing model.
[0055] In this embodiment, the segmented reward function can include a progressive reward function and a grabbing reward function, wherein the progressive reward function is used to represent whether the action has a better state change trend, and the grabbing reward function is used to represent the action grabbing object.
[0056] Specifically, the environment state information of the object to be grabbed by the robot arm is input into the progressive reward function to obtain progressive reward information, and the environment state information of the object to be grabbed by the robot arm is input into the grabbing reward function to obtain grabbing reward information, and then the reward information is determined according to the progressive reward information and the grabbing reward information, which realizes segmented determination of the reward information, can effectively avoid the problem of sparse reward, thereby improving the training effect of the robot arm grabbing model, and further improving the success rate of using the robot arm grabbing model to grab the object.
[0057] On the basis of the above embodiments, optionally, the progressive reward function can be:
[0058] r tendency = v(s t+1 ) - v(s t );
[0059] Wherein, r tendency represents the progressive reward function, v(s t ) represents the value function of the current environment state, and v(s t+1 ) represents the value function of the next environment state.
[0060] On the basis of the above embodiments, optionally, the grabbing reward function can be:
[0061]
[0062] wherein, r grasping represents a grasp reward function, v(s t ) represents a value function of a current environment state, grasp∩v(s t ) represents that the robot arm grasps the object and the value function of the current environment state is greater than a preset value threshold, and δ represents the preset value threshold.
[0063] It can be understood that if the robot arm does not grasp the object, the grasp reward information is 0; if the robot arm grasps the object and the value function of the current environment state does not exceed the preset value threshold, the grasp reward information is the value function of the current environment state minus a constant 100; if the robot arm grasps the object and the value function of the current environment state is greater than the preset value threshold, the grasp reward information is a constant 500, and the preset value threshold can be determined according to detection and grasping performance.
[0064] In some embodiments, the constant 500 and the constant 100 in the grasp reward function can be changed according to specific grasping requirements, for example, in the case that the robot arm grasps the object and the value function of the current environment state does not exceed the preset value threshold, the grasp reward information can be the value function of the current environment state minus a constant 150, and in the case that the robot arm grasps the object and the value function of the current environment state is greater than the preset value threshold, the grasp reward information can be a constant 600, which is not limited herein.
[0065] For example, the segmented reward function can be:
[0066] r=r tendency +r grasping ;
[0067] r tendency =v(s t+1 )-v(s t );
[0068]
[0069] wherein, r represents the determined reward information.
[0070] The technical scheme of the embodiment of the application inputs the environment state information of the object to be grasped by the robot arm into the progressive reward function to obtain progressive reward information, and inputs the environment state information of the object to be grasped by the robot arm into the grasp reward function to obtain grasp reward information, and then determines the reward information according to the progressive reward information and the grasp reward information, realizes segmented determination of the reward information, can effectively avoid the problem of sparse reward, and thus improves the training effect of the robot arm grasping model, and further improves the success rate of grasping the object using the robot arm grasping model.
[0071] Embodiment three
[0072] Figure 3 A flowchart of a training method of a mechanical arm grasping model is provided for Embodiment Three of the present application. The method of this embodiment can be combined with any of the optional solutions of the training method of the mechanical arm grasping model provided in the above embodiments. The training method of the mechanical arm grasping model provided in this embodiment is further optimized. Optionally, the to-be-trained model is trained based on the environment state information of the object to be grasped by the mechanical arm, the action information corresponding to the environment state information of the object to be grasped by the mechanical arm, and the reward information, to obtain the mechanical arm grasping model, including: determining the loss of the critic network layer based on the reward information, and updating the network parameters of the critic network layer based on the loss of the critic network layer; determining the loss of the actor network layer based on the environment state information of the object to be grasped by the mechanical arm, the action information corresponding to the environment state information of the object to be grasped by the mechanical arm, and the reward information, and updating the network parameters of the actor network layer based on the loss of the actor network layer; until the stop training condition of the to-be-trained model is met, the mechanical arm grasping model is obtained.
[0073] As shown in Figure 3 , the method includes:
[0074] S310, environment state information of an object to be grasped by a mechanical arm is acquired.
[0075] S320, the environment state information of the object to be grasped by the mechanical arm is input into a pre-constructed segmented reward function to obtain reward information.
[0076] S330, the loss of the critic network layer is determined based on the reward information, and the network parameters of the critic network layer are updated based on the loss of the critic network layer.
[0077] S340, the loss of the actor network layer is determined based on the environment state information of the object to be grasped by the mechanical arm, the action information corresponding to the environment state information of the object to be grasped by the mechanical arm, and the reward information, and the network parameters of the actor network layer are updated based on the loss of the actor network layer.
[0078] S350, until the stop training condition of the to-be-trained model is met, the mechanical arm grasping model is obtained.
[0079] In this embodiment, the to-be-trained model can include a shared network layer, an actor network layer, and a critic network layer. The shared network layer is used to extract position feature information of the object to be grasped by the mechanical arm. The actor network layer can select actions based on probability, the critic network layer can judge the score of the action based on the action of the actor network layer, and the actor network layer can modify the probability of the selected action according to the score of the critic network layer.
[0080] Specifically, the electronic device obtains the environment state information of the object to be grasped by the robot arm, and then inputs the environment state information of the object to be grasped by the robot arm into a pre-constructed segmented reward function to obtain reward information; then determines the loss of the critic network layer based on the reward information, updates the network parameters of the critic network layer based on the loss of the critic network layer, determines the loss of the actor network layer based on the environment state information of the object to be grasped by the robot arm, the action information corresponding to the environment state information of the object to be grasped by the robot arm, and the reward information, and updates the network parameters of the actor network layer based on the loss of the actor network layer until the stop training condition of the to-be-trained model is met, and the robot arm grasping model is obtained.
[0081] On the basis of the above embodiments, the loss of the critic network layer is determined based on the reward information, including: determining the discount reward corresponding to the reward information based on the reward information; determining the advantage function value based on the discount reward corresponding to the reward information; and determining the loss of the critic network layer based on the advantage function value.
[0082] For example, the reward information can be calculated by R = [R0, R1, R2, R3, …], where γ represents a discount factor, r t represents the reward information at time t, R represents the discount reward; then the advantage function value At = R - V' is calculated, where V' represents the state value corresponding to each environment state information, and At represents the advantage function value; and then the loss of the critic network layer is determined based on the pre-configured loss function c_loss = mean(square(At)).
[0083] On the basis of the above embodiments, the loss of the actor network layer is determined based on the environment state information of the object to be grasped by the robot arm, the action information corresponding to the environment state information of the object to be grasped by the robot arm, and the reward information, and the network parameters of the actor network layer are updated based on the loss of the actor network layer, including: determining the advantage function value based on the reward information; determining a first normal distribution model and a second normal distribution model based on the environment state information of the object to be grasped by the robot arm; inputting the action information corresponding to the environment state information of the object to be grasped by the robot arm into the first normal distribution model and the second normal distribution model respectively to obtain first probability information and second probability information; determining a ratio based on the first probability information and the second probability information; and determining the loss of the actor network layer based on the advantage function value and the ratio.
[0084] For example, the reward information can be calculated by Figure 4Fig. 1 is a structural schematic diagram of a to-be-trained network provided by the embodiment, the actor network layer includes an actor-new network and an actor-old network, the network structure of the actor-new network is the same as that of the actor-old network, the to-be-trained network can be trained in a PyBullet simulation environment and perform object grasping, so that the model converges, and specifically includes the following steps.
[0085] Step 1, input the environment state information to the convolutional neural network, and extract the position information features of the object to be grasped based on the convolutional neural network (CNN).
[0086] Step 2-1, input the position information features of the object to be grasped extracted by the convolutional neural network (CNN) to the actor-new network, obtain mu and sigma, then construct a normal distribution by taking mu and sigma as the mean and variance of the normal distribution respectively, the normal distribution is used to represent the distribution of action information, and then the action information is generated through the normal distribution, the action information is input into the environment to obtain reward information and next-step environment state information, the next-step environment state information is input into the actor-new network, and the next-step environment state information is also input into the critic network layer to obtain a value function, and store [(s, a, r), …], wherein s represents environment state information, a represents action information, and r represents reward information, and the step 2-1 is repeated until a preset number of [(s, a, r), …] is stored, and the actor-new network is not updated in the process. The reward function for determining the reward information is:
[0087] r=r tendency +r grasping ;
[0088] r tendency =v(s t+1 )-v(s t );
[0089]
[0090] Step 2-2, calculate the discounted reward, and obtain R=[R0, R1, R2, R3, …] through , wherein γ represents a discount factor, r t represents the reward information at time t, and R represents the discounted reward.
[0091] Step 2-3, calculate the advantage function, At=R–V’, wherein V’ represents the state value corresponding to each environment state information, and At represents the advantage function.
[0092] Step 2-4, determine the loss of the critic network layer based on the loss function c_loss = mean(square(At)), and then update the network parameters of the critic network layer by backpropagation.
[0093] Step 2-5, input all the stored environment state information into the actor-old and actor-new networks respectively to obtain the first and second normal distribution models respectively, and input all the stored action information into the first and second normal distribution models respectively to obtain the first and second probability information corresponding to the action information, and then divide the second probability information by the first probability information to obtain the ratio.
[0094] Step 2-6, determine the loss of the actor network layer based on the loss function a_loss = mean(min((ratio*At, clip(ratio, 1-ξ, 1+ξ)*At))), where ratio represents the ratio, ξ represents a constant set by default, and At represents the advantage function value; and then update the actor-new network based on the loss of the actor network layer.
[0095] Step 2-7, loop steps 2-2 to 2-6, after the preset number of loops, the loop ends, and the actor-old network is updated using the actor-new network weight.
[0096] Step 2-8, loop steps 2-1 to 2-7, stop training after 1000 times of training to obtain the robot arm grasping model.
[0097] Step 3, save the trained robot arm grasping model and deploy it to a real robot arm to complete the grasping action.
[0098] To verify the feasibility of the robot arm grasping model training method proposed in this embodiment, a simulation experiment is performed to compare the robot arm grasping model training method (PRPPO) proposed in this embodiment with the Trust Region Policy Optimization (TRPO) algorithm and the Proximal Policy Optimization (PPO) algorithm in the prior art. The simulation results are shown in Figure 5 Figure 5 It can be seen that the reward information obtained by the robot arm grasping model training method proposed in this embodiment is better than the TRPO algorithm, and the stability is also better than the PPO algorithm.
[0099] The technical scheme of the embodiment of the present application determines the loss of the critic network layer based on the reward information, updates the network parameters of the critic network layer based on the loss of the critic network layer, determines the loss of the actor network layer based on the environment state information of the object to be grabbed by the robot arm, the action information corresponding to the environment state information of the object to be grabbed by the robot arm, and the reward information, updates the network parameters of the actor network layer based on the loss of the actor network layer, completes the training of the robot arm grabbing model, and supplies the real robot arm for deployment.
[0100] Embodiment four
[0101] Figure 6 A structural schematic diagram of a robot arm grabbing model training device provided by the fourth embodiment of the present application is shown in FIG. 4. As shown in the figure, the device comprises: Figure 6
[0102] An environment state information acquisition module 410 is configured to acquire the environment state information of the object to be grabbed by the robot arm.
[0103] A reward information determination module 420 is configured to input the environment state information of the object to be grabbed by the robot arm into a pre-constructed segmented reward function to obtain reward information.
[0104] A grabbing model training module 430 is configured to train a to-be-trained model based on the environment state information of the object to be grabbed by the robot arm, the action information corresponding to the environment state information of the object to be grabbed by the robot arm, and the reward information, and obtain a robot arm grabbing model.
[0105] The technical scheme of the embodiment of the present application acquires the environment state information of the object to be grabbed by the robot arm, then inputs the environment state information of the object to be grabbed by the robot arm into a pre-constructed segmented reward function to obtain reward information, and then trains a to-be-trained model based on the environment state information of the object to be grabbed by the robot arm, the action information corresponding to the environment state information of the object to be grabbed by the robot arm, and the reward information, and obtains a robot arm grabbing model. The above technical scheme determines the reward information through a segmented reward function, which can effectively avoid the problem of sparse reward, thereby improving the training effect of the robot arm grabbing model and further improving the success rate of grabbing objects using the robot arm grabbing model.
[0106] In some optional embodiments, the segmented reward function comprises a progressive reward function and a grabbing reward function; and the reward information determination module 420 is specifically configured to:
[0107] input the environment state information of the object to be grabbed by the robot arm into the progressive reward function to obtain progressive reward information;
[0108] input the environment state information of the object to be grabbed by the robot arm into the grabbing reward function to obtain grabbing reward information.
[0109] determine reward information based on the progressive reward information and the grasp reward information.
[0110] In some optional embodiments, the progressive reward function is:
[0111] r tendency = v(s t+1 ) - v(s t );
[0112] wherein r tendency represents the progressive reward function, v(s t ) represents the value function of the current environment state, and v(s t+1 ) represents the value function of the next environment state.
[0113] In some optional embodiments, the grasp reward function is:
[0114]
[0115] wherein r grasping represents the grasp reward function, v(s t ) represents the value function of the current environment state, grasp∩v(s t ) represents that the robot arm grasps the object and the value function of the current environment state is greater than a preset value threshold, and δ represents the preset value threshold.
[0116] In some optional embodiments, the grasp model training module 430 comprises:
[0117] a critic network layer updating unit configured to determine a loss of a critic network layer based on the reward information, and update network parameters of the critic network layer based on the loss of the critic network layer;
[0118] an actor network layer updating unit configured to determine a loss of an actor network layer based on the environment state information of the object to be grasped by the robot arm, the action information corresponding to the environment state information of the object to be grasped by the robot arm, and the reward information, and update network parameters of the actor network layer based on the loss of the actor network layer;
[0119] a model training stopping unit configured to obtain a robot arm grasping model until a stopping training condition of the model to be trained is met.
[0120] In some optional embodiments, the critic network layer updating unit is specifically configured to:
[0121] determine a discounted reward corresponding to the reward information based on the reward information;
[0122] determine an advantage function value based on the reward information;
[0123] determine a critic network layer loss based on the advantage function value.
[0124] In some optional embodiments, the actor network layer updating unit is specifically configured to:
[0125] determine an advantage function value based on the reward information;
[0126] determine a first normal distribution model and a second normal distribution model based on the environment state information of the object to be grasped by the robot arm;
[0127] input the action information corresponding to the environment state information of the object to be grasped by the robot arm into the first normal distribution model and the second normal distribution model respectively to obtain first probability information and second probability information;
[0128] determine a ratio based on the first probability information and the second probability information;
[0129] determine an actor network layer loss based on the advantage function value and the ratio.
[0130] In some optional embodiments, the training device of the robot arm grasping model further comprises:
[0131] a real environment state information acquisition module configured to acquire real environment state information of an object to be grasped by a robot arm;
[0132] an object grasping module configured to input the real environment state information of the object to be grasped by the robot arm into the robot arm grasping model to complete a grasping action of the object to be grasped.
[0133] The training device of the robot arm grasping model provided in the embodiments of the present application can execute the training method of the robot arm grasping model provided in any of the embodiments of the present application, and has the corresponding functional modules and beneficial effects of the execution method.
[0134] Embodiment five
[0135] Figure 7A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0136] like Figure 7 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded into the RAM 13 from storage unit 18. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An I / O interface 15 is also connected to the bus 14.
[0137] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0138] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as a training method for a robotic arm grasping model, which includes:
[0139] Obtain environmental state information of the object to be grasped by the robotic arm;
[0140] The environment state information of the object to be grabbed by the mechanical arm is input to a pre-constructed segmented reward function to obtain reward information;
[0141] The training model is trained based on the environment state information of the object to be grabbed by the mechanical arm, the action information corresponding to the environment state information of the object to be grabbed by the mechanical arm, and the reward information, to obtain a mechanical arm grabbing model.
[0142] In some embodiments, the training method of the mechanical arm grabbing model can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the training method of the mechanical arm grabbing model described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the training method of the mechanical arm grabbing model by any other appropriate means, for example, by means of firmware.
[0143] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0144] Computer programs used to implement the methods of the application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program, when executed by the processor of the machine, implements the functions / acts specified in the flowcharts and / or block diagrams. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, and partially on a machine or a remote machine or a server.
[0145] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0146] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0147] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0148] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.
[0149] It should be understood that the various forms of flow shown above can be used to reorder, add or delete steps. For example, each step described in the present application can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions of the present application can be achieved, which is not limited herein.
[0150] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A training method for a robotic arm grasping model, characterized in that, include: Obtain environmental state information of the object to be grasped by the robotic arm; The environmental state information of the object to be grasped by the robotic arm is input into a pre-constructed piecewise reward function to obtain reward information; wherein, the piecewise reward function is a piecewise reward function designed according to the grasping distance; The robotic arm grasping model is trained based on the environmental state information of the object to be grasped by the robotic arm, the action information corresponding to the environmental state information of the object to be grasped by the robotic arm, and the reward information, so as to obtain the robotic arm grasping model. The segmented reward function includes an incremental reward function and a capture reward function; Accordingly, the step of inputting the environmental state information of the object to be grasped by the robotic arm into a pre-constructed piecewise reward function to obtain reward information includes: The environmental state information of the object to be grasped by the robotic arm is input into the progressive reward function to obtain progressive reward information; The environmental state information of the object to be grasped by the robotic arm is input into the grasping reward function to obtain grasping reward information; The reward information is determined based on the progressive reward information and the capture reward information; The progressive reward function is as follows: ; in, This represents the asymptotic reward function. The value function representing the current environmental state. The value function representing the next environmental state; The capture reward function is: ; in, This indicates the function for capturing rewards. The value function representing the current environmental state. This indicates that the value function of the robotic arm grasping an object and the current environmental state is greater than a preset value threshold. This indicates a preset value threshold.
2. The method according to claim 1, characterized in that, The process of training the model based on the environmental state information of the object to be grasped by the robotic arm, the corresponding action information of the environmental state information of the object to be grasped by the robotic arm, and the reward information to obtain the robotic arm grasping model includes: The loss of the commentator network layer is determined based on the reward information, and the network parameters of the commentator network layer are updated based on the loss of the commentator network layer. The loss of the actor network layer is determined based on the environmental state information of the object to be grasped by the robotic arm, the action information corresponding to the environmental state information of the object to be grasped by the robotic arm, and the reward information. The network parameters of the actor network layer are updated based on the loss of the actor network layer. The training continues until the stop training condition of the model to be trained is met, thus obtaining the robotic arm grasping model.
3. The method according to claim 2, characterized in that, Determining the loss of the commenter network layer based on the reward information includes: Determine the discount reward corresponding to the reward information based on the reward information; The advantage function value is determined based on the discount reward corresponding to the aforementioned reward information; The loss of the commentator network layer is determined based on the aforementioned advantage function value.
4. The method according to claim 2, characterized in that, The process of determining the loss of the actor network layer based on the environmental state information of the object to be grasped by the robotic arm, the corresponding action information of the environmental state information of the object to be grasped by the robotic arm, and the reward information, and updating the network parameters of the actor network layer based on the loss of the actor network layer, includes: The advantage function value is determined based on the reward information; Based on the environmental state information of the object to be grasped by the robotic arm, a first normal distribution model and a second normal distribution model are determined; The motion information corresponding to the environmental state information of the object to be grasped by the robotic arm is input into the first normal distribution model and the second normal distribution model respectively to obtain the first probability information and the second probability information. The ratio is determined based on the first probability information and the second probability information; The loss of the actor network layer is determined based on the advantage function value and the ratio.
5. The method according to claim 1, characterized in that, After training the robotic arm grasping model based on the environmental state information of the object to be grasped by the robotic arm, the corresponding action information of the environmental state information of the object to be grasped, and the reward information, the process further includes: Obtain the real environmental state information of the object to be grasped by the robotic arm; The real environmental state information of the object to be grasped by the robotic arm is input into the robotic arm grasping model to complete the grasping action of the object to be grasped.
6. A training device for a robotic arm to grasp a model, characterized in that, include: The environmental status information acquisition module is used to acquire the environmental status information of the object to be grasped by the robotic arm; The reward information determination module is used to input the environmental state information of the object to be grasped by the robotic arm into a pre-constructed piecewise reward function to obtain reward information; wherein, the piecewise reward function is a piecewise reward function designed according to the grasping distance; The grasping model training module is used to train the model to be trained based on the environmental state information of the object to be grasped by the robotic arm, the action information corresponding to the environmental state information of the object to be grasped by the robotic arm, and the reward information, so as to obtain the robotic arm grasping model. The segmented reward function includes a progressive reward function and a capture reward function; the reward information determination module is specifically used for: The environmental state information of the object to be grasped by the robotic arm is input into the progressive reward function to obtain progressive reward information; The environmental state information of the object to be grasped by the robotic arm is input into the grasping reward function to obtain grasping reward information; The reward information is determined based on the progressive reward information and the capture reward information; The progressive reward function is as follows: ; in, This represents the asymptotic reward function. The value function representing the current environmental state. The value function representing the next environmental state; The capture reward function is: ; in, This indicates the function for capturing rewards. The value function representing the current environmental state. This indicates that the value function of the robotic arm grasping an object and the current environmental state is greater than a preset value threshold. This indicates a preset value threshold.
7. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the training method for the robotic arm grasping model according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the training method for the robotic arm grasping model according to any one of claims 1-5.
Citation Information
Patent Citations
Robot stirring and grabbing combination method based on deep reinforcement learning
CN112102405A
Robot rapid assembly method and system based on near-end strategy optimization algorithm
CN113977583A