Robotic arm control method, device, electronic device and storage medium
By predicting the probability of each candidate action of the robot arm on a tightly packed object, determining the target action and controlling the execution of the robot arm, the accuracy and reliability of the robot arm when grabbing a tight object is solved, and the grab success rate is improved.
Patent Information
- Application Number
- CN202110626122.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-04
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-06-04
AI Technical Summary
When the robot arm grasps closely packed objects, it is difficult to find a suitable grasping position, and the collision between the robot arm claws and the object causes the grasping failure.
By obtaining the depth map and color map of the target object, predict the probability values of each candidate action of the robot arm, including the probability of separation of the target object from the adjacent object and the probability of successful capture, determine the target action and control the execution of the robot arm.
It improves the accuracy of target action determination, avoids collision of robotic arms, improves the reliability of robotic arms movement execution, and thus improves the success rate of object grabbing in complex environments.
Smart Images

Figure CN115431258B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, in particular to the field of computer vision technology, and specifically to a robot arm control method, device, electronic device and storage medium. Background Art
[0002] The field of robotics has important applications in classification robots, service robots, human-computer interaction and other scenarios, and has received increasing attention in recent years. However, it is still a challenge for robotic arms to automatically grasp tightly stacked objects. This is because directly grasping tightly stacked objects will not only cause the algorithm to find a suitable grasping posture for grasping, but also cause the robot arm's gripper to collide with the object, resulting in grasping failure. Summary of the invention
[0003] The present application provides a robot arm control method, device, electronic device and storage medium for improving the grasping success rate.
[0004] According to one aspect of the present application, a method for controlling a robot arm is provided, comprising:
[0005] Obtain a first depth map and a first color map of the target object;
[0006] According to the first depth map and the first color map, a first prediction value and a second prediction value of each candidate action of the robotic arm are predicted, wherein the first prediction value is the probability that the robotic arm performs the corresponding candidate action to separate the target object from the adjacent objects; the second prediction value is the probability that the robotic arm performs the corresponding candidate action to successfully grasp the target object;
[0007] Determining a target action according to the first prediction value and the second prediction value of each of the candidate actions;
[0008] Control the robotic arm to perform the target action.
[0009] According to another aspect of the present application, a robot arm control device is provided, comprising:
[0010] An acquisition module, used to acquire a first depth map and a first color map of a target object;
[0011] A prediction module, configured to predict, based on the first depth map and the first color map, a first prediction value and a second prediction value of each candidate action of the robotic arm, wherein the first prediction value is a probability that the robotic arm performs the corresponding candidate action to separate the target object from adjacent objects; and the second prediction value is a probability that the robotic arm performs the corresponding candidate action to successfully grasp the target object;
[0012] A determination module, configured to determine a target action according to the first prediction value and the second prediction value of each of the candidate actions;
[0013] A control module is used to control the robotic arm to perform the target action.
[0014] According to another aspect of the present application, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method described in the aforementioned aspect is implemented.
[0015] According to another aspect of the present application, a non-transitory computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method described in the aforementioned aspect is implemented.
[0016] According to another aspect of the present application, a computer program product is provided. When instructions in the computer program product are executed by a processor, the method described in the above aspect is performed.
[0017] The technical solution provided in the embodiments of the present application may have the following beneficial effects:
[0018] A first depth map and a first color map of the target object are obtained, and a first prediction value and a second prediction value of each candidate action of the robot arm are predicted based on the first depth map and the first color map, wherein the first prediction value is the probability that the robot arm performs the corresponding candidate action to separate the target object from adjacent objects, and the second prediction value is the probability that the robot arm performs the corresponding candidate action to successfully grasp the target object. According to the first prediction value and the second prediction value of each candidate action, the target action is determined, and the robot arm is controlled to perform the target action. In the present application, by predicting the first prediction value and the second prediction value of each candidate action, a collaborative analysis of the strategy of separating the target object from adjacent objects and the strategy of grasping the target object is achieved, so as to select the action corresponding to the largest prediction value as the target action, thereby improving the accuracy of determining the target action, avoiding collision of the robot arm, and improving the reliability of the robot arm's action execution.
[0019] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0021] Figure 1 A schematic diagram of a flow chart of a robot arm control method provided in an embodiment of the present application;
[0022] Figure 2 A schematic flow chart of another robotic arm control method provided in an embodiment of the present application;
[0023] Figure 3 A schematic diagram of the structure of a prediction network provided in an embodiment of the present application;
[0024] Figure 4 A schematic flow chart of another robotic arm control method provided in an embodiment of the present application;
[0025] Figure 5 A schematic flow chart of another robotic arm control method provided in an embodiment of the present application;
[0026] Figure 6 A schematic diagram of the structure of a classification network provided in an embodiment of the present application;
[0027] Figure 7 A schematic diagram of the structure of a robotic arm control device provided in an embodiment of the present application;
[0028] Figure 8 A block diagram of an exemplary computer device suitable for implementing embodiments of the present application is shown. DETAILED DESCRIPTION
[0029] Embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0030] The following describes the robot arm control method, device, electronic device and storage medium of the embodiments of the present application with reference to the accompanying drawings.
[0031] Figure 1 A schematic flow chart of a robotic arm control method provided in an embodiment of the present application.
[0032] like Figure 1 As shown, the method comprises the following steps:
[0033] Step 101: Obtain a first depth map and a first color map of a target object.
[0034] The first depth map and the first color map are depth maps and color maps in the robot arm coordinate system.
[0035] In this embodiment, by acquiring the first depth map and the first color map of the target object, the current depth information and color information of the target object can be determined.
[0036] In one implementation method of an embodiment of the present application, the collected original depth map and original color map are obtained, and the original depth map and original color map are converted to the robotic arm coordinate system to obtain the three-dimensional coordinate position corresponding to each pixel point in the original depth map and the original color map in the robotic arm coordinate system.
[0037] Specifically, the original depth map I is collected using an RGB-D camera depth and the original color map I color , where the resolution of the original depth map and the original color map is M*N, where M and N are the width and height of the original depth map and the original color map, respectively. pix Convert to the robot coordinate system to obtain the three-dimensional coordinate position of each pixel in the original depth map and the original color map in the robot coordinate system, denoted as P w :
[0038] P w =R×(K -1 ×z×P pix )+T
[0039] Among them, K -1 represents the inverse matrix of the intrinsic parameter matrix, R represents the extrinsic parameter rotation matrix, T represents the extrinsic parameter translation matrix, and z represents the depth in the camera coordinate system.
[0040] Then, each three-dimensional coordinate position is projected onto a set two-dimensional plane to obtain a two-dimensional coordinate point corresponding to each three-dimensional coordinate position. As an implementation method, the X-axis and Y-axis coordinates (x w ,y w ) is mapped to a two-dimensional plane of size H×W, denoted as (x s ,y s ):
[0041] (x s ,y s )=(floor((x w -x l ) / res),floor((y w -y l ) / res))
[0042] Among them, x l Indicates the minimum value of the working space in the X-axis direction of the robot coordinate system, y l It represents the minimum value of the working space in the Y-axis direction of the robot coordinate system, res represents the actual size of each pixel after mapping, and floor(·) represents the rounding down operation.
[0043] In an implementation of this embodiment, setting the two-dimensional plane is determined based on a minimum working distance along a set direction in the robot arm coordinate system.
[0044] Further, according to the depth of the pixel corresponding to each two-dimensional coordinate point in the original depth map, a first depth map is generated, and according to the color of the pixel corresponding to each two-dimensional coordinate point in the original color map, a first color map is generated. Specifically, the depth Z of the pixel corresponding to each two-dimensional coordinate point in the original depth map is transferred to the corresponding coordinate (x s ,y s ) to obtain a depth state map I with a size of H×W depth_map ; Transfer the color information (r, g, b) of the corresponding pixel point of each two-dimensional coordinate point in the original color map to the corresponding coordinate (x s ,y s ) to obtain a color state image I with a size of H×W color_map .
[0045] Step 102, based on the first depth map and the first color map, predict the first prediction value and the second prediction value of each candidate action of the robotic arm, wherein the first prediction value is the probability that the robotic arm performs the corresponding candidate action to separate the target object from the adjacent objects, and the second prediction value is the probability that the robotic arm performs the corresponding candidate action to successfully grasp the target object.
[0046] The candidate actions include an action for separating the target object from adjacent objects and an action for grabbing the target object.
[0047] In this embodiment, according to the depth information and color information carried in the first depth map and the first color map, the corresponding features of the first depth map and the features of the first color map are extracted, and the first prediction value and the second prediction value of each candidate action of the robotic arm are predicted based on the extracted features of the first depth map and the features of the first color map.
[0048] Step 103: Determine a target action according to the first prediction value and the second prediction value of each candidate action.
[0049] In this embodiment, the first prediction value and the second prediction value of each candidate action are compared, and the candidate action with the larger prediction value is used as the target action, wherein the target action is an action for separating the target object from the adjacent object, or an action for grabbing the target object.
[0050] Step 104: Control the robotic arm to perform the target action.
[0051] Furthermore, according to the determined target action, the robotic arm is controlled to perform the target action. In the embodiment of the present application, by predicting the first prediction value and the second prediction value of each candidate action, a collaborative analysis of the strategy of separating the target object from the adjacent objects and the strategy of grasping the target object is achieved, so as to select the action corresponding to the largest prediction value as the target action, thereby improving the accuracy of determining the target action, avoiding collision of the robotic arm, and improving the reliability of the execution of the robotic arm's actions.
[0052] In the robotic arm control method of the present embodiment, a first depth map and a first color map of the target object are obtained, and a first prediction value and a second prediction value of each candidate action of the robotic arm are predicted based on the first depth map and the first color map, wherein the first prediction value is the probability that the robotic arm performs the corresponding candidate action to separate the target object from adjacent objects, and the second prediction value is the probability that the robotic arm performs the corresponding candidate action to successfully grasp the target object. Based on the first prediction value and the second prediction value of each candidate action, the target action is determined, and the robotic arm is controlled to perform the target action. In the present application, by predicting the first prediction value and the second prediction value of each candidate action, a collaborative analysis of the strategy of separating the target object from adjacent objects and the strategy of grasping the target object is achieved, so as to select the action corresponding to the largest prediction value as the target action, thereby improving the accuracy of determining the target action, reducing the probability of robotic arm collision, and improving the reliability of robotic arm action execution, thereby improving the success rate of object grasping in complex environments.
[0053] Based on the previous embodiment, this embodiment provides another robot arm control method. Figure 2 A flowchart of another robotic arm control method provided in an embodiment of the present application.
[0054] like Figure 2 As shown, step 102 may include the following steps:
[0055] Step 201 : rotating a first depth map by a plurality of set angles along a set rotation direction to obtain a plurality of input depth maps.
[0056] In this embodiment, in order to construct depth maps under different scenarios and obtain more features of the depth maps, the state space of the reinforcement learning algorithm is formed with multiple color maps obtained in the following steps. The first depth map is rotated along a set rotation direction by multiple set angles to obtain multiple input depth maps.
[0057] For example, the first depth map is rotated in a set direction within a 360° circle with a rotation interval of Δθ, for example, counterclockwise, then d = 360° / Δθ times, and d groups of depth maps with different rotation angles are obtained. For example, the number d of the obtained multiple depth maps is 16.
[0058] Step 202: Rotate the first color map by a plurality of set angles along a set rotation direction to obtain a plurality of input color maps.
[0059] In this embodiment, in order to construct color maps under different scenes and obtain more features of the color maps, so as to constitute the state space of the reinforcement learning algorithm together with the multiple depth maps obtained above, the first color map is rotated by multiple set angles along a set rotation direction to obtain multiple input color maps.
[0060] For example, the first color map is rotated in a set direction within a 360° circle with a rotation interval of Δθ, for example, counterclockwise, then d = 360° / Δθ times, and d groups of color maps with different rotation angles are obtained. For example, the number d of the obtained multiple color maps is 16.
[0061] Step 203: Input multiple input depth maps and multiple input color maps into a first prediction network to predict and obtain a first prediction value for each candidate action.
[0062] Among them, the first prediction network consists of a feature extraction layer, a feature fusion layer, a prediction layer and a dynamic optimization layer.
[0063] For example, Figure 3 As shown in the figure, the feature extraction layer is a convolutional layer of a DenseNet network, for example, a DenseNet-121 network pre-trained on ImageNet. The feature fusion layer consists of a Batch Normalization layer, a Rectified Linear Unit activation layer, and a convolutional layer with a convolution kernel size of 3×3. The prediction layer is an upstate layer.
[0064] In one implementation of this embodiment, a feature extraction layer of a first prediction network is used to extract features from multiple input depth maps and multiple input color maps, and the features of the multiple input depth maps are fused with the features of the corresponding input color maps to obtain multiple first fused feature maps, and each first fused feature map is reversely rotated along a set rotation direction so that the direction of the reversely rotated first fused feature map is consistent with that of the first depth map or the first color map. Furthermore, the prediction layer of the first prediction network is used to perform action prediction on each reversely rotated first fused feature map to obtain a first prediction value for each candidate action.
[0065] For example, Figure 3 As shown, in this embodiment, multiple depth maps and multiple color maps are d=16 groups as an example for description, wherein each group includes a depth map and a color map. i group as an example, where the depth map and colormap Each of them is subjected to feature extraction operation through the convolutional layer of a DenseNet-121 network pre-trained on ImageNet to obtain a color feature map. and deep feature maps Then, the color feature map and deep feature maps Perform channel splicing operation to obtain the initial fusion feature map Then After two convolution groups with the same structure, the features are deeply fused to obtain the first fused feature map Further, the first fusion feature map Rotate clockwise so that the rotation matches the color map I color_map The same angle direction is used to obtain the second fused feature map after reverse rotation; similarly, the second fused feature map after reverse rotation corresponding to other groups can be obtained. Then, the prediction layer of the first prediction network is used to perform action prediction on the first fused feature map after reverse rotation to obtain the first prediction value of each candidate action. As a possible implementation method, in the prediction layer of the first prediction network, the first fused feature maps after reverse rotation are subjected to channel splicing operations to obtain the first prediction value Q1(s) of each candidate action with a d-dimensional size of H×W. t ,a;θ1), where the d groups of color images collected under the current time robot working range environment are and depth map Recorded as state s t , θ1 represents the first prediction network parameter; a represents the action space, which consists of the candidate action types of the robot arm, the execution position of the robot arm (x w ,y w ,z w ), the gripper rotation angle Θ and the gripper pushing length L are composed of four parts.
[0066] Step 204: input the multiple input depth maps and the multiple input color maps into a second prediction network to predict and obtain a second prediction value for each candidate action.
[0067] Among them, the first prediction network consists of a feature extraction layer, a feature fusion layer, a prediction layer and a dynamic optimization layer.
[0068] For example, Figure 3 As shown in the figure, the feature extraction layer is a convolutional layer of a DenseNet network, for example, a DenseNet-121 network pre-trained on ImageNet. The feature fusion layer consists of a Batch Normalization layer, a Rectified Linear Unit activation layer, and a convolutional layer with a convolution kernel size of 3×3. The prediction layer is an upstate layer.
[0069] In one implementation of the present embodiment, a feature extraction layer of a second prediction network is used to perform feature extraction on multiple input depth maps and multiple input color maps, and the features of the multiple input depth maps are fused with the features of the corresponding input color maps to obtain multiple second fused feature maps. Each second fused feature map is reversely rotated along a set rotation direction, and the prediction layer of the second prediction network is used to perform action prediction on each second fused feature map after reverse rotation to obtain a second prediction value of each candidate action.
[0070] For example, Figure 3 As shown, in this embodiment, multiple depth maps and multiple color maps are d=16 groups as an example for description, wherein each group includes a depth map and a color map. i group as an example, where the depth map and colormap Each of them is subjected to feature extraction operation through the convolutional layer of a DenseNet-121 network pre-trained on ImageNet to obtain a color feature map. and deep feature maps Then, the color feature map and deep feature maps Perform channel splicing operation to obtain the initial fusion feature map Then After two convolution groups with the same structure, the features are deeply fused to obtain the corresponding d i The second fusion feature map of the group Further, the second fusion feature map Rotate clockwise so that the rotated second fusion feature map is consistent with the color state map I color_map The same angle direction; similarly, the second fused feature maps corresponding to the other groups after reverse rotation can be obtained. Then, the prediction layer of the second prediction network is used to perform action prediction on each second fused feature map after reverse rotation to obtain the second prediction value of each candidate action. As a possible implementation method, in the prediction layer of the second prediction network, each second fused feature map after reverse rotation is subjected to channel splicing operation to obtain the second prediction value Q2(s t ,a;θ2), where the d groups of color images collected under the current robot arm working environment are and depth map Recorded as state s t , θ2 represents the second prediction network parameter; a represents the action space, which consists of the candidate action types of the robot arm, the execution position of the robot arm (x w ,y w ,z w), the gripper rotation angle Θ and the gripper pushing length L are composed of four parts.
[0071] In the robotic arm control method of the present embodiment, by rotating the first color map and the first depth map along a set rotation direction by multiple set angles, multiple input color maps and multiple input depth maps are obtained, thereby realizing the construction of color maps and depth maps in different scenes, so as to obtain more features of color maps and depth maps, so as to constitute the state space of the reinforcement learning algorithm, and then, based on the state space composed of multiple groups of depth maps and color maps, the first prediction network is used to predict the first prediction value of each candidate action and the second prediction value of each candidate action, and then based on the first prediction value and the second prediction value of each candidate action, a collaborative analysis of the strategy of separating the target object from the adjacent objects and the strategy of grasping the target object is realized, so as to select the action corresponding to the largest prediction value as the target action, thereby improving the accuracy of the target action determination, reducing the probability of robotic arm collision, and improving the reliability of the robotic arm action execution, thereby improving the success rate of object grasping in complex environments.
[0072] Based on the above embodiments, this embodiment provides an implementation method, which illustrates that the first prediction value of each candidate action and the second prediction value of each candidate action obtained by prediction are corrected to improve the accuracy of the first prediction value and the second prediction value, thereby improving the accuracy of determining the target action.
[0073] Figure 4 A flow chart of another robot arm control method provided in an embodiment of the present application is shown as follows: Figure 4 As shown, step 103 includes the following steps:
[0074] Step 401: modify the first prediction value of each candidate action according to the contour of the target object indicated by the first depth map.
[0075] As an implementation manner, a dynamic mask is calculated according to the outline of the target object indicated by the first depth map, and the dynamic mask is multiplied by the first prediction value of each candidate action to obtain a corrected first prediction value of each candidate action.
[0076] Step 402: Correct the second prediction value of each candidate action according to the center position of the target object indicated by the first depth map.
[0077] As an implementation manner, a dynamic mask is calculated according to the center position of the target object indicated by the first depth map, and the dynamic mask is multiplied by the second prediction value of each candidate action to obtain a corrected second prediction value of each candidate action.
[0078] Step 403: Select a target action from each candidate action according to the corrected first prediction value and the corrected second prediction value.
[0079] In one implementation of this embodiment, the maximum value among the revised first prediction values is determined based on the revised first prediction values of each candidate action, and the maximum value among the revised second prediction values is determined based on the revised second prediction values of each candidate action. Then, the maximum value of the revised first prediction value and the maximum value of the second prediction value are compared to determine the candidate action corresponding to the largest prediction value, and the candidate action corresponding to the largest prediction value is taken as the target action.
[0080] In one scenario, the candidate actions are "pushing action" and "grasping action", and the first corrected prediction value of "pushing action" and the first corrected prediction value of "grasping action" are determined, and then the maximum value among the corrected first prediction values is determined, for example, the maximum value is the corrected first prediction value of "pushing action", and the second corrected prediction value of "pushing action" and the second corrected prediction value of "grasping action" are determined, and then the maximum value among the corrected second prediction values is determined, for example, the maximum value is the corrected second prediction value of "grasping action". Further, the corrected first prediction value of "pushing action" and the corrected second prediction value of "grasping action" are compared, and the larger prediction value is selected, for example, the corrected first threshold value of "pushing action", and "pushing action" is used as the target action of the robot arm.
[0081] In the control method of the robotic arm of the present embodiment, the first prediction value and the second prediction value of each candidate action are corrected to improve the accuracy of the first prediction value and the second prediction value corresponding to each candidate action, and then the target action is determined according to the corrected first prediction value and the second prediction value of each candidate action, thereby realizing a collaborative analysis of the strategy of separating the target object from the adjacent objects and the strategy of grasping the target object, so as to select the action corresponding to the largest prediction value as the target action, thereby improving the accuracy of the target action determination, reducing the probability of robotic arm collision, and improving the reliability of the robotic arm action execution, thereby improving the success rate of object grasping in complex environments.
[0082] Based on the above embodiments, this embodiment provides a possible implementation method, which specifically describes how to train the first prediction network and the second prediction network.
[0083] Figure 5 A flow chart of another robot arm control method provided in an embodiment of the present application is shown as follows: Figure 5 As shown, step 104 includes the following steps:
[0084] Step 501: collect a second depth map of the target object after the target action is performed.
[0085] Among them, after controlling the robotic arm to perform the target action, the position distribution of the target object will change, and a second depth map of the target object after the target action is performed is obtained. Among them, the method for obtaining the second depth map can refer to the description in the previous embodiment and is no longer limited in this embodiment.
[0086] Step 502: Determine a first reward value for the target action using a classification network based on the second depth map and the first depth map, wherein the first reward value is used to indicate the effectiveness of the robot arm in performing the target action to separate the target object from adjacent objects.
[0087] The first depth map is the depth map obtained before the robot arm is controlled to perform the target action.
[0088] like Figure 6 As shown, Figure 6 A schematic diagram of a classification network structure provided in an embodiment of the present application is shown in FIG. Figure 6 As shown in the figure, the first depth map and the second depth map are used as the input of the classification network, and each is subjected to feature extraction through a convolutional layer of a pre-trained VGG16 network, and the features of the extracted first depth map and the features of the second depth map are subjected to channel concatenation to obtain a fused feature map. Then, the fused feature map is subjected to a convolutional layer with a convolution kernel size of 1×1 to obtain a convolution feature. Figure 1 , and then, the convolution feature Figure 1 Input the convolution layer with a convolution kernel size of 3×3 to obtain the convolution feature Figure 2 ; Finally, the convolution feature Figure 2 After three fully connected layers, the degree to which the target action separates the target objects is determined, and the first reward value of the target action is determined based on the degree of separation.
[0089] For example, the first reward value is r p ,
[0090]
[0091] Among them, output = 0 means that the target action can aggregate the target objects, which is not conducive to the successful grasping of the target objects. The first reward value of this type of target action is determined to be a penalty value of -0.5; output = 1 means that the target action can separate the objects, which is conducive to the successful grasping of the target objects. The first reward value of this type of target action is determined to be a positive reward value of 0.5.
[0092] Step 503: Determine a second reward value for the target action according to whether the robot arm successfully grasps the target object.
[0093] In one implementation of this embodiment, the second reward value r g Defined as:
[0094] If the robot arm successfully grasps the target object after executing the target action, the second reward value of the target action is determined to be 1.5. If the robot arm successfully grasps the target object after executing the target action, the second reward value of the target action is determined to be 0.
[0095] Step 504: training a first prediction network according to the first reward value, and training a second prediction network according to the second reward value.
[0096] In one implementation of this embodiment, the loss function of the first prediction network is:
[0097]
[0098] Among them, r p is the first reward value, s t is the state space composed of multiple depth maps and multiple color maps corresponding to time t, Q1(s t ,a t ; θ) is the value function of the predicted value of the target action of the first prediction network at time t, θ1 is the network parameter at the current time t, θ 1target is the target network parameter of the first prediction network, It is the value function of the predicted value of the target action of the first prediction network at time t+1, and γ represents the attenuation factor.
[0099] According to the loss function determined by the first reward value, the parameters of the first prediction network are continuously adjusted to train the first prediction network.
[0100] In one implementation of this embodiment, the loss function of the second prediction network is:
[0101]
[0102] Among them, rg is the second reward value, s t is the state space composed of multiple depth maps and multiple color maps corresponding to time t, Q2(s t ,a t ; θ2) is the value function of the predicted value of the target action of the second prediction network at time t, θ2 is the network parameter at the current time t, θ 2target is the target network parameter of the second prediction network, It is the value function of the predicted value of the target action of the second prediction network at time t+1, and γ represents the attenuation factor.
[0103] According to the loss function determined by the second reward value, the parameters of the second prediction network are continuously adjusted to train the second prediction network.
[0104] In the robotic arm control method of the present embodiment, by determining the first reward value and the second reward value, the first reward value is used to train the first prediction network so that the trained first prediction network learns the corresponding relationship between multiple depth references and multiple color images and the first prediction value of each candidate action, and the second reward value is used to train the second prediction network so that the trained second prediction network learns the corresponding relationship between multiple depth references and multiple color images and the second prediction value of each candidate action.
[0105] In order to implement the above embodiments, the present application also proposes a robot arm control device.
[0106] Figure 7 A schematic diagram of the structure of a robotic arm control device provided in an embodiment of the present application.
[0107] like Figure 7 As shown, the device comprises:
[0108] The acquisition module 71 is used to acquire a first depth map and a first color map of the target object.
[0109] The prediction module 72 is used to predict the first prediction value and the second prediction value of each candidate action of the robotic arm based on the first depth map and the first color map, wherein the first prediction value is the probability that the robotic arm performs the corresponding candidate action to separate the target object from the adjacent objects; the second prediction value is the probability that the robotic arm performs the corresponding candidate action to successfully grasp the target object.
[0110] The determination module 73 is used to determine the target action according to the first prediction value and the second prediction value of each candidate action.
[0111] The control module 74 is used to control the robotic arm to perform the target action.
[0112] Furthermore, in a possible implementation of the embodiment of the present application, the prediction module 72 is used to:
[0113] Rotating the first depth map by a plurality of set angles along a set rotation direction to obtain a plurality of input depth maps;
[0114] Rotating the first color map along the set rotation direction by the multiple set angles to obtain multiple input color maps;
[0115] Inputting the multiple input depth maps and the multiple input color maps into a first prediction network to predict and obtain the first prediction value of each of the candidate actions;
[0116] The multiple input depth maps and the multiple input color maps are input into a second prediction network to predict the second prediction value of each of the candidate actions.
[0117] In a possible implementation of the embodiment of the present application, the device further includes:
[0118] A collection module, used for collecting a second depth map of the target object after the target action is performed;
[0119] a processing module, configured to determine a first reward value for the target action using a classification network according to the second depth map and the first depth map, wherein the first reward value is used to indicate an effectiveness degree of the robot arm performing the target action to separate the target object from adjacent objects;
[0120] The determination module is used to determine a second reward value for the target action according to whether the robot arm successfully grasps the target object;
[0121] A training module is used to train the first prediction network according to the first reward value, and to train the second prediction network according to the second reward value.
[0122] In a possible implementation of the embodiment of the present application, the processing module is used to:
[0123] Using a feature extraction layer of a classification network, respectively extracting features from the first depth map and the second depth map;
[0124] Fusing the features of the first depth map with the features of the second depth map to obtain a first fused feature;
[0125] The classification layer of the classification network is used to perform classification prediction on the first fusion feature to obtain a first reward value for the target action.
[0126] In a possible implementation of the embodiment of the present application, the prediction module 72 is specifically configured to:
[0127] Using the feature extraction layer of the first prediction network, extracting features from the multiple input depth maps and the multiple input color maps, and fusing features of the multiple input depth maps with features of corresponding input color maps to obtain multiple first fused feature maps;
[0128] Rotating each of the first fused feature images in the opposite direction along the set rotation direction;
[0129] The prediction layer of the first prediction network is used to perform action prediction on each of the first fused feature maps after reverse rotation to obtain a first prediction value for each of the candidate actions.
[0130] In a possible implementation of the embodiment of the present application, the prediction module 72 is specifically configured to:
[0131] Using the feature extraction layer of the second prediction network, extracting features from the multiple input depth maps and the multiple input color maps, and fusing features of the multiple input depth maps with features of corresponding input color maps to obtain multiple second fused feature maps;
[0132] Rotating each of the second fused feature images in the opposite direction along the set rotation direction;
[0133] The prediction layer of the second prediction network is used to perform action prediction on each of the second fused feature maps after reverse rotation to obtain a second prediction value for each of the candidate actions.
[0134] In a possible implementation of the embodiment of the present application, the determination module 73 is specifically configured to:
[0135] Modifying the first prediction value of each of the candidate actions according to the contour of the target object indicated by the first depth map;
[0136] Correcting the second prediction value of each of the candidate actions according to the center position of the target object indicated by the first depth map;
[0137] The target action is selected from the candidate actions according to the corrected first prediction value and the corrected second prediction value.
[0138] In a possible implementation of the embodiment of the present application, the acquisition module 71 is specifically configured to:
[0139] Obtain the collected original depth map and original color map;
[0140] Convert the original depth map and the original color map to the robot coordinate system to obtain the three-dimensional coordinate position of each pixel point in the original depth map and the original color map in the robot coordinate system;
[0141] Projecting each of the three-dimensional coordinate positions onto a set two-dimensional plane to obtain a two-dimensional coordinate point corresponding to each of the three-dimensional coordinate positions;
[0142] Generate the first depth map according to the depth of the pixel point corresponding to each of the two-dimensional coordinate points in the original depth map;
[0143] The first color map is generated according to the colors of the pixels corresponding to the two-dimensional coordinate points in the original color map.
[0144] In a possible implementation of the embodiment of the present application, the set two-dimensional plane is determined according to the minimum working distance along the set direction in the robot arm coordinate system.
[0145] It should be noted that the above explanation of the method embodiment is also applicable to the device of this embodiment, and will not be repeated here.
[0146] In the robotic arm control device of the present embodiment, a first depth map and a first color map of the target object are obtained, and a first prediction value and a second prediction value of each candidate action of the robotic arm are predicted based on the first depth map and the first color map, wherein the first prediction value is the probability that the robotic arm performs the corresponding candidate action to separate the target object from the adjacent objects, and the second prediction value is the probability that the robotic arm performs the corresponding candidate action to successfully grasp the target object. Based on the first prediction value and the second prediction value of each candidate action, the target action is determined, and the robotic arm is controlled to perform the target action. In the present application, by predicting the first prediction value and the second prediction value of each candidate action, a collaborative analysis of the strategy of separating the target object from the adjacent objects and the strategy of grasping the target object is achieved, so as to select the action corresponding to the largest prediction value as the target action, thereby improving the accuracy of determining the target action, reducing the probability of robotic arm collision, and improving the reliability of robotic arm action execution, thereby improving the success rate of object grasping in complex environments.
[0147] In order to implement the above embodiments, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method described in the above method embodiment is implemented.
[0148] In order to implement the above-mentioned embodiments, the embodiments of the present application provide a non-temporary computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the method described in the above-mentioned method embodiments is implemented.
[0149] In order to implement the above embodiment, the embodiment of the present application provides a computer program product. When the instructions in the computer program product are executed by a processor, the method described in the above method embodiment is executed.
[0150] Figure 8 A block diagram of an exemplary computer device suitable for implementing embodiments of the present application is shown. Figure 8 The computer device 12 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0151] like Figure 8 As shown, the computer device 12 is in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, a memory 28, and a bus 18 that connects various system components (including the memory 28 and the processing unit 16).
[0152] The bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor or a local bus using any of a variety of bus structures. For example, these architectures include but are not limited to Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus and Peripheral Component Interconnection (PCI) bus.
[0153] The computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the computer device 12, including volatile and non-volatile media, removable and non-removable media.
[0154] The memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be used to read and write non-removable, non-volatile magnetic media ( Figure 8 not shown, usually called a "hard drive"). Although Figure 8 Not shown, a disk drive for reading and writing a removable non-volatile disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing a removable non-volatile optical disk (e.g., a compact disc read only memory (CD-ROM), a digital versatile disc read only memory (DVD-ROM), or other optical media) may be provided. In these cases, each drive may be connected to the bus 18 via one or more data medium interfaces. The memory 28 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the various embodiments of the present application.
[0155] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in the memory 28, such program modules 42 including, but not limited to, an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment. The program modules 42 generally perform the functions and / or methods of the embodiments described herein.
[0156] The computer device 12 may also communicate with one or more external devices 14 (e.g., keyboard, pointing device, display 24, etc.), one or more devices that enable a user to interact with the computer device 12, and / or any device that enables the computing device 12 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication may be performed via an input / output (I / O) interface 22. In addition, the computer device 12 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 20. As shown, the network adapter 20 communicates with other modules of the computer device 12 via a bus 18. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the computer device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0157] The processing unit 16 executes various functional applications and data processing by running the programs stored in the memory 28, such as implementing the methods mentioned in the above embodiments.
[0158] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0159] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0160] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0161] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute the instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.
[0162] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0163] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0164] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0165] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A robot arm control method, characterized in that: The following steps are involved: Obtain a first depth map and a first color map of the target object; According to the first depth map and the first color map, a first prediction value and a second prediction value of each candidate action of the robotic arm are predicted, wherein the first prediction value is the probability that the robotic arm performs the corresponding candidate action to separate the target object from the adjacent objects; the second prediction value is the probability that the robotic arm performs the corresponding candidate action to successfully grasp the target object; Modifying the first prediction value of each of the candidate actions according to the contour of the target object indicated by the first depth map; Correcting the second prediction value of each of the candidate actions according to the center position of the target object indicated by the first depth map; Selecting a target action from the candidate actions according to the corrected first prediction value and the corrected second prediction value; Control the robotic arm to perform the target action.
2. The robot arm control method according to claim 1, characterized in that: The step of predicting a first prediction value and a second prediction value of each candidate action of the robot arm according to the first depth map and the first color map includes: Rotating the first depth map by a plurality of set angles along a set rotation direction to obtain a plurality of input depth maps; Rotating the first color map along the set rotation direction by the multiple set angles to obtain multiple input color maps; Inputting the multiple input depth maps and the multiple input color maps into a first prediction network to predict and obtain the first prediction value of each of the candidate actions; The multiple input depth maps and the multiple input color maps are input into a second prediction network to predict the second prediction value of each of the candidate actions.
3. The robot arm control method according to claim 2, characterized in that: The method further comprises: Acquiring a second depth map of the target object after the target action is performed; Determine, using a classification network, a first reward value for the target action according to the second depth map and the first depth map, wherein the first reward value is used to indicate the effectiveness of the robot arm in performing the target action to separate the target object from adjacent objects; Determining a second reward value for the target action according to whether the robot arm successfully grasps the target object; The first prediction network is trained according to the first reward value, and the second prediction network is trained according to the second reward value.
4. The robot arm control method according to claim 3, characterized in that: The step of determining a first reward value of the target action using a classification network according to the second depth map and the first depth map includes: Using a feature extraction layer of a classification network, respectively extracting features from the first depth map and the second depth map; Fusing the features of the first depth map with the features of the second depth map to obtain a first fused feature; The classification layer of the classification network is used to perform classification prediction on the first fusion feature to obtain a first reward value for the target action.
5. The robot arm control method according to claim 2, characterized in that: The step of inputting the plurality of input depth maps and the plurality of input color maps into a first prediction network to predict and obtain the first prediction value of each of the candidate actions comprises: Using the feature extraction layer of the first prediction network, extracting features from the multiple input depth maps and the multiple input color maps, and fusing features of the multiple input depth maps with features of corresponding input color maps to obtain multiple first fused feature maps; Reversely rotating each of the first fused feature images along the set rotation direction; The prediction layer of the first prediction network is used to perform action prediction on each of the first fused feature maps after reverse rotation to obtain a first prediction value for each of the candidate actions.
6. The robot arm control method according to claim 2, characterized in that: The step of inputting the plurality of input depth maps and the plurality of input color maps into a second prediction network to predict and obtain the second prediction value of each of the candidate actions comprises: Using the feature extraction layer of the second prediction network, extracting features from the multiple input depth maps and the multiple input color maps, and fusing features of the multiple input depth maps with features of corresponding input color maps to obtain multiple second fused feature maps; Rotating each of the second fused feature images in the opposite direction along the set rotation direction; The prediction layer of the second prediction network is used to perform action prediction on each of the second fused feature maps after reverse rotation to obtain a second prediction value for each of the candidate actions.
7. The robot arm control method according to any one of claims 1 to 6, characterized in that: The step of obtaining a first depth map and a first color map of the target object includes: Obtain the collected original depth map and original color map; Convert the original depth map and the original color map to the robot coordinate system to obtain the three-dimensional coordinate position of each pixel point in the original depth map and the original color map in the robot coordinate system; Projecting each of the three-dimensional coordinate positions onto a set two-dimensional plane to obtain a two-dimensional coordinate point corresponding to each of the three-dimensional coordinate positions; Generate the first depth map according to the depth of the pixel point corresponding to each of the two-dimensional coordinate points in the original depth map; The first color map is generated according to the colors of the pixels corresponding to the two-dimensional coordinate points in the original color map.
8. The robot arm control method according to claim 7, characterized in that: The set two-dimensional plane is determined according to the minimum working distance along the set direction in the robot arm coordinate system.
9. A robot arm control device, characterized in that: include: An acquisition module, used to acquire a first depth map and a first color map of a target object; A prediction module, configured to predict, based on the first depth map and the first color map, a first prediction value and a second prediction value of each candidate action of the robotic arm, wherein the first prediction value is a probability that the robotic arm performs the corresponding candidate action to separate the target object from adjacent objects; and the second prediction value is a probability that the robotic arm performs the corresponding candidate action to successfully grasp the target object; A determination module, configured to modify a first prediction value of each of the candidate actions according to the outline of the target object indicated by the first depth map; modify a second prediction value of each of the candidate actions according to the center position of the target object indicated by the first depth map; and select a target action from each of the candidate actions according to the modified first prediction value and the modified second prediction value; A control module is used to control the robotic arm to perform the target action.
10. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 8 is implemented.
11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
12. A computer program product, characterized in that When the instructions in the computer program product are executed by a processor, the method according to any one of claims 1 to 8 is performed.
Citation Information
Patent Citations
Mechanical arm pushing and grabbing cooperation method suitable for dense environment
CN112643668A