Humanoid robot grabbing method and device, computer equipment and storage medium
By acquiring multimodal sensor data and using the crawling model to generate execution strategies, the problem of insufficient stability and security of robot crawling in the existing technology in complex environments is solved, and a higher crawling accuracy and success rate is achieved.
Patent Information
- Application Number
- CN202510320988.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-05-09
AI Technical Summary
The robot grasping methods in the prior art are difficult to meet the requirements in complex environments, especially when there are various types of objects, irregular shapes, many changes in dynamic scenes, and uncertain factors such as lighting changes and occlusion exist.
By obtaining multimodal sensor data, including object feature data and environmental feature data, the grab model is used to generate a grab strategy, including theoretical joint angle, position and torque, and the robot can accurately grasp it by executing the strategy. This method combines the convolutional layer, DDF layer and multi-head output layer to dynamically adjust the focus of the network through the multi-head self-attention mechanism to improve the robustness of the system and the crawling success rate.
It improves the accuracy and real-time grasping of humanoid robots, enhances stability and security in complex environments, improves the success rate of grasping, and enables the robot to adapt to more application occasions.
Smart Images

Figure CN119952718A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robots, and in particular to a humanoid robot grasping method, device, computer equipment and storage medium. Background Art
[0002] With the rapid development of artificial intelligence and robotics technology, humanoid robots are being used more and more widely in industries, services, medical care and other fields.
[0003] In complex environments, robot grasping tasks face many challenges, such as the variety of objects, irregular shapes, changing dynamic scenes, and uncertain factors such as lighting changes and occlusions in the environment. These all place higher demands on the robot's perception capabilities and decision-making mechanisms.
[0004] Traditional robotic grasping methods mainly rely on single-modality data, such as vision, force or touch. Single-modality data has significant limitations when facing complex tasks, such as visual data lacking information in occluded scenes and force data having a delayed response in dynamic environments.
[0005] It can be seen that the robot grasping method in the existing technology cannot meet the requirements in terms of stability, security and timeliness. Summary of the invention
[0006] In order to solve the above technical problems or at least partially solve the above technical problems, the present invention provides a humanoid robot grasping method, device, computer equipment and storage medium.
[0007] In a first aspect, the present invention provides a grasping method of a humanoid robot, wherein the humanoid robot comprises a mechanical arm, and a mechanical hand is disposed at the end of the mechanical arm, and the method comprises:
[0008] Acquire sensor data, wherein the sensor data includes object feature data and environment feature data, wherein the object feature data includes target object color information, target object depth image, and target object material information;
[0009] Acquire a grasping strategy according to the sensor data and the grasping model, wherein the grasping strategy includes a theoretical joint angle, a theoretical joint position, and a theoretical joint torque;
[0010] According to the grasping strategy, an execution strategy is acquired, wherein the execution strategy includes an execution joint angle, an execution joint position and an execution joint torque.
[0011] Optionally, before acquiring the grasping strategy according to the sensor data and the grasping model, the method further includes:
[0012] Collecting raw data, the raw data includes raw object feature data, raw environment feature data and raw robot data, the raw robot data includes: raw robot arm end posture data, raw robot hand joint angle, raw robot hand torque, raw robot hand tactile feedback data;
[0013] Acquire training data according to the original data;
[0014] The original network is trained according to the training data to obtain the grasping model.
[0015] Optionally, obtaining a grasping strategy according to the sensor data and the grasping model includes:
[0016] The input layer receives input data, wherein the input data includes the object feature data and the environment feature data;
[0017] The encoding layer extracts and outputs high-level features of the input data;
[0018] The decoding layer generates the capture strategy after upsampling the high-level features to the size of the input data;
[0019] The output layer outputs the grasping strategy.
[0020] Optionally, the encoding layer includes a convolution layer, a DDF layer and a multi-head output layer, and the DDF layer includes a compressed convolution layer, an expanded convolution layer and a dense residual connection layer.
[0021] The encoding layer extracts and outputs high-level features of the input data, including:
[0022] The convolution layer performs a convolution operation on the input data to obtain spatial features;
[0023] The DDF layer compresses, expands and reuses the spatial features;
[0024] The multi-head output layer captures global dependencies through a multi-head self-attention mechanism, splices the output data of the DDF layer into the high-level features according to the global dependencies, and outputs the high-level features;
[0025] The DDF layer compresses, expands and reuses the spatial features, including:
[0026] The compressed convolution layer reduces the number of channels of the spatial feature according to the first convolution kernel to obtain a compressed feature.
[0027] The dilated convolution layer applies a second convolution kernel with different dilation rates on the compressed features to capture multi-scale features.
[0028] The dense residual connection layer connects the outputs of different layers to the dense residual connection layer to achieve feature reuse.
[0029] Optionally, the convolution layer performs a convolution operation on the input data to obtain spatial features in the following manner:
[0030] F conv =X*W conv +b conv
[0031] Among them, F conv is the spatial feature, X is the input data of the convolutional layer, and W conv is the convolution kernel of the convolution layer, b conv is the bias term;
[0032] The compressed convolution layer reduces the number of channels of the spatial feature according to the first convolution kernel to obtain the compressed feature in the following manner:
[0033] F sq =F conv *W sq +b sq
[0034] Among them, F sq is the compression feature, F conv is the spatial feature, W sq is the first convolution kernel, b sq is the compression bias term;
[0035] The dilated convolution layer applies a second convolution kernel with different dilation rates on the compressed features to capture multi-scale features in the following manner:
[0036] F dil =F sq *rW dil +b dil
[0037] F dil is a multi-scale feature, F sq is the compression characteristic, r is the expansion rate, W dil is the second convolution kernel, b dil is the expansion bias term;
[0038] The dense residual connection layer connects the outputs of different layers to the dense residual connection layer to achieve feature reuse in the following way:
[0039] F (l) =H (l) ([F (0) ,F (1) ,...,F (l-1) ])
[0040] Among them, H (l) represents the nonlinear transformation function of the lth layer, [] represents the cascade of features, F (0) ……F (l-1) It is a multi-scale feature;
[0041] The multi-head output layer captures global dependencies through a multi-head self-attention mechanism, splices the output data of the DDF layer into the high-level features according to the global dependencies, and outputs the high-level features in the following manner:
[0042] MultiHead(F in )=Concat(Z1,Z2,...,Z h )W O
[0043] Z m =AV m
[0044]
[0045] in, is the mth query weight matrix, is the mth key weight matrix, is the mth value weight matrix, F in is the input data, d k is the dimension of the key vector, W O is the output weight matrix, h is the number of attention heads, MultiHead means concatenation, Concat means linear transformation, and m=1,2…h.
[0046] Optionally, after the decoding layer upsamples the high-level features to the size of the input data, generating the capture strategy includes:
[0047] The deconvolution layer upsamples the high-level features to the size of the input data;
[0048] The fusion connection layer fuses the corresponding features in the encoding layer with the features in the decoding layer to generate the capture strategy.
[0049] Optionally, the deconvolution layer upsamples the high-level features to the size of the input data in the following manner:
[0050] F deconv =F′ in*T W deconv +b deconv
[0051] Among them, F′ in is a high-level feature, W deconv is the deconvolution kernel, bdeconv is the deconvolution bias;
[0052] The fusion connection layer fuses the corresponding features in the encoding layer with the features in the decoding layer to generate the capture strategy in the following manner:
[0053]
[0054] in, represents the pixel-by-pixel sum of features, is the output feature of the encoder layer l, Output features for the decoder layer l.
[0055] Optionally, the grasping model further includes multimodal contrast loss, posture stability and grasping force coordination loss, and dynamic grasping success rate loss.
[0056] The multimodal contrast loss is obtained as follows:
[0057] L′ MMC =L MMC +λ
[0058]
[0059] λ = ReLU(s ivis ·s itor -∈)
[0060] Among them, L′ MMC is the multimodal contrast loss, N is the number of samples, M is the number of modalities, and the modalities include visual modality and moment modality, s ij is the true feature of the i-th sample in the h-th mode, is the prediction feature, s ivis is the visual modality feature, s itor is the moment modal characteristic, ∈ is the preset threshold;
[0061] The posture stability and grasping force coordination loss are obtained in the following manner:
[0062] L′ SSC =L SSC +γ
[0063]
[0064] γ=ReLU((y ist -θ st )·(y iforce -θ force ))
[0065] Among them, L′ SSCis the coordinated loss of posture stability and grasping strength, y ist is the true posture stability score of the i-th sample, y iforce Score the true grasping strength of the i-th sample, is the model prediction posture stability score for the i-th sample, The model predicts the grasping strength score for the i-th sample, θ st is the attitude stability threshold, θ force is the grasping force threshold, N is the number of samples;
[0066] The dynamic crawling success rate is obtained in the following manner:
[0067] L′ DSC =δ·L DSC
[0068]
[0069] Among them, L′ DSC is the dynamic crawling success rate, y isucc is the true successful capture label of the i-th sample, is the model prediction probability, is the prediction error, and N is the number of samples.
[0070] In a second aspect, a humanoid robot grasping device is provided, which is applied to a humanoid robot, wherein the humanoid robot comprises a mechanical arm, and a mechanical hand is arranged at the end of the mechanical arm, and the device comprises:
[0071] A sensor for acquiring sensor data, wherein the sensor data includes object feature data and environment feature data, wherein the object feature data includes target object color information, target object depth image, and target object material information;
[0072] A model unit, used for acquiring a grasping strategy according to the sensor data and the grasping model, wherein the grasping strategy includes a theoretical joint angle, a theoretical joint position and a theoretical joint torque;
[0073] The execution unit is used to obtain an execution strategy according to the grasping strategy, and the execution strategy includes an execution joint angle, an execution joint position and an execution joint torque.
[0074] According to a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above methods when executing the computer program.
[0075] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the method as described in any one of the above items is implemented.
[0076] The present invention provides a humanoid robot grasping method, device, computer equipment and storage medium, wherein the humanoid robot includes a mechanical arm, and a mechanical hand is arranged at the end of the mechanical arm. The method includes: acquiring sensor data, wherein the sensor data includes object feature data and environmental feature data, wherein the object feature data includes target object color information, target object depth image, and target object material information; acquiring a grasping strategy according to the sensor data and a grasping model, wherein the grasping strategy includes theoretical joint angles, theoretical joint positions, and theoretical joint torques; acquiring an execution strategy according to the grasping strategy, wherein the execution strategy includes execution joint angles, execution joint positions, and execution joint torques. The acquired sensor data includes object feature data and environmental feature data, wherein the object feature data includes target object color information, target object depth image, and target object material information, that is, the data collected in the embodiment of the present invention includes data of multiple modalities, providing the system with all-round environmental perception capabilities. The output grasping strategy includes the execution joint torque, which allows the humanoid robot to adjust the execution joint torque according to more material information of the target object when grasping, that is, adjust the grasping force and grasping point, etc., making the humanoid robot's grasping more targeted and adaptable to more application occasions, and also making the humanoid robot's grasping more accurate and real-time, and improving the grasping success rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0078] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0079] Figure 1 The figure shows an application environment diagram of the humanoid robot grasping method according to an embodiment of the present invention;
[0080] Figure 2 It is a schematic diagram of the process of the humanoid robot grasping method according to an embodiment of the present invention;
[0081] Figure 3 The figure is a structural block diagram of a humanoid robot grasping device according to an embodiment of the present invention;
[0082] Figure 4 FIG. 2 is a diagram showing the internal structure of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION
[0083] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0084] Figure 1 FIG. 1 is an application environment diagram of a humanoid robot grasping method in an embodiment. Figure 1 , the humanoid robot grasping method is applied to a humanoid robot grasping system. The humanoid robot grasping method includes a terminal 110 and / or a server 120. The terminal 110 and the server 120 are connected via a network. The terminal 110 may be a desktop terminal or a mobile terminal, and the mobile terminal may be at least one of a mobile phone, a tablet computer, a laptop computer, etc. The server 120 may be implemented as an independent server or a server cluster consisting of multiple servers.
[0085] The humanoid robot grasping method of the present invention is applied to the terminal 110 and / or the server 120 .
[0086] like Figure 2 As shown, in one embodiment, a humanoid robot grasping method is provided. This embodiment mainly applies the method to the above Figure 1 The server 120 in FIG. 1 is used as an example.
[0087] In the embodiment of the present invention, the humanoid robot comprises a mechanical arm, and a mechanical hand is arranged at the end of the mechanical arm.
[0088] Reference Figure 2 , the humanoid robot grasping method comprises:
[0089] Step 210, acquiring sensor data, wherein the sensor data includes object feature data and environment feature data, wherein the object feature data includes target object color information, target object depth image, and target object material information;
[0090] Step 220, obtaining a grasping strategy according to the sensor data and the grasping model, wherein the grasping strategy includes a theoretical joint angle, a theoretical joint position, and a theoretical joint torque;
[0091] Step 230: acquiring an execution strategy according to the grasping strategy, wherein the execution strategy includes an execution joint angle, an execution joint position and an execution joint torque.
[0092] The method of the embodiment of the present invention obtains sensor data including object feature data and environmental feature data, wherein the object feature data includes color information of the target object, depth image of the target object, and material information of the target object, that is, the data collected in the embodiment of the present invention includes data of multiple modes, providing the system with all-round environmental perception capabilities. The output grasping strategy includes the execution joint torque, which enables the humanoid robot to adjust the execution joint torque according to more material information of the target object when grasping, that is, adjust the grasping force and grasping landing point, etc., so that the grasping of the humanoid robot is more targeted and can adapt to more application occasions, and also makes the grasping of the humanoid robot more accurate and real-time, and improves the grasping success rate.
[0093] In the embodiment of the present invention, before step 220, before acquiring the grasping strategy according to the sensor data and the grasping model, the method further includes:
[0094] Collecting raw data, the raw data includes raw object feature data, raw environment feature data and raw robot data, the raw robot data includes: raw robot arm end posture data, raw robot hand joint angle, raw robot hand torque, raw robot hand tactile feedback data;
[0095] Acquire training data according to the original data;
[0096] The original network is trained according to the training data to obtain the grasping model.
[0097] In the embodiment of the present invention, the step of acquiring the grasping strategy according to the sensor data and the grasping model includes:
[0098] The input layer receives input data, wherein the input data includes the object feature data and the environment feature data;
[0099] The encoding layer extracts and outputs high-level features of the input data;
[0100] The decoding layer generates the capture strategy after upsampling the high-level features to the size of the input data;
[0101] The output layer outputs the grasping strategy.
[0102] In the embodiment of the present invention, the encoding layer includes a convolution layer, a DDF layer and a multi-head output layer, and the DDF layer includes a compressed convolution layer, an expanded convolution layer and a dense residual connection layer.
[0103] The encoding layer extracts and outputs high-level features of the input data, including:
[0104] The convolution layer performs a convolution operation on the input data to obtain spatial features;
[0105] The DDF layer compresses, expands and reuses the spatial features;
[0106] The multi-head output layer captures global dependencies through a multi-head self-attention mechanism, splices the output data of the DDF layer into the high-level features according to the global dependencies, and outputs the high-level features;
[0107] The DDF layer compresses, expands and reuses the spatial features, including:
[0108] The compressed convolution layer reduces the number of channels of the spatial feature according to the first convolution kernel to obtain a compressed feature.
[0109] The dilated convolution layer applies a second convolution kernel with different dilation rates on the compressed features to capture multi-scale features.
[0110] The dense residual connection layer connects the outputs of different layers to the dense residual connection layer to achieve feature reuse.
[0111] In the embodiment of the present invention, the convolution layer performs a convolution operation on the input data to obtain spatial features in the following manner:
[0112] F conv =X*W conv +b conv
[0113] Among them, F conv is the spatial feature, X is the input data of the convolutional layer, and W conv is the convolution kernel of the convolution layer, b conv is the bias term;
[0114] The compressed convolution layer reduces the number of channels of the spatial feature according to the first convolution kernel to obtain the compressed feature in the following manner:
[0115] F sq =F conv *W sq +b sq
[0116] Among them, F sq is the compression feature, F conv is the spatial feature, W sq is the first convolution kernel, b sq is the compression bias term;
[0117] The dilated convolution layer applies a second convolution kernel with different dilation rates on the compressed features to capture multi-scale features in the following manner:
[0118] F dil =Fsq *rW dil +b dil
[0119] F dil is a multi-scale feature, F sq is the compression characteristic, r is the expansion rate, W dil is the second convolution kernel, b dil is the expansion bias term;
[0120] The dense residual connection layer connects the outputs of different layers to the dense residual connection layer to achieve feature reuse in the following way:
[0121] F (l) =H (l) ([F (0) ,F (1) ,...,F (l-1) ])
[0122] Among them, H (l) represents the nonlinear transformation function of the lth layer, [] represents the cascade of features, F (0) ……F (l-1) It is a multi-scale feature;
[0123] The multi-head output layer captures global dependencies through a multi-head self-attention mechanism, splices the output data of the DDF layer into the high-level features according to the global dependencies, and outputs the high-level features in the following manner:
[0124] MultiHead(F ib )=Concat(Z1,Z2,...,Z h )W O
[0125] Z m =AV m
[0126]
[0127]
[0128] in, is the mth query weight matrix, is the Mth bond weight matrix, is the mth value weight matrix, F in is the input data, d k is the dimension of the key vector, W O is the output weight matrix, h is the number of attention heads, MultiHead means concatenation, Concat means linear transformation, and m=1,2…h.
[0129] In the embodiment of the present invention, after the decoding layer upsamples the high-level features to the size of the input data, the capture strategy is generated, including:
[0130] The deconvolution layer upsamples the high-level features to the size of the input data;
[0131] The fusion connection layer fuses the corresponding features in the encoding layer with the features in the decoding layer to generate the capture strategy.
[0132] In an embodiment of the present invention, the deconvolution layer upsamples the high-level features to the size of the input data in the following manner:
[0133] F deconv =F′ in*T W deconv +b deconv
[0134] Among them, F′ in is a high-level feature, W deconv is the deconvolution kernel, b deconv is the deconvolution bias;
[0135] The fusion connection layer fuses the corresponding features in the encoding layer with the features in the decoding layer to generate the capture strategy in the following manner:
[0136]
[0137] in, represents the pixel-by-pixel sum of features, is the output feature of the encoder layer l, Output features for the decoder layer l.
[0138] In the embodiment of the present invention, the grasping model further includes multimodal contrast loss, posture stability and grasping force coordination loss, and dynamic grasping success rate loss.
[0139] The multimodal contrast loss is obtained as follows:
[0140] L′ MMC =L MMC +λ
[0141]
[0142] λ = ReLU(s ivis ·s itor -∈)
[0143] Among them, L′ MMCis the multimodal contrast loss, N is the number of samples, M is the number of modalities, and the modalities include visual modality and moment modality, s ij is the true feature of the i-th sample under the j-th mode, is the prediction feature, s ivis is the visual modality feature, s itor is the moment modal characteristic, ∈ is the preset threshold;
[0144] The posture stability and grasping force coordination loss are obtained in the following manner:
[0145] L′ SSC =L SSC +γ
[0146]
[0147] γ=ReLU((y ist -θ st )·(y iforce -θ force ))
[0148] Among them, L′ SSC is the coordinated loss of posture stability and grasping strength, y ist is the true posture stability score of the i-th sample, y iforce Score the true grasping strength of the i-th sample, is the model prediction posture stability score for the i-th sample, The model predicts the grasping strength score for the i-th sample, θ st is the attitude stability threshold, θ force is the grasping force threshold, N is the number of samples;
[0149] The dynamic crawling success rate is obtained in the following manner:
[0150] L′ DSC =δ·L DSC
[0151]
[0152] Among them, L′ DSC is the dynamic crawling success rate, y isucc is the true successful capture label of the i-th sample, is the model prediction probability, is the prediction error, and N is the number of samples.
[0153] The grasping model of the embodiment of the present invention adopts an attention mechanism, which can dynamically adjust the "focus" of the network by simulating the selective attention function of the human visual system, so that the robot can focus on the key areas of the target object in the visual data, such as the shape features and surface material information near the grasping point. At the same time, in the process of multimodal data fusion, the attention mechanism can dynamically assign weights of different modalities according to task requirements, such as relying more on force feedback when the lighting changes drastically, and giving priority to visual information when the force signal is weak. This dynamic adjustment capability greatly improves the robustness of the system and the success rate of grasping.
[0154] In the embodiment of the present invention, a variety of loss functions are included, which combine the grasping point positioning loss, force feedback loss and posture stability loss to comprehensively improve the accuracy and real-time performance of the grasping strategy. At the same time, these loss functions improve the stability and safety of the grasping action by combining the characteristics of the attention mechanism.
[0155] like Figure 3 As shown, the present invention also provides a humanoid robot grasping device, which is applied to a humanoid robot, wherein the humanoid robot comprises a mechanical arm, and a mechanical hand is arranged at the end of the mechanical arm, and the device comprises:
[0156] Sensor 310, used to obtain sensor data, the sensor data includes object feature data and environment feature data, the object feature data includes target object color information, target object depth image, target object material information;
[0157] A model unit 320, for acquiring a grasping strategy according to the sensor data and the grasping model, wherein the grasping strategy includes a theoretical joint angle, a theoretical joint position and a theoretical joint torque;
[0158] The execution unit 330 is used to obtain an execution strategy according to the grasping strategy, and the execution strategy includes an execution joint angle, an execution joint position and an execution joint torque.
[0159] In the embodiment of the present invention, the model unit 320 is further used for:
[0160] Collecting raw data, the raw data includes raw object feature data, raw environment feature data and raw robot data, the raw robot data includes: raw robot arm end posture data, raw robot hand joint angle, raw robot hand torque, raw robot hand tactile feedback data;
[0161] Acquire training data according to the original data;
[0162] The original network is trained according to the training data to obtain the grasping model.
[0163] In the embodiment of the present invention, the model unit 320 is further used for:
[0164] The input layer receives input data, wherein the input data includes the object feature data and the environment feature data;
[0165] The encoding layer extracts and outputs high-level features of the input data;
[0166] The decoding layer generates the capture strategy after upsampling the high-level features to the size of the input data;
[0167] The output layer outputs the grasping strategy.
[0168] In the embodiment of the present invention, the encoding layer includes a convolution layer, a DDF layer and a multi-head output layer, and the DDF layer includes a compressed convolution layer, an expanded convolution layer and a dense residual connection layer.
[0169] The model unit 320 is also used for:
[0170] The convolution layer performs a convolution operation on the input data to obtain spatial features;
[0171] The DDF layer compresses, expands and reuses the spatial features;
[0172] The multi-head output layer captures global dependencies through a multi-head self-attention mechanism, splices the output data of the DDF layer into the high-level features according to the global dependencies, and outputs the high-level features;
[0173] The model unit 320 is also used for:
[0174] The compressed convolution layer reduces the number of channels of the spatial feature according to the first convolution kernel to obtain a compressed feature,
[0175] The expanded convolution layer applies a second convolution kernel with a different dilation rate on the compressed feature to capture multi-scale features,
[0176] The dense residual connection layer connects outputs of different layers to the dense residual connection layer to achieve feature reuse.
[0177] In the embodiment of the present invention, the model unit 320 is further configured to enable the convolution layer to perform a convolution operation on the input data to obtain spatial features in the following manner:
[0178] F conv =X*W conv +b conv
[0179] Among them, F conv is the spatial feature, X is the input data of the convolutional layer, and Wconv is the convolution kernel of the convolution layer, b conv is the bias term;
[0180] In the embodiment of the present invention, the capture model is further used to enable the compressed convolution layer to reduce the number of channels of the spatial feature according to the first convolution kernel in the following manner to obtain the compressed feature:
[0181] F sq =F conv *W sq +b sq
[0182] Among them, F sq is the compression feature, F conv is the spatial feature, W sq is the first convolution kernel, b sq is the compression bias term;
[0183] In the embodiment of the present invention, the model unit 320 is further configured to enable the dilated convolution layer to apply a second convolution kernel with a different dilation rate to the compressed feature in the following manner to capture multi-scale features:
[0184] F dsl =F sq *rW dil +b dil
[0185] F dil is a multi-scale feature, F sq is the compression characteristic, r is the expansion rate, W dil is the second convolution kernel, b dil is the expansion bias term;
[0186] In the embodiment of the present invention, the model unit 320 is further configured to enable the dense residual connection layer to connect outputs of different layers to the dense residual connection layer in the following manner to achieve feature reuse:
[0187] F (l) =H (l) ([F (0) ,F (1) ,...,F (l-1) ])
[0188] Among them, H (l) represents the nonlinear transformation function of the lth layer, [] represents the cascade of features, F (0) ……F (l-1) It is a multi-scale feature;
[0189] In the embodiment of the present invention, the model unit 320 is further used to enable the multi-head output layer to capture the global dependency through the multi-head self-attention mechanism, splice the output data of the DDF layer into the high-level features according to the global dependency, and output the high-level features in the following manner:
[0190] MultiHead(F in )=Concat(Z1,Z2,...,Z h )W O
[0191] Z m =AV m
[0192]
[0193] in, is the mth query weight matrix, is the mth key weight matrix, is the mth value weight matrix, F in is the input data, d k is the dimension of the key vector, W O is the output weight matrix, h is the number of attention heads, MultiHead means concatenation, Concat means linear transformation, and m=1,2…h.
[0194] In the embodiment of the present invention, the model unit 320 is further used for:
[0195] Enable the deconvolution layer to upsample the high-level features to the size of the input data;
[0196] The fusion connection layer is made to fuse the corresponding features in the encoding layer with the features in the decoding layer to generate the capture strategy.
[0197] In the embodiment of the present invention, the model unit 320 is further configured to enable the deconvolution layer to upsample the high-level features to the size of the input data in the following manner:
[0198] F deconv =f′ in*T W deconv +b deconv
[0199] Among them, F′ in is a high-level feature, W deconv is the deconvolution kernel, b deconv is the deconvolution bias;
[0200] In the embodiment of the present invention, the model unit 320 is further configured to enable the fusion connection layer to fuse the corresponding features in the encoding layer with the features in the decoding layer to generate the capture strategy in the following manner:
[0201]
[0202] in, represents the pixel-by-pixel sum of features, is the output feature of the encoder layer l, Output features for the decoder layer l.
[0203] In the embodiment of the present invention, the grasping model further includes multimodal contrast loss, posture stability and grasping force coordination loss, and dynamic grasping success rate loss.
[0204] The model unit 320 is also used to obtain the multimodal contrast loss in the following manner:
[0205] L′ MMC =L MMC +λ
[0206]
[0207] λ = ReLU(s ivis ·s itor -∈)
[0208] Among them, L′ MMC is the multimodal contrast loss, N is the number of samples, M is the number of modalities, and the modalities include visual modality and moment modality, s ij is the true feature of the i-th sample under the j-th mode, is the prediction feature, s ivis is the visual modality feature, s itor is the moment modal characteristic, ∈ is the preset threshold;
[0209] The model unit 320 is also used to obtain the posture stability and grasping force coordinated loss in the following manner:
[0210] L′ SSC =L SSC +γ
[0211]
[0212] γ=ReLU((y ist -θ st )·(y iforce -θ force ))
[0213] Among them, L′ SSCis the coordinated loss of posture stability and grasping strength, y ist is the true posture stability score of the i-th sample, y iforce Score the true grasping strength of the i-th sample, is the model prediction posture stability score for the i-th sample, The model predicts the grasping strength score for the i-th sample, θ st is the attitude stability threshold, θ force is the grasping force threshold, N is the number of samples;
[0214] The model unit 320 is also used to obtain the dynamic crawling success rate in the following manner:
[0215] L′ DSC =δ·L DSC
[0216]
[0217] Among them, L′ DSC is the dynamic crawling success rate, y isucc is the true successful capture label of the i-th sample, is the model prediction probability, is the prediction error, and N is the number of samples.
[0218] The embodiments of the present invention enable the humanoid robot to grasp more accurately and in real time, and also improve the grasping success rate
[0219] An embodiment of the present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following method when executing the computer program: acquiring sensor data, wherein the sensor data includes object feature data and environmental feature data, wherein the object feature data includes target object color information, target object depth image, and target object material information; acquiring a grasping strategy based on the sensor data and a grasping model, wherein the grasping strategy includes a theoretical joint angle, a theoretical joint position, and a theoretical joint torque; acquiring an execution strategy based on the grasping strategy, wherein the execution strategy includes an execution joint angle, an execution joint position, and an execution joint torque.
[0220] An embodiment of the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the following method: acquiring sensor data, wherein the sensor data includes object feature data and environmental feature data, wherein the object feature data includes target object color information, target object depth image, and target object material information; acquiring a grasping strategy based on the sensor data and a grasping model, wherein the grasping strategy includes a theoretical joint angle, a theoretical joint position, and a theoretical joint torque; acquiring an execution strategy based on the grasping strategy, wherein the execution strategy includes an execution joint angle, an execution joint position, and an execution joint torque.
[0221] The above-mentioned humanoid robot grasping method achieves the beneficial effect of being able to solve the technical problems raised in the background technology.
[0222] Figure 2 FIG. 1 is a flow chart of a humanoid robot grasping method in one embodiment. It should be understood that although Figure 2 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 2 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0223] Figure 4 The internal structure diagram of a computer device in one embodiment is shown. The computer device may specifically be Figure 1 The server 120 in FIG. Figure 4 As shown, the computer device includes a processor, a memory, a network interface, an input device and a display screen connected through a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor can implement the humanoid robot grasping method. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor can execute the humanoid robot grasping method. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a button, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.
[0224] Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0225] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0226] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0227] The foregoing is merely a specific embodiment of the present invention, which enables those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A humanoid robot grasping method, characterized in that: The humanoid robot comprises a mechanical arm, a mechanical hand is arranged at the end of the mechanical arm, and the method comprises: Acquire sensor data, wherein the sensor data includes object feature data and environment feature data, wherein the object feature data includes target object color information, target object depth image, and target object material information; Acquire a grasping strategy according to the sensor data and the grasping model, wherein the grasping strategy includes a theoretical joint angle, a theoretical joint position, and a theoretical joint torque; According to the grasping strategy, an execution strategy is acquired, wherein the execution strategy includes an execution joint angle, an execution joint position and an execution joint torque.
2. The method according to claim 1, characterized in that Before acquiring the grasping strategy according to the sensor data and the grasping model, the method further includes: Collecting raw data, the raw data includes raw object feature data, raw environment feature data and raw robot data, the raw robot data includes: raw robot arm end posture data, raw robot hand joint angle, raw robot hand torque, raw robot hand tactile feedback data; Acquire training data according to the original data; The original network is trained according to the training data to obtain the grasping model.
3. The method according to claim 2, characterized in that The method of obtaining a grasping strategy according to the sensor data and the grasping model includes: The input layer receives input data, wherein the input data includes the object feature data and the environment feature data; The encoding layer extracts and outputs high-level features of the input data; The decoding layer generates the capture strategy after upsampling the high-level features to the size of the input data; The output layer outputs the grasping strategy.
4. The method according to claim 3, characterized in that The encoding layer includes a convolution layer, a DDF layer and a multi-head output layer, and the DDF layer includes a compressed convolution layer, an expanded convolution layer and a dense residual connection layer. The encoding layer extracts and outputs high-level features of the input data, including: The convolution layer performs a convolution operation on the input data to obtain spatial features; The DDF layer compresses, expands and reuses the spatial features; The multi-head output layer captures global dependencies through a multi-head self-attention mechanism, splices the output data of the DDF layer into the high-level features according to the global dependencies, and outputs the high-level features; The DDF layer compresses, expands and reuses the spatial features, including: The compressed convolution layer reduces the number of channels of the spatial feature according to the first convolution kernel to obtain a compressed feature. The dilated convolution layer applies a second convolution kernel with different dilation rates on the compressed features to capture multi-scale features. The dense residual connection layer connects the outputs of different layers to the dense residual connection layer to achieve feature reuse.
5. The method according to claim 4, characterized in that The convolution layer performs a convolution operation on the input data to obtain spatial features in the following manner: F conv =X*W conv +b conv Among them, F conv is the spatial feature, X is the input data of the convolutional layer, and W conv is the convolution kernel of the convolution layer, b conv is the bias term; The compressed convolution layer reduces the number of channels of the spatial feature according to the first convolution kernel to obtain the compressed feature in the following manner: F sq =F conv *W sq +b sq Among them, F sq is the compression feature, F conv is the spatial feature, W sq is the first convolution kernel, b sq is the compression bias term; The dilated convolution layer applies a second convolution kernel with different dilation rates on the compressed features to capture multi-scale features in the following manner: F dil =F sq *rW dil +b dil F dil is a multi-scale feature, F sq is the compression characteristic, r is the expansion rate, W dil is the second convolution kernel, b dil is the expansion bias term; The dense residual connection layer connects the outputs of different layers to the dense residual connection layer to achieve feature reuse in the following way: F (l) =H (l) ([F (0) ,F (1) ,...,F (l-1) ]) Among them, H (l) represents the nonlinear transformation function of the lth layer, [] represents the cascade of features, F (0) ……F (l-1) It is a multi-scale feature; The multi-head output layer captures global dependencies through a multi-head self-attention mechanism, splices the output data of the DDF layer into the high-level features according to the global dependencies, and outputs the high-level features in the following manner: MultiHead(F in )=Concat(Z1,Z2,...,Z h )IN O Z m =OFF m in, is the mth query weight matrix, is the mth key weight matrix, is the mth value weight matrix, F in is the input data, d k is the dimension of the key vector, W O is the output weight matrix, h is the number of attention heads, MultiHead means concatenation, Concat means linear transformation, and m=1,2…h.
6. The method according to claim 4, characterized in that After the decoding layer upsamples the high-level features to the size of the input data, the capture strategy is generated, including: The deconvolution layer upsamples the high-level features to the size of the input data; The fusion connection layer fuses the corresponding features in the encoding layer with the features in the decoding layer to generate the capture strategy.
7. The method according to claim 6, characterized in that The deconvolution layer upsamples the high-level features to the size of the input data in the following way: F deconv =F′ in*T W deconv +b deconv Among them, F′ in is a high-level feature, W deconv is the deconvolution kernel, b deconv is the deconvolution bias; The fusion connection layer fuses the corresponding features in the encoding layer with the features in the decoding layer to generate the capture strategy in the following manner: in, represents the pixel-by-pixel sum of features, is the output feature of the encoder layer l, Output features for the decoder layer l.
8. The method according to claim 1, characterized in that The grasping model also includes multimodal contrast loss, posture stability and grasping strength coordination loss, and dynamic grasping success rate loss. The multimodal contrast loss is obtained as follows: L′ MMC =L MMC +λ λ=ReLU(s ivis ·s itor -∈) Among them, L′ MMC is the multimodal contrast loss, N is the number of samples, M is the number of modalities, and the modalities include visual modality and moment modality, s ij is the true feature of the i-th sample under the j-th mode, is the prediction feature, s ivis is the visual modality feature, s itor is the moment modal feature, ∈ is the preset threshold; The posture stability and grasping force coordination loss are obtained in the following manner: L′ SSC =L SSC +g γ=ReLU((y ist -θ st )·(y iforce -θ force )) Among them, L′ SSC is the coordinated loss of posture stability and grasping strength, y ist is the true posture stability score of the i-th sample, y iforce Score the true grasping strength of the i-th sample, is the model prediction posture stability score for the i-th sample, The model predicts the grasping strength score for the i-th sample, θ st is the attitude stability threshold, θ force is the grasping force threshold, N is the number of samples; The dynamic crawling success rate is obtained in the following manner: L′ DSC =δ·L DSC Among them, L′ dSG is the dynamic crawling success rate, y isucc is the true successful capture label of the i-th sample, is the model prediction probability, is the prediction error, and N is the number of samples.
9. A humanoid robot grasping device, characterized in that: Applied to a humanoid robot, the humanoid robot comprises a mechanical arm, a mechanical hand is provided at the end of the mechanical arm, and the device comprises: A sensor for acquiring sensor data, wherein the sensor data includes object feature data and environment feature data, wherein the object feature data includes target object color information, target object depth image, and target object material information; A model unit, used for acquiring a grasping strategy according to the sensor data and the grasping model, wherein the grasping strategy includes a theoretical joint angle, a theoretical joint position and a theoretical joint torque; The execution unit is used to obtain an execution strategy according to the grasping strategy, and the execution strategy includes an execution joint angle, an execution joint position and an execution joint torque.
10. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.