Uncalibrated grabbing method and system for irregular parts

By using an Oriented R-CNN structure for rotating bounding box object detection and image visual servo control, the adaptability and calibration problems of grasping irregular parts in traditional methods are solved, achieving high-precision and robust grasping results.

CN120901953APending Publication Date: 2025-11-07BEIBU GULF UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511152477.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Traditional automation methods are difficult to adapt to tasks involving the grasping of irregularly shaped parts placed arbitrarily, and require recalibration after changes in the relative hand-eye position, increasing the time cost of system deployment.

Method used

A rotating bounding box grasping target detection model based on Oriented R-CNN structure is adopted. Combined with a depth camera and robot end effector, uncalibrated grasping is achieved through image visual servo control. The robot joint velocity is calculated by rotating bounding box detection and image Jacobian matrix, and the posture is iteratively adjusted.

Benefits of technology

It enables precise gripping of irregular parts, avoids tedious hand-eye calibration steps, improves gripping and positioning accuracy and success rate, and adapts to different placement postures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120901953A_ABST
    Figure CN120901953A_ABST
Patent Text Reader

Abstract

The invention discloses an irregular part uncalibrated grabbing method and system, and belongs to the technical field of mechanical control, and the method comprises the steps: image collection: based on a depth camera installed on an end effector of a robot, capturing a real-time RGB image and a depth image, and outputting image data; rotating target detection and positioning: building a rotating frame grabbing target detection model and training, and executing rotating frame grabbing target detection; servo control of image vision: extracting current image features, defining expected image features, designing an image vision servo control law, and obtaining the joint speed of the robot; grabbing is performed through robot motion execution and iteration. A closed-loop visual servo mode directly calculates errors in an image space and generates a control instruction, so that not only is a tedious hand-eye calibration step in a traditional method avoided, but also an end effector of a robot can be continuously guided to accurately and robustly approach a target grabbing pose, system errors and environmental disturbance are effectively compensated, and grabbing positioning accuracy and success rate are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of mechanical control, and particularly relates to an irregular part non-calibration grasping method and system. BACKGROUND

[0002] In modern manufacturing, robot automation sorting technology is a key means to improve production efficiency and optimize operating costs. Especially in the toy manufacturing, electronic product and other industries, there are a large number of irregular parts with different shapes and sizes that need to be quickly and accurately grasped and sorted from the disordered state. However, the traditional automation method and device have significant limitations in performing sorting tasks of multiple types of irregular targets: the traditional teaching method is difficult to adapt to any placed task scene; the method based on three-dimensional visual positioning, although it can handle scenes with arbitrary changes in pose, the grasping effect is highly dependent on the accuracy of hand-eye parameter calibration. In addition, the robot grasping system needs to be re-calibrated after the relative position of the camera and the robot changes, further increasing the time cost of system deployment. SUMMARY

[0003] The present application proposes an irregular part non-calibration grasping method and system to solve the problem that the prior art has significant limitations, is not suitable for any placed task scene, and needs to be re-calibrated after the relative position of the hand-eye changes, further increasing the time cost of system deployment.

[0004] To achieve the above purpose, the technical solution adopted by the present application is:

[0005] An irregular part non-calibration grasping method, comprising: image acquisition: based on a depth camera installed on the end effector of a robot, capturing real-time RGB images and depth images containing target parts to be grasped, and outputting current frame image data; rotating target detection and positioning: building a rotating frame grasping target detection model, training the rotating frame grasping target detection model, and performing rotating frame grasping target detection; image vision-based servo control: current image feature extraction, desired image feature definition, design of image vision servo control law, and acquisition of robot joint speed; and grasping is performed through robot motion execution and iteration.

[0006] The building rotating frame grabbing target detection model includes a network model based on an Oriented R-CNN structure, and the rotating frame grabbing target detection model is built.

[0007] The backbone network adopts EfficientnetV2 combined with a feature pyramid network, and the directional region proposal network includes a 3x3 convolution layer Two parallel 1x1 convolution layers are connected in sequence, and the convolution layers are respectively used for target classification and directional proposal box regression; the target proposal box adopts a midpoint offset representation method to describe (x, y, w, h, Delta alpha, Delta beta). Wherein (x, y) is the center point coordinate of the target horizontal outer boundary box, w and h are respectively the width and height of the horizontal outer rectangle, Delta alpha represents the horizontal offset of the top vertex of the target after rotation relative to the top midpoint of the horizontal outer boundary box, and Delta beta represents the vertical offset of the right vertex of the target after rotation relative to the right midpoint of the horizontal outer boundary box; the network structure of the rotating region of interest alignment layer is an alignment and pooling operation unit, and does not contain trainable parameters, and the core is bilinear interpolation and maximum pooling; the head network includes a full connection layer group, performs flattening processing on the fixed size feature map, and outputs features, and then the features are respectively sent into two parallel classification layers and regression layers, the classification layer processes to classify the predefined part category to which each proposal box belongs, and the regression layer regresses the parameters of the proposal box to output the parameterized offset of the final rotating boundary box description.

[0008] The training of the rotating bounding box grasping target detection model comprises: data set making: collecting a data set containing images of parts to be grasped, and performing rotating bounding box labeling on each sample instance in the data set; loss function design: adopting a multi-task loss function, wherein a total loss function comprises a directional region proposal network loss and a head loss, both of which adopt a classification loss of a cross-entropy loss function and a regression loss of a Smooth L1 loss function, wherein the classification loss of the directional region proposal network loss is used to distinguish foreground and background anchors according to an intersection over union threshold of anchors and a real target horizontal outer rectangle, and the regression loss of the directional region proposal network loss is used to learn the offset of a six-parameter directional proposal relative to an anchor; the classification loss of the head loss is used for multi-class part classification on features after rotating region of interest alignment processing, and the regression loss of the head loss is used to further optimize the accuracy of five parameters of the rotating bounding box; network training: adopting AdamW as an optimizer, calculating the gradient of the loss function on the network parameters through a back propagation algorithm during the training process, and updating the network weights; the learning rate adjustment strategy adopts a Flat-Cosine learning rate adjustment strategy, and at the last training round, the learning rate will gradually decrease in the form of a cosine function to optimize the stability in the later training period.

[0009] The current image feature extraction comprises determining and extracting four corner point coordinates of a bounding box based on parameters of the rotating bounding box; the expected image feature definition comprises predefining coordinates of corner points of a rotating bounding box of a target part in a camera image when the target part is in an ideal grasping pose, to form an expected image feature vector; the design of the image visual servo control law comprises calculating an image feature error vector based on the current image feature and the expected image feature, to construct an image Jacobian matrix, and calculating a camera speed based on the image feature error and the image Jacobian matrix, to realize convergence of the image error to the zero axis direction; the acquisition of the robot joint speed comprises converting the calculated speed command in the camera coordinate system into a speed command in the robot joint space by using a robot Jacobian matrix.

[0010] The robot motion execution and iteration comprise: sending the speed command to a robot controller to drive the robot to move for a small time step according to the speed command, and then re-executing image acquisition to obtain a new image; if a feature error norm between the new image and the expected image is greater than a preset convergence threshold, the speed command is recalculated based on the new image and executed to complete one iteration; the iteration process continues, the robot gradually adjusts the pose under the guidance of visual feedback, and the current image feature continuously approaches the expected image feature.

[0011] The execution of the grabbing comprises: in each iteration process, calculating the size of the image feature error norm, and when the error norm is less than the convergence threshold, the visual servo control process is terminated, and then the robot executes the grabbing.

[0012] An irregular part calibration-free grabbing system comprises: an image acquisition module, configured to capture real-time RGB images and depth images containing target parts to be grabbed based on a depth camera installed on a robot end effector, and output current frame image data; a rotating target detection and positioning module, configured to build a rotating box target detection model, train the rotating box target detection model, and perform rotating box target detection; an image visual servo control module, configured to extract current image features, define expected image features, design an image visual servo control law, and obtain robot joint speeds; and a control module, configured to perform grabbing through robot motion execution and iteration.

[0013] Thanks to the above technical solutions, the present application has the following advantages:

[0014] 1. The present application performs grabbing through image acquisition, rotating target detection and positioning, image visual servo control, and robot motion execution and iteration. The closed-loop visual servo method directly calculates errors in the image space and generates control instructions, thereby avoiding the cumbersome hand-eye calibration step in the traditional method, and continuously guiding the robot end effector to accurately and robustly approach the target grabbing pose, effectively compensating for system errors and environmental disturbances, and improving the grabbing positioning accuracy and success rate. BRIEF DESCRIPTION OF DRAWINGS

[0015] Figure 1 A schematic diagram of an irregular part calibration-free grabbing method proposed by the present application;

[0016] Figure 2 A schematic diagram of identifying white toy parts placed on a wood grain background proposed by the present application;

[0017] Figure 3 An image error convergence curve experimental graph in the process of grabbing by a mechanical arm proposed by the present application;

[0018] Figure 4 An overall working schematic diagram proposed by the present application. DETAILED DESCRIPTION

[0019] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0020] An irregular part calibration-free grasping method, comprising: image acquisition: based on a depth camera installed on the end effector of a robot, capturing real-time RGB images and depth images containing target parts to be grasped, and outputting current frame image data; performing rotation target detection and positioning: building a rotation box grasping target detection model, training the rotation box grasping target detection model, and performing rotation box grasping target detection; image vision-based servo control: current image feature extraction, desired image feature definition, image vision servo control law design, and robot joint speed acquisition; and grasping is performed through robot motion execution and iteration.

[0021] The building of the rotation box grasping target detection model comprises a network model based on an Oriented R-CNN structure, and the rotation box grasping target detection model comprises: a backbone network for extracting multi-level deep feature maps from input images; a directional region proposal network for obtaining the deep feature maps and generating target proposal boxes with rotation angles; a rotated region of interest alignment layer for obtaining the target proposal boxes with rotation angles and each layer feature map, processing to output fixed-size feature maps; and a head network for obtaining the fixed-size feature maps, processing to predict: a predefined part category to which each proposal box belongs, and a parameterized offset of a final rotation bounding box description.

[0022] The backbone network adopts EfficientnetV2 combined with a feature pyramid network; the directional region proposal network comprises a 3x3 convolution layer followed by two parallel 1x1 convolution layers, which are respectively used for target classification and directional proposal box regression; the target proposal box adopts a midpoint offset representation method to describe: (x, y, w, h, Delta alpha, Delta beta). Wherein (x, y) is the center point coordinate of the horizontal outer bounding box of the target, w and h are respectively the width and height of the horizontal outer rectangle, Delta alpha represents the horizontal offset of the top vertex of the target after rotation relative to the top midpoint of the horizontal outer bounding box, and Delta beta represents the vertical offset of the right vertex of the target after rotation relative to the right midpoint of the horizontal outer bounding box; the network structure of the rotated region of interest alignment layer is an alignment and pooling operation unit, which does not contain trainable parameters, and the core is bilinear interpolation and maximum pooling; the head network comprises a full connection layer group, which performs flattening processing on the fixed-size feature maps, and then outputs the features into two parallel classification layers and regression layers, the classification layer processes to classify the predefined part category to which each proposal box belongs, and the regression layer regresses the parameters of the proposal box to output the parameterized offset of the final rotation bounding box description.

[0023] The training of the rotating bounding box grasping target detection model comprises: data set making: collecting a data set containing images of parts to be grasped, and performing rotating bounding box labeling on each sample instance in the data set; loss function design: adopting a multi-task loss function, wherein a total loss function comprises a directional region proposal network loss and a head loss, and both the classification loss of the cross-entropy loss function and the regression loss of the Smooth L1 loss function are adopted, wherein the classification loss of the directional region proposal network loss is used to distinguish foreground and background anchors according to an intersection over union threshold of anchors and a real target horizontal outer rectangle, and the regression loss of the directional region proposal network loss is used to learn the offset of a six-parameter directional proposal relative to an anchor; the classification loss of the head loss is used for multi-class part classification on the features after rotating region of interest alignment processing, and the regression loss of the head loss is used to further optimize the accuracy of the five parameters of the rotating bounding box; network training: adopting AdamW as an optimizer, calculating the gradient of the loss function on the network parameters through a back propagation algorithm during the training process, and updating the network weights; the learning rate adjustment strategy adopts a Flat-Cosine learning rate adjustment strategy, and at the last training round, the learning rate will gradually decrease in the form of a cosine function to optimize the stability in the later training period.

[0024] The current image feature extraction comprises determining the four corner point coordinates of the bounding box based on the parameters of the rotating bounding box; the expected image feature definition comprises predefining the coordinates of the corner points of the rotating bounding box of the target part in the camera image when the target part is in an ideal grasping pose, to form an expected image feature vector; the design of the image visual servo control law comprises calculating an image feature error vector based on the current image feature and the expected image feature, to construct an image Jacobian matrix, and calculating a camera speed based on the image feature error and the image Jacobian matrix, to realize the convergence of the image error to the zero axis direction; the acquisition of the robot joint speed comprises converting the speed command in the camera coordinate system into a speed command in the robot joint space by using a robot Jacobian matrix.

[0025] The robot motion execution and iteration comprise: sending the speed command to a robot controller to drive the robot to move for a small time step according to the speed command, and then re-executing image acquisition to obtain a new image; if the feature error norm between the new image and the expected image is greater than a preset convergence threshold, the speed command is recalculated based on the new image and executed to complete one iteration; the iteration process continues, and the robot gradually adjusts the pose under the guidance of visual feedback, and the current image feature continuously approaches the expected image feature.

[0026] The execution of the grabbing comprises: in each iteration process, calculating the size of the image feature error norm, and when the error norm is less than the convergence threshold, the visual servo control process is terminated, and then the robot executes the grabbing.

[0027] An irregular part calibration-free grabbing system comprises: an image acquisition module for capturing real-time RGB images and depth images containing target parts to be grabbed based on a depth camera installed on the end effector of a robot, and outputting current frame image data; a rotating target detection and positioning module for building a rotating box grabbing target detection model, training the rotating box grabbing target detection model, and performing rotating box grabbing target detection; an image visual servo control module for current image feature extraction, expected image feature definition, image visual servo control law design, and robot joint speed acquisition; and a control module for robot motion execution and iteration to execute grabbing.

[0028] The present application provides a kind of robot irregular part grabbing method without hand-eye calibration, utilizes rotating target detection network to accurately identify part and output rotating boundary box, can accurately obtain the position and arbitrary rotation angle of part in image plane, improve the adaptability of system to part posture. By taking the corner point of rotating box as the general visual servo key point, the error term between the current image feature and the expected image feature corresponding to the target grabbing pose is constructed, and the image Jacobian matrix and PD control algorithm are combined, and the error is mapped into the motion speed instruction of robot in real time. This closed-loop visual servo mode directly calculates error in image space and generates control instruction, not only avoids the cumbersome hand-eye calibration step in traditional method, but also can continuously guide the end effector of robot to accurately and robustly approach target grabbing pose, effectively compensates system error and environmental disturbance, improves grabbing positioning precision and success rate.

[0029] Specifically, as Figure 1 The robot irregular part grabbing method without hand-eye calibration comprises:

[0030] S1: RGB-D image acquisition

[0031] Start the depth camera (RGB-D camera) arranged near the end effector of the robot or fixedly installed, capture real-time RGB images and depth images containing target parts to be grabbed, and transmit the current frame image data to the processing unit.

[0032] S2: rotating target detection and positioning

[0033] S2.1: build a rotating box grabbing target detection model

[0034] The rotating box grabbing target detection model is specifically a network model based on Oriented R-CNN structure, which is composed of the following modules:

[0035] The backbone network adopts EfficientnetV2 combined with Feature Pyramid Network (FPN) to extract multi-level deep feature maps P = {P2, P3, P4, P5} from the input image, and FPN can fuse features of different levels to adapt to target detection of different sizes.

[0036] The oriented region proposal network (Oriented RPN) acts on each layer of feature maps P output by the FPN i , which contains a 3x3 convolution layer followed by two parallel 1x1 convolution layers for target classification and oriented proposal box regression, respectively. Define the shared intermediate feature The classification branch outputs the foreground score map where, The output channel number of is A, that is, the number of anchors at each anchor position. The regression branch outputs a six-parameter offset map where, The output channel number of is 6A. Each value in is passed through the Sigmoid activation to obtain the foreground probability p obj , The six values in correspond to the six-parameter offset of each anchor point δ = (δ x , δ y , δ w , δ h , δ Δα , δ Δβ ).

[0037] The RPN network module receives the feature maps from the FPN to generate target proposal boxes with rotation angles. The target proposal box uses the midpoint offset representation method to describe: (x, y, w, h, Δα, Δβ). Where (x, y) is the center point coordinate of the horizontal bounding box of the target, w and h are the width and height of the horizontal bounding box, respectively, Δα represents the horizontal offset of the top point of the target after rotation relative to the top midpoint of the horizontal bounding box, and Δβ represents the vertical offset of the right top point of the target after rotation relative to the right midpoint of the horizontal bounding box. The above six parameters can uniquely determine the four vertices of the rotated bounding box.

[0038] Rotated Region of Interest Alignment Layer: This module receives the target proposal box R prop with a rotation angle generated by the Oriented RPN and the feature map P i. Its network structure is an alignment and pooling operation unit, which does not contain trainable parameters, and its core is bilinear interpolation and max pooling. It can accurately extract fixed-size feature maps from irregularly shaped orientation proposal regions. pool while preserving the rotation invariance. This is achieved by converting the six-parameter proposals output by the RPN into standard five-parameter rotated box representations R obb = (x c , y c , w obb , h obb , θ obb ) and then performing an alignment pooling operation. For each cell (m, n) and each channel c of the output feature map F pool , its value is calculated as follows:

[0039]

[0040] where Pool max represents the max pooling operation, and SampleGrid(m, n, R obb ) represents a set of sampling points (u i , v s ) generated for the output cell (m, n) on the input feature map P s . These sampling points are uniformly distributed according to the position, size, and angle θ obb of the rotated box R obb . represents the coordinate transformation considering the rotation θ s for the sampling points (u s , v obb ), and bilinear interpolation is used to obtain sub-pixel accuracy feature values from P i .

[0041] Head network: receives the fixed-size feature map F pool output by the rotated region of interest alignment layer. The head network consists of a fully connected layer FC1, which performs flattening processing on F pool , and outputs features that are then sent to two parallel classification layers FC cls and regression layers FC reg respectively. The classification layer predicts which predefined part category each proposal box belongs to, mathematically represented as follows:

[0042]

[0043] where the output dimension of FC cls is C+1, i.e., C part categories plus 1 background category.

[0044] The final class probability is calculated by the Softmax function:

[0045]

[0046] wherein, represents the probability that the proposal box belongs to the kth class.

[0047] The regression layer also receives the output features from FC1, and regresses the parameters of the proposal box, and outputs the parameterized offset Δ' of the final rotating bounding box description param , which is mathematically expressed as follows:

[0048] Δ' param = FC reg (Flatten(F pool ));

[0049] wherein, the output dimension of FC reg is 5xC, that is, Δ' param = (Δ' x , Δ' y , Δ' w , Δ' h , Δ' θ ). The final rotating bounding box parameters (x, y, w, h, θ) are obtained by the following operation:

[0050]

[0051] wherein, % is the modulus operation, (x, y) is the center point pixel coordinate of the rotating bounding box in the image coordinate system, w and h are the width and height of the rotating bounding box respectively, and θ is the angle of the long side of the rotating bounding box relative to the positive direction of the x-axis of the image coordinate system.

[0052] S2.2: Training of the rotating box grasping target detection model

[0053] The rotating box grasping target detection network built in S2.1 is trained, and the specific steps are as follows:

[0054] (1) Data set preparation: a data set containing images of parts to be grasped is prepared. Each sample instance in the data set is labeled with a rotating bounding box, including the center point coordinates (x gt , y gt ) of the rotating box, the width w gt , the height h gt , and the rotation angle θ gt .

[0055] (2) Loss function design: in the present application, the training of the rotating box grasping target detection model adopts a multi-task loss function, and the total loss function L totalThe oriented region proposal network loss is composed of two parts: a head loss. The oriented region proposal network loss includes: 1) a classification loss L cls_rpn , which adopts a cross-entropy loss function and distinguishes foreground and background anchors according to an intersection-over-union threshold of the anchor and the real target horizontal outer rectangle, and in the present application, an intersection-over-union greater than 0.7 is set as a positive sample and an intersection-over-union less than 0.3 is set as a negative sample; and 2) a regression loss λ1L reg_rpn , which adopts a Smooth L1 loss function and is used to learn the offset of the predicted six-parameter oriented proposal (δx, δy, δw, δh, δα, δβ) relative to the anchor. The head loss also includes: 1) a classification loss L cls_head , which adopts a cross-entropy loss function and performs multi-class part classification on the features of the aligned rotated region of interest; and 2) a regression loss L reg_head , which adopts a Smooth L1 loss function and is used to further optimize the accuracy of the five parameters of the rotated bounding box, and is expressed as:

[0056]

[0057] The final total loss function L total is a weighted sum of the above four loss parts, and the specific form is as follows:

[0058] L total =L cls_rpn +λ1L reg_rpn +L cls_head +λ2L reg_head ;

[0059] Wherein, λ1 and λ2 are preset weight coefficients for balancing the contribution of different task losses.

[0060] (3) Network training: AdamW is used as the optimizer, and in the training process, the gradient of the loss function to the network parameters is calculated through the back propagation algorithm, and the network weights are updated. The network is trained for 500 rounds, the initial learning rate is set to 0.001, the optimizer momentum is set to 0.9, and the weight decay is set to 0.05. The batch size of the training is set to 64, and 1000 warm-up iterations are set at the beginning of the training to stabilize the initial training process. The learning rate adjustment strategy adopts the Flat-Cosine learning rate adjustment strategy, and the learning rate will gradually decrease according to the cosine function form from the last 50 rounds of training, in order to optimize the stability in the later training.

[0061] S2.3: Perform rotated box grasp target detection

[0062] The real-time RGB image frame captured in the S1 step is input into the trained rotated box grasp target detection network to obtain the category of the part and the prediction information (x, y, w, h, θ) of the rotated bounding box of the part in the image plane.

[0063] As Figure 2 The rotated bounding box in white detects a white toy part placed on a wood grain background. The white oblique rectangular box in the figure is the rotated bounding box predicted by the model, which closely fits the outline of the target part and accurately indicates the specific position, size and rotation angle of the part in the image.

[0064] S3: Image visual servoing control. Based on the depth information of S1 and the detection results of S2, an image visual servoing control module for a 6-DOF serial robot is constructed to realize the transformation of target image point coordinates-robot joint speed. Specifically:

[0065] S3.1: Current image feature extraction

[0066] Based on the output target rotated bounding box information (x, y, w, h, θ) in S2, the four corner coordinates (x1, y1, x2, y2, x3, y3, x4, y4) of the bounding box are extracted, as follows:

[0067]

[0068] S3.2: Desired image feature definition

[0069] The coordinates of the corner points of the rotated bounding box of the target part in the camera image when the target part is in the ideal grasping pose are defined in advance to form the desired image feature vector This desired feature s * represents the desired image s of the target in the camera when the robot end effector reaches the ideal pre-grasping state relative to the target part. * Through one-time teaching during system deployment, the robot is manually moved to the ideal grasping position of the part, the image is captured and the rotated box corner coordinates at this time are extracted as s * , which remains unchanged thereafter. *

[0070] S3.3: Design of image visual servoing control law

[0071] Based on the current image feature s and the desired image feature s * , the image feature error vector e = s-s * is calculated.

[0072] The image Jacobian matrix L s is constructed as follows:

[0073]

[0074] ​where f is the camera focal length provided by the camera manufacturer, and Z is the distance from the camera to the target in the camera coordinate system, which is directly acquired by the depth camera. The image Jacobian matrix L s The camera velocity v in the camera coordinate system is described as c The linear relationship between the image feature rate of change , that is

[0075] Based on the image feature error e and the estimated image Jacobian matrix L s , the camera velocity is solved using the proportional control rate to achieve the convergence of the image error e to the zero axis direction, and the formula is as follows:

[0076] where λ = 0.3 is a positive scalar gain, is the pseudo-inverse of L s , and is the image feature error rate of change.

[0077] S3.4: Obtain the robot joint velocity. The calculated velocity command v in the camera coordinate system c is converted into the velocity command robot in the robot joint space using the robot Jacobian matrix J

[0078] S4: Robot motion execution and iteration

[0079] The robot velocity command calculated in step S3.4 is sent to the robot controller. The robot controller drives the robot to move according to the velocity command for a small time step. After the movement, return to step S1 to collect a new image and calculate the error norm ||e|| between the new image feature and the expected image feature. When the error norm is greater than the preset convergence threshold ∈ = 0.01, repeat the closed-loop control process of S2 to S4. This iterative process continues, and the robot will gradually adjust the pose under the guidance of visual feedback, so that the current image feature s continuously approaches the expected image feature s * .

[0080] As Figure 3 is the image error convergence curve experimental diagram in the robot grasping process. The figure shows the convergence curve of the image feature error in the visual servo process. The multiple curves (x1, y1 to x4, y4) in the figure represent the coordinate errors of each feature point. Under the action of the control law, all errors are smoothly and quickly converged to zero from the initial value, and are stably around zero in about 10 seconds. This indicates that the robot end has accurately reached the target pose, verifying that the visual servo control method proposed in the invention has good convergence, stability and high positioning accuracy.

[0081] S5: Capture and Execute

[0082] In each iteration of S4, the magnitude of the image feature error norm ||e|| is calculated. When the error norm is less than the preset convergence threshold ∈ = 0.01, it is considered that the robot end effector has reached the ideal grasping pose that is sufficiently close to the target part, and the visual servo control process terminates. At this time, a grasping command is sent to the robot system to control the gripper to close and complete the grasping operation of the target part.

[0083] like Figure 4 The overall working diagram shown includes the cabinet, i.e. the processing system, the robotic arm (the end of the robotic arm is equipped with RGB-D and a clamping structure), and the worktable.

[0084] The above description is a detailed description of the preferred embodiments of the present invention. However, the embodiments are not intended to limit the scope of the patent application of the present invention. All equivalent changes or modifications made under the technical spirit of the present invention should fall within the patent scope covered by the present invention.

Claims

1. An irregular part calibration-free grasping method, characterized in that, The method comprises the following steps: Image acquisition: based on a depth camera installed on the end effector of a robot, real-time RGB images and depth images containing target parts to be grabbed are captured, and current frame image data is output; Rotating target detection and positioning: a rotating box grabbing target detection model is built, the rotating box grabbing target detection model is trained, and rotating box grabbing target detection is performed; Image vision-based servo control: current image feature extraction, desired image feature definition, image vision servo control law design, and robot joint speed acquisition; The method is executed by robot motion execution and iteration to perform grabbing.

2. The irregular part calibration-free grasping method according to claim 1, characterized in that, The building of the rotating box grabbing target detection model comprises a network model based on an Oriented R-CNN structure. The rotating box grabbing target detection model comprises: a backbone network for extracting multi-level deep feature maps from input images; a directional region proposal network for obtaining the deep feature maps and generating target proposal boxes with rotation angles; a rotated region of interest alignment layer for obtaining the target proposal boxes with rotation angles and the feature maps of each layer, and processing to output fixed-size feature maps; a head network for obtaining the fixed-size feature maps and processing to predict: the predefined part category to which each proposal box belongs, and the parameterized offset of the final rotating bounding box description.

3. The irregular part calibration-free grasping method according to claim 2, wherein, The backbone network adopts EfficientnetV2 combined with a feature pyramid network; The orientation region proposal network comprises a 3x3 convolution layer followed by two parallel 1x1 convolution layers for target classification and orientation proposal box regression, respectively; The target proposal box adopts a midpoint offset representation method to describe: (x, y, w, h, Δα, Δβ). Wherein (x, y) is the center point coordinate of the horizontal bounding box of the target, w and h are the width and height of the horizontal bounding box respectively, Δα represents the horizontal offset of the top vertex of the target after rotation relative to the top midpoint of the horizontal bounding box, and Δβ represents the vertical offset of the right vertex of the target after rotation relative to the right midpoint of the horizontal bounding box; The network structure of the rotating region of interest alignment layer is an alignment and pooling operation unit, which does not contain trainable parameters, and its core is bilinear interpolation and maximum pooling; The head network comprises a fully connected layer group, which performs flattening processing on the fixed-size feature maps, and outputs features which are then sent into two parallel classification layers and regression layers respectively. The classification layer processes to classify the predefined part category to which each proposal box belongs, and the regression layer regresses the parameters of the proposal box to output the parameterized offset of the final rotating bounding box description.

4. The irregular part calibration-free grasping method according to claim 3, wherein, The training of the rotating box grabbing target detection model comprises: Dataset preparation: a dataset containing images of parts to be grabbed is collected, and each sample instance in the dataset is labeled with a rotating bounding box; Loss function design: a multi-task loss function is adopted, and a total loss function includes a directional region proposal network loss and a head loss, both of which adopt a classification loss of a cross-entropy loss function and a regression loss of a Smooth L1 loss function, wherein the classification loss of the directional region proposal network loss is used to distinguish foreground and background anchors according to an intersection over union threshold of an anchor point and a real target horizontal outer rectangle, and the regression loss of the directional region proposal network loss is used to learn an offset of a six-parameter directional proposal relative to an anchor point; the classification loss of the head loss is used for multi-class part classification of features after rotation region of interest alignment processing, and the regression loss of the head loss is used to further optimize the accuracy of five parameters of the rotated bounding box; Network training: AdamW is used as an optimizer, in the training process, the gradient of the loss function to the network parameters is calculated by the back propagation algorithm, and the network weight is updated; the learning rate adjustment strategy adopts a Flat-Cosine learning rate adjustment strategy, and in the last training round, the learning rate will gradually decrease in the form of a cosine function, so as to optimize the stability in the later training.

5. The irregular part calibration-free grasping method according to claim 4, wherein, The current image feature extraction includes determining and extracting four corner point coordinates of the bounding box based on the parameters of the rotated bounding box; The expected image feature definition includes predefining the coordinates of the corner points of the rotated bounding box of the target part in the camera image when the target part is in an ideal grasping pose, to form an expected image feature vector; The image visual servo control law design includes calculating an image feature error vector based on the current image feature and the expected image feature, constructing an image Jacobian matrix, and calculating a camera speed using a proportional control rate based on the image feature error and the image Jacobian matrix, so as to realize the convergence of the image error to the zero axis direction; The robot joint speed acquisition includes converting the calculated speed command in the camera coordinate system into a speed command in the robot joint space by using a robot Jacobian matrix.

6. The irregular part calibration-free grasping method according to claim 5, wherein, The robot motion execution and iteration include: sending the speed command to a robot controller to drive the robot to move for a small time step according to the speed command, and then re-executing image acquisition to obtain a new image; if the feature error norm between the new image and the expected image is greater than a preset convergence threshold, the speed command is recalculated based on the new image and executed to complete one iteration; the iteration process continues, and the robot gradually adjusts the pose under the guidance of visual feedback, and the current image feature continuously approaches the expected image feature.

7. The irregular part calibration-free grasping method according to claim 6, wherein, The execution of the grasp includes: in each iteration process, the size of the image feature error norm is calculated, and when the error norm is less than the convergence threshold, the visual servo control process is terminated, and then the robot executes the grasp.

8. An irregular part no calibration grasping system, characterized in that, The image acquisition module is configured to capture real-time RGB images and depth images containing a target part to be grasped based on a depth camera installed on a robot end effector, and output current frame image data. The image acquisition module is configured to capture real-time RGB images and depth images containing a target part to be grasped based on a depth camera installed on a robot end effector, and output current frame image data. A rotating target detection and positioning module is configured to build a rotating box target detection model, train the rotating box target detection model, and execute rotating box target detection. An image visual servo control module is configured to extract current image features, define expected image features, design an image visual servo control law, and obtain robot joint speeds. A control module is configured to execute grasping through robot motion execution and iteration.

Citation Information

Cited By

  • Robot action planning method and device, terminal and medium

    CN121374638A