Photovoltaic junction box grabbing pose estimation method and system based on instance segmentation

Through the example segmentation method, the lightweight grab instance segmentation network model FG-YOLO is trained, combined with depth images and confidence screening candidate mask maps, the two-dimensional grab point and curvature direction are extracted, the three-dimensional grab pose matrix is ​​calculated and optimized, which solves the problem of low success rate in the photovoltaic junction box scenario, and realizes efficient and accurate grab pose estimation.

CN120219743APending Publication Date: 2025-06-27CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510311491.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The traditional three-dimensional grasping algorithm has a low success rate in photovoltaic junction box grabbing scenarios, especially when objects are highly overlapped and severely obstructed, it is difficult to accurately estimate the grasping position.

Method used

Using an instance segmentation method, the instance mask diagram and its confidence are obtained by collecting RGB images and depth images of the photovoltaic junction box, constructing a data set and training a lightweight grabbing instance segmentation network model FG-YOLO. Then, the candidate mask map is screened using depth images and confidence, and the two-dimensional grab points and curvature direction are extracted by a refinement algorithm, the three-dimensional grab pose matrix is ​​calculated, and the position of the photovoltaic junction box in the material box is optimized.

Benefits of technology

The success rate of photovoltaic junction box grabbing is improved, and the problem of inaccurate estimation of traditional methods in complex scenarios is solved, thereby achieving accurate and efficient collision-free grabbing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219743A_ABST
    Figure CN120219743A_ABST
Patent Text Reader

Abstract

The invention provides a photovoltaic junction box grabbing pose estimation method and system based on instance segmentation, and relates to the field of robot control, and the method comprises the steps: collecting an RGB image and a depth image, and constructing a data set; constructing a lightweight capture instance segmentation network model FG-YOLO and performing training to obtain an instance mask map and a confidence coefficient thereof; screening the instance mask patterns, and determining candidate mask patterns; extracting two-dimensional capture points and curvature directions of the candidate mask patterns by using a refinement algorithm; a three-dimensional grabbing pose matrix is calculated according to the two-dimensional grabbing points and the curvature directions of the two-dimensional grabbing points; the three-dimensional grabbing pose matrix is optimized according to the position of the photovoltaic junction box in the material frame; and controlling the robot to grab the photovoltaic junction box through the optimized three-dimensional grabbing pose matrix. According to the technical scheme, the optimal grabbing posture can be selected under the condition that the photovoltaic junction boxes are stacked in a scattered mode, and the grabbing success rate is greatly increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of robot control, and in particular to a method and system for estimating the grasping posture of a photovoltaic junction box based on instance segmentation. Background Art

[0002] In recent years, with the rapid development of industrial automation and intelligent robotics, the automatic pick-and-place operations of vision-guided robots have been widely used in material handling and assembly tasks in the manufacturing industry, greatly improving production efficiency and precision. However, for photovoltaic junction boxes such as photovoltaic junction boxes in unstructured environments, providing robots with accurate grasping posture information to ensure that they can complete tasks efficiently and accurately is still a problem that needs to be studied and solved in depth. There are two types of grasping pose estimation: grasping pose estimation based on 2D vision and grasping pose estimation based on 3D vision. The grasping pose estimation method based on 2D vision uses techniques such as template matching and target detection to guide the robot to complete the grasping operation by calculating the 2D coordinate information of the target in the grayscale image or RGB image. This method is mainly used for vertical grasping operations in simple scenes. Although grasping pose estimation based on 2D vision performs well in segmenting and locating workpieces, its application range is limited to simple robot grasping tasks due to the lack of depth information.

[0003] The grasping pose estimation method based on 3D vision usually uses a structured light camera to collect 3D point clouds, segments and clusters the scene point clouds through a point cloud segmentation algorithm, and then aligns them with the point cloud template of the workpiece to achieve grasping pose estimation.

[0004] Traditional grasping pose estimation methods based on point cloud segmentation and registration have shown excellent performance in grasping rigid objects, but they still have limitations in more complex scenarios. First, the high overlap and severe occlusion of objects will significantly weaken the feature details of the 3D point cloud, resulting in poor results in the point cloud processing algorithm. Second, the deformability of photovoltaic junction boxes (such as photovoltaic junction boxes) further increases the difficulty of point cloud registration, making it difficult to apply in the automatic feeding link of photovoltaic junction boxes. Summary of the invention

[0005] The purpose of the present invention is to provide a method and system for photovoltaic junction box grasping pose estimation based on instance segmentation in order to solve the problem of low success rate of traditional three-dimensional grasping algorithms in photovoltaic junction box grasping scenarios.

[0006] The above-mentioned purpose of the present application is achieved through the following technical solutions: S1: Collect RGB images and depth images of photovoltaic junction boxes in different stacking conditions; construct a data set through RGB images; S2: Construct a lightweight grasping instance segmentation network model FG-YOLO and train it with a dataset to obtain instance mask maps and their confidence levels; S3: Screen the instance mask maps based on the depth image and the confidence level to determine candidate mask maps; S4: Use a refinement algorithm to extract the 2D grasping points of the candidate mask maps and their curvature directions; S5: Calculate the 3D grasping pose matrix from the 2D grasping points and their curvature directions; optimize the 3D grasping pose matrix according to the position of the photovoltaic junction box in the material box; control the robot to grasp the photovoltaic junction box by using the optimized 3D grasping pose matrix.

[0007] Optionally, step S2 includes: Use a structured light camera integrated with an RGB camera to collect the RGB image and depth image of the photovoltaic junction box; the structured light camera includes: an RGB camera and a depth camera; Map and perform coordinate transformation on the depth image through the external parameters of the structured light camera to align the depth image to the RGB camera coordinate system; Obtain the RGB images of the photovoltaic junction box under different stacking conditions, and use the annotation tool Labelme to annotate the graspable areas of the photovoltaic junction box in the RGB images; Perform data augmentation on the annotated RGB images to obtain the final dataset; The data augmentation includes: scaling, rotation and translation, and contrast transformation.

[0008] Optionally, step S2 includes: The lightweight grasping instance segmentation network model FG-YOLO includes: a YOLOv8 backbone network module, a lightweight fusion LOD module, and a segmentation head module adjusted based on a dynamic mechanism; The YOLOv8 backbone network module is connected to the lightweight fusion LOD module; The lightweight fusion LOD module is connected to the segmentation head module adjusted based on a dynamic mechanism; The YOLOv8 backbone network module is used to extract the main features of the image; The lightweight fusion LOD module is used to fuse the extracted main features; Use the ODConv module to replace the C2f module and the ordinary Conv module in the neck of the YOLOv8 model to construct the lightweight fusion LOD module; The segmentation head module adjusted based on a dynamic mechanism includes: a DMA module and a Scale layer.

[0009] Optionally, the loss function of the lightweight grasping instance segmentation network model FG-YOLO includes: The target box regression loss function is as follows:

[0010]

[0011]

[0012] Wherein, represents the target box regression loss; represents the Inner-SIoU loss based on the auxiliary bounding box; represents the smooth L1 loss function; and represent the scaling factor; represents the SIoU loss; represents the intersection over union; represents the intersection over union inside the auxiliary bounding box; represents the distance loss; represents the shape loss; The smooth L1 loss function is defined as follows:

[0013] Wherein, x is the coordinate deviation between the ground truth box and the predicted box; The calculation formula of

[0014]

[0015]

[0016]

[0017]

[0018] Wherein, A represents the predicted box, B represents the ground truth box; represents the scale factor; ( , ) represents A the center coordinates; ( , ) represents A the width and height of , ) represents A the upper left corner coordinates; ( , ) represents A the lower right corner coordinates; ( , ) represents B the central coordinates; ( , ) represents B the width and height of; ( , ) represents B the upper left corner coordinates; ( , ) represents B the lower right corner coordinates; Inter represents A and B the intersection of, Union represents A and B the union of.

[0019] Optionally, step S3 includes: S31: Remove the instance mask images with confidence lower than the preset threshold; S32: Use the depth information of the depth image to calculate the comprehensive score of each instance mask image, as follows:

[0020] where S represents the comprehensive score, C represents the confidence, H represents the depth information of the mask, u and v represent the corresponding weights; S33: Select a preset number of masks as candidate mask images through the comprehensive scores of each instance mask image.

[0021] Optionally, step S4 includes: S41: Extract the central skeleton of the candidate mask image through the refinement algorithm; S42: Sort the central skeleton through the nearest neighbor algorithm to obtain an ordered skeleton; S43: In the ordered skeleton, combine the preset step size to extract multiple equally spaced pixels as candidate grasping points; calculate the curvature direction of each candidate grasping point in the local skeleton; The curvature direction is the closing direction of the robot gripper in the pixel coordinate system; S44: Obtain the size parameters of the robot gripper; calculate the rectangular area passed by the two fingers of the gripper closing in the depth image through the candidate grasping points, closing direction, and size parameters; S45: Judge the interference degree of the rectangular area through the depth of the grasping point and the diameter of the photovoltaic junction box; calculate the score of each candidate grasping point according to the interference degree of the rectangular area; S46: Compare the scores of each candidate grasping point and select the candidate grasping point with the highest score as the two-dimensional grasping point of the candidate mask image.

[0022] Optionally, step S5 includes: S51: Calculate the three-dimensional coordinates of the two-dimensional grasping point mapped into the RGB camera coordinate system; Map the unit vector along the curvature direction with the two-dimensional grasping point as the starting point into a three-dimensional vector in the RGB camera coordinate system; Let the two-dimensional grasping point coordinates be and the depth value corresponding to the coordinates in the depth image be h, and the three-dimensional coordinates of the two-dimensional grasping point are calculated as follows:

[0023]

[0024]

[0025] where and represent the pixel focal lengths in the x and y directions, and represent the principal point coordinates; Let the unit vector along the curvature direction be , and construct a three-dimensional vector parallel to the camera plane as follows:

[0026]

[0027]

[0028] This three-dimensional vector represents the direction in which the gripper closes in the RGB camera coordinate system; S52: Cross-multiply the three-dimensional vector with the three-dimensional unit vector parallel to the z-axis of the RGB camera coordinate system to obtain the three-dimensional vector , and form the following rotation matrix from the three three-dimensional vectors:

[0029] Together with the rotation matrix and the three-dimensional coordinates of the two-dimensional grasping point, form the grasping pose matrix in the camera coordinate system:

[0030] S53: Through morphological processing, obtain the edge information of the material frame in the RGB image; Based on the edge information, calculate the distance between the two-dimensional grasping point and the nearest edge; If the distance is less than the preset threshold, optimize the calculated grasping pose matrix; Let the unit vector starting from the two-dimensional grasping point and pointing to the center of the material box be ; Construct a three-dimensional vector parallel to the camera plane , combined with the three-dimensional vector parallel to the camera plane, construct a pose matrix in the RGB camera coordinate system ; The homogeneous transformation matrix tilted towards the center of the material box is:

[0031] Among them, is the tilt angle; The final optimized grasping pose matrix of the gripper in the RGB camera coordinate system is: ; S54: Convert the grasping pose matrix from the RGB camera coordinate system to the robot coordinate system; Generate a transition pose matrix at a preset distance above along the Z direction of the grasping pose matrix; Control the robot to grasp the photovoltaic junction box through the grasping pose matrix and the transition pose matrix after coordinate transformation.

[0032] A photovoltaic junction box grasping pose estimation system based on instance segmentation, the system includes: a structured light camera, a computer module, and a robot; The robot includes: an electric parallel gripper, a robot control cabinet, and a six-axis robotic arm; The structured light camera is connected to the computer module; The computer module is connected to the robot; The structured light camera is used to collect RGB images and depth images of the photovoltaic junction box under different stacking conditions; The computer module is used to construct a dataset through the RGB images; The computer module is also used to construct a lightweight grasping instance segmentation network model FG-YOLO and train it through the dataset to obtain an instance mask image and its confidence; The computer module is also used to screen the instance mask image through the depth image and the confidence to determine the candidate mask image; The computer module is also used to use a refinement algorithm to extract the two-dimensional grasping point and its curvature direction of the candidate mask image; The computer module is also used to calculate a three-dimensional grasping pose matrix from the two-dimensional grasping point and its curvature direction; optimize the three-dimensional grasping pose matrix according to the position of the photovoltaic junction box in the material box; The computer module is also used to generate a pose control instruction by using the optimized three-dimensional grasping pose matrix, and control the robot to grasp the photovoltaic junction box; The robot is used to receive the pose control instruction and grasp the photovoltaic junction box.

[0033] The beneficial effects brought by the technical solution provided by this application are as follows: 1. Design a lightweight and efficient instance segmentation network model FG-YOLO, which adopts a lightweight fusion LOD module and a segmentation head module with dynamic mechanism adjustment, and improves the target box regression loss function, making the image segmentation speed more efficient and accurate. This model is improved on the basis of YOLOv8 and can efficiently and accurately identify and locate photovoltaic junction boxes in different states in RGB images.

[0034] 2. According to the morphological characteristics of the photovoltaic junction box mask and combined with the depth image, a highly robust grasping point selection and grasping pose generation method is designed. According to the segmentation result and the characteristics of the junction box, a suitable image processing method is used to select the best 2D grasping point, and then a 3D grasping pose is constructed. And according to the position relationship between the grasping point and the material box, the grasping pose is optimized to achieve accurate and efficient collision-free grasping. Solve the problem that it is difficult to calculate the best grasping pose by traditional methods when the photovoltaic junction boxes are stacked in a highly scattered manner.

[0035] 3. Optimize the three-dimensional grasping pose matrix according to the position of the photovoltaic junction box in the material box, solve the potential collision problem between the robot gripper and the material box for placing the photovoltaic junction box in the actual scenario, and can effectively avoid the collision between the robot and the material box. Description of the Drawings

[0036] The following will further illustrate this application in conjunction with the drawings. In the drawings: Figure 1 is the step diagram in the embodiment of this application; Figure 2 is the structure diagram in the embodiment of this application; Figure 3 is the model structure diagram of the lightweight grasping instance segmentation network model FG-YOLO in the embodiment of this application; Figure 4 is the loss parameter diagram in the embodiment of this application; Figure 5 is the 2D grasping point effect diagram in the embodiment of this application; Figure 6 is the schematic diagram of the refinement algorithm in the embodiment of this application. Detailed Embodiment

[0037] To have a clearer understanding of the technical features, objectives, and effects of this application, the specific implementation manners of this application will now be described in detail with reference to the accompanying drawings.

[0038] An embodiment of this application provides a method for estimating the grasping pose of a photovoltaic junction box based on instance segmentation.

[0039] Please refer to Figure 1 , Figure 1 which is a step diagram of a method for estimating the grasping pose of a photovoltaic junction box based on instance segmentation in an embodiment of this application, including: S1: Collect RGB images and depth images of the photovoltaic junction box under different stacking conditions; construct a dataset through the RGB images; S2: Construct a lightweight grasping instance segmentation network model FG-YOLO and train it through the dataset to obtain an instance mask map and its confidence; S3: Screen the instance mask map through the depth image and the confidence to determine the candidate mask map; S4: Use a refinement algorithm to extract the two-dimensional grasping points of the candidate mask map and their curvature directions; S5: Calculate the three-dimensional grasping pose matrix from the two-dimensional grasping points and their curvature directions; optimize the three-dimensional grasping pose matrix according to the position of the photovoltaic junction box in the material box; control the robot to grasp the photovoltaic junction box by using the optimized three-dimensional grasping pose matrix.

[0040] Step S2 includes: Use a structured light camera integrated with an RGB camera to collect RGB images and depth images of the photovoltaic junction box; the structured light camera includes: an RGB camera and a depth camera; Map and perform coordinate transformation on the depth image through the external parameters of the structured light camera to align the depth image to the coordinate system of the RGB camera; Obtain RGB images of the photovoltaic junction box under different stacking conditions, and use the annotation tool Labelme to annotate the graspable areas of the photovoltaic junction box in the RGB images; Perform data augmentation on the annotated RGB images to obtain the final dataset; The data augmentation includes: scaling, rotation and translation, and contrast transformation.

[0041] Step S2 includes: The lightweight grasping instance segmentation network model FG-YOLO includes: a YOLOv8 backbone network module, a lightweight fusion LOD module, and a segmentation head module based on dynamic mechanism adjustment; The YOLOv8 backbone network module is connected to the lightweight fusion LOD module; The lightweight fusion LOD module is connected to the segmentation head module based on a dynamic mechanism adjustment; The YOLOv8 backbone network module is used to extract the main features of the image; The lightweight fusion LOD module is used to fuse the extracted main features; The ODConv module is used to replace the C2f module and the ordinary Conv module in the neck of the YOLOv8 model to construct the lightweight fusion LOD module; The segmentation head module based on a dynamic mechanism adjustment includes: the DMA module and the Scale layer.

[0042] The present application provides an embodiment as follows. The neck of YOLOv8 is mainly composed of the PAN+FPN module, which performs sampling connection operations on the features extracted from the backbone network to achieve image feature fusion. FG-YOLO proposes the fusion LOD module, and the main change is to use ODConv to replace the C2f module and the ordinary Conv in the neck of YOLOv8. The advantage of ODConv (dynamic convolution module) is that it can comprehensively consider information in multiple dimensions such as spatial position, input channels, output channels, and the number of convolution kernels during the convolution process, so as to achieve all-round dynamic adjustment of the convolution kernel. This meticulous adjustment method enables the improved neck to more accurately extract feature flow information and gradient information, enhancing the learning and expression capabilities of the network; The present application provides an embodiment as follows. A segmentation head module adjusted by a dynamic mechanism is introduced: The segmentation head part of YOLOv8 mainly adopts a decoupled head structure, separating classification and detection, and outputting classification information and segmentation masks. Before outputting the neck feature fusion result to the decoupled head, FG-YOLO introduces the DMA module. The difference between GnConv (recursive gated convolution) and ordinary Conv in this module is that GnConv uses group normalization, which has been proven to enhance the localization and segmentation performance of the detection head. By using shared convolution, the number of parameters can be significantly reduced, making the model more lightweight, especially outstanding on devices with limited resources. While using shared convolution, considering the differences in target scales faced by different detection heads, the Scale layer is introduced to adjust the scale of features, aiming to ensure that the detection head reduces the number of parameters and computational complexity while minimizing the accuracy loss as much as possible.

[0043] An embodiment provided by the present application is as follows: a lightweight FG-YOLO network model is constructed for the segmentation of photovoltaic junction boxes. In the actual production line, there are certain rhythm requirements for the feeding of photovoltaic junction boxes, and it is necessary to ensure that the grasping method has sufficient efficiency and success rate. As an important part of the method, instance segmentation requires the network model to balance segmentation speed and accuracy. To sum up, a lightweight and fast instance segmentation network model FG-YOLO is proposed to achieve fast and accurate segmentation of photovoltaic junction boxes in different poses. The model structure is shown in Figure 3 .

[0044] The loss function of the lightweight grasping instance segmentation network model FG-YOLO includes: The target box regression loss function is as follows:

[0045]

[0046]

[0047] Among them, represents the target box regression loss; represents the Inner-SIoU loss based on the auxiliary bounding box; represents the smooth L1 loss function; and represent the scaling factor; represents the SIoU loss; represents the intersection over union; represents the internal intersection over union of the auxiliary bounding box; represents the distance loss; represents the shape loss; The definition of the smooth L1 loss function is as follows:

[0048] Among them, x is the coordinate deviation between the ground truth box and the predicted box; The calculation formula of

[0049]

[0050]

[0051]

[0052]

[0053] Among them, AIndicates the predicted bounding box, B Indicates the ground truth bounding box; Indicates the scale factor; ( , ) indicates A The center coordinates; ( , ) indicates A The width and height of; ( , ) indicates A The upper left corner coordinates; ( , ) indicates A The lower right corner coordinates; ( , ) indicates B The center coordinates; ( , ) indicates B The width and height of; ( , ) indicates B The upper left corner coordinates; ( , ) indicates B The lower right corner coordinates; Inter Indicates A And B The intersection of, Union Indicates A And B The union of.

[0054] As an example, The idea of can be seen in Figure 4 . Improve the target bounding box regression loss function: Incorporate SIoU and Inner-IoU into the loss function of target bounding box regression to ensure the quality of target bounding box regression. The initial value of this loss function is relatively large at the beginning of training, and the loss gradient may be too large, affecting the stability of model training. In this regard, it is hoped that the loss gradient gradually decreases as training progresses to the later stage to ensure that the model can converge to the global or local minimum. And ensure that the gradient is small enough when the difference is small.

[0055] Step S3 includes: S31: Remove the instance mask images with confidence lower than the preset threshold; S32: Use the depth information of the depth image to calculate the comprehensive score of each instance mask image, as follows:

[0056] Among them, S represents the comprehensive score, C represents the confidence, H represents the depth information of the mask, u And v Indicates the corresponding weights; S33: Select a preset number of masks as candidate mask images based on the comprehensive scores of each instance mask image.

[0057] As an embodiment, for the multiple instance masks obtained by segmentation, select the best 2D grasping points. For the implementation effect, see Figure 5 .

[0058] Step S4 includes: S41: Extract the central skeleton of the candidate mask image through a thinning algorithm; As an embodiment, for each of the selected masks, use the thinning algorithm to extract a central skeleton. Using the points on the skeleton as the grasping points can ensure that when grasping subsequently, the jaws grasp at a position closer to the center of the junction box, reducing the situation of grasping failure. Taking Figure 6 as an example, where is the central pixel, are its eight-neighbor pixels, and the discrimination of the thinning algorithm has the following four conditions: a)

[0059] b)

[0060] c) Stage 1: (Stage 2 is ) d) Stage 1: (Stage 2 is ) Among them, represents the number of non-zero pixels in the eight-neighbor of represents from to and then to the number of times of changing from zero pixels to non-zero pixels in order. The thinning process is carried out in two stages. The first two discrimination conditions are the same in both stages, and the last two conditions are used to discriminate the pixel points in different orientations in the mask in the two stages respectively. Set the pixels that meet all the above conditions to zero, and continuously iterate the two stages in turn until no pixels are deleted, then the thinning process ends and the central skeleton is obtained; S42: Sort the central skeleton through the nearest neighbor algorithm to obtain an ordered skeleton; As an embodiment, the extracted skeleton is in an unordered state, and it is sorted through the nearest neighbor algorithm. In the sorting process, first discriminate all the pixel points on the skeleton. If there is only one non-zero pixel in its eight-neighbor, set this pixel as the starting point and delete it from the skeleton. Then calculate the distance between other points and the starting point in the skeleton, find the nearest point as the second point, and delete it from the skeleton. Repeat the above operations to find the point closest to the previous point until all points are deleted and the sorting is completed; S43: In the ordered skeleton, in combination with a preset step size, extract multiple equidistant pixels as candidate grasping points; calculate the curvature direction of each candidate grasping point in the local skeleton; The curvature direction is the closing direction of the robot gripper in the pixel coordinate system; S44: Obtain the size parameters of the robot gripper; calculate the rectangular area passed by the closing of the two fingers of the gripper in the depth image through the candidate grasping points, the closing direction, and the size parameters; S45: Judge the interference degree of the rectangular area through the depth of the grasping point and the diameter of the photovoltaic junction box; calculate the score of each candidate grasping point according to the interference degree of the rectangular area; As an embodiment, considering the stability during grasping, the top of the gripper needs to extend to a depth position that completely covers the photovoltaic junction box. Therefore, the depth threshold for judging the interference degree is the sum of the depth of the grasping point and the diameter of the photovoltaic junction box. In the area corresponding to the depth map of the two fingers of the gripper, pixels smaller than the depth threshold are regarded as interference points, and collisions may occur when the gripper descends, resulting in grasping failure. The scoring criteria are as follows: both regions with the number of interference points less than 5 get 3 points, only one region with the number of interference points less than 5 gets 2 points, both regions with the number of interference points greater than 5 get 1 point, and grasping points with too little depth information in the mask get 0 points; S46: Compare the scores of each candidate grasping point, and select the candidate grasping point with the highest score as the two-dimensional grasping point of the candidate mask image.

[0061] As an embodiment, the structured light camera takes pictures from top to bottom. In the depth map, the smaller the depth value of the pixel, the closer it is to the camera and the higher its position in the actual scene. Therefore, when sorting the grasping points of each mask, it is carried out in ascending order of depth, and the grasping point with the highest score and relatively smaller depth value is used as the final grasping point. Step S5 includes: S51: Calculate the three-dimensional coordinates of the two-dimensional grasping point mapped to the RGB camera coordinate system; map the unit vector along the curvature direction with the two-dimensional grasping point as the starting point to a three-dimensional vector in the RGB camera coordinate system; Let the two-dimensional grasping point coordinates be and the depth value corresponding to the coordinates in the depth image be h, the three-dimensional coordinates of the two-dimensional grasping point are calculated as follows:

[0062]

[0063]

[0064] Among them, and Indicates the pixel focal lengths in the x and y directions, and indicates the principal point coordinates; Let the unit vector along the curvature direction be , and construct a three-dimensional vector parallel to the camera plane The calculation formula is as follows:

[0065]

[0066]

[0067] This three-dimensional vector represents the direction of gripper closing in the RGB camera coordinate system; S52: Cross-multiply the three-dimensional vector with the three-dimensional unit vector parallel to the z-axis of the RGB camera coordinate system to obtain the three-dimensional vector , and form the following rotation matrix from the three three-dimensional vectors:

[0068] Together with the rotation matrix and the three-dimensional coordinates of the two-dimensional grasping point form the grasping pose matrix in the camera coordinate system:

[0069] S53: Through morphological processing, obtain the edge information of the material frame in the RGB image; Based on the edge information, calculate the distance between the two-dimensional grasping point and the nearest edge; If the distance is less than the preset threshold, optimize the calculated grasping pose matrix; Let the unit vector starting from the two-dimensional grasping point and pointing to the center of the material frame be ; Construct a three-dimensional vector parallel to the camera plane , and combine it with the three-dimensional vector parallel to the camera plane to construct the pose matrix in the RGB camera coordinate system ; The homogeneous transformation matrix for tilting towards the center of the material frame is:

[0070] where, is the tilting angle; The final optimized grasping pose matrix of the gripper in the RGB camera coordinate system is: ; S54: Convert the grasping pose matrix from the RGB camera coordinate system to the robot coordinate system; Generate a transition pose matrix at a preset distance above along the Z direction of the grasping pose matrix; Control the robot to grasp the photovoltaic junction box through the grasping pose matrix and the transition pose matrix after coordinate transformation.

[0071] As an embodiment, convert the two pose matrices (the grasping pose matrix and the transition pose matrix) into the rotation angle type acceptable to the robot, and send the transition pose and the grasping pose to the robot client in sequence to guide the robot to move to the grasping point, realizing collision-free grasping of the photovoltaic junction box.

[0072] Please refer to Figure 2 , Figure 2 which is the system structure diagram of a photovoltaic junction box grasping pose estimation system based on instance segmentation in the embodiments of the present application, including: A structured light camera, a computer module, and a robot; The robot includes: an electric parallel gripper, a robot control cabinet, and a six-axis robotic arm; The structured light camera is connected to the computer module; The computer module is connected to the robot; The structured light camera is used to collect RGB images and depth images of the photovoltaic junction box under different stacking conditions; The computer module is used to construct a data set through the RGB images; The computer module is also used to construct a lightweight grasping instance segmentation network model FG-YOLO and train it through the data set to obtain an instance mask map and its confidence; The computer module is also used to screen the instance mask map through the depth image and the confidence to determine the candidate mask map; The computer module is also used to use a refinement algorithm to extract the two-dimensional grasping points and their curvature directions of the candidate mask map; The computer module is also used to calculate a three-dimensional grasping pose matrix from the two-dimensional grasping points and their curvature directions; optimize the three-dimensional grasping pose matrix according to the position of the photovoltaic junction box in the material box; The computer module is also used to generate a pose control instruction by the optimized three-dimensional grasping pose matrix to control the robot to grasp the photovoltaic junction box; The robot is used to receive the pose control instruction and grasp the photovoltaic junction box.

[0073] The above are only exemplary embodiments of the present disclosure and should not be used to limit the scope of the present disclosure. That is, all equivalent changes and modifications made according to the teachings of the present disclosure still fall within the scope covered by the present disclosure.

[0074] This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include well-known knowledge or conventional technical means in the technical field not recorded in the present disclosure. The description and examples are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.

Claims

1. A photovoltaic junction box grasping pose estimation method based on instance segmentation, characterized in that: The method comprises the following steps: S1: Collect RGB images and depth images of photovoltaic junction boxes in different stacking conditions; construct a data set through RGB images; S2: Build a lightweight instance segmentation network model FG-YOLO and train it through the dataset to obtain the instance mask map and its confidence; S3: Screen the instance mask map through the depth image and confidence level to determine the candidate mask map; S4: Using the refinement algorithm, extract the two-dimensional grasping points and curvature directions of the candidate mask image; S5: Calculate the three-dimensional grasping posture matrix based on the two-dimensional grasping points and their curvature directions; optimize the three-dimensional grasping posture matrix according to the position of the photovoltaic junction box in the material frame; and control the robot to grasp the photovoltaic junction box by using the optimized three-dimensional grasping posture matrix.

2. A photovoltaic junction box grasping pose estimation method based on instance segmentation as claimed in claim 1, characterized in that: Step S1 includes: Use a structured light camera with an integrated RGB camera to collect RGB images and depth images of the photovoltaic junction box; the structured light camera includes: an RGB camera and a depth camera; Through the external parameters of the structured light camera, the depth image is mapped and the coordinates are converted to align the depth image to the bottom of the RGB camera coordinate system; Obtain RGB images of photovoltaic junction boxes under different stacking conditions, and use the labeling tool Labelme to label the graspable areas of the photovoltaic junction boxes in the RGB images; Perform data enhancement on the annotated RGB images to obtain the final data set; Data augmentation includes: scaling, rotation and translation, and contrast transformation.

3. The method for estimating the pose of a photovoltaic junction box based on instance segmentation according to claim 1, characterized in that: Step S2 includes: The lightweight capture instance segmentation network model FG-YOLO includes: a YOLOv8 backbone network module, a lightweight fusion LOD module, and a segmentation head module based on dynamic mechanism adjustment; The YOLOv8 backbone network module is connected to the lightweight fusion LOD module; The lightweight fusion LOD module is connected to the segmentation head module based on dynamic mechanism adjustment; The YOLOv8 backbone network module is used to extract the main features of the image; The lightweight fusion LOD module is used to fuse the extracted main features; Use the ODConv module to replace the C2f module and the ordinary Conv module in the neck of the YOLOv8 model to build a lightweight fusion LOD module; The segmentation header module adjusted based on the dynamic mechanism includes: a DMA module and a Scale layer.

4. The method for estimating the pose of a photovoltaic junction box based on instance segmentation according to claim 1, characterized in that: The loss function of the lightweight capture instance segmentation network model FG-YOLO includes: The target box regression loss function is as follows: in, represents the target box regression loss; Indicates that the auxiliary border is based on Inner-SIoU loss; represents the smooth L1 loss function; and represents the scaling factor; express ikB loss; represents intersection and union ratio; Indicates the internal intersection ratio of the auxiliary border; Indicates distance loss; represents shape loss; The smooth L1 loss function is defined as follows: Where x is the coordinate deviation between the real box and the predicted box; The calculation formula is as follows: in, A represents the prediction box, B represents the real frame; represents the scale factor; ( , )express A Center coordinates; ( , )express A The width and height of , )express A Upper left corner coordinates; ( , )express A Coordinates of the lower right corner; ( , )express B Center coordinates; ( , )express B The width and height of , )express B Upper left corner coordinates; ( , )express B Coordinates of the lower right corner; Inter express A and B The intersection of Union express A and B The union of .

5. The method for estimating the pose of a photovoltaic junction box based on instance segmentation according to claim 1, characterized in that: Step S3 includes: S31: removing instance mask images whose confidence level is lower than a preset threshold; S32: Using the depth information of the depth image, calculate the comprehensive score of each instance mask map, the input is as follows: Among them, S represents the comprehensive score, C represents the confidence, and H represents the depth information of the mask. u and v represents the corresponding weight; S33: Selecting a preset number of masks as candidate mask images based on the comprehensive score of each instance mask image.

6. The method for estimating the pose of a photovoltaic junction box based on instance segmentation according to claim 1, characterized in that: Step S4 includes: S41: extracting the central skeleton of the candidate mask image through a refinement algorithm; S42: sorting the central skeleton by a nearest neighbor algorithm to obtain an ordered skeleton; S43: extracting a plurality of equally spaced pixels from the ordered skeleton in combination with a preset step size as candidate grasping points; and calculating the curvature direction of each candidate grasping point in the local skeleton; The curvature direction is the closing direction of the robot gripper in the pixel coordinate system; S44: Obtaining the size parameters of the robot gripper; calculating the rectangular area where the two fingers of the gripper close and pass through in the depth image through the candidate grasping points, closing direction and size parameters; S45: judging the interference degree of the rectangular area according to the depth of the grasping point and the diameter of the photovoltaic junction box; and calculating the score of each candidate grasping point according to the interference degree of the rectangular area; S46: Compare the scores of each candidate grasping point, and select the candidate grasping point with the highest score as the two-dimensional grasping point of the candidate mask image.

7. The method for estimating the pose of a photovoltaic junction box based on instance segmentation according to claim 1, characterized in that: Step S5 includes: S51: Calculate the three-dimensional coordinates of the two-dimensional grasping point mapped to the RGB camera coordinate system; take the two-dimensional grasping point as the starting point, and map the unit vector along the curvature direction to the three-dimensional vector in the RGB camera coordinate system; Assume the coordinates of the two-dimensional grab point are The depth value of the corresponding coordinate in the depth image is h, and the three-dimensional coordinates of the two-dimensional grasping point are The calculation formula is as follows: in, and represents the pixel focal length in the x and y directions, and represents the coordinates of the principal point; Let the unit vector along the curvature direction be , construct a 3D vector parallel to the camera plane The calculation formula is as follows: This three-dimensional vector represents the direction of the gripper closing in the RGB camera coordinate system; S52: Convert three-dimensional vector The three-dimensional unit vector parallel to the z-axis of the RGB camera coordinate system Cross product to get a three-dimensional vector , the following rotation matrix is ​​formed by three three-dimensional vectors: The three-dimensional coordinates of the two-dimensional grasping point through the rotation matrix Together they form the grasping pose matrix in the camera coordinate system: S53: Obtain edge information of the material frame in the RGB image through morphological processing; Calculate the distance between the two-dimensional grasping point and the nearest edge through edge information; If the distance is less than the preset threshold, the calculated grasping pose matrix is ​​optimized; Assume that the unit vector pointing to the center of the material frame with the two-dimensional grabbing point as the starting point is ; Construct a 3D vector parallel to the camera plane , combined with the 3D vector parallel to the camera plane , construct the pose matrix in the RGB camera coordinate system ; The homogeneous transformation matrix tilted toward the center of the material frame is: in, is the tilt angle; The final optimized gripper's grasping pose matrix in the RGB camera coordinate system is: ; S54: convert the grasping pose matrix from the RGB camera coordinate system to the robot coordinate system; Generate a transition pose matrix at a preset distance above the grasped pose matrix along the Z direction; The robot is controlled to grasp the photovoltaic junction box through the grasping pose matrix and the transition pose matrix after coordinate system conversion.

8. A photovoltaic junction box grasping pose estimation system based on instance segmentation, used to implement a photovoltaic junction box grasping pose estimation method based on instance segmentation as described in any one of claims 1 to 7, characterized in that: The system comprises: a structured light camera, a computer module and a robot; The robot includes: electric parallel gripper, robot control cabinet and six-axis robot arm; The structured light camera is connected to a computer module; The computer module is connected to the robot; The structured light camera is used to collect RGB images and depth images of photovoltaic junction boxes in different stacking conditions; The computer module is used to construct a data set through RGB images; The computer module is also used to construct a lightweight capture instance segmentation network model FG-YOLO and train it through a data set to obtain an instance mask image and its confidence; The computer module is also used to screen the instance mask image through the depth image and the confidence level to determine the candidate mask image; The computer module is also used to extract two-dimensional grab points and curvature directions of the candidate mask image using a thinning algorithm; The computer module is also used to calculate the three-dimensional grasping posture matrix from the two-dimensional grasping points and their curvature directions; and optimize the three-dimensional grasping posture matrix according to the position of the photovoltaic junction box in the material frame; The computer module is also used to generate a posture control instruction by using the optimized three-dimensional grasping posture matrix to control the robot to grasp the photovoltaic junction box; The robot is used to receive posture control instructions and grasp the photovoltaic junction box.