Intelligent and rapid positioning method for spatial target grasping points based on dual-arm collaborative vision joint perception
Through the dual-arm collaborative visual joint perception method, intelligent autonomous recognition and rapid and stable grasping of non-cooperative targets in space are achieved, which solves the accuracy and robustness problems existing in traditional methods and improves the intelligent autonomy level and grasping accuracy of spatial operations.
Patent Information
- Application Number
- CN202211299547.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-23
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2042-10-23
AI Technical Summary
Traditional target recognition and positioning algorithms suffer from the problems of lack of prior information in complex spatial environments, slow target capture area segmentation, poor accuracy and weak robustness, making it difficult to achieve intelligent and autonomous recognition and rapid and stable capture of non-cooperative targets in space.
The dual-arm collaborative visual joint perception method is adopted. Through the intelligent autonomous recognition and instance segmentation of global targets, combined with the fast and accurate segmentation of the local contour of the target by visual joint perception, the intelligent autonomous recognition, segmentation and fast and stable acquisition of the target grasping position of the spatial target docking ring of the dual-arm collaborative visual joint perception are realized.
It improves the intelligent autonomy level of spatial operations, enhances the positioning accuracy and robustness of the target grasping position, and ensures the safe arrival and fine operation control of the robotic arm.
Smart Images

Figure CN115723123B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of space on-orbit service and control technology, and in particular to a method for intelligently and rapidly locating a space target grasping point using dual-arm collaborative visual joint perception. Background Art
[0002] With the increasing diversity and complexity of on-orbit manipulation tasks, manipulation technologies primarily based on space robots have become a hot topic. Manipulation targets are transitioning from stable, cooperative targets to rotating, non-cooperative targets such as failed spacecraft and space debris. Space manipulators can approach and capture non-cooperative targets, enabling tasks such as inspection, repair, and replacement of faulty spacecraft components and debris removal. Identifying, locating, and grasping non-cooperative targets are key aspects of space manipulators' on-orbit servicing and manipulation missions. To enhance the intelligent autonomy and reliability of non-cooperative target capture technology, a method for intelligent and rapid location of space target grasping points using dual-arm collaborative visual perception is developed. This method uses a dual-arm collaborative visual perception system to achieve intelligent and autonomous recognition, segmentation, and rapid and stable acquisition of the target grasping position of the non-cooperative target docking loop. This method comprehensively enhances the intelligent autonomy of space operations and improves the positioning accuracy of the non-cooperative target grasping position, laying the foundation for subsequent precise control of manipulator operations such as safe arrival, capture, and component removal. Summary of the Invention
[0003] Technical problems to be solved
[0004] In order to solve the problems in traditional target recognition and positioning algorithms that rely on direct image processing methods such as feature extraction and contour segmentation of targets with specific structural shapes to achieve recognition and acquisition of grasped targets, such as weak generalization of target recognition in complex spatial environments where prior information is missing, slow processing speed of target grasping area segmentation, poor grasping point positioning accuracy and weak robustness, the present invention provides an intelligent and rapid positioning of spatial target grasping points with dual-arm collaborative visual joint perception, which can realize intelligent and autonomous recognition, segmentation and rapid and stable acquisition of target grasping positions of spatial non-cooperative target docking rings based on dual-arm collaborative visual joint perception.
[0005] Technical Solution
[0006] A method for intelligently and rapidly locating a spatial target grasping point using dual-arm collaborative vision joint perception is characterized by the following steps:
[0007] S1, intelligent autonomous recognition of global targets and instance segmentation;
[0008] S2, fast and accurate segmentation of the local contour of the target based on joint visual perception;
[0009] S3. Dynamic and precise positioning of the dual-arm collaborative target grasping point guided by vision, realizing the intelligent autonomous recognition, segmentation and rapid and stable acquisition of the target grasping position of the spatial target docking ring of the dual-arm collaborative vision joint perception.
[0010] A further technical solution of the present invention: Step S1 specifically comprises:
[0011] S11. Collect images of non-cooperative space targets during the robotic arm's on-orbit servicing and manipulation missions, annotate the samples, and complete the construction of a labeled training dataset for key components of non-cooperative space targets.
[0012] S12. Use the training dataset constructed in S11 as the input for training the Mask R-CNN-based network model for rapid target detection and instance segmentation. Use the images acquired in real time by a global large-field-of-view RGB-D camera to input the model to achieve intelligent autonomous recognition and instance segmentation of non-cooperative target docking rings during the end-of-line approach of the robotic arm grasping task.
[0013] A further technical solution of the present invention: Step S2 specifically comprises:
[0014] S21. Based on the results of the global target intelligent autonomous recognition and instance segmentation in S1, the target docking ring components are screened, and the position of the surrounding detection frame of the target docking ring and the target pixel segmentation mask are obtained;
[0015] S22, synchronously mapping the target detection and segmentation results and the target docking ring center position to the depth map correspondingly obtained by the global large-field-of-view RGB-D camera to obtain the target docking ring center, radius, detection frame, and boundary three-dimensional information;
[0016] S23, guiding the target docking ring to the center position of the image obtained by the global large-field-of-view RGB-D camera, and based on the three-dimensional information of the target docking ring center, detection frame, and boundary obtained at the current moment, achieving the positioning and acquisition of the initial position of the target docking ring under the high-precision hand-eye depth camera of the dual-arm collaborative end;
[0017] S24. Based on the detection border, the local window area of the target docking ring under the hand-eye depth camera at the end of the dual robotic arm is obtained, and the edge detection and segmentation of the local target in the depth image of this area are performed. Combined with the target instance segmentation pixel boundary results, the local contour of the target is quickly and accurately segmented by visual joint perception.
[0018] A further technical solution of the present invention: Step S3 specifically comprises:
[0019] S31. Based on the target edge information obtained by rapid and accurate segmentation of the local contour of the target based on joint visual perception, a piecewise combined third-order Bezier curve is used to continuously fit the local edge arc in the double-arm hand-eye depth image, and the farthest tangent point of the edge arc segment is dynamically obtained as the current target capture point;
[0020] S32: When the distance between the two grasping points meets a certain condition, the grasping process is executed, and the position information of the grasping points is sent to the end robot arm controller;
[0021] S33. A weighted speed optimization adjustment method based on segmented distance is used to enable the two arms to synchronously and collaboratively approach the target docking ring capture point in segments during movement, ensuring that the target is captured synchronously.
[0022] A further technical solution of the present invention: The condition in S32 is specifically: calculating the distance d between the two capture points normalized to the right camera coordinate system. If d is greater than 4 / 5 times the diameter of the docking ring, the capture requirements are met.
[0023] Beneficial effects
[0024] The present invention provides a method for intelligently and rapidly locating spatial target grasping points using dual-arm collaborative visual joint perception. Compared to traditional target recognition and positioning algorithms that rely on direct image processing methods such as feature extraction and contour segmentation of targets with specific structural shapes to achieve recognition and acquisition of grasped targets, the present invention adopts a method for intelligently and rapidly locating spatial target grasping points using dual-arm collaborative visual joint perception. Through global target intelligent recognition and segmentation and local contour fine acquisition and joint perception, the method accurately locates the target grasping points based on dual-arm collaboration, realizes intelligent and autonomous recognition, segmentation, and rapid and stable acquisition of the target grasping position of the spatial non-cooperative target docking ring, can comprehensively enhance the intelligent and autonomous level of spatial operations, improve the positioning accuracy of the spatial non-cooperative target grasping position, and lay the foundation for subsequent precise operational control of the robotic arm, such as safe arrival, capture, and component disassembly. It has the advantages of high target recognition generalization, fast grasping area segmentation speed, high grasping point positioning accuracy, and strong robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings are only for the purpose of illustrating particular embodiments and are not to be considered limiting of the present invention. Like reference symbols denote like parts throughout the drawings.
[0026] Figure 1 is a flow chart of the method of the present invention;
[0027] Figure 2 Flowchart for implementing step S2 in the method of the present invention;
[0028] Figure 3 Flowchart for implementing step S23 in the method of the present invention;
[0029] Figure 4 Flowchart for implementing step S3 in the method of the present invention;
[0030] Figure 5 Flowchart of implementation of step S31 in the method of the present invention; DETAILED DESCRIPTION
[0031] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only intended to illustrate the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.
[0032] like Figure 1 As shown, the present invention proposes a method for intelligent rapid positioning of spatial target grasping points using dual-arm collaborative visual joint perception, which is used for intelligent identification, segmentation and precise positioning of the grasping position of spatial non-cooperative target docking rings. This method uses a dual-arm collaborative visual joint perception system to achieve intelligent autonomous identification, segmentation and rapid and stable acquisition of the spatial non-cooperative target docking ring and target grasping position. The method specifically includes the following steps:
[0033] S1, intelligent autonomous recognition of global targets and instance segmentation;
[0034] S2, fast and accurate segmentation of the local contour of the target based on joint visual perception;
[0035] S3. Dynamic and precise positioning of the dual-arm collaborative target grasping point guided by vision, realizing the intelligent autonomous recognition, segmentation and rapid and stable acquisition of the target grasping position of the spatial target docking ring of the dual-arm collaborative vision joint perception.
[0036] In the above-mentioned method for intelligent and rapid positioning of spatial target grasping points by dual-arm collaborative vision joint perception, step S1 specifically includes:
[0037] S11. Collect images of non-cooperative space targets for on-orbit servicing and manipulation tasks such as on-orbit capture and component disassembly by robotic arms, and annotate samples to construct a labeled training dataset for key components of non-cooperative space targets.
[0038] S12. The constructed training dataset is used to complete the training of the Mask R-CNN-based network model for rapid target detection and instance segmentation. The images acquired in real time by the global large-field-of-view RGB-D camera are used to complete the intelligent autonomous recognition and instance segmentation of the non-cooperative target docking ring during the end-of-the-line approach of the robotic arm grasping task.
[0039] The images of non-cooperative space targets should be rich and diverse, and simulate as realistically as possible the non-cooperative space targets of various types and motion states at multiple angles under different missions, different lighting, and different background conditions in a real space environment. Specifically, these images should include images of non-cooperative space targets under conditions such as stray light interference, complex and changing backgrounds, extremely harsh lighting conditions, occlusion of space targets, and tumbling or high-speed spinning of space targets.
[0040] The sample annotation includes marking the instance segmentation boundaries and detection boxes of the target key components in each original image in the training sample and creating corresponding labels;
[0041] The sample annotation is performed by using Labelme to annotate each collected original image. Specifically, the Create Polygons function in the Labelme tool is used to mark the instance segmentation boundaries of the target key components and create labels. Then, the Create Rectangle function is used on this basis to mark the detection box corresponding to the target key components and select the corresponding label;
[0042] The key components of the space non-cooperative target include a dish antenna, a thruster, a ground observation camera, a solar panel, a solar sensor, an infrared horizon, a rigid arm, a feed source, a flexible arm, a solar panel bracket, a conical spiral antenna, a satellite body, a star sensor, a satellite-rocket docking ring, a docking surface, a laser data transmission, a nozzle, a measurement and control antenna, an end effector, and a GPS antenna.
[0043] The global large field of view RGB-D camera is installed on the service satellite body, located at the center of the connection line of the dual robotic arm bases, and can realize the acquisition of global information of the non-cooperative target docking ring of the service satellite at the end of the robotic arm grasping task, which is 2 meters away from the operating satellite;
[0044] The intelligent autonomous recognition and instance segmentation of the non-cooperative target docking ring adopts the Mask R-CNN algorithm to obtain the position of the target initial state positioning detection frame, target pixel level classification and accurate segmentation of target boundaries and other information parameters in real time.
[0045] Specifically, the Mask R-CNN implementation steps are as follows:
[0046] (1) Image feature extraction based on ResNeXt-101 network;
[0047] (2) Generate candidate regions of interest (ROIs) through the RPN target estimation network for the obtained feature map;
[0048] (3) The obtained ROI is mapped into a fixed-dimensional feature vector through the ROI Align layer with bicubic linear interpolation to maintain the fineness of the feature pixels. Two branches are processed by Faster RCNN to complete the target classification and bounding box regression, and the other branch is processed by the fully convolutional neural network (FCN) for upsampling to obtain pixel-level instance segmentation of the target object.
[0049] The instance segmentation branch uses an FCN with 16x upsampling, resulting in more refined segmentation results and greater robustness for smaller objects. Furthermore, the FCN adds five convolutional layers and three deconvolutional layers to the con5_3 layer of the VGG-Net, enabling parallel processing of object recognition and detection branches, reducing overall task execution time.
[0050] The multi-task loss function used in training continuously reduces the value of the loss function through learning until the global optimal solution is obtained. The formula of the loss function is shown in Equation (1), where the three terms are classification error, bounding box error, and segmentation error.
[0051] L=L cls +L box +L mask (1)
[0052] The improved Mask R-CNN algorithm will get three outputs: the target category, the position of the bounding box, and the mask of the object.
[0053] In the above-mentioned method for intelligent and rapid positioning of spatial target grasping points by dual-arm collaborative vision joint perception, step S2 specifically includes:
[0054] S21. Based on the results of the global target intelligent autonomous recognition and instance segmentation in S1, the target docking ring components are screened, and the position of the surrounding detection frame of the target docking ring and the target pixel segmentation mask are obtained;
[0055] S22, synchronously mapping the target detection and segmentation results and the target docking ring center position to the depth map correspondingly obtained by the global large-field-of-view RGB-D camera, and completing the acquisition of the target docking ring center, radius, detection frame, and boundary three-dimensional information;
[0056] S23, guiding the target docking ring to the center position of the image obtained by the global large-field-of-view RGB-D camera, and based on the three-dimensional information of the target docking ring center, radius, detection frame, and boundary obtained at the current moment, achieving the positioning and acquisition of the initial position of the target docking ring under the local high-precision hand-eye depth camera;
[0057] S24. Based on the detection frame, the local window area of the target docking ring under the hand-eye depth camera at the end of the dual robotic arm is acquired, and the edge detection and segmentation of the local target in the depth image of the area are performed. Combined with the pixel boundary results of the target instance segmentation, the local contour of the target is quickly and accurately segmented by the joint visual perception.
[0058] The screening of the target docking ring components in step S21 is based on the results of the global target intelligent autonomous recognition and instance segmentation in step S1, and the target docking ring of interest is quickly screened out from multiple components, thereby realizing the rapid intelligent recognition and segmentation of the target docking ring components under different tasks based on the global depth camera, which has the advantages of high speed and high generalization.
[0059] In step S23, the global large field of view RGB-D camera is located at the center of the dual robotic arm base at the end of the service satellite, and the local high-precision hand-eye depth camera is installed at the end of the robotic arm;
[0060] The specific process of locating and acquiring the initial position of the target docking ring is as follows:
[0061] (1) Dynamically adjust the positional relationship between the service satellite and the target docking ring according to the center coordinates of the target docking ring until the target docking ring is located at the center of the field of view of the global large-field-of-view RGB-D camera;
[0062] (2) Start the local hand-eye camera at the end of the dual robotic arms and complete the calibration of the global large-field-of-view RGB-D camera and the high-precision hand-eye camera at the end of the service satellite robotic arm;
[0063] (3) The three-dimensional information of the target docking ring center, radius, detection frame and boundary obtained by the global RGB-D camera at the current moment is mapped to the depth image observed by the high-precision hand-eye camera at the end of the dual robotic arm, and the initial position of the target docking ring is located and acquired under the local high-precision hand-eye depth camera.
[0064] The process of fast and accurate segmentation of the local contour of the target by visual joint perception is as follows:
[0065] Based on the above S1 global target intelligent autonomous recognition and instance segmentation, the four vertices A, B, C, D of the target docking ring's enclosing detection frame and the target pixel segmentation mask Mask boundary point N are obtained. i The coordinates are:
[0066] A(x a ,y a , z a ),B(x b ,y b , z b ),C(x c ,y c , zc ),D(x d ,y d , z d ) and N i (x i ,y i , z i )(i=1...n), the center coordinate of the target docking ring O(x O ,y O , z O ).
[0067] The three-dimensional information of the target docking ring center, radius, detection frame and boundary obtained by the global RGB-D camera is mapped to the depth map of the high-precision hand-eye camera at the end of the dual robotic arm as follows Figure 3 As shown:
[0068] The coordinates of the target docking ring center, radius, detection frame and boundary three-dimensional information of the target docking ring initially positioned by the local high-precision hand-eye depth camera are as follows:
[0069] The vertex coordinates of the target docking ring detection frame in the depth image of the left robotic arm's hand-eye depth camera:
[0070] A1(x a1 ,y a1 , z a1 ), B1(x b1 ,y b1 , z b1 ), M1(x m1 ,y m1 , z m1 ), N1(x n1 ,y n1 , z n1 );
[0071] The virtual center coordinates of the target docking ring in the depth image of the left robotic arm hand-eye depth camera O1(x O1 ,y O1 , z O1 );
[0072] Coordinates of the target docking ring boundary point in the depth image of the left robotic arm's hand-eye depth camera: N 1i (x 1i ,y 1i , z 1i )(i=1...t);
[0073] The coordinates of the target docking ring boundary endpoints in the depth image of the left robotic arm's hand-eye depth camera:
[0074] P1(x p1 ,y p1 , z p1 ), Q1(xq1 ,y q1 , z q1 );
[0075] The vertex coordinates of the target docking ring detection frame in the depth image of the right robotic arm's hand-eye depth camera:
[0076] C2(x c2 ,y c2 , z c2 ), D2(x d2 ,y d2 , z d2 ), M2(x m2 ,y m2 , z m2 ), N2(x n2 ,y n2 , z n2 );
[0077] Coordinates of the virtual center of the target docking ring in the depth image of the right robotic arm's hand-eye depth camera:
[0078] O2(x O2 ,y O2 , z O2 );
[0079] Coordinates of the target docking ring boundary points in the depth image of the right robotic arm's hand-eye depth camera:
[0080] N 2i (x 2i ,y 2i , z 2i )(i=1...s);
[0081] The coordinates of the target docking ring boundary endpoints in the depth image of the right robotic arm's hand-eye depth camera:
[0082] P2(x p2 ,y p2 , z p2 ), Q2(x q2 ,y q2 , z q2 ).
[0083] in:
[0084] s+t≤n (2)
[0085] In step S24, the acquisition of the local window area of the target docking ring under the hand-eye depth camera at the end of the dual manipulator is completed based on the detection frame. Figure 3 As shown, the local area observed by the hand-eye depth camera at the end of the left / right robotic arm is the area composed of C2, D2, N2, M2 and A1, B1, N1, M1.
[0086] In step S24, it is only necessary to complete the detection and segmentation of the edge of the target in the local area depth image composed of C2, D2, N2, M2 and A1, B1, N1, M1 observed by the hand-eye depth camera at the end of the left / right robotic arm, and obtain the segmentation edge point T 2j (x 2j ,y 2j , z 2j )(j=1...n 20 ) and T 1j (x 1j ,y 1j , z 1j )(j=1...n 10 );
[0087] In step S24, the pixel boundary results of the target instance segmentation are combined to complete the rapid and accurate segmentation of the local contour of the target in the joint visual perception. Specifically, the target edge points of the target instance segmentation are screened and the detection and segmentation edge map information of the local edge of the target in the depth image under the small window area is obtained to obtain the final target depth map edge extraction result, which is as follows:
[0088] Traverse each edge point of the depth image obtained by the hand-eye depth camera at the end of the left / right robotic arm and retain the edge point J that meets the following conditions at the same time 1k (x 1k ,y 1k , z 1k )(k=1...m) and J 2k (x 2k ,y 2k , z 2k )(k=1...n);
[0089]
[0090]
[0091] The described rapid and accurate segmentation of the target local contour by visual joint perception adopts the method of globally processing only the initial position image, and subsequently locally processing the local image of the scene containing the target or the image sequence to achieve continuous segmentation of the local features of the remaining sequence images. There is no need to perform edge detection, feature extraction and target segmentation on the entire satellite image, which has the advantage of good real-time performance. The method of visual joint perception by global intelligent segmentation and local edge detection segmentation can effectively improve the accuracy of target local contour area segmentation.
[0092] In the above-mentioned method for intelligent and rapid positioning of grabbing points of a space target docking ring by visual joint perception and dual-arm collaboration, step S3 specifically includes:
[0093] S31. Rapidly and accurately segment the target's local contour based on joint visual perception to obtain target edge information. A segmented combined third-order Bezier curve is used to continuously fit the local edge arc in the hand-eye depth image of both arms. The farthest tangent point of the edge arc segment is dynamically obtained to complete the acquisition of the target capture point.
[0094] S32: Unify the grabbing points of both arms into the same coordinate system, execute the grabbing process when the distance between the two grabbing points meets certain conditions, and send the position information of the grabbing points to the end robot arm controller;
[0095] S33. A weighted speed optimization adjustment method based on segmented distance is adopted to enable the two arms to synchronously and collaboratively approach the target docking ring capture point in segments during the movement process, ensuring the synchronous completion of the target capture. It has the advantages of high target capture point positioning accuracy and strong robustness.
[0096] The acquisition of the target capture point in step S31 is performed by using a segmented combined third-order Bezier curve to complete the continuous fitting of the local edge arc in the double-arm hand-eye depth image, and dynamically obtain the farthest tangent point of the edge arc segment;
[0097] The process of implementing the continuous fitting of the local edge arcs in the double-arm hand-eye depth image using the segmented combined third-order Bezier curve is as follows:
[0098] (1) The current arc segment is divided into N equal parts, N = 10, and the angle corresponding to each arc segment
[0099] (2) Using the third-order Bezier curve to complete the continuous fitting of the arc segments in the local small area for each small arc segment;
[0100] (3) Merge N arc segments after fitting the third-order Bezier curve to complete the acquisition of the entire arc segment;
[0101] The arc segment fitting process of the third-order Bezier curve is as follows:
[0102] like Figure 5 As shown, for each small arc segment A i-1 ,M,N,A i There are four control points, among which A i-1 As the starting point, A i is the end point, and the intermediate control points are M and N. The radius of the arc is r.
[0103] P(t)=A i (1-t) 3 +M·3(1-t) 2 +N·3(1-t)t 2 +A i-1 ·t 3 , t=0...1 (5)
[0104]
[0105] The coordinates of the four control points are:
[0106] A i-1 (r,0), M(r,h), N(rcosθ1+hsinθ1,rsinθ1-hcosθ1), A i (rcosθ1,rsinθ1).
[0107] The dynamic acquisition of the farthest tangent point of the edge arc segment is specifically performed by traversing each point on the curve and selecting the leftmost and rightmost points corresponding to the left and right robotic arms as the capture points;
[0108] In step S32, when the distance between the two capture points meets certain conditions, the capture process is executed by calculating the distance d between the two capture points normalized to the right camera coordinate system. If d is greater than 4 / 5 times the diameter of the docking ring, the capture requirement is met; if the distance requirement is not met, the right capture point is used as a reference, the right capture point and the virtual circle center coordinates are connected, and the intersection of the extension line and the left robotic arm depth image is obtained, and the capture point is searched within a certain area r = 0.001m around the intersection to ensure that the two arms can capture the target area far enough at the same time and can stably capture the same target object.
[0109] In step S33, a weighted speed optimization adjustment method based on segmented distance is adopted to enable the two arms to synchronously and collaboratively approach the target docking ring capture point in segments during the movement;
[0110] The implementation process of the segmented distance weighted speed optimization adjustment method is as follows:
[0111] (1) Obtain the distances d1(0) and d2(0) between the grasping points of the left and right manipulators under the hand-eye depth camera at the initial 2m position;
[0112] (2) Control the left and right robotic arms to synchronize at a distance of 0.2 m until they stop at the closest point of 0.2 m. The number of synchronization steps is N = 9. The movement speed of the left and right robotic arms at the current distance step is
[0113]
[0114] The segmented distance-weighted speed optimization adjustment method is described to achieve the acquisition of the target grasping point position of the dual-arm collaborative operation. Specifically, the movement speed of the dual robotic arms is adjusted according to the percentage weight of the distance information of each arm from the grasping point currently obtained during the movement of the dual robotic arms. A segmented synchronous collaborative approach to the target docking ring grasping point strategy is adopted to ensure that the left and right hand-eye cameras synchronously acquire the position of the target grasping point. It has the advantages of high target grasping point positioning accuracy and strong robustness.
[0115] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present invention, and these modifications or replacements should all be included in the scope of protection of the present invention.
Claims
1. A method for intelligent and rapid positioning of spatial target grasping points using dual-arm collaborative vision joint perception, characterized by Here are the steps: S1, intelligent autonomous recognition of global targets and instance segmentation; S11. Collect images of non-cooperative space targets during the robotic arm's on-orbit servicing and manipulation missions, annotate the samples, and complete the construction of a labeled training dataset for key components of non-cooperative space targets. S12. Use the training dataset constructed in S11 as input for a Mask R-CNN-based network model for rapid object detection and instance segmentation. Use real-time images acquired by a global, large-field-of-view RGB-D camera to input the model to achieve intelligent autonomous recognition and instance segmentation of non-cooperative target docking loops during the end-of-line approach of a robotic arm grasping task. The images of non-cooperative space targets should be rich and diverse, and simulate as realistically as possible the non-cooperative space targets of various types and motion states at multiple angles under different missions, different lighting, and different background conditions in a real space environment. Specifically, these images should include images of non-cooperative space targets under conditions such as stray light interference, complex and changing backgrounds, extremely harsh lighting conditions, occlusion of space targets, and tumbling or high-speed spinning of space targets. The sample annotation includes marking the instance segmentation boundaries and detection boxes of the target key components in each original image in the training sample and creating corresponding labels; The sample annotation is performed by using Labelme to annotate each collected original image. Specifically, the Create Polygons function in the Labelme tool is used to mark the instance segmentation boundaries of the target key components and create labels. Then, the Create Rectangle function is used on this basis to mark the detection box corresponding to the target key components and select the corresponding label; The key components of the space non-cooperative target include a dish antenna, a thruster, an earth observation camera, a solar panel, a solar sensor, an infrared horizon, a rigid arm, a feed source, a flexible arm, a solar panel bracket, a conical spiral antenna, a satellite body, a star sensor, a satellite-rocket docking ring, a docking surface, a laser data transmission, a nozzle, a measurement and control antenna, an end effector, and a GPS antenna; The global large field of view RGB-D camera is installed on the service satellite body, located at the center of the connection line of the dual robotic arm bases, and can realize the acquisition of global information of the non-cooperative target docking ring of the service satellite at the end of the robotic arm grasping task, which is 2 meters away from the operating satellite; The intelligent autonomous recognition and instance segmentation of the non-cooperative target docking ring adopts the Mask R-CNN algorithm to obtain the position of the target initial state positioning detection frame, target pixel level classification and accurate segmentation of target boundaries and other information parameters in real time; Specifically, the Mask R-CNN implementation steps are as follows: (1) Image feature extraction based on ResNeXt-101 network; (2) Generate candidate regions of interest (ROIs) through the RPN target estimation network for the obtained feature maps; (3) The obtained ROI is mapped into a fixed-dimensional feature vector through the ROI Align layer with bicubic linear interpolation to maintain the fineness of the feature pixels; two branches are processed by Faster RCNN to complete the target classification and bounding box regression, and the other branch is processed by the fully convolutional neural network FCN for upsampling to obtain pixel-level instance segmentation of the target object; The instance segmentation branch uses an FCN with 16x upsampling, which results in finer segmentation results and is more robust to smaller objects. Furthermore, the FCN adds five convolutional layers and three deconvolutional layers to the con5_3 layer of VGG-Net, enabling parallel processing of object recognition and detection branches, reducing overall task execution time. The multi-task loss function used in the training continuously reduces the value of the loss function through learning until the global optimal solution is obtained; the formula of the loss function is as shown in formula (1), where the three terms are classification error, bounding box error, and segmentation error: (1) After the improved Mask R-CNN algorithm, three outputs are obtained: the target category, the position of the bounding box, and the mask of the object; S2, fast and accurate segmentation of the local contour of the target based on joint visual perception; S21. Based on the results of the global target intelligent autonomous recognition and instance segmentation in S1, the target docking ring components are screened, and the position of the surrounding detection frame of the target docking ring and the target pixel segmentation mask are obtained; S22, synchronously mapping the target detection and segmentation results and the target docking ring center position to the depth map correspondingly obtained by the global large-field-of-view RGB-D camera to obtain the target docking ring center, radius, detection frame, and boundary three-dimensional information; S23, guiding the target docking ring to the center position of the image obtained by the global large-field-of-view RGB-D camera, and based on the three-dimensional information of the target docking ring center, detection frame, and boundary obtained at the current moment, achieving the positioning and acquisition of the initial position of the target docking ring under the high-precision hand-eye depth camera of the dual-arm collaborative end; S24. Based on the detected bounding box, the local window area of the target docking ring under the hand-eye depth camera at the end of the dual manipulator is obtained, and the edge detection and segmentation of the local target in the depth image of the area are performed. Combined with the pixel boundary results of the target instance segmentation, the local contour of the target is quickly and accurately segmented by the joint visual perception. The screening of the target docking ring components in step S21 is based on the results of the global target intelligent autonomous recognition and instance segmentation in step S1, and the target docking ring of interest is quickly screened out from multiple components, thereby realizing the rapid intelligent recognition and segmentation of the target docking ring components under different tasks based on the global depth camera, which has the advantages of high speed and high generalization. In step S23, the global large field of view RGB-D camera is located at the center of the dual robotic arm base at the end of the service satellite, and the local high-precision hand-eye depth camera is installed at the end of the robotic arm; The specific process of locating and acquiring the initial position of the target docking ring is as follows: (1) Dynamically adjust the positional relationship between the service satellite and the target docking ring according to the center coordinates of the target docking ring until the target docking ring is located at the center of the field of view of the global large-field-of-view RGB-D camera; (2) Start the local hand-eye camera at the end of the dual robotic arms and complete the calibration of the global large-field-of-view RGB-D camera and the high-precision hand-eye camera at the end of the service satellite robotic arm; (3) Map the three-dimensional information of the target docking ring center, radius, detection frame and boundary obtained by the global RGB-D camera at the current moment to the depth image observed by the high-precision hand-eye camera at the end of the dual manipulator, and complete the positioning and acquisition of the initial position of the target docking ring under the local high-precision hand-eye depth camera; The process of fast and accurate segmentation of the local contour of the target by visual joint perception is as follows: Based on the above S1 global target intelligent autonomous recognition and instance segmentation, the four vertices of the surrounding detection box of the target docking ring are obtained and target pixel segmentation mask Mask boundary point The coordinates are: and , Center coordinates of the target docking ring ; The three-dimensional information of the target docking ring center, radius, detection frame and boundary acquired by the global RGB-D camera is mapped to the depth map of the high-precision hand-eye camera at the end of the dual robotic arm; The three-dimensional information coordinates of the target docking ring center, radius, detection frame and boundary of the target docking ring initially positioned by the local high-precision hand-eye depth camera are as follows: The vertex coordinates of the target docking ring detection frame in the depth image of the left robotic arm's hand-eye depth camera: ; Coordinates of the virtual center of the target docking ring in the depth image of the left robotic arm's hand-eye depth camera ; Coordinates of the target docking ring boundary points in the depth image of the left robotic arm's hand-eye depth camera: ; The coordinates of the target docking ring boundary endpoints in the depth image of the left robotic arm's hand-eye depth camera: ; The vertex coordinates of the target docking ring detection frame in the depth image of the right robotic arm's hand-eye depth camera: ; Coordinates of the virtual center of the target docking ring in the depth image of the right robotic arm's hand-eye depth camera: ; Coordinates of the target docking ring boundary points in the depth image of the right robotic arm's hand-eye depth camera: ; The coordinates of the target docking ring boundary endpoints in the depth image of the right robotic arm's hand-eye depth camera: ; in: (2) In step S24, the acquisition of the local window area of the target docking ring under the hand-eye depth camera at the end of the dual manipulator is completed based on the detection frame. The local area observed by the hand-eye depth camera at the end of the left / right manipulator is and the area it comprises; In step S24, only the left / right robotic arm end hand-eye depth camera observation is required. and Detection and segmentation of the edge of the target in the local area depth image composed of the obtained segmentation edge points and ; In step S24, the pixel boundary results of the target instance segmentation are combined to complete the rapid and accurate segmentation of the local contour of the target in the joint visual perception. Specifically, the target edge points of the target instance segmentation are screened and the detection and segmentation edge map information of the local edge of the target in the depth image under the small windowed area is taken to obtain the final target depth map edge extraction result, which is as follows: Traverse each edge point of the depth image obtained by the hand-eye depth camera at the end of the left / right robotic arm and retain the edge points that meet the following conditions at the same time and ; (3) (4) The rapid and accurate segmentation of the local contour of the target by the joint visual perception adopts the method of only globally processing the initial position image, and then locally processing the local image of the scene containing the target or the image sequence to achieve continuous segmentation of the local features of the remaining sequence images. It does not need to perform edge detection, feature extraction and target segmentation on the entire satellite image, and has the advantage of good real-time performance. The joint visual perception method of global intelligent segmentation and local edge detection segmentation can effectively improve the accuracy of the target local contour area segmentation. S3, dynamic and precise positioning of the dual-arm collaborative target grasping point guided by vision, realizing the intelligent autonomous recognition and segmentation of the spatial target docking loop and the rapid and stable acquisition of the target grasping position by the dual-arm collaborative vision joint perception; S31. Based on the target edge information obtained by rapid and accurate segmentation of the local contour of the target based on joint visual perception, a piecewise combined third-order Bezier curve is used to continuously fit the local edge arc in the double-arm hand-eye depth image, and the farthest tangent point of the edge arc segment is dynamically obtained as the current target capture point; S32: When the distance between the two grasping points meets a certain condition, the grasping process is executed, and the position information of the grasping points is sent to the end robot arm controller; S33. Using a segmented distance-based weighted speed optimization adjustment method, the two arms can synchronously approach the target docking ring capture point in segments during movement, ensuring that the target is captured synchronously. The acquisition of the target capture point in step S31 is performed by using a segmented combined third-order Bezier curve to complete the continuous fitting of the local edge arc in the double-arm hand-eye depth image, and dynamically obtain the farthest tangent point of the edge arc segment; The segmented combined third-order Bezier curve completes the continuous fitting of the local edge arc in the double-arm hand-eye depth image: The dynamic acquisition of the farthest tangent point of the edge arc segment is specifically performed by traversing each point on the curve and selecting the leftmost and rightmost points corresponding to the left and right robotic arms as the capture points; In step S32, when the distance between the two capture points meets certain conditions, the capture process is executed by calculating the distance d between the two capture points normalized to the right camera coordinate system. If d is greater than 4 / 5 times the diameter of the docking ring, the capture requirement is met; if the distance requirement is not met, the right capture point is used as a reference, the right capture point is connected with the virtual circle center coordinate, and the intersection of the extension line and the left robotic arm depth image is obtained. The capture point is searched within a certain area r=0.001m around the intersection to ensure that the two arms can capture the target area far enough at the same time and can stably capture the same target object. In step S33, a weighted speed optimization adjustment method based on segmented distance is adopted to enable the two arms to synchronously and collaboratively approach the target docking ring capture point in segments during the movement; The segmented distance weighted speed optimization adjustment method: The segmented distance-weighted speed optimization adjustment method is described to achieve the acquisition of the target grasping point position of the dual-arm collaborative operation. Specifically, the movement speed of the dual robotic arms is adjusted according to the percentage weight of the distance information of each arm from the grasping point currently obtained during the movement of the dual robotic arms. A segmented synchronous collaborative approach to the target docking ring grasping point strategy is adopted to ensure that the left and right hand-eye cameras synchronously acquire the position of the target grasping point. It has the advantages of high target grasping point positioning accuracy and strong robustness.
Citation Information
Patent Citations
Double-mechanical-arm grabbing system control method based on multi-view vision
CN115194774A