A Grasping Perception Method for a Space Manipulator in Complex Scenario Tasks

By adopting a hybrid algorithm of long-distance grayscale feature detection, lightweight network recognition and segmentation network segmentation on the space robot arm, the problem of illumination changes and occlusion interference in the capture task in the microgravity environment is solved, and high-precision target recognition and attitude estimation are achieved, reducing the computational complexity.

CN116476059BActive Publication Date: 2025-07-29CHINA ORDNANCE SCI INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310441415.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-23
Publication Date
2025-07-29
Estimated Expiration
2043-04-23

AI Technical Summary

Technical Problem

When performing the capture task in a microgravity environment, the space robotic arm faces problems such as hardware resource limitations, lighting changes, occlusion interference, complex backgrounds and target diversity, which makes it difficult for traditional image processing methods to accurately segment and identify targets, and the deep detection network has a large amount of computation and strong inadaptability.

Method used

A hybrid algorithm of long-distance grayscale feature detection, lightweight deep convolution network recognition and lightweight segmentation network segmentation is adopted, and target position and posture information is obtained by combining time-division three-objective positioning, including BP neural network, lightweight YoloTinyV2 and DDRNet-SLim network, target recognition and segmentation through multi-angle observation.

Benefits of technology

The target recognition accuracy has been improved to 98.7%, greatly reducing the calculation amount, significantly improving adaptability and accuracy, and is suitable for the capture tasks of space robot arms in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116476059B_ABST
    Figure CN116476059B_ABST
Patent Text Reader

Abstract

The present invention discloses a grasping perception method for a space manipulator facing complex scenario tasks. When at a long distance, an image of the extravehicular operation scene area is captured by an extravehicular operation monitoring camera, and the presence or absence of the target is detected through a BP neural network using the gray-scale distribution characteristics of the image to obtain the effective target area where the target is located. During the approaching process, through a lightweight deep convolutional detection and recognition network, continuous target recognition, classification, and position calculation are performed to obtain a set of effective target areas after recognition and classification. When at a close distance, through the observation images from multiple angles of the camera, a lightweight segmentation convolutional network is used to perform semantic segmentation on the target to find the area of the target in the image, and binary data of the target shape from multiple angles is obtained. According to the obtained binary data of the target shape from multiple angles, time-division three-eye positioning is used to obtain the position and attitude information of the target, and the feedback control central controller is used to complete the grasping task, which not only improves the accuracy but also greatly reduces the calculation speed through a hierarchical strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of space manipulators, and more specifically, to a space manipulator grasping perception method for complex scenario tasks. Background Art

[0002] With the continuous construction of space stations and the proposal of lunar research station tasks, the autonomous repair and maintenance of equipment equipped on extraterrestrial vehicles and experimental modules have become a research hotspot. Equipping manipulators and mobile robots with end-effectors and supporting relevant visual perception systems to complete fine operations such as grasping, placing, and interacting with astronauts outside the cabin has become an important direction for intelligent space services. In this process, high-precision position and attitude perception of cooperative and non-cooperative components and operating tools is the basis for the normal operation of the entire system.

[0003] There are mainly the following three types of problems in space target position and attitude perception compared with the ground: During the extravehicular repair and maintenance process, due to space, lighting, and camera performance limitations, the imaging signal-to-noise ratio is low and the target features are not obvious, resulting in inaccurate extraction results after processing; during the movement of the manipulator, the lighting and occlusion state of the target imaging area change, and it is difficult for the algorithm to adapt, resulting in an increase in the error of the time-series result fusion process; in the microgravity environment, objects such as components and tools undergo spatial displacement and rotation, resulting in the invalidation of static analysis results.

[0004] In response to the need for high-precision attitude estimation during the grasping process, traditional manipulator grasping perception is mainly divided into two categories. One is feature point detection based on traditional image processing methods, and the other is recognition relying on deep convolutional networks. The former has a small computational load and poor adaptability, while the latter has strong adaptability but a large computational load. Moreover, for space manipulators, the control algorithm is still different from that on the ground. How to operate well under microgravity and floating states, and whether the recognized position and attitude are suitable for the application of the control algorithm, requires systematic adaptation design for the algorithm, which belongs to both system design problems and the applicability problems of perception algorithms.

[0005] Therefore, in view of the problems of limited hardware resources, lighting changes, occlusion interference, complex backgrounds, and target diversity existing in the space manipulator grasping process, how to provide a space manipulator grasping perception method to replace traditional image processing methods to solve the following problems: First, the light and dark differences caused by the target and complex background structures make it difficult for traditional image processing methods to accurately segment the foreground and background; second, the weak light of the optical fiber affects the recognition of certain specific reflection difference angles by traditional image processing methods; third, the traditional image processing threshold segmentation algorithm cannot effectively segment low signal-to-noise ratio and low-luminance targets; fourth, the general depth detection and segmentation networks have a large computational load and need to be combined with traditional image processing algorithms to form an effective engineering application. Summary of the Invention

[0006] In view of this, the present invention provides a grasping perception method for a space manipulator facing complex scenario tasks to solve the problems mentioned in the background art.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] A grasping perception method for a space manipulator facing complex scenario tasks, comprising the following steps:

[0009] S1. When at a long distance, an image of the extravehicular operation scenario area is captured by an extravehicular operation monitoring camera, and the presence or absence of a target is detected through a BP neural network using the gray-scale distribution characteristics of the image to obtain the effective target area where the target is located;

[0010] S2. During the approach process, the obtained effective target area is input into a lightweight deep convolutional detection and recognition network, and continuous target recognition classification and position calculation are performed to obtain a set of effective target areas after recognition classification;

[0011] S3. When at a short distance, through the observation images of the camera at multiple angles, a lightweight segmentation convolutional network is used to perform semantic segmentation on the target, find the area of the target in the image, and obtain the specific position and pose shape binary data of the target at multiple angles;

[0012] S4. According to the obtained specific position and pose shape binary data of the target at multiple angles, time-division three-eye positioning is used to obtain the position and pose information of the target, and the feedback control central controller is used to complete the grasping task.

[0013] Preferably, in S1, the specific content of detecting the presence or absence of a target through a BP neural network using the gray-scale distribution characteristics of the image to obtain the area where the target is located is as follows: a sliding window is used to calculate the gray-scale histogram of the overlapping area, the image is split into block areas, the gray-scale histogram of each block area is statistically obtained to obtain the area gray-scale histogram information, and a BP neural network is used to perform inference calculation on whether the gray-scale histogram information of each area is valid to obtain the area where the target is located, that is, the set of effective target areas.

[0014] Preferably, the BP neural network adopts a three-layer BP neural network, the activation function is the LeakyRelu function, and the calculated output result y (i,j) = 0, 1, when it is 1, it means that the block area Q (i,j) is an effective target area, otherwise it is an invalid area, and the set of effective target areas is where k = 0, 1,..., K - 1, and K represents the number of effective target areas.

[0015] Preferably, the lightweight deep convolutional detection and recognition network adopts a cropped YoloTinyV2 network and removes 8 BN layers.

[0016] Preferably, the set of effective target regions after recognition is where k' = 0, 1, …, K'-1, and K' represents the total number of effective target regions with non-zero numbers of recognized classification targets.

[0017] Preferably, in S3, the lightweight segmentation convolutional network adopts the lightweight depth double-resolution network DDRNet-SLim.

[0018] Preferably, the position and pose information of the target in S4 includes the target azimuth pitch angle, position, and distance information.

[0019] Preferably, the specific content of S4 includes:

[0020] S41. The camera is placed at the center directly above the operating table. After imaging, the stent target is segmented to obtain the centroid position, and the distances in the horizontal and vertical directions from the centroid to the edge are calculated to obtain the value of the horizontal direction azimuth angle.

[0021] S42. The camera is respectively moved to the left rear, right rear azimuths, and 30 degrees of pitch of the operating table. According to the positions of the vertices after segmentation and the number of pixels in the vertical direction starting from the vertices, the pitch angle of the stent target is judged. Another perspective is used for fusion, and the left rear data is switched to when the illumination condition is poor at the right rear.

[0022] S43. According to the calculated target azimuth pitch angle, position, and distance information, it is sent to the central controller of the whole arm for the execution of the grasping task.

[0023] Preferably, before S1, offline simulation training and testing are also included for the BP neural network, lightweight depth convolutional detection and recognition network, and lightweight segmentation convolutional network.

[0024] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a grasping perception method for a space manipulator facing complex scenario tasks, which replaces the traditional image processing method and adopts a hybrid algorithm strategy of long-distance gray feature search, lightweight detection network classification, and lightweight segmentation network to obtain the centroid, position, and attitude. When the end is far from the target, within the field of view, for high-resolution images, a simple BP neural network is used to perform target presence detection through sliding window histogram feature screening, quickly determining the area where the target of interest is located. The effective slices obtained only account for 1% of the original image, greatly reducing the subsequent calculation amount. The area with the target is extracted, and during the approaching movement process, a lightweight deep convolutional detection and recognition network is used to continuously identify and classify the target, determine the key parts to be grasped, and determine whether the target is a valid target. After secondary screening by the lightweight detection and recognition network, the false alarm rate is further reduced. At the same time, the real target is framed, and the framed part is further sliced, further reducing the calculation amount and improving the accuracy of the target presence in the slices. When reaching a relatively close distance, a lightweight segmentation convolutional network is used to perform semantic segmentation on the target to obtain shape binary data at multiple angles, so as to obtain the position and attitude information of the target according to binocular stereo vision. Finally, the target recognition and classification accuracy reaches 98.7%. Compared with previous designs, it not only improves the accuracy but also greatly reduces the calculation amount in high-resolution large-pixel images through a hierarchical strategy. The simulation verification test of the data shows the effectiveness and adaptability of the method. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0026] Figure 1 The drawings are the schematic flowcharts of the space manipulator grasping perception method provided by the present invention;

[0027] Figure 2 The drawings are the schematic diagrams of the gray feature statistics of the extravehicular operation scenario area image provided by the present invention;

[0028] Figure 3 The drawings are the schematic diagrams of the effective target area calculation process of the gray histogram based on the BP neural network provided by the present invention;

[0029] Figure 4 The drawings are the schematic diagrams of the cropped YoloTinyV2 network structure provided by the present invention;

[0030] Figure 5The attached figure is a schematic diagram of the recognition and classification results of the operating platform at different distances and observation postures provided by the present invention;

[0031] Figure 6 The attached figure is a schematic diagram of the DDRNet-Slim framework of the target segmentation network provided by the present invention;

[0032] Figure 7 The attached figure is a schematic diagram of the target time-division three-eye positioning provided by the present invention;

[0033] Figure 8 The attached figure is a schematic diagram of the target recognition and classification results under low signal-to-noise ratio and low light intensity provided by the present invention;

[0034] Figure 9 The attached figure is a schematic diagram of the target segmentation results at different postures at close range provided by the present invention. Detailed implementation manners

[0035] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the attached figures in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0036] The embodiment of the present invention discloses a grasping perception method for a space manipulator facing complex scene tasks, such as Figure 1 , including the following steps:

[0037] S1. Detection of effective target area: At a long distance, an image of the extravehicular operation scene area is captured by an extravehicular operation monitoring camera, and the presence or absence of the target is detected through a BP neural network using the gray-scale distribution characteristics of the image to obtain the effective target area where the target is located;

[0038] S2. Target recognition and classification: During the approach process, the obtained effective target area is input into a lightweight deep convolutional detection and recognition network, and continuous target recognition, classification, and position calculation are performed to obtain a set of effective target areas after recognition and classification;

[0039] S3. Close-range target segmentation: At a close range, through the observation images of the camera at multiple angles, a lightweight segmentation convolutional network is used to perform semantic segmentation on the target, find the area of the target in the image, and obtain the binary data of the specific position and posture shape of the target at multiple angles;

[0040] S4. According to the binary data of the specific position and posture shape of the target obtained at multiple angles, the time-division three-eye positioning is used to obtain the position and posture information of the target, and the feedback control central controller is used to complete the grasping task.

[0041] In order to further implement the above technical solution, S1 uses the grayscale distribution characteristics of the image to detect the presence of the target through the BP neural network. The specific content of obtaining the target area is as follows: using a sliding window to calculate the grayscale histogram of the overlapping area, splitting the image into block areas, performing grayscale histogram statistics on each block area to obtain regional grayscale histogram information, and using the BP neural network to perform inference calculation to determine whether the grayscale histogram information of each area is valid, and obtaining the target area, that is, the valid target area set, such as Figure 2 .

[0042] In this embodiment, the image in the detection process is defined as I N , where N = 2048, and a window of M = 256 is used to calculate the grayscale histogram with an overlapping area of P = 64, then I N It is divided into (N+P+M / 2) / [M / 2+P]=11 rows / columns, a total of 121 blocks, Q (i,j) Represents a block area, where i,j∈[0,M-1], i,j∈Z+.

[0043] The scenes outside the space capsule mainly include three categories: the surface of the spacecraft, the deep space background, and the external payload. Among them, the surface of the spacecraft and the deep space background are both grayscale monotonous areas. The external payload may have a grayscale distribution similar to that of the manipulator console and the operating area, but it is generally installed in a specific area and position and can be masked out by using the predicted information. Therefore, the preliminary elimination and confirmation of the effective area can be directly performed by the grayscale histogram statistics method. (i,j) Perform statistics to obtain the regional 8-bit grayscale histogram information h (i,j) ,Effective slices are generally more regular, with grayscale concentrated in a few areas, while invalid areas will have more grayscale distribution interference caused by natural background or single background.

[0044] In order to further implement the above technical solutions, Figure 3 , BP neural network uses a three-layer BP neural network, the activation function is LeakyRelu function, and the calculated output result y (i,j) =0,1, when it is 1, it indicates the block area Q (i,j) is a valid target area, otherwise it is an invalid area. The set of valid target areas is Wherein, k = 0, 1, ..., K-1, K represents the number of valid target areas.

[0045] In order to further implement the above technical solutions, Figure 4 ,The lightweight deep convolution detection and recognition network adopts the cropped YoloTinyV2 network and removes 8 layers of BN layers. In this process, false alarms in the effective target area are eliminated, and the target is classified and located at the same time.

[0046] The network description after cropping optimization is as follows in the table:

[0047] 1) The first layer takes a 224×224 pixel image as input, performs calculations using 16 3×3 convolutional kernels, and outputs 16 112×112 pixel images;

[0048] 2) The second layer takes the 16 112×112 images from the first layer as input, inputs them into 32 16×3×3 convolutional kernels for calculation, and obtains 32 56×56 pixel images;

[0049] 3) The third layer takes the 32 56×56 images from the second layer as input, inputs them into 64 32×3×3 convolutional kernels for calculation, and obtains 64 28×28 pixel images;

[0050] 4) The fourth layer takes the 64 28×28 images from the third layer as input, inputs them into 128 64×3×3 convolutional kernels for calculation, and obtains 128 14×14 pixel images;

[0051] 5) The fifth layer takes the 128 14×14 images from the fourth layer as input, inputs them into 256 128×3×3 convolutional kernels for calculation, and obtains 256 7×7 pixel images;

[0052] 6) The sixth layer takes the 256 7×7 images from the fifth layer as input, inputs them into 512 256×3×3 convolutional kernels for calculation, and obtains 512 7×7 pixel images;

[0053] 7) The seventh layer takes the 512 7×7 images from the sixth layer as input, inputs them into 1024 512×3×3 convolutional kernels for calculation, and obtains 1024 7×7 pixel images;

[0054] 8) The seventh layer takes the 1024 7×7 images from the sixth layer as input, inputs them into 125 1024×1×1 convolutional kernels for calculation, and obtains data of 1024×125;

[0055] 9) The eighth layer takes the result 1024×125 from the seventh layer as input, inputs it into the Yolo layer, and obtains a 7×7×5×20 calculation output.

[0056] Number of Layers Input Output Number of Convolutions Remarks 1 224×224 112×112×16 16 3×3 Convolution 2 112×112×16 56×56×32 16×32 3×3 Convolution 3 56×56×32 28×28×64 32×64 3×3 Convolution 4 28×28×64 14×14×128 64×128 3×3 Convolution 5 14×14×128 7×7×256 128×256 3×3 Convolution 6 7×7×256 7×7×512 256×512 3×3 Convolution 7 7×7×512 7×7×1024 512×1024 3×3 Convolution 8 7×7×1024 7×7×125 1024×125 1×1 Convolution 9 7×7×125 7×7×5×20 - -

[0057] In this embodiment, for any Q′ k, first, perform zero-padding calculation on the pixels, and then input them into the cropped YoloTinyV2 network. YoloTinyV2 consists entirely of 9 layers. The input data is 416×416×3 pixels. Modify the input data to 224×224 pixels and remove the BN layers of the other 8 layers to obtain the recognition and classification results of the pitch angle of -45°, the pitch angle of 10°, and the close-range pitch angle of 10°, such as Figure 5 。

[0058] To further implement the above technical solution, the set of effective target regions after recognition is where k' = 0, 1, …, K'-1, and K' represents the total number of effective target regions with non-zero recognition and classification target numbers.

[0059] To further implement the above technical solution, in S3, the lightweight segmentation convolutional network adopts the lightweight deep double-resolution network DDRNet-SLim, such as Figure 6 。

[0060] To further implement the above technical solution, in S4, the position and attitude information of the target includes the target azimuth pitch angle, position, and distance information.

[0061] To further implement the above technical solution, such as Figure 7 , the specific content of S4 includes:

[0062] S41. Place the camera at the center directly above the operating table. After imaging, perform segmentation on the stent target to obtain the centroid position, calculate the distances in the horizontal and vertical directions from the centroid to the edge, and obtain the value of the horizontal direction azimuth angle;

[0063] S42. Move the camera to the left rear, right rear azimuth, and 30 degrees of pitch of the operating table respectively. According to the position of the vertices after segmentation and the number of pixels in the vertical direction starting from the vertices, judge the pitch angle of the stent target. Another perspective is used for fusion. At the same time, when the lighting condition is poor at the right rear, switch to the left rear data;

[0064] S43. Send the calculated target azimuth pitch angle, position, and distance information to the central controller of the entire arm for the execution of the grasping task.

[0065] To further implement the above technical solution, before S1, it also includes offline simulation training and testing of the BP neural network, lightweight deep convolutional detection and recognition network, and lightweight segmentation convolutional network.

[0066] In this embodiment, the experimental test data are all obtained by simulation using the model library, and simulated data are used for training in the offline stage. Different scenarios of data at different distances (the target size ranges from 50 to 500 pixels from far to near) are tested, and are divided into different postures, illumination angle changes, superposed noise and other conditions. And the detection of the effective target area, the target recognition and classification, and the close-range target segmentation are respectively tested.

[0067] Specifically, the test of the effective target area detection:

[0068] The training and test data include three categories. The first category is high signal-to-noise ratio data with a signal-to-noise ratio of more than 6, the second category is low signal-to-noise ratio data with a signal-to-noise ratio between 3 and 6, and the third category is images with a gray level of 10% to 20% of the full amplitude under low illumination. First, 30 images of each of the three categories are split, obtaining 3630 images of 256×256 pixels, which are input into a BP neural network for training. The total parameters obtained from training are 264 KB.

[0069] Comparison of the test results of the effective target area detection for different types of data:

[0070]

[0071]

[0072] For the split test of another 30 images of the three categories, it can be seen that the first category can reach an accuracy rate of 97.3%, but at the same time the false alarm rate is as high as 12.6%, indicating that there is a gray level distribution in some areas outside the cabin that is relatively consistent with the target under the noise background; the detection accuracy rate of the second category reaches 96.5%, and the false alarm rate reaches 29.7%, which also shows that the accuracy rate will be greatly affected when the noise is high. This is mainly because the gray level of the target itself is close to that of the background and the cabin body, and it is difficult to distinguish the statistical distribution of the foreground and background due to the high noise; for the data of the third category under low light, the detection rate reaches more than 95%, and the false alarm rate is as high as 48.3%, which also shows that the illumination intensity has a greater impact on the detection accuracy rate and false alarm rate.

[0073] At a detection accuracy rate of 95.1%, considering that the operating platform can span at most four sliding window areas, 95% of the data for subsequent processing can be reduced, greatly reducing the huge computational amount brought by the subsequent CNN application. At the same time, it also shows that it is difficult to suppress the false alarm rate by simply statistically analyzing the regional features through the gray level histogram, and subsequent accurate recognition is still needed.

[0074] Target recognition and classification test:

[0075] In this embodiment, 500 images of 256×256 pixels with an operating table and a target at different signal-to-noise ratios are used for training, and another 500 images are used for testing. The training and test data sets both include data of the target at various angles.

[0076] For data with a gray value of not less than 100, target recognition and classification are performed to obtain the test results of recognition and classification accuracy at different signal-to-noise ratios (target luminosity: gray value not less than 120) and the test results of recognition and classification accuracy under different light intensities (target signal-to-noise ratio: SNR = 7):

[0077]

[0078] When the signal-to-noise ratio is not greater than 3, the detection accuracy is below 90%. When the signal-to-noise ratio is not less than 7, the detection accuracy is better than 95%.

[0079] For the image data with a signal-to-noise ratio equal to 7, the recognition results are calculated for different low-luminance grayscales. Among them, when the gray value is less than 40, the detection accuracy drops significantly to 88.3%. When the image gray value increases to more than 60, the accuracy is higher than 91.1%.

[0080] The results of recognition and classification are used to confirm whether there is a target in the image slice, and at the same time, to further determine the position of the target in the image, facilitating the approach of the end of the robotic arm, such as Figure 8 。

[0081] Close-range target segmentation test:

[0082] During the close-range imaging process, target segmentation is adopted. At this time, the end camera is close enough to the operating table, and it can be considered that the operating table occupies most of the field of view. At this time, since the end generally has an illumination LED lamp, the image signal-to-noise ratio is generally not low at close range. The biggest problem of target segmentation at this time is to cope with the segmentation accuracy under different targets and different postures, as well as the resulting calculation errors of the target centroid, azimuth / pitch estimation angles.

[0083] After testing, four types of targets, namely the camera, the bracket, the protective cover, and the screwdriver, are tested in different postures. The centroid estimation error is better than 1.3 pixels, and the posture estimation is better than 0.3 degrees, such as Figure 9 。

[0084] Using a computer with an i7-11800 model CPU, a GTX3060 graphics card, and 16GB of memory for calculation, combined with the calculation results of effective area detection, target recognition and classification, and target segmentation, the final classification accuracy of the target is shown in the following table. The recognition accuracy is improved by 1.4%, and the calculation speed is reduced by 58%.

[0085]

[0086] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.

[0087] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A grasping perception method for a space manipulator facing complex scenario tasks, characterized in that, It includes the following steps: S1. At a long distance, the camera for out-of-cabin operation monitors the image of the out-of-cabin operation scene area, and uses the gray-scale distribution characteristics of the image to detect the presence or absence of the target through a BP neural network, and obtains the effective target area where the target is located; S2. During the approach process, the obtained effective target area is input into the lightweight deep convolutional detection and recognition network, and the target recognition classification and position calculation are continuously performed to obtain the set of effective target areas after recognition classification; S3. At a short distance, through the observation images of the camera at multiple angles, the lightweight segmentation convolutional network is used to perform semantic segmentation on the target, find the area of the target in the image, and obtain the binary data of the specific position, attitude and shape of the target at multiple angles; S4. According to the obtained binary data of the specific position, attitude and shape of the target at multiple angles, the time-division three-eye positioning is used to obtain the position and attitude information of the target, and the feedback control central controller is used to complete the grasping task.

2. The grasping perception method of a space manipulator for complex-scene tasks according to claim 1, characterized in that The specific content of using the gray-scale distribution characteristics of the image through the BP neural network to detect the presence or absence of the target in S1 and obtain the area where the target is located is as follows: The sliding window is used to calculate the gray-scale histogram of the overlapping area, the image is split into block areas, the gray-scale histogram of each block area is statistically obtained to obtain the area gray-scale histogram information, and the BP neural network is used to perform inference calculation to determine whether the gray-scale histogram information of each area is valid, and the area where the target is located, that is, the set of effective target areas, is obtained.

3. A grasping perception method for a space manipulator facing complex scene tasks according to claim 2, characterized in that, The BP neural network adopts a three-layer BP neural network, the activation function is the Leaky Relu function, and the calculated output result is y (i,j) = 0, 1. When it is 1, it represents the block area Q (i,j) is the effective target area, otherwise it is the invalid area. The set of effective target areas is where k = 0, 1, …, K - 1, and K represents the number of effective target areas.

4. A grasping perception method for a space manipulator facing complex scenario tasks according to claim 1, characterized in that, The lightweight deep convolutional detection and recognition network uses the cropped Yolo Tiny V2 network and removes 8 BN layers.

5. A grasping perception method for a space manipulator facing complex scenario tasks according to claim 3, characterized in that, The set of valid target regions after recognition is where k' = 0, 1, …, K'-1, and K' represents the total number of valid target regions with non-zero recognized classification targets.

6. A grasping perception method for a space manipulator facing complex scenario tasks according to claim 1, characterized in that, In S3, the lightweight segmentation convolutional network uses the lightweight deep double-resolution network DDRNet-SLim.

7. A grasping perception method for a space manipulator facing complex scenario tasks according to claim 1, characterized in that, In S4, the position and attitude information of the target includes the azimuth pitch angle, position and distance information of the target.

8. A grasping perception method for a space manipulator facing complex scenario tasks according to claim 7, characterized in that, The specific content of S4 includes: S41. The camera is placed at the center directly above the operation table. After imaging, the bracket target is segmented to obtain the centroid position, and the distances in the horizontal and vertical directions from the centroid to the edge are calculated to obtain the value of the azimuth angle in the horizontal direction; S42. The camera is respectively moved to the left rear, right rear azimuth and 30-degree pitch positions of the operation table. According to the position of the vertex after segmentation and the number of pixels in the vertical direction starting from the vertex, the pitch angle of the bracket target is judged. Another perspective is used for fusion, and when the illumination condition is poor at the right rear, the data from the left rear is used instead; S43. According to the calculated azimuth pitch angle, position and distance information of the target, it is sent to the central controller of the whole arm for the execution of the grasping task.

9. A grasping perception method for a space manipulator facing complex scenario tasks according to claim 1, characterized in that Before S1, it also includes offline simulation training and testing of the BP neural network, lightweight deep convolutional detection and recognition network, and lightweight segmentation convolutional network.