Space station D-shaped hole bolt target detection method based on YOLOv5 improvement
By improving the YOLOv5 algorithm and combining MobileNetv3 and CBAM attention mechanisms, the precise identification and positioning of the D-shaped bolts in the space station in a narrow environment is solved, efficient target detection and ranging functions are realized, and automated operation of the space station is supported.
Patent Information
- Application Number
- CN202510325264.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-18
AI Technical Summary
The existing space dexterity devices have insufficient precise operation and autonomous decision-making capabilities in narrow environments, especially in identifying and positioning the D-hole bolts of the space station.
Using the improved object detection algorithm based on YOLOv5, by introducing the MobileNetv3 structure and CBAM attention mechanism, combined with the distance measurement algorithm, an object detection model for D-type bolts is constructed to achieve accurate identification and distance calculation of D-type bolts.
It realizes high-precision identification and distance measurement of D-hole bolts in orbit environments, improves the accuracy and efficiency of automated operations, and is suitable for target detection tasks in the narrow space of the space station.
Smart Images

Figure CN120339576A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of target detection methods, and specifically relates to a method for detecting D-hole bolts on a space station. Background Art
[0002] As the space competition becomes a key indicator for evaluating a country's comprehensive national strength, the advancement of space on-orbit manipulation technology has become a strategic focus for all countries. The space dexterous device integrates technologies such as target detection and can perform tasks such as on-orbit equipment maintenance and target capture, improving space operation capabilities. In 2010, the European Robotic Arm (ERA), jointly developed by the European Space Agency and the Russian Space Agency, was transported to the Russian segment of the International Space Station and put into use. The ERA is approximately 11.3 m long, weighs 630 kg, and has an accuracy of up to 3 mm. The work content of the ERA robotic arm is to install and maintain the solar panels of the Russian segment, can walk autonomously outside the cabin, and assist astronauts in completing some tasks. In 2015, in order to maintain the Hubble Space Telescope (HST), NASA proposed the Hubble Robotic Servicing Vehicle, which includes a deorbit module and a pop-up module equipped with a robotic arm. The former is used to provide services for the HST, and the latter is used to push the HST out of orbit when it retires. In 2020, MDA Space signed a development contract for the Canadarm 3 for NASA's Lunar Gateway project. The Canadarm 3 robotic arm is an 8.5 m long dexterous robotic arm that integrates technologies such as machine vision and planning software. In addition to regular operations, it mainly cooperates with NASA to accelerate the lunar exploration program. Domestically, the Extravehicular Mobile Robot (EMR) developed by the Beijing Institute of Control Engineering is an early representative achievement. It has 9 joints and 14 sensors, has the ability to walk and operate, and can complete tasks such as the installation and removal of screws, the plugging and unplugging of plugs, and the picking up of floating objects. In 2016, the space laboratory "Tiangong-2" successfully launched a robotic arm system integrating a six-degree-of-freedom lightweight robotic arm and a five-finger anthropomorphic dexterous hand, equipped with global stereo vision and a monocular hand-eye camera. In November 2022, the robotic arm of the core module "Tianhe" of the space station and the robotic arm of the experimental module "Wentian" have completed on-orbit combined tests. Its visual monitoring and measurement system includes an Ethernet switch, wrist and elbow cameras, etc., realizing functions such as video monitoring, real-time target recognition, and pose measurement, marking an important progress in China's space robotic arm technology.
[0003] Although existing technical solutions have made significant progress in the research and development of dexterous space devices, they still face limitations such as limited on-orbit operation accuracy and insufficient autonomous capabilities, especially in precise operations and autonomous decision-making in confined environments. Summary of the invention
[0004] In view of the limitations of on-orbit control and perception recognition research in the prior art, the present invention discloses a set of target detection algorithms based on visual sensing, which can accurately identify and locate the D-hole bolts of the Alpha solar orientation device of the Chinese space station, navigate the end tool to the vicinity of the D-hole bolts, and lay the foundation for subsequent automated operations.
[0005] The present invention is achieved in that:
[0006] A method for detecting D-type hole bolts in a space station based on an improved YOLOv5, the method comprising: step one, collecting original images of D-type hole bolts in the space station; step two, labeling the images to obtain a D-type hole bolt image dataset; step three, optimizing and improving the YOLOv5 model to obtain an optimized model for D-type hole bolt image detection; step four, using the optimized model obtained in step three to train the dataset in step two to obtain a trained detection model; step five, using the trained detection model to detect D-type hole bolt images.
[0007] Further, the step 1 is specifically as follows:
[0008] The method of directly holding a mobile phone to shoot simulates the running trajectory of the camera on the robot arm in actual operation. A total of 277 images were collected from different angles and directions through this method; during the shooting process, additional interference objects similar in appearance to D-type hole bolts were introduced as negative samples;
[0009] Further, the step 2 is specifically as follows:
[0010] After completing the image sampling, the original image dataset was divided into a training set, a test set, and a validation set according to the ratio of 6:2:2. Then, the labelimg tool was used to annotate the images in the training set and the test set, where the D-type hole bolts were simply labeled as d and J599 was labeled as j.
[0011] Further, the step three is specifically as follows:
[0012] 3.1. Improvement of image recognition algorithm based on YOLOv5:
[0013] Based on the improvement of the YOLOv5 algorithm, the Conv, C3, and FPP structures in the Backbone network are replaced with the MobileNetv3 structure; MobileNetv3 uses depthwise separable convolutions to replace traditional convolutions (Conv); at the same time, the attention mechanism CBAM algorithm is introduced to replace the original C3 module in the Head; CBAM integrates a channel attention module and a spatial attention module. The channel attention module can enhance the feature expression of channels, and the spatial attention module enhances the recognition accuracy of the model for small targets, which can improve the classification performance of the image recognition algorithm.
[0014] 3.2. Introduction of the ranging algorithm:
[0015] In the image detection algorithm, the distance between the image and the camera center can be calculated by formula (1);
[0016]
[0017] In formula (1): D is the distance from the target to the camera, F is the pixel focal length of the camera, W is the width or height of the target, and P is the number of pixels occupied by the target in the image.
[0018] The pixel focal length of the camera is calculated according to the size and resolution of the camera's photosensitive element.
[0019] The formula for converting the known focal length (pixels) to millimeters is:
[0020]
[0021] In formulas (2) and (3): dx and dy are the values of the pixel focal length in two directions, S is the size of the camera CCD, and x and y are the number of pixels of the image in two directions.
[0022] When the camera lens is parallel to the plane where the D-shaped hole is located, the distance of the target object can be calculated through the relevant parameters of the camera and the size of the D-shaped hole bolt.
[0023]
[0024] In formula (4): the actual distance between the two centers is d r1 , the number of pixels between the two centers in the picture is d p1 , the diameter of the D-shaped hole is d r , the number of pixels occupied by the height of the D-shaped hole is d p ;
[0025] Without considering lens distortion, the horizontal and vertical distances between the center of the D-shaped hole bolt and the camera center are approximately calculated by formula (4).
[0026] Furthermore, the step four is:
[0027] Training results of the D-shaped hole detection model improved based on YOLOv5:
[0028] After completing the collection and preprocessing process of image data, the model was trained using computer resources, and finally a simplified model for D-shaped hole detection was constructed;
[0029] To deeply evaluate the accuracy of the model, a two-dimensional analysis tool, the confusion matrix, was introduced; from the confusion matrix, we can clearly observe that the model has extremely high accuracy when predicting D-shaped holes. When the actual object is a D-shaped hole, the probability that the model predicts it as a D-shaped hole is as high as 0.99; similarly, when predicting J599, the model also shows extremely high accuracy, and the probability that the actual J599 is correctly predicted by the model reaches 1;
[0030] Check the overfitting situation of the model through the F1-score-confidence curve;
[0031] Further, step five is as follows:
[0032] When only the edges or corners of the D-shaped hole and J599 are exposed in the field of view, according to the test results, the D-shaped hole bolt detection model can still perform excellently in identifying the target.
[0033] To further test the recognition performance of the model, simulate the problems of D-shaped hole bolt wear, partial occlusion, and interference of the photographing device by cosmic radiation that may exist in the real situation. Add noise and perform random occlusion processing on some pictures to simulate the damaged situation of D-shaped hole bolts during on-orbit service; after obtaining the image recognition results, use the pictures with drawn rectangular frames for ranging; the target detection model with both D-shaped hole recognition and ranging functions will provide data support for the next path planning algorithm.
[0034] The beneficial effects of the present invention compared with the prior art are as follows:
[0035] The present invention proposes an innovative method for object detection for the automatic recognition and operation task of D-shaped hole bolts in the narrow space of the Chinese space station. Based on the YOLOv5 algorithm, the algorithm is lightweighted by introducing the MobileNetv3 structure, and the CBAM attention mechanism is added to improve the detection performance. Experimental results show that the optimized object detection algorithm has an identification accuracy of 0.99, can quickly and accurately identify D-shaped hole bolts, and calculate the distance information of the target object.
[0036] The improved YOLOv5 object detection algorithm of the present invention can be transferred to the recognition tasks of other small parts on the space station. By replacing the training dataset, it can quickly adapt to different targets and improve the generalization ability of on-orbit maintenance. In the future, it can be expanded from a multi-modal perspective. Combining with a force sensor, force / position hybrid control can be developed to improve the working accuracy of the end device. Considering the limited computing resources of space equipment, the number of model parameters will be further compressed, and an embedded deployment solution will be developed. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 is a schematic flowchart of the method of the present invention;
[0038] Figure 2 is a schematic diagram of the operating space of the D-type hole bolt screwing device in the present invention;
[0039] Figure 3 is a comparison diagram of the YOLOv5 network structures before and after optimization;
[0040] Figure 4 is the overall structural diagram of the algorithm of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] To make the objectives, technical solutions and effects of the present invention clearer and more definite, the following examples are listed to further illustrate the present invention in detail. It should be noted that the specific implementations described herein are only used to explain the present invention and are not used to limit the present invention.
[0042] The specific design of the object detection algorithm improved based on YOLOv5 of the present invention is as follows:
[0043] 1. Establishment of the training environment for the object detection algorithm
[0044] When the end tool of the Chinese space station performs on-orbit operation tasks, it is necessary to perform precise information perception, recognition and operation tasks on the target parts. After actual measurement, the available space range of the end tool is limited to 80mm×40mm×300mm, and the operation area is a narrow and "L"-shaped space. The vision sensor is small in size and can ensure sufficient information feedback in a narrow space, which is suitable for realizing precise target information perception and recognition work and can provide assistance for space dexterous manipulation, such as Figure 1 shown.
[0045] In the present invention, the target detection algorithm uses Python 3.8 as the main programming language and performs deep learning training with the help of the PyTorch computing library. To ensure the independence and stability of the project environment, Conda is used to create a virtual environment, and various dependent libraries required for the project are installed in the virtual environment, including NumPy (Numerical Python), OpenCV (OpenSource Computer Vision Library), etc. The target detection task based on the YOLOv5 algorithm is implemented in this virtual environment; a CPU-based computing platform is used for model training, and the specific hardware configuration is a device equipped with the i7-1260 processor in the 12th generation Intel(R) Core(TM) i7 series.
[0046] 2. Introduction to the Target Detection Algorithm
[0047] 2.1 Image Recognition Algorithm
[0048] YOLOv5 is an open-source project in the prior art. As a highly mature image recognition algorithm model, its structure mainly includes four core components: the input end, the Backbone network (main network), the Neck network (neck network), and the Head network (detection head).
[0049] The Focus structure and the CSP structure are integrated in the Backbone network, which is responsible for processing the data at the input end and generating three key feature layers. The Neck network consists of the SPP structure, the pooling structure, and the PAN structure, aiming to strengthen the data of these feature layers and perform effective feature fusion. The Head network performs calculations, generates prediction results, and conducts loss analysis. The prediction results are presented in the form of bounding boxes, and NMS (Non-Maximum Suppression) is used to filter out the prediction boxes with lower scores to improve the accuracy; the loss function uses IoU loss to accurately evaluate the difference between the predicted bounding box and the true bounding box.
[0050] The present invention has made several improvements based on the YOLOv5 algorithm, replacing the Conv, C3, and FPP structures in the Backbone network with the MobileNetv3 structure.
[0051] MobileNetv3 uses depthwise separable convolutions to replace traditional convolutions (Conv). Compared with the traditional convolutions in YOLOv5, depthwise separable convolutions reduce the number of parameters and the dimension of the input channels, thus reducing the computational load. MobileNetv3 also uses NAS platform-aware search for the global network structure and combines the NetAdapt algorithm to find the most effective optimized model for a given hardware platform. These features make MobileNetv3 lighter and easier to deploy on spaceborne embedded devices.
[0052] At the same time, the attention mechanism CBAM algorithm is introduced to replace the original C3 module in the Head. CBAM integrates a channel attention module and a spatial attention module. The channel attention module can enhance the feature expression of channels, and the spatial attention module can enhance the recognition accuracy of the model for small targets, which can improve the classification performance of the image recognition algorithm.
[0053] It can be seen from Figure 2 that (2a is the localization loss rate of the training set, 2b is the localization loss rate of the validation set, 2c is the confidence loss rate of the training set, 2d is the confidence loss rate of the validation set, 2e is the classification loss rate of the training set, 2f is the classification loss rate of the validation set, 2g is the mean average precision with an IoU threshold of 0.5, 2h is the mean average precision with an IoU threshold from 0.5 to 0.95, 2i is the precision, and 2j is the recall rate). When the training batch reaches more than 80, the localization loss, confidence loss, and classification loss of the model on the training set and validation set are all close to 0, and the precision, recall rate, and mean average precision are all close to 1, demonstrating extremely high performance.
[0054] Table 1 Comparison of the average time used to process a single image before and after optimization
[0055]
[0056] It can be seen from Table 1 that the model trained by the improved algorithm significantly reduces the inference time, and the time to process a single image is reduced by 27.2%.
[0057] It can be seen from Figure 3 that when the number of training times reaches about 50, the classification loss rate of the training set is close to 0, and the precision, mean average precision with an IoU threshold of 0.5, and recall rate are all close to 1, which is similar to the model before the improvement of the YOLOv5 algorithm, demonstrating extremely high performance.
[0058] The experimental results show that the introduction of MobileNetv3 and the CBAM algorithm not only maintains a high level of accuracy but also reduces the time spent processing an image by 27.2%, improving the overall efficiency of the model.
[0059] 2.2 Ranging algorithm:
[0060] In the image detection algorithm, the distance between the image and the camera center can be calculated by formula (1).
[0061]
[0062] In formula (1): D is the distance from the target to the camera, F is the pixel focal length of the camera, W is the width or height of the target, and P is the number of pixels occupied by the target in the image.
[0063] The pixel focal length of the camera can be calculated based on the size and resolution of the camera's photosensitive element.
[0064] The formula for converting the known focal length (pixels) to millimeters is:
[0065]
[0066] In formulas (2) and (3): dx and dy are the values of the pixel focal length in two directions respectively, S is the size of the camera's CCD, and x and y are the number of pixels of the image in two directions respectively.
[0067] When the camera lens is parallel to the plane where the D-shaped hole is located, the distance of the target object can be calculated through the relevant parameters of the camera and the size of the D-shaped hole bolt.
[0068]
[0069] In formula (4): the actual distance between the two centers is d r1 , the number of pixels difference between the two centers in the picture is d p1 , the diameter of the D-shaped hole is d r , the number of pixels occupied by the height of the D-shaped hole is d p .
[0070] Without considering lens distortion, the horizontal and vertical distances between the center of the D-shaped hole bolt and the camera center are approximately calculated through formula (4).
[0071] 3. Implementation process of the D-shaped hole detection function
[0072] 3.1 Image collection and preprocessing
[0073] In order to obtain the original image data for model training and verification, the method of directly holding a mobile phone to take pictures was adopted, simulating the running trajectory of the camera on the robotic arm. A large number of images were collected from different angles and orientations through this method.
[0074] During the shooting process, interference objects similar to the appearance of the D-shaped hole bolt were additionally introduced (as shown in Figure 4 ), and these were used as negative samples to enhance the model's ability to identify target features.
[0075] After completing the image sampling, the original image dataset is divided into a training set, a test set, and a validation set according to a ratio of 6:2:2. Subsequently, the labelimg tool is used to annotate the pictures in the training set and the test set, where the D-type hole bolts are uniformly and simply marked as "d", and J599 is marked as "j".
[0076] 3.2 Training Results of the D-Type Hole Detection Model Improved Based on YOLOv5
[0077] After completing the collection and preprocessing process of the image data, the model is trained using computer resources, and a simplified model for D-type hole detection is finally constructed.
[0078] To deeply evaluate the accuracy of the model, we introduce a two-dimensional analysis tool, the confusion matrix (as Figure 4 shown). From the confusion matrix, we can clearly observe that the model has a very high accuracy when predicting D-type holes. When the actual object is a D-type hole, the probability that the model predicts it as a D-type hole is as high as 0.99; similarly, when predicting J599, the model also shows very high accuracy, and the probability that the actual J599 is correctly predicted by the model reaches 1.
[0079] By checking the overfitting situation of the model through the F1-score-confidence curve, the model does not show obvious overfitting phenomena, which further verifies the stability and reliability of the model.
[0080] When only the edges or corners of the D-type hole and J599 are exposed in the field of view, according to the test results, the D-type hole bolt detection model can still perform excellently in identifying the target. This result indicates that the model has high accuracy when dealing with detection tasks under partial occlusion or limited viewing angle conditions.
[0081] To further test the recognition performance of the model, problems that may exist in real situations, such as wear of D-type hole bolts, partial occlusion, and interference of the photographing device by cosmic radiation, are simulated. Some pictures are processed by inserting noise and random occlusion to simulate the damaged situation of D-type hole bolts during on-orbit service. The experimental results prove that the model can still correctly identify the D-type hole bolts.
[0082] After obtaining the image recognition results, the pictures with drawn rectangular frames are used for distance measurement. Since the work only requires distance measurement of the D-type hole, the distance of the J599 parts is not detected to simplify the code.
[0083] The target detection model with both D-type hole recognition and distance measurement functions will provide data support for the next path planning algorithm.
[0084] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, several improvements can be made without departing from the principle of the present invention, and these improvements should also be regarded as within the protection scope of the present invention.
Claims
1. A method for detecting the bolts of D-shaped holes in a space station improved based on YOLOv5, characterized in that, The method described is as follows: Step 1: Collect the original images of the D-shaped hole bolts on the space station; Step 2: Perform annotation processing on the images to obtain the D-shaped hole bolt image dataset; Step 3: Optimize and improve the YOLOv5 model to obtain an optimized model for D-shaped hole bolt image detection; Step 4: Use the optimized model obtained in Step 3 to train the dataset in Step 2 to obtain a trained detection model; Step 5: Use the trained detection model to detect the D-shaped hole bolt images.
2. The method for detecting the D-type hole bolts of the space station improved based on YOLOv5 according to claim 1, characterized in that, The specific content of Step 1 is as follows: Adopt the method of shooting with a micro camera to simulate the running trajectory of the camera on the robotic arm during actual operation; Through this method, a large number of images are collected from different angles and orientations; During the shooting process, interference objects similar to the appearance of the D-shaped hole bolts are additionally introduced as negative samples.
3. The method for detecting the bolts of D-shaped holes in the space station improved based on YOLOv5 according to claim 1, wherein, The specific content of Step 2 is as follows: After completing the image sampling, according to the ratio of 6:2:2, the original image dataset is divided into a training set, a test set, and a validation set; Subsequently, use the labelimg tool to label the pictures in the training set and the test set, where the D-shaped hole bolts are uniformly simply labeled as d, and J599 is labeled as j.
4. A method for bolt target detection of D-shaped holes in a space station improved based on YOLOv5 according to claim 1, characterized in that, The specific content of Step 3 is as follows: 3.1 Improvement of the image recognition algorithm based on YOLOv5: Replace the Conv, C3, and FPP structures in the Backbone network of the YOLOv5 algorithm with the MobileNetv3 structure; Compared with the YOLOv5 algorithm, MobileNetv3 uses depthwise separable convolution to replace the traditional convolution (Conv), reducing the computational amount; Introduce the attention mechanism CBAM algorithm and replace the original C3 module in the Head with it; CBAM integrates the channel attention module and the spatial attention module. The channel attention module can enhance the feature expression of the channels, and the spatial attention module enhances the recognition accuracy of the model for small targets, which can improve the classification performance of the image recognition algorithm; 3.2 Introduction of the ranging algorithm: In the image detection algorithm, the distance between the image and the camera center can be calculated by formula (1); In formula (1): D is the distance from the target to the camera, F is the pixel focal length of the camera, W is the width or height of the target, and P is the number of pixels occupied by the target in the image; The pixel focal length of the camera is calculated according to the size and resolution of the camera's photosensitive element; The formula for converting the known focal length (pixels) to millimeters is: In formulas (2)(3): dx and dy are the values of the pixel focal length in two directions respectively, S is the size of the camera CCD, and x and y are the number of pixels of the image in two directions respectively; When the camera lens is parallel to the plane where the D-shaped hole is located, the distance of the target object can be calculated through the relevant parameters of the camera and the size of the D-shaped hole bolt; In formula (4): The actual distance between the two centers is d r1 , the number of pixels between the two centers in the picture is d p1 , the diameter of the D-shaped hole is d r , the number of pixels occupied by the height of the D-shaped hole is d p ; Without considering lens distortion, the horizontal and vertical distances between the center of the D-shaped hole bolt and the camera center are approximately calculated by formula (4).
5. A method for detecting bolts with D-shaped holes in a space station improved based on YOLOv5 according to claim 1, characterized in that, The specific content of Step 4 is as follows: After completing the collection and preprocessing process of the image data, use computer resources to train the model, and finally construct a simplified model for D-shaped hole detection; To deeply evaluate the accuracy of the model, a two-dimensional analysis tool, the confusion matrix, is introduced; from the confusion matrix, we can clearly observe that the model has a very high accuracy in predicting D-shaped holes. When the actual hole is a D-shaped hole, the probability that the model predicts it as a D-shaped hole is as high as 0.99; similarly, when predicting J599, the model also shows extremely high accuracy, and the probability that the actual J599 is correctly predicted by the model reaches 1. The overfitting situation of the model is checked through the F1 score-confidence curve.
6. The method for detecting the bolts of the D-shaped holes in the space station improved based on YOLOv5 according to claim 1, characterized in that, The specific content of step five is as follows: Use the trained model to detect the pictures that have not been used for training the model. When only the edges or corners of the D-shaped holes and J599 are exposed in the field of view, or when partial occlusion and noise are added, according to the test results, the D-shaped hole bolt detection model can still excellently identify the target.
7. A method for detecting D-shaped hole bolts on a space station based on the improvement of YOLOv5 according to claim 1, characterized in that To further test the recognition performance of the model, simulate the problems of wear, local occlusion, and interference of the photographing device by cosmic radiation that may exist in the real situation for some pictures, and perform noise insertion and random occlusion processing to simulate the damaged situation of D-shaped hole bolts during on-orbit service; after obtaining the image recognition results, use the pictures with drawn rectangular frames for ranging; the target detection model with both D-shaped hole recognition and ranging functions will provide data support for the next path planning algorithm.
Citation Information
Cited By
Intelligent identifying and counting method, system and equipment for industrial hoister and medium
CN121214161A