Failure spacecraft component detection network and method based on YOLOv8 network

By introducing the HCE module into the backbone network of the YOLOv8 network and adding R-GAM to the neck network, the illuminance and motion state problems in the detection method of failed spacecraft components in optical images are solved, and the detection accuracy and effectiveness are significantly improved.

CN120070965AActive Publication Date: 2025-05-30XIAN INST OF OPTICS & PRECISION MECHANICS CHINESE ACAD OF SCI
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510120896.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-30
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

In the prior art, the detection method for detecting failed spacecraft components in optical images leads to missed or missed detection due to illumination problems and motion state problems, and the detection accuracy and effectiveness are insufficient.

Method used

The failed spacecraft component detection network based on the YOLOv8 network is adopted. By introducing HCE modules into the backbone network, combining CBAM and ECA mechanisms, the ability to extract complex image noise characteristics is enhanced; R-GAM is added to the neck network, and the residual network and attention mechanism are used to optimize the feature map to improve the detection accuracy of motion blur images.

Benefits of technology

It significantly improves the recognition of target information in the image, overcomes the negative impact of illumination problems on detection, improves the detection accuracy of motion blur images, enhances the anti-interference ability of the network, and improves the detection accuracy and effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070965A_ABST
    Figure CN120070965A_ABST
Patent Text Reader

Abstract

The invention relates to a target detection network and method, in particular to a failure spacecraft component detection network and method based on a YOLOv8 network, and aims to overcome the defect of missing detection or false detection of key components of a failure spacecraft due to the illumination problem and the motion state problem in a failure spacecraft component detection method in an optical image at the present stage. According to the failure spacecraft component detection network based on the YOLOv8 network, HCE is introduced into a YOLOv8 backbone network to replace C2f, the extraction capability of complex image noise features is effectively enhanced by combining CBAM and ECA mechanisms, target features are further accurately extracted, the identification degree of target information in an image is improved, R-GAM is added in a YOLOv8 neck network, and the detection efficiency of the failure spacecraft component detection network based on the YOLOv8 network is improved. Information can be ensured to be efficiently transmitted among different layers through leap connection of a residual network, a feature map is optimized and updated in combination with an attention mechanism, and more attention is focused on key component features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a target detection network and method, and particularly to a detection network and method for failed spacecraft components based on the YOLOv8 (You Only Look Once version 8) network. Background Art

[0002] Artificial satellites in orbit may fail due to natural impacts, accidents, or fuel exhaustion. Such failed satellites, as non-cooperative targets, not only waste orbital resources but may also disintegrate, posing a threat to space security. Therefore, capturing or maintaining failed satellites has become a crucial task. The core of the capture task lies in identifying key components of the satellite, such as solar panels and radar antennas; while the focus of on-orbit maintenance is to identify the overall structure and key interfaces of the satellite, such as the satellite body and docking surfaces.

[0003] Optical imaging technology has shown excellent applicability in the task of capturing failed spacecraft components due to its intuitiveness, high resolution, and rich information in target detection. The detection of spacecraft components based on optical images belongs to the target detection direction in the field of computer vision. Currently, such detection methods are mainly divided into two categories: traditional target detection methods and detection methods based on convolutional neural networks (CNNs).

[0004] Traditional target detection methods are usually achieved through feature fitting (such as points, lines, circles), but the fitting parameters need to be adjusted under different lighting conditions and target types, with poor adaptability. In addition, such methods rely on complex image preprocessing processes. With the development of deep learning technology, detection methods based on CNNs have shown remarkable results in local component detection, such as high-voltage insulator detection and material surface defect detection.

[0005] Currently, the detection and identification of key spacecraft components face two major challenges:

[0006] (1) Illumination problem: Images taken in the deep space background often have problems such as insufficient brightness, low contrast, and a large number of noise points, seriously affecting the image quality and visual effect, resulting in the target information being hidden in the noise, thus reducing the performance of the detector.

[0007] (2) Motion state problem: When photographing a moving spacecraft, jitter and blurring are likely to occur. These problems hinder the accurate identification of key components and significantly restrict the detection task.

[0008] The above problems often lead to missed detection or false detection of key components of failed spacecraft. Therefore, it is of great significance to develop an efficient and robust detection method for spacecraft components to cope with the complex space environment and improve the detection accuracy. Summary of the Invention

[0009] The object of the present invention is to solve the deficiencies of the current detection methods for failed spacecraft components in optical images, which may lead to missed or false detections of critical components of failed spacecraft due to illumination problems and motion state problems, and to provide a detection network and method for failed spacecraft components based on the YOLOv8 network.

[0010] To solve the above-mentioned deficiencies of the prior art, the present invention provides the following technical solutions:

[0011] A detection network for failed spacecraft components based on the YOLOv8 network, characterized in that it includes a backbone network, a neck network, and a head network;

[0012] The backbone network is used to extract features from the input image, and it includes Conv (convolution), Conv, HCE (Hybrid CBAM and ECA, hybrid attention module), Conv, HCE, Conv, HCE, Conv, HCE, SPPF (Spatial Pyramid Pooling-Fast, spatial pyramid pooling module) connected in sequence from the output to the input;

[0013] The neck network is used to further process the feature map output by the backbone network, and it includes Unsample (upsampling), Concat (concatenation), C2f (CSP Bottleneck with 2 Convolutions), Unsample, Concat, C2f, R-GAM (Residual Global Attention Mechanism, residual global attention mechanism), Conv, Concat, C2f, R-GAM, Conv, Concat, C2f, R-GAM connected in sequence from the output to the input; the head network is used to complete the classification and localization of the target according to the feature map output by the neck network, and it includes three Detect.

[0014] The second output end of the HCE in the backbone network is also connected to the second input end of the Concat in the neck network, the third output end of the HCE in the backbone network is also connected to the first input end of the Concat in the neck network, and the output end of the SPPF in the backbone network is respectively connected to the first input end of the Unsample in the neck network and the fourth input end of the Concat in the neck network; the first output end of the R-GAM in the neck network is also connected to the first Detect, the second output end of the R-GAM is also connected to the second Detect, and the third output end of the R-GAM is connected to the first Detect;

[0015] Each of the HCEs includes a Conv, a Split, n BottleneckCBAMs (Bottleneck with CBAM), a Concat, and a Conv connected in sequence from output to input, where n≥2. The output end of each BottleneckCBAM is connected to the input end of the Concat through a BottleneckECA (Bottleneck with ECA); the output end of the Split is also connected to the input end of the Concat;

[0016] Each of the R-GAMs includes an MLP (Multi-Layer Perceptron), a Sigmoid, a multiplication module, a Depthwise Conv, a Pointwise Conv, a Sigmoid, a multiplication module, and a Residual Connection connected in sequence from output to input. The input end of the MLP is connected to the output end of the corresponding C2f, and the output end of the corresponding C2f is also connected to the second input end of the first multiplication module. The output end of the first multiplication module is also connected to the second input end of the second multiplication module. The MLP is used to transform the output of the corresponding C2f from the C×W×H dimension to the W×H×C dimension, where C is the number of channels, and W and H are the width and height of the feature map. The first Sigmoid is used to generate a weight matrix from the output of the MLP. The first multiplication module is used to perform element-wise multiplication on the output of the corresponding C2f and the weight matrix to generate a second feature map. The Depthwise Conv is used to perform independent convolution operations on each channel of the second feature map to generate intermediate features. The Pointwise Conv is used to transform the intermediate features from the W×H×C dimension to the W×H×C′ dimension, where C′ is the number of output channels of the Pointwise Conv. The second Sigmoid is used to generate a new weight matrix from the output of the Pointwise Conv. The second multiplication module is used to perform element-wise multiplication on the intermediate features and the new weight matrix to generate a third feature map. The Residual Connection is used to generate the output features of the R-GAM from the third feature map.

[0017] Further, the BottleneckCBAM includes a Conv, a Conv, and a CBAM (Convolutional Block Attention Module) connected in sequence from output to input. The first Conv is used to extract preliminary features from the input of the BottleneckCBAM and transform them into deeper features. The second Conv is used to extract higher-level features from the output of the first Conv. The CBAM is used to generate a spatial attention map based on the output of the second Conv;

[0018] If shortcut = True, the output of BottleneckCBAM is the sum of the input of the first Conv and the spatial attention map; if shortcut = False, the output of BottleneckCBAM is the said spatial attention map.

[0019] Furthermore, the BottleneckECA includes a Conv, a Conv, and an ECA (Efficient Channel Attention) connected in sequence to the output input; the first Conv is used to extract preliminary features from the input of BottleneckECA and convert them into features of a lower dimension; the second Conv is used to continue extracting features from the output of the first Conv to maintain the depth of the feature map; the ECA is used to calculate channel attention weights through local convolution operations and apply the channel attention weights to the output of the second Conv;

[0020] If shortcut = True, the output of BottleneckECA is the sum of the input of the first Conv and the output of the ECA; if shortcut = False, the output of BottleneckECA is the output of the ECA.

[0021] Meanwhile, the present invention also provides a method for detecting failed spacecraft components based on the YOLOv8 network, which is characterized in that it includes the following steps:

[0022] Step 1, within the backbone network of YOLOv8, each Bottleneck of each C2f is replaced by BottleneckCBAM, and a BottleneckECA is used to connect the output end of each BottleneckCBAM to the Concat input end, that is, each C2f is replaced by HCE;

[0023] Within the neck network of YOLOv8, at the output ends of the second C2f and the third C2f, they are respectively connected to the corresponding Conv input end within the neck network and the corresponding Detect input end within the head network through an R-GAM, and at the output end of the fourth C2f, it is connected to the corresponding Detect input end within the head network through an R-GAM to obtain the initial detection network for failed spacecraft components based on the YOLOv8 network;

[0024] Step 2: Use the model of the failed spacecraft to be detected to collect images of the failed spacecraft, perform data augmentation, and then divide them into a training set, a validation set, and a test set. After image annotation of the training set and the validation set, input them into the initial failed spacecraft component detection network obtained in Step 1 for training, and use the test set for evaluation to obtain a trained failed spacecraft component detection network based on the YOLOv8 network;

[0025] Step 3: Input the real image of the failed spacecraft to be detected into the trained failed spacecraft component detection network to obtain the detection results of the failed spacecraft components, and complete the detection of the failed spacecraft components.

[0026] Further, in Step 1, the working process of each HCE is as follows:

[0027] Step A1: Perform a convolution operation on the input of HCE through the first Conv of HCE to extract local features and input them into Split;

[0028] Step A2: Split the output of the first Conv through Split to obtain multiple segmented features, and input them into the first BottleneckCBAM and Concat respectively;

[0029] Step A3: The first Conv of the i-th BottleneckCBAM extracts preliminary features from the segmented features, converts them into deeper features, and then the second Conv extracts higher-level features from the output of the first Conv. Then, CBAM generates a spatial attention map according to the output of the second Conv; i ∈ [1, n - 1]; when shortcut = true, add the input of the first Conv to the spatial attention map and use it as the output of the i-th BottleneckCBAM, and input it into the i-th BottleneckECA and the (i + 1)-th BottleneckCBAM respectively; when shortcut = false, use the spatial attention map as the output of the i-th BottleneckCBAM, and input it into the i-th BottleneckECA and the (i + 1)-th BottleneckCBAM respectively;

[0030] The first Conv of the i-th Bottleneck ECA extracts preliminary features from the output of the i-th Bottleneck CBAM, converts them into features of a lower dimension, and then the second Conv further extracts features from the output of the first Conv to maintain the depth of the feature map. Then, ECA calculates the channel attention weights through local convolution operations and applies the channel attention weights to the output of the second Conv; when shortcut = true, the input of the first Conv is added to the output of ECA as the output of the i-th Bottleneck ECA and input to Concat; when shortcut = false, the output of ECA is used as the output of the i-th Bottleneck ECA and input to Concat;

[0031] The first Conv of the n-th Bottleneck CBAM extracts preliminary features from the segmented features, converts them into deeper features, and then the second Conv extracts higher-level features from the output of the first Conv. Then, CBAM generates a spatial attention map based on the output of the second Conv; when shortcut = true, the input of the first Conv is added to the spatial attention map and used as the output of the n-th Bottleneck CBAM, which is respectively input to Concat and the n-th Bottleneck ECA; when shortcut = false, the spatial attention map is used as the output of the n-th Bottleneck CBAM and input to Concat and the n-th Bottleneck ECA respectively;

[0032] Follow the above process until the output of the n-th Bottleneck ECA is input to Concat, and then execute step A4;

[0033] Step A4: Concat concatenates the segmented features output by step A2 Split, the output of the n-th Bottleneck CBAM in step A3, and the output of the n Bottleneck ECAs to obtain the concatenated features, which are input to the second Conv of HCE for convolution processing to generate the output of HCE.

[0034] Furthermore, in step 1, the working process of each of the R-GAMs is as follows:

[0035] Step B1: The outputs corresponding to C2f are respectively input to the MLP and the first multiplication module of the R-GAM;

[0036] Step B2: The MLP transforms the output corresponding to C2f from the C×W×H dimension to the W×H×C dimension, inputs the first Sigmoid, and then the first Sigmoid generates a weight matrix and inputs it into the first multiplication module;

[0037] Step B3: The first multiplication module performs an element-wise multiplication on the output corresponding to C2f and the weight matrix to generate a second feature map, which is respectively input into the depth convolution module and the second multiplication module;

[0038] Step B4: The depth convolution module performs an independent convolution operation on each channel of the second feature map to generate intermediate features and inputs them into the point convolution module;

[0039] Step B5: The point convolution module transforms the intermediate features from the W×H×C dimension to the W×H×C′ dimension, inputs the second Sigmoid, and then generates a new weight matrix and inputs it into the second multiplication module;

[0040] Step B6: The second multiplication module performs an element-wise multiplication on the second feature map and the new weight matrix to generate a third feature map and inputs it into the residual connection module, and then the residual connection module generates the output of the R-GAM.

[0041] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0042] (1) The present invention provides a detection network for failed spacecraft components based on the YOLOv8 network. HCE is introduced in the YOLOv8 backbone network to replace C2f, that is, each Bottleneck of each C2f is replaced by BottleneckCBAM, and a BottleneckECA is used to connect the output end of each BottleneckCBAM to the Concat input end; by combining the CBAM and ECA mechanisms, the ability to extract complex image noise features is effectively enhanced, thereby accurately extracting target features, improving the recognition rate of target information in images, and overcoming the negative impact of image quality defects caused by illumination problems on detection, ensuring the accuracy and effectiveness of detection.

[0043] (2) The present invention adds R-GAM to the YOLOv8 neck network. The skip connections of the residual network ensure that information can be efficiently transmitted between different layers. By combining the attention mechanism to optimize and update the feature map, irrelevant information generated by motion blur can be effectively ignored, and more attention can be focused on the key component features. At the same time, R-GAM strengthens the key features of the target from both the channel and spatial dimensions, enhances the anti-interference ability of the network, significantly improves the detection accuracy of the network for targets in motion-blurred images, and reduces the detection error caused by the unstable motion state of the satellite.

[0044] (3) The method for detecting failed spacecraft components based on the YOLOv8 network of the present invention can effectively improve the mAP (mean Average Precision) of failed spacecraft components (an increase of 2.57%) compared to the original YOLOv8 model. Especially when dealing with challenges such as image noise and jitter blur, the present invention has good detection effects. Description of the Drawings

[0045] Figure 1 It is a schematic structural diagram of a YOLOv8 network;

[0046] Figure 2 is Figure 1 the schematic structural diagram of C2f in

[0047] Figure 3 It is the schematic structural diagram of the initial failed spacecraft component detection network in step 1 of the embodiment of the method for detecting failed spacecraft components based on the YOLOv8 network of the present invention;

[0048] Figure 4 is Figure 3 the schematic structural diagram of HCE in

[0049] Figure 5 is Figure 4 the schematic structural diagram of BottleneckCBAM in

[0050] Figure 6 is Figure 4 the schematic structural diagram of BottleneckECA in

[0051] Figure 7 is Figure 3 the schematic structural diagram of R-GAM in Detailed Embodiment

[0052] The present invention will be further described below with reference to the drawings and exemplary embodiments.

[0053] Referring to Figures 1 - 2 , the YOL0v8 network includes a backbone network, a neck network, and a head network; the backbone network is responsible for extracting features from the input image, and it includes Conv, Conv, C2f, Conv, C2f, Conv, C2f, Conv, C2f, SPPF connected in sequence from the output to the input; as Figure 2As shown, in the backbone network, C2f includes Conv, Split, n Bottlenecks (n≥2), Concat, and Conv connected in sequence from output to input. The output ends of Split and n Bottlenecks are also connected to the input end of Concat. C2F is a module based on the CSP (Cross Stage Partial) architecture, which gradually refines and fuses features by repeating n Bottleneck modules; the neck network is used to aggregate features from the backbone network, and it includes Unsample, Concat, C2f, Unsample, Concat, C2f, Conv, Concat, C2f, Conv, Concat, C2f connected in sequence from output to input; the head network receives the refined features from the neck network and outputs detection results, including predictions of categories and bounding boxes.

[0054] The present invention discloses a method for detecting failed spacecraft components based on the YOLOv8 network, which includes the following steps:

[0055] Step 1, in Figure 1 the backbone network of the shown YOLOv8, replace each C2f with an HCE. Specifically, replace each Bottleneck in each C2f with BottleneckCBAM, and use a BottleneckECA to connect the output end of each BottleneckCBAM to the input end of Concat within C2f;

[0056] In the neck network of YOLOv8, connect the output ends of the second C2f and the third C2f to the corresponding Conv input end within the neck network and the corresponding Detect input end within the head network respectively through an R-GAM, and connect the output end of the fourth C2f to the corresponding Detect input end within the head network through an R-GAM to obtain an initial detection network for failed spacecraft components based on the YOLOv8 network, as Figure 3 shown;

[0057] The detection network for failed spacecraft components includes a backbone network, a neck network, and a head network; the backbone network is used to extract features from the input image, and it includes Conv, Conv, HCE, Conv, HCE, Conv, HCE, Conv, HCE, and SPPF connected in sequence from output to input; the neck network is used to further process the feature maps output by the backbone network, and it includes Unsample, Concat, C2f, Unsample, Concat, C2f, R-GAM, Conv, Concat, C2f, R-GAM, Conv, Concat, C2f, R-GAM connected in sequence from output to input; the head network is used to complete the classification and localization of the target according to the feature maps output by the neck network; the head network includes three Detect;

[0058] Refer to Figure 4 , each HCE includes Conv, Split, n BottleneckCBAMs, Concat, and Conv connected in sequence from output to input, where n≥2; the output end of each BottleneckCBAM is connected to the input end of Concat through a BottleneckECA; the output end of Split is also connected to the input end of Concat;

[0059] Refer to Figure 5 , BottleneckCBAM includes Conv, Conv, and CBAM connected in sequence from output to input. The first Conv is used to extract preliminary features and transform them into deeper features. The second Conv is used to extract higher-level features from the output of the first Conv. CBAM is used to generate a spatial attention map according to the output of the second Conv. If shortcut = True (residual connection), the output of BottleneckCBAM is the sum of the input of the first Conv and the spatial attention map; if shortcut = False (no residual connection), the output of BottleneckCBAM is the spatial attention map; shortcut (shortcut connection) refers to a connection method that directly adds the input feature map to the output of the subsequent layer;

[0060] Refer to Figure 6, BottleneckECA includes Conv, Conv, and ECA connected in sequence of output and input. The first Conv is used to extract the preliminary features corresponding to the output of BottleneckCBAM and convert them into features of a lower dimension; the second Conv is used to continue extracting features from the output of the first Conv to maintain the depth of the feature map. ECA is used to calculate the channel attention weights through local convolution operations and apply the channel attention weights to the output of the second Conv. If shortcut = True, the output of BottleneckECA is the sum of the input of the first Conv and the output of ECA; if shortcut = False, the output of BottleneckECA is the output of ECA;

[0061] The working process of each HCE is as follows:

[0062] Step A1: Perform a convolution operation on the input of HCE through the first Conv of HCE to extract local features and input them into Split;

[0063] Step A2: Split the output of the first Conv through Split to obtain multiple segmented features, which are respectively input into the first BottleneckCBAM and Concat;

[0064] Step A3: The first Conv of the i-th BottleneckCBAM extracts preliminary features from the segmented features and converts them into deeper features. Then, the second Conv extracts higher-level features from the output of the first Conv. Then, CBAM generates a spatial attention map according to the output of the Conv; i ∈ [1, n - 1]. When shortcut = true, the sum of the input of the first Conv and the spatial attention map is used as the output of the i-th BottleneckCBAM, which is respectively input into the i-th BottleneckECA and the (i + 1)-th BottleneckCBAM. When shortcut = false, the spatial attention map is used as the output of the i-th BottleneckCBAM, which is respectively input into the i-th BottleneckECA and the (i + 1)-th BottleneckCBAM;

[0065] The first Conv of the i-th Bottleneck ECA extracts preliminary features from the output of the i-th Bottleneck CBAM, converts them into features of a lower dimension, and then the second Conv continues to extract features from the output of the first Conv to maintain the depth of the feature map. Then, ECA calculates the channel attention weights through local convolution operations and applies the channel attention weights to the output of the second Conv; when shortcut = true, the input of the first Conv is added to the output of ECA as the output of the i-th Bottleneck ECA, and the input is Concat; when shortcut = false, the output of ECA is used as the output of the i-th Bottleneck ECA, and the input is Concat;

[0066] The first Conv of the n-th Bottleneck CBAM extracts preliminary features from the segmented features, converts them into deeper features, and then the second Conv extracts higher-level features from the output of the first Conv. Then, CBAM generates a spatial attention map based on the output of the second Conv; when shortcut = true, the input of the first Conv is added to the spatial attention map and then used as the output of the n-th Bottleneck CBAM, and the input is respectively Concat and the n-th Bottleneck ECA; when shortcut = false, the spatial attention map is used as the output of the n-th Bottleneck CBAM, and the input is respectively Concat and the n-th Bottleneck ECA;

[0067] Follow the above process until the output of the n-th Bottleneck ECA is input to Concat, and then execute step A4;

[0068] Step A4: Concat splices the segmented features output by step A2 Split, the output of the n-th Bottleneck CBAM in step A3, and the output of n Bottleneck ECAs to obtain the spliced features, and inputs them to the second Conv of HCE for convolution processing to generate the output of HCE;

[0069] Refer to Figure 7 , each R-GAM includes an MLP, a Sigmoid, a multiplication module, a depth convolution module, a point convolution module, a Sigmoid, a multiplication module, and a residual connection module connected in sequence at the output and input; the input end of the MLP is connected to the output end of the corresponding C2f, and the output end of the corresponding C2f is also connected to the second input end of the first multiplication module, and the output end of the first multiplication module is also connected to the second input end of the second multiplication module;

[0070] The MLP is used to transform the output corresponding to C2f from the C×W×H dimension to the W×H×C dimension (only changing the order of dimension arrangement), where C is the number of channels, and W and H are the width and height of the feature map; the first Sigmoid is used to generate a weight matrix from the output of the MLP; the first multiplication module is used to perform element-wise multiplication on the output corresponding to C2f and the weight matrix to generate a second feature map; the depth convolution module is used to perform independent convolution operations on each channel of the second feature map to generate intermediate features; the point convolution module is used to transform the intermediate features from the W×H×C dimension to the W×H×C′ dimension, where C′ is the number of output channels of the point convolution module; the first Sigmoid is used to generate a new weight matrix from the output of the point convolution module; the second multiplication module is used to perform element-wise multiplication on the intermediate features and the new weight matrix to generate a third feature map; the residual connection module is used to generate the output features of R-GAM from the third feature map;

[0071] The working process of each R-GAM is as follows:

[0072] Step B1: Input the output corresponding to C2f into the MLP and the first multiplication module of R-GAM respectively;

[0073] Step B2: The MLP transforms the output corresponding to C2f from the C×W×H dimension to the W×H×C dimension, inputs it into the first Sigmoid, and then the first Sigmoid generates a weight matrix and inputs it into the first multiplication module;

[0074] Step B3: The first multiplication module performs element-wise multiplication on the output corresponding to C2f and the weight matrix to generate a second feature map, and inputs it into the depth convolution module and the second multiplication module respectively;

[0075] Step B4: The depth convolution module performs independent convolution operations on each channel of the second feature map to generate intermediate features and inputs them into the point convolution module;

[0076] Step B5: The point convolution module transforms the intermediate features from the W×H×C dimension to the W×H×C′ dimension, inputs it into the second Sigmoid, and then generates a new weight matrix and inputs it into the second multiplication module;

[0077] Step B6: The second multiplication module performs element-wise multiplication on the second feature map and the new weight matrix to generate a third feature map and inputs it into the residual connection module, and then the residual connection module generates the output of R-GAM;

[0078] Step 2: Use the model of the failed spacecraft to be detected to collect images of the failed spacecraft, perform data augmentation, and then divide them into a training set, a validation set, and a test set. After image annotation of the training set and the validation set, input them into the initial failed spacecraft component detection network obtained in Step 1 for training, and use the test set for evaluation to obtain a trained failed spacecraft component detection network based on the YOLOv8 network;

[0079] The specific process of Step 2 can be referred to in Chinese Patent CN119048869A;

[0080] Step 3: Input the real image of the failed spacecraft to be detected into the trained failed spacecraft component detection network to obtain the detection results of the failed spacecraft components, and complete the detection of the failed spacecraft components.

[0081] To evaluate the effects of the embodiments of the present invention, ablation experiments are respectively carried out on the backbone network and the neck network of the YOLOv8 network.

[0082] Experiments are respectively carried out on the "HCE", "2×HCE", "3×HCE", "4×HCE", "5×HCE", "2×HCE+R-GAM", and "4×HCE+R-GAM of the embodiment of the present invention" modules. To better show the influence of the embodiments of the present invention on the detection accuracy of the failed spacecraft components, the AP (Average Precision), mAP, and Recall of each failed spacecraft component are listed for evaluation, and the results are shown in Table 1:

[0083] Table 1

[0084]

[0085] In Table 1, when the added module is HCE, it means that the first C2f of the YOLOv8 backbone network is replaced by HCE; when the added module is 2×HCE, it means that the first to second C2f of the YOLOv8 backbone network are replaced by HCE; when the added module is 3×HCE, it means that the first to third C2f of the YOLOv8 backbone network are replaced by HCE; when the added module is 4×HCE, it means that the first to fourth C2f of the YOLOv8 backbone network are replaced by HCE; when the added module is 5×HCE, it means that the first to fourth C2f of the YOLOv8 backbone network are replaced by HCE, and an additional HCE is added between the fourth HCE and the SPPF; when the added module is 2×HCE+R-GAM, it means that the first to second C2f of the YOLOv8 backbone network are replaced by HCE, and the same R-GAM as in the embodiment of the present invention is added to the YOLOv8 neck network;

[0086] As can be seen from Table 1, compared with the original YOLOv8 backbone network, adding modules of HCE, 2×HCE, 3×HCE, 4×HCE, and 5×HCE increased the mAP by 0.48%, 0.83%, 1.15%, 1.30%, and decreased it by 0.43% respectively, indicating that adding too many HCE modules to the YOLOv8 backbone network may lead to performance degradation; when the added module is 2×HCE+R-GAM, the mAP increased by 0.56% compared with the added module of 2×HCE, and the mAP increased by 1.39% compared with the original YOLOv8 backbone network; while the AP of the solar panel in the embodiment of the present invention increased the most, by 2.72%. Generally speaking, by adding HCE to the YOLOv8 backbone network and R-GAM to the neck network for improvement, not only the overall detection accuracy of the detection model is improved, but also the detection accuracy in the face of target image noise and jitter blur is improved. The improvement strategy proposed by the present invention is effective.

Claims

1. A failed spacecraft component detection network based on the YOLOv8 network, characterized by: Includes backbone network, neck network and head network; The backbone network is used to extract features from an input image, and includes Conv, Conv, HCE, Conv, HCE, Conv, HCE, Conv, HCE, SPPF, where output and input are sequentially connected; The neck network is used to further process the feature map output by the backbone network, and includes Unsample, Concat, C2f, Unsample, Concat, C2f, R-GAM, Conv, Concat, C2f, R-GAM, Conv, Concat, C2f, R-GAM, which are sequentially connected in output and input; the head network is used to complete the classification and positioning of the target according to the feature map output by the neck network, and includes three Detect; The second HCE output of the backbone network is also connected to the second Concat input of the neck network, the third HCE output of the backbone network is also connected to the first Concat input of the neck network, and the SPPF output of the backbone network is respectively connected to the first Unsample input of the neck network and the fourth Concat input of the neck network; the first R-GAM output of the neck network is also connected to the first Detect, the second R-GAM output is also connected to the second Detect, and the third R-GAM output is connected to the first Detect; Each of the HCEs includes a Conv, a Split, n BottleneckCBAMs, a Concat, and a Conv, whose output and input are connected in sequence, where n≥2, and each of the BottleneckCBAM outputs is connected to the Concat input via a BottleneckECA; the Split output is also connected to the Concat input; Each of the R-GAMs includes an MLP, a Sigmoid, a multiplication module, a depth convolution module, a point convolution module, a Sigmoid, a multiplication module, and a residual connection module, the input of which is connected in sequence; the input of the MLP is connected to the corresponding C2f output, the corresponding C2f output is also connected to the second input of the first multiplication module, and the output of the first multiplication module is also connected to the second input of the second multiplication module; the MLP is used to transform the output of the corresponding C2f from the C×W×H dimension to the W×H×C dimension, where C is the number of channels, and W and H are the width and height of the feature map; the first Sigmoid is used to generate a weight matrix from the output of the MLP ; The first multiplication module is used to perform element-level multiplication on the output of the corresponding C2f and the weight matrix to generate a second feature map; the deep convolution module is used to perform an independent convolution operation on each channel of the second feature map to generate intermediate features; the point convolution module is used to transform the intermediate features from W×H×C dimensions to W×H×C′ dimensions, where C′ is the number of output channels of the point convolution module; the second Sigmoid is used to generate a new weight matrix from the output of the point convolution module; the second multiplication module is used to perform element-level multiplication on the intermediate features and the new weight matrix to generate a third feature map; the residual connection module is used to generate the output features of R-GAM from the third feature map.

2. According to claim 1, a failed spacecraft component detection network based on a YOLOv8 network is characterized in that: The BottleneckCBAM includes Conv, Conv and CBAM whose output and input are connected in sequence; the first Conv is used to extract preliminary features from the input of the BottleneckCBAM and convert it into deeper features; the second Conv is used to extract higher-level features from the output of the first Conv; the CBAM is used to generate a spatial attention map according to the output of the second Conv; If shortcut=True, the output of BottleneckCBAM is the sum of the input of the first Conv and the spatial attention map; if shortcut=False, the output of BottleneckCBAM is the spatial attention map.

3. A failed spacecraft component detection network based on a YOLOv8 network according to claim 1 or 2, characterized in that: The BottleneckECA includes a Conv, a Conv and an ECA whose output and input are connected in sequence; the first Conv is used to extract preliminary features from the input of the BottleneckECA and convert it into features of lower dimensions; the second Conv is used to continue extracting features from the output of the first Conv and maintain the depth of the feature map; the ECA is used to calculate the channel attention weight through a local convolution operation and apply the channel attention weight to the output of the second Conv; If shortcut = True, the output of BottleneckECA is the sum of the input of the first Conv and the output of ECA; if shortcut = False, the output of BottleneckECA is the output of ECA.

4. A method for detecting failed spacecraft components based on a YOLOv8 network, characterized in that: The steps include: Step 1. In the backbone network of YOLOv8, replace each Bottleneck of each C2f with BottleneckCBAM, and use a BottleneckECA to connect each BottleneckCBAM output end with the Concat input end, that is, replace each C2f with HCE; In the neck network of YOLOv8, the output ends of the second C2f and the third C2f are respectively connected to the corresponding Conv input end in the neck network and the corresponding Detect input end in the head network through an R-GAM, and the output end of the fourth C2f is connected to the corresponding Detect input end in the head network through an R-GAM, so as to obtain the initial failed spacecraft component detection network based on the YOLOv8 network as described in claim 1; Step 2: Use the model of the failed spacecraft to be detected to collect images of the failed spacecraft, perform data augmentation, and then divide them into a training set, a validation set, and a test set; after annotating the images of the training set and the validation set, input them into the initial failed spacecraft component detection network obtained in step 1 for training, and use the test set for evaluation to obtain the trained failed spacecraft component detection network based on the YOLOv8 network as described in claim 1; Step 3: Input the real image of the failed spacecraft to be detected into the trained failed spacecraft component detection network to obtain the detection result of the failed spacecraft component and complete the failed spacecraft component detection.

5. The method for detecting failed spacecraft components based on a YOLOv8 network according to claim 4, characterized in that: In step 1, the working process of each HCE is as follows: Step A1, perform convolution operation on the input of HCE through the first Conv of HCE, extract local features and input them into Split; Step A2: Segment the output of the first Conv by Split to obtain multiple segmented features, which are input into the first BottleneckCBAM and Concat respectively; Step A3: The first Conv of the i-th Bottleneck CBAM extracts preliminary features from the segmented features and converts them into deeper features. The second Conv then extracts higher-level features from the output of the first Conv. Then, the CBAM generates a spatial attention map based on the output of the second Conv. i∈[1, n-1]; when shortcut=true, the input of the first Conv is added to the spatial attention map as the output of the i-th BottleneckCBAM, and input into the i-th BottleneckECA and the i+1-th BottleneckCBAM respectively; when shortcut=false, the spatial attention map is used as the output of the i-th BottleneckCBAM, and input into the i-th BottleneckECA and the i+1-th BottleneckCBAM respectively; The first Conv of the i-th BottleneckECA extracts preliminary features from the output of the i-th BottleneckCBAM and converts them into features of lower dimensions. The output of the first Conv is then extracted by the Conv to continue extracting features, maintaining the depth of the feature map. The ECA then calculates the channel attention weights through local convolution operations and applies the channel attention weights to the output of the Conv. When shortcut = true, the input of the first Conv is added to the output of the ECA as the output of the i-th BottleneckECA and input to Concat. When shortcut = false, the output of the ECA is used as the output of the i-th BottleneckECA and input to Concat. The first Conv of the nth BottleneckCBAM extracts preliminary features from the segmented features and converts them into deeper features. The second Conv extracts higher-level features from the output of the first Conv, and then CBAM generates a spatial attention map based on the output of the second Conv. When shortcut = true, the input of the first Conv is added to the spatial attention map as the output of the nth BottleneckCBAM, and input into Concat and the nth BottleneckECA respectively. When shortcut = false, the spatial attention map is used as the output of the nth BottleneckCBAM, and input into Concat and the nth BottleneckECA respectively. Follow the above process until the output of the nth BottleneckECA is input into Concat, and then execute step A4; Step A4: Concat the segmented features output by the Split in step A2, the output of the nth BottleneckCBAM in step A3, and the output of n BottleneckECAs to obtain the concatenated features, which are input into the second Conv of HCE for convolution processing to generate the output of HCE.

6. A method for detecting failed spacecraft components based on a YOLOv8 network according to claim 4 or 5, characterized in that: In step 1, the working process of each R-GAM is as follows: Step B1, input the output corresponding to C2f into the MLP and the first multiplication module of R-GAM respectively; Step B2: The output of the corresponding C2f is transformed from the C×W×H dimension to the W×H×C dimension by the MLP, and then input into the first Sigmoid, and then the weight matrix generated by the first Sigmoid is input into the first multiplication module; Step B3: The first multiplication module performs element-wise multiplication on the output of the corresponding C2f and the weight matrix to generate a second feature map, which is input into the deep convolution module and the second multiplication module respectively; Step B4: The deep convolution module performs an independent convolution operation on each channel of the second feature map to generate an intermediate feature input point convolution module; Step B5: The point convolution module transforms the intermediate features from W×H×C dimensions to W×H×C′ dimensions, inputs the second Sigmoid, and then generates a new weight matrix and inputs it into the second multiplication module; Step B6: The second multiplication module performs element-wise multiplication on the second feature map and the new weight matrix to generate a third feature map which is input into the residual connection module, and then the residual connection module generates the output of R-GAM.

Citation Information

Patent Citations

  • Failure spacecraft part detection method based on improved YOLOv5s

    CN119048869A

  • Wafer surface defect mode detection method based on deep attention network

    CN113362320A

  • Cloth flaw detection method based on improved Yolov4 network

    CN114240885A

  • Improved yolov5-based aerial insulator orientation identification method

    CN115690542A

  • Unmanned aerial vehicle aerial photography target detection method based on PCRS-YOLO network

    CN118799766A