Earth-moon space non-cooperative target identification method and system based on improved deep learning model
By introducing self-attention mechanism, spatial attention module and multimodal data fusion technology on the YOLOv5 model, combined with self-supervised comparative learning and lightweight design, the identification and matching problems of African cooperation goals in earth-moon space in complex contexts is solved, and high-precision and real-time recognition effects are achieved.
Patent Information
- Application Number
- CN202510179130.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-18
AI Technical Summary
Under complex, sparse and low signal-to-noise observation conditions, it is difficult to effectively distinguish and match non-cooperative targets in earth-moon space, such as space debris and abandoned satellites, especially when background and target similarity are high.
The improved deep learning model YOLOv5 is adopted, combining the self-attention mechanism, spatial attention module, multimodal data fusion and cross-modal attention mechanism, and through self-supervised comparative learning and lightweight model design, an optimized spatial non-cooperational goal matching recognition model is formed.
It significantly improves the recognition and matching accuracy of space debris in complex backgrounds, enhances the robustness of the model under low signal-to-noise ratio conditions, and realizes the ability of real-time monitoring and early warning.
Smart Images

Figure CN120047794A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to space monitoring and deep learning technologies, particularly in the field of identification and matching for non-cooperative space objects. Specifically, the present invention relates to an improved deep learning method that can efficiently identify and match non-cooperative targets (such as space debris, abandoned satellites, etc.) in the Earth-Moon space region, providing support for spacecraft orbit avoidance and mission planning. Background Art
[0002] With the continuous increase in human space activities, especially the growing number of space missions in regions such as low Earth orbit and medium Earth orbit, the problem of space debris has become an increasingly serious challenge. Space debris mainly refers to abandoned satellites, rocket remnants, and debris generated by events such as collisions. These debris move at high speeds in Earth's orbit, posing a huge threat to existing spacecraft. According to data from the National Aeronautics and Space Administration (NASA) and other space agencies, the number of space debris in Earth's orbit has reached millions, and most of them are non-cooperative targets (Non-Cooperative Targets, NCTs), that is, targets that cannot be interacted with through traditional communication or control means. In the space field, non-cooperative targets usually refer to space debris that does not operate according to predetermined rules, abandoned satellites, uncalibrated small spacecraft, or even extraterrestrial objects. In the research on the identification and orbit determination technology of non-cooperative targets in the Earth-Moon space, a prominent technical difficulty lies in: how to effectively distinguish and associate-match these unlabeled and non-active-signal non-cooperative targets from numerous background stars under complex, sparse, and low signal-to-noise ratio observation conditions.
[0003] Currently, the detection and identification of space debris mainly rely on radar, optical sensors, and other observation devices. However, these methods still have some deficiencies in practical applications: 1. Detection accuracy issue: Due to the small size, complex motion trajectories, and high speeds of space debris, traditional observation means are difficult to provide sufficiently accurate positioning and identification in a dynamically changing environment.
[0004] 2. Complex background interference: Space debris usually exists in a high background noise environment. For example, when sunlight shines strongly, or affected by other satellites and celestial bodies, it is easily interfered, resulting in a relatively high false identification rate of traditional methods.
[0005] 3. Real-time issue: The rapid movement of space debris requires the monitoring system to have real-time capabilities, while the existing technologies mostly adopt periodic scanning and offline processing methods, which are difficult to meet the requirements for efficient and real-time processing.
[0006] In recent years, with the development of deep learning technology, especially the successful application of deep learning models such as convolutional neural networks (CNNs) in the field of image recognition, researchers have begun to attempt to use deep learning methods to improve the recognition accuracy and real-time performance of space debris. Deep learning can overcome the limitations of manual feature design by automatically learning and extracting features, and can handle recognition problems in complex backgrounds. However, the existing deep learning-based space debris recognition methods still face the following challenges: 1. Feature extraction problem: Existing deep learning models have limited ability to extract features of space debris and cannot fully identify debris information in complex backgrounds, especially in low signal-to-noise ratio or high dynamic range situations.
[0007] 2. Efficient matching problem: The matching task of space debris not only requires high-precision feature extraction but also efficient computational methods for comparing debris with the target database. Existing technologies often face problems of low computational efficiency and poor real-time performance.
[0008] Therefore, how to use deep learning technology to overcome these challenges, improve the recognition and matching accuracy of space debris, and achieve real-time monitoring and early warning has become an important research direction in the current space monitoring field. Summary of the Invention
[0009] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a method and system for identifying non-cooperative targets in the Earth-Moon space based on an improved deep learning model. By improving the deep learning model YOLOv5 and introducing self-supervised learning, efficient and accurate recognition and matching of space debris in complex deep space environments, especially in situations where the similarity between the background and the target is relatively high, are achieved.
[0010] According to one aspect of the specification of the present invention, a method for identifying non-cooperative targets in the Earth-Moon space based on an improved deep learning model is provided, including: Obtain a space debris image; Input the obtained space debris image into a trained space non-cooperative target matching and recognition model to output the recognized non-cooperative target; wherein, the training of the space non-cooperative target matching and recognition model includes: Construct a space debris data set; On the basis of the deep learning model YOLOv5, introduce a self-attention mechanism and a spatial attention module, and also introduce multi-modal data fusion and cross-modal attention mechanisms to form an improved space non-cooperative target matching and recognition model; Introduce self-supervised contrast learning to perform efficient feature extraction and lightweight model design on the improved space non-cooperative target matching and recognition model to form an optimized space non-cooperative target matching and recognition model; According to the designed loss function, the optimized spatial non-cooperative target matching and recognition model is trained on the constructed space debris dataset to obtain the trained spatial non-cooperative target matching and recognition model.
[0011] As a further technical solution, a space debris dataset is constructed, including: The orbital dynamics model is used to simulate the target trajectories in the Earth-Moon space, considering the multi-source gravitational field and the effects of the non-spherical symmetric gravity field, and high-precision position information and velocity distribution for the orbital evolution of the target under different initial conditions are provided through numerical integration. Highly realistic target optical imaging is generated in the simulation, including simulating the three-dimensional shape and surface reflection characteristics of the target, and batch data generation for different observation periods, observation angles, and sensor parameters is performed through a parallel rendering pipeline to form the original simulation dataset. Using the known non-cooperative target trajectories and imaging geometric relationships, the target positions in the original simulation dataset are accurately labeled to form the space debris dataset.
[0012] As a further technical solution, the self-attention mechanism and the spatial attention module are introduced, and it also includes: The correlations of each position in the feature map are calculated through the self-attention mechanism to dynamically adjust the feature weights and enhance the representation ability of the key regions; the extraction of target features is enhanced through the spatial attention module to strengthen the model's attention to the target region, making the recognition of small targets more accurate in complex environments.
[0013] As a further technical solution, multi-modal data fusion is introduced, and it also includes: A two-stream network structure is constructed to process optical imaging data and laser ranging data respectively, and data fusion is performed through the feature pyramid network and the path aggregation network to improve the integration ability of multi-modal data.
[0014] As a further technical solution, the cross-modal attention mechanism is introduced, and it also includes: The weights of different modal data are dynamically adjusted to enhance the adaptability of the model under different sensor data.
[0015] As a further technical solution, self-supervised contrastive learning is introduced for efficient feature extraction and lightweight model design of the improved spatial non-cooperative target matching and recognition model, including: Self-supervised contrastive learning is introduced to minimize the feature differences between different targets by enhancing the feature similarity of the same target under different perspectives; Depthwise separable convolutions are used to reduce the computational complexity and the number of parameters of the model; Optimize the YOLOv5 model using model pruning techniques and lightweight network architecture design so that it can run in real time in resource-constrained environments.
[0016] As a further technical solution, the designed loss function also includes: Introduce focal loss to reduce the weight of easily classified samples and focus on optimizing difficult-to-classify samples; introduce complete IoU loss to consider the overlap degree of bounding boxes, the distance of center points, and the consistency of aspect ratios, and improve the accuracy of target boundary regression.
[0017] According to one aspect of the specification of the present invention, there is provided a lunar-earth space non-cooperative target recognition system based on an improved deep learning model, including: An image acquisition module for acquiring space debris images; An image recognition module for inputting the acquired space debris images into a trained space non-cooperative target matching and recognition model and outputting the recognized non-cooperative targets; wherein, the training of the space non-cooperative target matching and recognition model includes: Construct a space debris data set; Based on the deep learning model YOLOv5, introduce a self-attention mechanism and a spatial attention module, and also introduce multi-modal data fusion and cross-modal attention mechanisms to form an improved space non-cooperative target matching and recognition model; Introduce self-supervised contrast learning to perform efficient feature extraction and lightweight model design on the improved space non-cooperative target matching and recognition model to form an optimized space non-cooperative target matching and recognition model; According to the designed loss function, train the optimized space non-cooperative target matching and recognition model on the constructed space debris data set to obtain a trained space non-cooperative target matching and recognition model.
[0018] According to one aspect of the specification of the present invention, there is provided a lunar-earth space non-cooperative target recognition device including a memory and a processor, the memory stores program instructions executed by the processor, and the processor calls the program instructions to execute the steps of the lunar-earth space non-cooperative target recognition method based on the improved deep learning model.
[0019] According to one aspect of the specification of the present invention, there is provided a non-transitory computer-readable storage medium, the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the steps of the lunar-earth space non-cooperative target recognition method based on the improved deep learning model.
[0020] Compared with the prior art, based on a high-fidelity dataset, the present invention trains a deep learning object detection and recognition model that still has strong robustness in the scenario of high similarity between non-cooperative targets and background stars. The main innovation points and advantages of the present invention are embodied in the following aspects: 1. Combination of deep learning and non-cooperative targets: The present invention innovatively combines deep learning with non-cooperative targets and proposes a new method for identifying non-cooperative targets. By introducing a self-attention mechanism and a spatial attention module into the model, the ability to identify weak targets is enhanced. Especially under low-contrast and complex background conditions, high-precision object detection can be maintained.
[0021] 2. Multi-modal data fusion and cross-modal attention mechanism: The present invention proposes a method for multi-modal data fusion. By designing a dual-stream network structure, optical imaging data and laser ranging data are processed separately, and a feature pyramid network and a path aggregation network are used for efficient fusion. Combined with a cross-modal attention mechanism, the feature weights of different modalities are dynamically adjusted, thereby improving the model's adaptability to different sensor data and ensuring efficient object recognition under various observation conditions. The advantage of this method lies in its powerful data processing ability, which can flexibly cope with complex and changeable observation environments.
[0022] 3. Construction and application of a high-fidelity simulation dataset: Based on an accurate orbital dynamics model and the optical imaging law of non-cooperative targets, the present invention constructs a high-fidelity simulation dataset. This dataset covers a variety of target types, lighting conditions, noise interferences, and complex backgrounds, greatly improving the training effect of the deep learning model and providing rich data support for subsequent algorithm optimization and performance evaluation. This dataset is not only crucial for object recognition and trajectory prediction tasks but also provides an important reference for practical applications in deep space environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings used in the description of the embodiments or the prior art. Obviously, the following-described drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0024] Figure 1 It is a schematic flowchart of a method for identifying non-cooperative targets in the Earth-Moon space based on an improved deep learning model provided by an embodiment of the present invention.
[0025] Figure 2 This is the training flow chart of the spatial non - cooperative target matching and recognition model provided by the embodiments of the present invention.
[0026] Figure 3 This is the flow chart for constructing the spatial debris data set provided by the embodiments of the present invention.
[0027] Figure 4 This is a partial sample diagram of the spatial non - cooperative target data set provided by the embodiments of the present invention.
[0028] Figure 5 This is the visualization diagram of partial results of the evaluation of the high - fidelity spatial non - cooperative target test set provided by the embodiments of the present invention. Detailed implementation manners
[0029] The present invention proposes a method for identifying non - cooperative targets in the Earth - Moon space based on an improved deep - learning model, specifically targeting the detection and matching problems of non - cooperative targets (such as space debris, retired spacecraft, etc.) in complex backgrounds under long exposure times and star - tracking modes. This method is based on the YOLOv5 model and combines the characteristics of the deep - space environment and multiple innovative technologies, significantly improving the accuracy and efficiency of identifying and matching non - cooperative targets in complex backgrounds, mainly reflected in: First, the present invention first constructs a high - fidelity simulation data set to accurately simulate non - cooperative targets in the Earth - Moon space environment. Based on the orbital dynamics model and spatial environmental characteristics of the Earth - Moon space, it simulates the target observation processes of multiple types of sensors (such as ground - based optical telescopes, on - orbit cameras, laser rangefinders, etc.) under different observation conditions. And in the simulation, factors such as the optical imaging law of the target, the distribution of the background star field, the motion blur effect, and the density of the star field are considered to ensure that the appearance characteristics of the target and the background similarity reach a high degree of realism in the real scene. In addition, through an automated annotation mechanism, accurate trajectories, appearance attributes, and orbital parameters of the target are generated for each piece of data to ensure the high quality of the training data.
[0030] Second, the present invention has developed a single-stage deep learning detection framework. In view of the high visual feature similarity between non-cooperative targets and background stars, the YOLOv5 model is adopted and deeply customized and optimized on this basis. Specific innovations include: the introduction of the Self-Attention Mechanism and the Spatial Attention Module, enabling the model to more precisely capture the features of weak targets and improving the sensitivity of target detection. Through the dual-stream network structure, optical imaging data and laser ranging data are processed separately, and a feature fusion layer is set between the Feature Pyramid Network (FPN) and the Path Aggregation Network (PANet) to enhance the integration ability of different modality data. The Cross-Modal Attention mechanism is introduced, which can dynamically adjust the weights of different modality data, enhance the adaptability of the model to multiple data sources, and improve the detection accuracy of non-cooperative targets. In addition, self-supervised methods such as Contrastive Learning are used to enhance the robustness of the model under weak feature and low signal-to-noise ratio conditions, enabling the model to distinguish targets from the background relying on deep features in an environment where visual discrimination is difficult.
[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Additionally, the technical features in each embodiment or individual embodiment provided by the present invention can be arbitrarily combined with each other to form a new technical solution. This combination is not restricted by the order of steps and / or the structural composition mode, but must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0032] Please refer to Figure 1 , the embodiments of the present invention provide a method for identifying non-cooperative targets in the Earth-Moon space based on an improved deep learning model. First, a space debris image is acquired; then, the acquired space debris image is input into a trained spatial non-cooperative target matching and recognition model, and the recognized non-cooperative target is output.
[0033] Further, input the space debris image into the space non-cooperative target matching and recognition model, extract features from the optical imaging data and laser ranging data in the space debris image respectively, and introduce the self-attention mechanism and the spatial attention module during feature extraction; then perform multi-modal data fusion on the spatial attention images on the two branches, and introduce the cross-modal attention mechanism during multi-modal data fusion, and finally output the recognized non-cooperative target.
[0034] Please refer to Figure 2 , the training and evaluation of the space non-cooperative target matching and recognition model described in the embodiments of the present invention include: Step 1: Construct and preprocess the space debris dataset, including obtaining different types of space debris image data, and performing image synthesis and data annotation; Step 2: Improve the YOLOv5 model, introduce the self-attention mechanism (Self-Attention Mechanism) and the spatial attention module (Spatial Attention Module) into the YOLOv5 model, and innovatively introduce the multi-modal data fusion technology and the cross-modal attention mechanism (Cross-Modal Attention); Step 3: Customize the loss function of the improved YOLOv5 to better address the class imbalance and target boundary blur problems in the deep space environment; Step 4: For the optimized YOLOv5 model, introduce the self-supervised contrastive learning (Contrastive Learning) method, and perform efficient feature extraction and lightweight model design; Step 5: Train the optimized YOLOv5 model and evaluate its accuracy.
[0035] Please refer to Figure 3 , the construction and preprocessing of the space debris dataset in Step 1 further includes: 1. Dataset construction and target simulation Orbit simulation: First, according to the distribution and motion characteristics of non-cooperative targets (such as space debris, retired spacecraft, and small unknown targets) in the Earth-Moon space, establish a dynamic model that includes multi-source gravitational fields and non-spherical symmetric gravity field effects of the Earth, Moon, Sun, etc. This model provides orbital evolution information of the target under different initial conditions through numerical integration and parameter iteration.
[0036] Imaging simulation: Simulate the imaging process of actual observation sensors (such as ground-based optical telescopes and on-orbit cameras) through a high-precision optical imaging simulation framework, considering factors such as atmospheric disturbance, optical system aperture, detector noise, thermal noise, and photon counting noise.
[0037] Background star field generation: Use star catalog data to generate a highly realistic starry sky background, simulating star scenes with different magnitudes and densities to ensure a high degree of visual similarity between the background and the target.
[0038] Target material and shape modeling: Through 3D CAD modeling and the optical BRDF (Bidirectional Reflectance Distribution Function) model, comprehensively consider the material, shape, and reflection characteristics of the target to ensure that the target presents a star-like appearance in the image.
[0039] In the above steps, first, for the distribution and motion characteristics of non-cooperative targets in the Earth-Moon space environment, a dynamic model is established that includes multi-source gravitational fields such as the Earth, Moon, and Sun, and the effects of non-spherical symmetric gravity fields. Through numerical integration and parameter iteration, this model provides high-precision position information and velocity distribution for the orbital evolution of the target under different initial conditions. The key point of this step is to ensure the diversity and authenticity of the target trajectory: According to various possible target types (space debris, decommissioned spacecraft, small unknown targets), we set different orbital inclinations, semi-major axes, eccentricities, and relative illumination conditions, making the target as close as possible to the actual deep space environment in the simulation. After obtaining the trajectory information of the non-cooperative target, we use the optical imaging simulation framework to simulate the imaging process of actual observation sensors (such as ground-based medium-sized optical telescopes, on-orbit cameras) under specific attitudes, detection sensitivities, and field-of-view conditions. Here, a high-precision optical transmission model and an atmospheric disturbance estimation model are introduced, considering the comprehensive effects of the optical system aperture size, detector pixel scale, readout noise, photon counting noise, thermal noise, and observation conditions (such as sky background brightness, moonlight interference, detection band, etc.). At the same time, we use star catalog data to generate a vast and realistic starry sky background, simulating star scenes with different magnitude levels and density distributions, making the appearance characteristics of the background light source and potential targets highly close to the real observation data in terms of brightness, star image size, and gray scale distribution. This not only increases the difficulty of data generation and processing but also significantly improves the simulation degree and scientific value of the data set. To highlight the proximity of non-cooperative targets and background stars in terms of visual characteristics, we comprehensively consider the target material, shape structure, and surface reflection characteristics in the target simulation. By constructing a 3D CAD model of the target and fitting the corresponding optical BRDF (Bidirectional Reflectance Distribution Function) model, we conduct lighting analysis and image rendering for targets with different shapes, sizes, and materials. Due to the lack of specific identification, the visible image points of the target are similar to point sources or near-point source star images, and their gray scale distribution characteristics and apparent sizes are very similar to those of the surrounding stars. This refined modeling process greatly increases the difficulty and workload of data set construction but also makes the obtained data more challenging and representative in subsequent deep learning training.
[0040] 2. Data augmentation and annotation Use a parallel rendering pipeline to batch generate images with different observation time periods, angles, and sensor parameters, simulating various possible interferences and noises. Accurately label the targets in the images, including target positions, border coordinates, relative brightness information, etc. Ensure the accuracy of the labeled information to provide high-quality data for subsequent model training.
[0041] Specifically, after completing the trajectory calculation and optical simulation, we use a parallel rendering pipeline to batch generate data for different observation time periods, observation angles, and observation sensor parameters. During this process, various noise models, lighting changes, scattering effects, and possible motion blurs are also introduced. To ensure data diversity, we have carried out large-scale randomization in parameter selection, including but not limited to random perturbations of target distance distribution, relative azimuth change, background star field density and luminosity, and moderate changes in the observation time period (dawn, dusk, night).
[0042] After the above multi-step, high-precision, and high-computation simulation and data synthesis, we finally generated a large-scale and highly realistic deep-space non-cooperative target dataset, as Figure 4 shown. This dataset consists of 6,000 optical observation images simulating deep-space environments. In each image, the presentation position, brightness characteristics, and star image size of the non-cooperative targets are carefully designed to be highly similar to background stars in terms of visual features. In the dataset, the basic components of each image include: 1) Background star field: Distributed with star point sources of different magnitudes, and the brightness and density ranges are generated according to real star catalog data. 2) Non-cooperative targets: One to several targets are scattered in the image, with appearance characteristics similar to star point sources, but can be distinguished by motion characteristics in images collected at multiple consecutive moments. 3) Observation noise and imaging distortion: Sensor noise, optical system aberration, field-of-view edge light reduction, and changes in the stellar point spread function (PSF) are simulated to provide a realistic actual perturbation scenario for subsequent algorithm research and development.
[0043] After constructing the original simulation data, we utilized the known non-cooperative target trajectories and imaging geometric relationships to accurately label the target positions in the dataset. Through automated annotation tools and subsequent inspections, we generated standardized annotation information for the non-cooperative targets in each image (such as the bounding box coordinates of the target position, circular ROI, relative brightness information of the target, and its orbital parameter markings). The annotation process was rigorous and cumbersome, ensuring that there were no missing, incorrect, or ambiguous annotation results, providing a high-quality reference standard for the supervised training of the deep learning model. After completing the annotation, we made a reasonable division of the dataset suitable for the deep learning training process. Randomly select 1000 images from 6000 images as the test set, and the remaining 5000 images as the training and validation set. During the training process, a part can be further divided from the training set as the validation set for parameter tuning and overfitting detection in model training. This data division method ensures that the model can fully learn the subtle difference features between non-cooperative targets and background stars in large-scale, complex, and diverse training data, and at the same time objectively examines the generalization performance of the final model through an independent test set.
[0044] 3. Dataset Division and Preparation Divide the 6000-image dataset into a training set, a validation set, and a test set. Among them, 1000 images are used as an independent test set, and the remaining 5000 images are used for training and validation. Use automated annotation tools and manual inspections to ensure that the annotation results are free of mislabeling, missing labels, or ambiguities, and use the data division for the training and validation of the deep learning model.
[0045] The improvement of the YOLOv5 model design in step two further includes: Introduce the self-attention mechanism and the spatial attention module. The self-attention mechanism enhances the model's ability to focus on the features of the target area and improves the detection ability of weak targets. This mechanism dynamically adjusts the feature weights by calculating the correlations between positions in the feature map, enhancing the representation ability of key regions; the spatial attention module further enhances the extraction of target features. By strengthening the model's attention to the target area, it makes the recognition of small targets more accurate in complex environments.
[0046] Introduce multi-modal data fusion and cross-modal attention mechanism, design a two-stream network structure to process optical imaging data and laser ranging data respectively. Through the feature fusion layers in the Feature Pyramid Network (FPN) and the Path Aggregation Network (PANet), efficient integration of different modal features is achieved. Introduce the cross-modal attention mechanism to dynamically adjust the weights of different modal features, further improving the accuracy and robustness of target detection.
[0047] In step two, for model design and training, a YOLOv5 model architecture suitable for space debris recognition is selected. First, we deeply customized the YOLOv5 model to adapt to the high visual feature similarity between non-cooperative targets and background stars in the deep space environment. The traditional YOLOv5 model is mainly oriented towards object detection in ground or low-orbit environments, and it lacks performance when dealing with scenes with high noise, low contrast, and high similarity between targets and backgrounds. Therefore, we introduced a self-attention mechanism and a spatial attention module in the feature extraction part of the model, enhancing the model's ability to capture weak target features and improving the fineness of feature expression. Second, we innovatively introduced a multi-modal data fusion technology, effectively combining optical imaging data and laser ranging data. By designing a two-stream network structure to process different modal data respectively and setting a feature fusion layer between the feature pyramid network (FPN) and the path aggregation network (PANet), the efficient integration of different modal features is achieved. Furthermore, the application of the cross-modal attention mechanism enables the model to dynamically adjust the weights of each modal feature, improving the accuracy and robustness of object detection.
[0048] Data augmentation and feature extraction: In the deep space environment, non-cooperative targets and background stars are highly similar in visual features (such as brightness, shape, gray-scale distribution, etc.), and single-modal optical data faces significant detection challenges. To effectively distinguish these highly similar targets and backgrounds, we designed and implemented a number of innovative optical data augmentation and feature extraction methods on the basis of the YOLOv5 basic model. These innovative methods significantly improved the detection accuracy and robustness of the model in the complex deep space environment, solving the limitations of traditional methods in dealing with low signal-to-noise ratio and high similarity backgrounds.
[0049] To enhance the model's ability to capture the features of weak targets, we introduce the Self-Attention Mechanism and the Spatial Attention Module in the feature extraction part of YOLOv5. The self-attention mechanism is a technique widely used in deep learning, especially when dealing with sequential data. It allows the model to dynamically consider the information of other parts when processing a certain part of the input, thus enhancing the model's expressive power. The self-attention mechanism is one of the cores of the Transformer model and is widely used in fields such as natural language processing (NLP) and computer vision. The core idea of the self-attention mechanism is to enable each element of the input data to adjust its representation according to the relationships of other elements. Specifically, each element in the input sequence is compared with other elements, and information is weighted and combined based on this comparison. This enables the model to capture the dependencies between different positions in the input sequence.
[0050] The self-attention mechanism is usually implemented through the following steps: The input is a sequence of length n. Assuming the dimension of each element is d, the input sequence can be represented as an n×d matrix. Each input element will generate a query (Q), a key (K), and a value (V) through three different linear transformations (matrix multiplications). These transformations are learned through training. For the i-th input element x i , its query, key, and value are respectively: (1) where Q, K, and V are the query, key, and value matrices respectively, and WQ, W K and W V are weight matrices to be learned. The dot product of the query and the key is used to calculate the attention scores, which reflect the correlations between the input elements: (2) where, is the dimension of the key vector, and usually the scores are scaled to avoid excessive numerical values. The self-attention mechanism enhances the model's attention to key regions through global feature interaction and improves the detection ability for weak targets. The self-attention mechanism is a powerful tool that can help the model learn the dependencies between different parts of the input data and is widely used in multiple fields such as natural language processing and computer vision. Its advantages include capturing global dependencies, strong parallelization ability, and high processing flexibility, making it an important innovation in deep learning.
[0051] The Spatial Attention Module (SAM) is a variant based on the self-attention mechanism, mainly used in image processing tasks. By focusing on the weight distribution of spatial positions, it enhances the information of specific regions. The goal of the Spatial Attention Module is to perform attention weighting in the spatial dimension, enabling the model to focus on the most important regions in the image and ignore irrelevant parts, thereby improving the performance of the model.
[0052] In many computer vision tasks, the important regions in an image are usually unevenly distributed. The Spatial Attention Module can help the model automatically discover these key regions, similar to how the human eye often focuses on the most visually prominent parts when observing an image, thus enhancing the expressive ability of image features. The Spatial Attention Module is usually used in combination with the Channel Attention Module to form a more refined attention mechanism. The basic principle of the Spatial Attention Module is as follows: Input feature map: Given a feature map F ∈ R H×W×C , where H is the height of the image, W is the width of the image, and C is the number of channels. The purpose of the Spatial Attention Module is to weight the spatial dimension (i.e., H × W) of this feature map.
[0053] Generate the spatial attention map: The core idea of the Spatial Attention Module is to learn a spatial attention map based on the features at different positions in the image. This process usually includes the following steps: Compress the input feature map along the channel dimension through some operations (such as max pooling, average pooling, etc.) to generate a two-dimensional spatial feature map (i.e., a feature map of H × W). The purpose of this step is to extract the global features of each spatial position.
[0054] Merge the above-generated spatial feature maps (e.g., concatenation or weighting), and then generate a spatial attention map through a convolutional layer (usually a 1×1 convolution). The size of the spatial attention map is H × W, and the value at each position represents the importance of that position.
[0055] Specifically, assume that the feature maps obtained through max pooling and average pooling are Fmax and Favg respectively, and the spatial attention map S is obtained through convolutional calculation after merging:
[0056] where, is an activation function (such as sigmoid), Conv1 is a 1×1 convolutional operation, and Concat represents a concatenation operation. After obtaining the spatial attention map, it is multiplied element-wise with the original input feature map. In this way, the model can weight the input feature map according to the weights at each spatial position, thereby strengthening or suppressing the information in specific regions.
[0057]
[0058] Among them, Fout is the weighted output feature map, and S is the spatial attention map. The spatial attention module helps the model focus on important regions in the image by learning and weighting the spatial information in the image, thereby enhancing the feature representation and improving the performance of computer vision tasks. When used in combination with the channel attention module, the spatial attention module can more comprehensively enhance the attention mechanism of the model, making it more flexible and efficient in processing images.
[0059] To enable the model to dynamically adjust the weights of various modality features and improve the accuracy and robustness of object detection, further, a cross-modal attention mechanism (Cross-Modal Attention) is introduced. The cross-modal attention mechanism (Cross-Modal Attention) is a mechanism for processing the relationships between multi-modal data (such as images, text, audio, etc.). Its goal is to better fuse multi-modal information and improve the performance of multi-modal tasks (such as visual question answering, image captioning, speech recognition, etc.) by establishing mutual dependencies and relationships between different modalities.
[0060] The core idea of the cross-modal attention mechanism is to perform information transfer and fusion between data of different modalities through the attention mechanism. Specifically, assume we have two different modality data (for example, an image and text). We hope to capture the correlation information between the image and text through the attention mechanism and fuse them into a joint representation.
[0061] Input feature representation: Assume we have two modality data: one is the image feature V ∈ R Hv×Wv×Cv (possibly a feature map extracted by a convolutional neural network), and the other is the text feature T ∈ R Nt×Dt (which can be word embeddings or sentence features extracted by models such as RNN, Transformer, etc.). Here, Hv, Wv are the height and width of the image, Cv is the number of channels of the image feature map, Nt is the length of the text, and Dt is the dimension of the text feature.
[0062] Generating Queries, Keys, and Values: To enable cross-modal attention, first, we use the features of one modality (e.g., text features) as queries, and the features of another modality (image features) as keys and values. For image features V, we usually generate keys and values through some linear transformations. Assuming that each pixel (or region) of the image feature map represents information at a certain position, we hope to find relevant information from the text. For text features T, we generate queries through linear transformations to obtain relevant information from the image features.
[0063] Calculating Cross-modal Attention Weights: Assume that we use text features as queries and image features as keys and values. First, calculate the similarity between text queries and image keys to obtain the relationship between each text element and image element:
[0064] where Q is the text feature, K is the image feature, and d k is the dimension of the key vector. Then, convert this similarity into attention weights through the softmax function:
[0065] This weight reflects the correlation between the i-th word in the text and the j-th region in the image. Weighted Summation for Information Fusion: After obtaining the attention weights, perform a weighted sum of the image values using these weights to obtain a new representation of the image features, which serves as a supplement to the text with image information:
[0066] Similarly, we can also calculate the relationship between image features and text features by using the image as a query and the text as keys and values.
[0067] Multi-modal Fusion: The weighted features calculated through the cross-modal attention mechanism will contain information from both modalities. Usually, these features will be further fused (such as concatenation, addition, etc.) and used for subsequent tasks (such as classification, generation, etc.).
[0068] The cross-modal attention mechanism is a very powerful tool that can help establish effective associations and information transfer between data of different modalities. It is widely used in various tasks such as visual question answering, image caption generation, multi-modal sentiment analysis, etc., significantly improving the performance of multi-modal learning. By establishing interdependent relationships between different modalities, the cross-modal attention mechanism not only enhances the accuracy of information fusion but also improves the model's ability to handle complex tasks.
[0069] In step 3, the loss function of the improved YOLOv5 is customized, which further includes: For the customized loss function design of the model, to address the class imbalance and target boundary blur problems in the deep space environment, a loss function is customized. The Focal Loss is introduced, which reduces the weight of easy-to-classify samples and focuses on optimizing difficult-to-classify samples, thereby improving the detection accuracy; the Complete IoU Loss (CIoU Loss) is introduced, which considers the overlap degree of bounding boxes, the distance between center points, and the consistency of aspect ratios, effectively improving the accuracy of target boundary regression.
[0070] In step 3, the loss function of YOLOv5 is customized to better address the class imbalance and target boundary blur problems in the deep space environment. Customized loss function design: For the class imbalance and target boundary blur problems in deep space non-cooperative target detection, we customize the loss function of YOLOv5, introducing the Focal Loss and the Complete IoU Loss (CIoU Loss) to better meet the special requirements of the deep space environment.
[0071] The expression of the Focal Loss is:
[0072] where, is the predicted probability of the model for the true class, and are hyperparameters. The Focal Loss focuses on optimizing difficult-to-classify samples by reducing the loss weight of easy-to-classify samples, significantly reducing the false detection rate.
[0073] The expression of the Complete IoU Loss (CIoU Loss) is:
[0074] where, represents the Euclidean distance between the center points, c is the diagonal length of the smallest closed box enclosing the two boxes, v is the aspect ratio consistency metric, is the weight coefficient. The CIoU Loss not only considers the overlap degree of the bounding boxes but also incorporates the distance between the center points and the aspect ratio consistency, making the bounding box regression more accurate and further improving the detection accuracy.
[0075] The expression of the comprehensive loss function is:
[0076] where, is the weight coefficient of each loss term. Through the customized design of the loss function, the learning ability of the model in the complex deep space environment is optimized, and the detection performance is significantly improved.
[0077] In step four, for the optimized YOLOv5 model, a self-supervised contrastive learning method is introduced, and efficient feature extraction and lightweight model design are carried out, which further includes: Introduce the self-supervised contrastive learning method. By enhancing the feature similarity of the same target under different perspectives and minimizing the feature differences between different targets, the discrimination ability of the model in complex environments is enhanced. Use depthwise separable convolution to reduce the computational complexity and the number of parameters of the model, and improve the feature extraction efficiency. Adopt model pruning technology and lightweight network architecture design to optimize the YOLOv5 model, enabling it to run in real time in resource-constrained environments such as deep space detectors or satellites.
[0078] In step four, a self-supervised contrastive learning method is introduced, and efficient feature extraction and lightweight model design are carried out. Self-supervised contrastive learning: In view of the sparsity and diversity of deep space observation data, a series of data augmentation techniques are adopted, and a self-supervised contrastive learning method is introduced to improve the feature expression ability and generalization performance of the model.
[0079]
[0080] Among them, represents the similarity of the features of two samples, τ is the temperature parameter, and N is the batch size. By maximizing the feature similarity of the same target under different perspectives and minimizing the feature differences between different targets, the discrimination ability of the model in the complex deep space environment is enhanced.
[0081] Efficient feature extraction and lightweight model design: To meet the high requirements of deep space target detection for real-time performance and computing resources, we have carried out efficient feature extraction and lightweight design on the YOLOv5 model, optimizing the computational efficiency and inference speed of the model. Depthwise separable convolution:
[0082] By decomposing the standard convolution into depthwise convolution and pointwise convolution, the number of model parameters and computational complexity are significantly reduced, and the feature extraction efficiency is improved. Model Pruning: Through pruning techniques, redundant convolutional kernels and neurons in the model are removed, further reducing the computational resource requirements of the model and improving the inference speed. Lightweight Network Architecture Design: Adopting a bottleneck structure similar to MobileNet, the structural layout of the Backbone network is optimized, improving the efficiency and effect of feature extraction, and ensuring that the model can run in real time on resource-constrained deep space probes or satellites.
[0083] In step five, model training and model evaluation further include: 1. Model Training The improved YOLOv5 model is trained using a high-fidelity simulation dataset. The training process uses a batch size of 16, the initial learning rate is set to 1e-3, and the cosine annealing learning rate scheduling strategy is combined. To prevent overfitting, the Early Stopping strategy is adopted. When the F1 score on the validation set has not improved for 10 consecutive epochs, the training is terminated early.
[0084] 2. Model Evaluation and Accuracy Verification After training is completed, the model is evaluated using an independent test set. As Figure 5 shown, the test results show that the improved YOLOv5 model performs excellently in terms of the recognition accuracy and recall rate of space debris in the deep space environment: the recall rate reaches 96%, and the precision rate reaches 93%.
[0085] In step five, the optimized YOLOv5 model is trained and its accuracy is evaluated. In the present invention, based on the YOLOv5 model and the constructed high-fidelity simulation deep space non-cooperative target dataset, experiments on the recognition and matching of deep space non-cooperative targets are carried out to verify the effectiveness and robustness of the developed deep learning algorithm. During the experiment, first, 1000 out of 6000 simulation images are randomly selected as the independent test set, and the remaining 5000 images are used for model training and validation. The model training uses a batch size of 16, the initial learning rate is set to 1e -3 , and the cosine annealing learning rate scheduling strategy is combined to ensure that the model can converge effectively during the training process. To prevent overfitting, we introduce the Early Stopping strategy. When the F1 score on the validation set has not improved for 10 consecutive epochs, the training is terminated early. During the training process, the customized loss function design, including Focal Loss and Complete IoU Loss (CIoULoss), plays a key role in solving the problems of class imbalance and target boundary blur, significantly improving the detection accuracy of the model.
[0086] Table 1 Evaluation Results of the High-Fidelity Space Non-Cooperative Target Test Set
[0087] After completing the model training, we evaluated the model on an independent test set. The results are shown in Table 1. The recall rate (recognition rate) of the model reached 96%, and the precision rate (matching rate) was 93%. This high recall rate indicates that the model can effectively identify the vast majority of real deep-space non-cooperative targets, greatly reducing the possibility of missed detections. This is crucial for the monitoring and management of the Earth-Moon space environment, ensuring that the vast majority of space debris and abandoned spacecraft can be detected in a timely manner and reducing the risk of in-orbit spacecraft collisions. At the same time, the 93% precision rate means that 93% of the identified targets are truly non-cooperative targets, and the false detection rate is only 7%. This high precision rate effectively reduces the resource waste and subsequent processing burden caused by false detections, improving the overall efficiency and reliability of the system.
[0088] Based on the foregoing training process, the present invention can achieve efficient and accurate identification and matching of space debris in a complex deep-space environment, especially in the case of a high similarity between the background and the target. The present invention not only improves the detection accuracy of space debris but also reduces the false detection rate, providing strong technical support for deep-space exploration and space debris management.
[0089] The present invention is developed based on the deep-space non-cooperative target recognition and matching algorithm of YOLOv5. By introducing and implementing a number of innovative technologies, it successfully solves the problem of the high visual feature similarity between the target and background stars in the deep-space environment, significantly improving the detection accuracy, robustness, and real-time performance of the model. This innovative algorithm system not only provides an efficient and reliable technical means for the automatic detection and matching of non-cooperative targets in the Earth-Moon space but also has important scientific significance and broad application prospects in promoting the application and development of deep learning technology in the field of space science. In the future, we will further optimize the algorithm structure to improve its adaptability and real-time performance in actual deep-space observation data, and explore more fusion methods of deep learning and orbital dynamics models to continuously improve the recognition and orbit determination capabilities of non-cooperative targets, providing more reliable technical guarantees for the safety and sustainable development of the space environment. Through the above steps, the present invention provides a method for matching and identifying space debris based on the YOLOv5 model, which can effectively improve the automatic recognition and orbit prediction accuracy of space debris, has high real-time performance and accuracy, and is applicable to multiple fields such as deep-space exploration and satellite collision avoidance.
[0090] The implementation basis of each embodiment of the present invention is achieved through programmed processing by a device with processor functions. Therefore, in engineering practice, the technical solutions and functions of each embodiment of the present invention are encapsulated into various modules. Based on this actual situation, on the basis of the above embodiments, an embodiment of the present invention provides a lunar-earth space non-cooperative target recognition system based on an improved deep learning model, which is used to execute the lunar-earth space non-cooperative target recognition method based on the improved deep learning model in the above method embodiments.
[0091] The system includes: an image acquisition module for acquiring space debris images; an image recognition module for inputting the acquired space debris images into a trained space non-cooperative target matching and recognition model and outputting the recognized non-cooperative targets; wherein the training of the space non-cooperative target matching and recognition model includes: constructing a space debris data set; introducing a self-attention mechanism and a spatial attention module on the basis of the deep learning model YOLOv5, and also introducing multi-modal data fusion and cross-modal attention mechanism to form an improved space non-cooperative target matching and recognition model; introducing self-supervised contrast learning to perform efficient feature extraction and lightweight model design on the improved space non-cooperative target matching and recognition model to form an optimized space non-cooperative target matching and recognition model; training the optimized space non-cooperative target matching and recognition model on the constructed space debris data set according to the designed loss function to obtain a trained space non-cooperative target matching and recognition model.
[0092] The lunar-earth space non-cooperative target recognition system based on the improved deep learning model provided by the embodiment of the present invention aims at the current demand in the aerospace monitoring field to improve the recognition and matching accuracy of space debris and achieve real-time monitoring and early warning. By using the foregoing several modules and taking YOLOv5 as the basic model, through model improvement and high-fidelity data set construction, it solves the problems of detection and matching of non-cooperative targets (such as space debris, retired aircraft, etc.) under long exposure time and star tracking mode in a complex background, and improves the accuracy and efficiency of recognition and matching of non-cooperative targets in a complex background.
[0093] It should be noted that the system embodiment provided by the present invention, in addition to being used to implement the method in the above method embodiment, is also used to implement the methods in other method embodiments provided by the present invention. The difference is only in setting corresponding functional modules, and its principle is basically the same as that of the above system embodiment provided by the present invention. As long as those skilled in the art, on the basis of the above system embodiment, refer to the specific technical solutions in other method embodiments, obtain corresponding technical means by combining technical features, and the technical solutions composed of these technical means, and on the premise of ensuring the practicability of the technical solutions, improve the modules in the above system embodiment to obtain corresponding system-like embodiments for implementing the methods in other method-like embodiments.
[0094] Based on the same inventive concept as the foregoing embodiments, an embodiment of the present invention further provides a non-cooperative target recognition device for the Earth-Moon space based on an improved deep learning model, including a memory and a processor. The memory stores program instructions executed by the processor, and the processor invokes the program instructions to execute the steps of the non-cooperative target recognition method for the Earth-Moon space based on the improved deep learning model.
[0095] Based on the same inventive concept as the foregoing embodiments, an embodiment of the present invention further provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the steps of the non-cooperative target recognition method for the Earth-Moon space based on the improved deep learning model.
[0096] In summary of the above embodiments, the present invention discloses a non-cooperative target recognition method for the Earth-Moon space based on an improved deep learning model, which is particularly applied to the detection and matching of non-cooperative targets such as space debris and abandoned satellites in a complex background. Aiming at the current situation that the detection and recognition of non-cooperative targets in low-Earth orbit and the Earth-Moon space face challenges such as strong background interference and low signal-to-noise ratio, the present invention proposes an efficient target matching method based on the YOLOv5 model. By introducing a self-attention mechanism, a spatial attention module, and a multi-modal data fusion technology, the recognition and matching accuracy of non-cooperative targets in a complex background is significantly improved. Specifically, the present invention first constructs a high-fidelity simulation dataset of deep-space non-cooperative targets by simulating the observation data of multiple sensors in the Earth-Moon space environment and accurately annotating the data. Secondly, based on the YOLOv5 model, by introducing a self-attention mechanism and a spatial attention module, aiming at the similarity problem between non-cooperative targets and background stars, the present invention processes optical imaging data and laser ranging data through a two-stream network architecture and introduces a cross-modal attention mechanism to optimize the feature extraction and target matching process and enhance the robustness of the model under low signal-to-noise ratio conditions. By enhancing the feature learning ability through self-supervised contrast learning, the present invention can effectively distinguish non-cooperative targets from the background in a weak signal environment and improve the real-time monitoring and early warning ability. The technical solution of the present invention can provide efficient and accurate support for spacecraft orbit avoidance and mission planning, and has broad application prospects and good market prospects.
[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.
Claims
1. A non-cooperative target recognition method in Earth-Moon space based on an improved deep learning model, characterized in that: include: Acquire images of space debris; Inputting the acquired space debris image into the trained space non-cooperative target matching and recognition model, and outputting the recognized non-cooperative target; wherein the training of the space non-cooperative target matching and recognition model includes: Constructing a space debris dataset; Based on the deep learning model YOLOv5, the self-attention mechanism and spatial attention module are introduced, as well as multimodal data fusion and cross-modal attention mechanism to form an improved spatial non-cooperative target matching and recognition model; Self-supervised contrastive learning is introduced to perform efficient feature extraction and lightweight model design on the improved spatial non-cooperative target matching and recognition model, thus forming an optimized spatial non-cooperative target matching and recognition model; According to the designed loss function, the optimized spatial non-cooperative target matching and recognition model is trained on the constructed space debris dataset to obtain a trained spatial non-cooperative target matching and recognition model.
2. The method for identifying non-cooperative targets in Earth-Moon space based on an improved deep learning model according to claim 1, characterized in that: Construct a space debris dataset, including: Use orbital dynamics models to simulate target trajectories in Earth-Moon space, consider multi-source gravitational fields and non-spherically symmetric gravitational field effects, and provide high-precision position information and velocity distribution for the orbital evolution of targets under different initial conditions through numerical integration; Generate highly realistic target optical imaging in simulation, including simulating the target's three-dimensional morphology and surface reflection characteristics, and generate batch data for different observation periods, observation angles and sensor parameters through a parallel rendering pipeline to form an original simulation data set; By using the known non-cooperative target trajectories and imaging geometric relationships, the target positions in the original simulation data set are accurately marked to form a space debris data set.
3. The method for identifying non-cooperative targets in Earth-Moon space based on an improved deep learning model according to claim 1, characterized in that: Introducing the self-attention mechanism and spatial attention module, also includes: The self-attention mechanism is used to calculate the correlation of each position in the feature map, dynamically adjust the feature weights, and enhance the representation ability of key areas. The spatial attention module is used to enhance the extraction of target features and strengthen the model's attention to the target area, making the recognition of small targets in complex environments more accurate.
4. The method for identifying non-cooperative targets in Earth-Moon space based on an improved deep learning model according to claim 3, characterized in that: Introducing multimodal data fusion, also includes: A dual-stream network structure is constructed to process optical imaging data and laser ranging data respectively, and data fusion is performed through feature pyramid network and path aggregation network to improve the integration capability of multimodal data.
5. The method for identifying non-cooperative targets in the terrestrial and lunar space based on an improved deep learning model according to claim 4, characterized in that: Introducing the cross-modal attention mechanism, also includes: Dynamically adjust the weights of different modal data to enhance the adaptability of the model under different sensor data.
6. The method for identifying non-cooperative targets in Earth-Moon space based on an improved deep learning model according to claim 1, characterized in that: Self-supervised contrastive learning is introduced to perform efficient feature extraction and lightweight model design on the improved spatial non-cooperative target matching and recognition model, including: Introducing self-supervised contrastive learning to minimize the feature differences between different targets by enhancing the feature similarities of the same target under different perspectives; Use depth-wise separable convolution to reduce the computational complexity and number of parameters of the model; Model pruning technology and lightweight network architecture design are used to optimize the YOLOv5 model so that it can run in real time in a resource-constrained environment.
7. The method for identifying non-cooperative targets in Earth-Moon space based on an improved deep learning model according to claim 1, characterized in that: Designing loss functions also includes: The focal loss is introduced to reduce the weight of easy-to-classify samples and focus on optimizing difficult-to-classify samples; the full IoU loss is introduced to consider the overlap degree, center point distance and aspect ratio consistency of the bounding box to improve the accuracy of target boundary regression.
8. The non-cooperative target recognition system in the Earth-Moon space based on the improved deep learning model is characterized by: include: An image acquisition module, used for acquiring space debris images; The image recognition module is used to input the acquired space debris image into the trained space non-cooperative target matching and recognition model, and output the recognized non-cooperative target; wherein the training of the space non-cooperative target matching and recognition model includes: Constructing a space debris dataset; Based on the deep learning model YOLOv5, the self-attention mechanism and spatial attention module are introduced, as well as multimodal data fusion and cross-modal attention mechanism to form an improved spatial non-cooperative target matching and recognition model; Self-supervised contrastive learning is introduced to perform efficient feature extraction and lightweight model design on the improved spatial non-cooperative target matching and recognition model, thus forming an optimized spatial non-cooperative target matching and recognition model; According to the designed loss function, the optimized spatial non-cooperative target matching and recognition model is trained on the constructed space debris dataset to obtain a trained spatial non-cooperative target matching and recognition model.
9. A non-cooperative target recognition device in Earth-Moon space based on an improved deep learning model, characterized in that: It includes a memory and a processor, the memory stores program instructions executed by the processor, and the processor calls the program instructions to execute the steps of the method for identifying non-cooperative targets in the Earth-Moon space based on an improved deep learning model as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores computer instructions, which enable the computer to execute the steps of the method for identifying non-cooperative targets in the Earth-Moon space based on an improved deep learning model as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Space non-cooperative target capturing method based on deep enhancement learning
CN109625333A
Spatial non-cooperative target component identification method based on lightweight and attention mechanism
CN115205467A
Traffic sign recognition algorithm based on multi-head attention mechanism
CN115909280A
Non-cooperative spacecraft high-reliability detection method based on deep learning
CN117011661A
Non-cooperative target detection method in weak light environment
CN117372764A