Artificial intelligence detection method and system for small space target

The multi-scale feature extraction and fusion method enhances detection of small space targets by integrating adaptive convolutional embedding and attention mechanisms, addressing detection inefficiencies and complexity in existing AI models, thereby improving accuracy and adaptability.

CN120318495APending Publication Date: 2025-07-15GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510441905.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The prior art has problems of missed detection, insufficient accuracy and complex environmental interference in space small target detection, making it difficult to effectively identify multi-scale targets, and has high computational complexity, making it difficult to deploy in spacecraft with resource-constrained.

Method used

Multi-scale features are extracted using YOLO backbone network, deeper feature maps are generated through adaptive convolution embedding, feature fusion is combined with spatial and channel feature extraction branches, and global attention and lightweight design is used to achieve efficient fusion of multi-scale features.

Benefits of technology

It improves the accuracy and sensitivity of identification of tiny targets, reduces the computational complexity, adapts to complex space environments, enhances the adaptability and robustness of the model, and is suitable for spacecraft with resource constrained.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318495A_ABST
    Figure CN120318495A_ABST
Patent Text Reader

Abstract

The invention discloses an artificial intelligence detection method and system for a space small target, and the method comprises the steps: extracting the multi-scale features of an input image through a preset YOLO backbone network, and obtaining a multi-scale feature map; generating a deeper-level feature map introduced by a first scale, a deeper-level feature map introduced by a second scale and a deeper-level feature map introduced by a third scale from the multi-scale feature map through convolution embedding operation of a self-adaptive layer number; splicing the deeper feature maps introduced by different scales; key feature extraction is carried out on the deeper feature map obtained through splicing through a spatial feature extraction branch and a channel feature extraction branch, and an effective deep scale feature map is obtained; performing multi-scale fusion on the effective deep scale feature map and an original feature map extracted by a YOLO backbone network to obtain a multi-scale fusion feature map; and outputting a target prediction result from the multi-scale fusion feature map through a detection head.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of small target detection, and specifically, to an artificial intelligence detection method and system for small space targets. Background Art

[0002] Small space target detection refers to the identification and tracking of small-sized orbital debris, defunct satellites, and microspacecraft in the low Earth orbit and deep space environment. According to NASA, there are currently over 34,000 objects with a diameter greater than 10 cm in the low Earth orbit, as well as an estimated 900,000 debris with a diameter between 1 cm and 10 cm, and over 128 million smaller debris. These tiny objects moving at high speeds pose a serious threat to on-orbit spacecraft. The energy generated by a single collision is equivalent to that of a grenade explosion, which has become a major challenge in the field of space security.

[0003] In recent years, with the explosive growth of commercial space activities, the number of satellite launches globally has increased sharply every year, resulting in an exponential increase in the congestion index of the low Earth orbit. At the same time, the risk of the "Kessler effect" caused by space debris collisions has intensified, making the real-time detection and accurate early warning of small targets one of the core technologies for ensuring spacecraft safety. Traditional space situational awareness (SSA) systems mainly rely on ground-based radars and optical telescopes for monitoring. This traditional method has limited detection capabilities for small-sized targets and has a data delay of several hours. Therefore, developing an artificial intelligence recognition algorithm that can accurately detect small space targets is of great significance for ensuring the safety of space spacecraft.

[0004] With the rapid development of space photography equipment and artificial intelligence technology, it is possible to provide new methods for detecting small space targets. By deploying lightweight imaging devices on spacecraft, the environmental conditions around the spacecraft can be clearly observed, providing a solid data foundation for the identification of abnormal targets. At the same time, using advanced artificial intelligence technologies such as deep learning, the captured space images can be directly analyzed and processed automatically, thereby extracting key feature information in the images and realizing the detection of tiny abnormal targets. This method not only improves the accuracy and objectivity of detection but also significantly reduces the detection time compared to traditional radar scanning, providing strong support for the detection of abnormal targets and the formulation of subsequent risk avoidance strategies. In addition, the small space target detection algorithm based on artificial intelligence technology has stronger adaptability and robustness. It can continuously learn and optimize its detection algorithm to adapt to the complex and changing space environment. Facing the diversity and uncertainty of space debris, the algorithm can improve the recognition ability of small target features by training a large amount of space image data, effectively reducing false alarms and missed detections. This means that spacecraft can obtain more reliable information about the surrounding environment, make more accurate decisions, and ensure the safe execution of tasks.

[0005] Current mainstream research methods mainly focus on space target detection, but generally ignore the problem of missed detection caused by too small target sizes. The model design that does not pay attention to the monitoring of small abnormal targets makes it difficult to build a complete space situation awareness (SSA) system, and may miss small spacecraft or space debris with potential threats. Secondly, the existing research on small space target detection mainly aims to improve the positioning accuracy and detection sensitivity of targets, but generally lacks the ability to finely identify target categories, restricting the comprehensive analysis of the complex situation in the space environment and making it difficult to support the formulation of subsequent space behavior prediction and response strategies.

[0006] Current methods mainly adopt a feature extraction framework with three scales or an additional shallow scale, and there are technical bottlenecks in dealing with space multi-scale target detection tasks. Such models are easily affected by many interference factors that affect target detection brought about by the complex physical environment in space, which has an impact on the accuracy of target detection. They are even difficult to effectively capture the detailed features of small satellites or space debris, resulting in incomplete feature representation of multi-scale targets. When encountering dense distribution of target groups or sudden changes in lighting conditions, feature confusion and false alarm phenomena are likely to occur, restricting the real-time decision-making ability and situation awareness accuracy of the monitoring system.

[0007] The prior art with patent number CN202410117710.6 uses feature information of three scales, large, medium, and small, aiming to improve the accuracy of target detection. However, under the special physical conditions in space, there are usually interference factors such as stray light, cosmic ray noise, and stellar background in the image background. The existence of these interference factors makes it insufficient to use only the information of three scales in detecting small targets. In addition, the FPN (Feature Pyramid Network) structure adopted by this technology can only achieve the upward transfer of high-level semantic feature information, but lacks the downward transfer of low-level detail information, resulting in inevitable loss of feature information in the image, thus having a negative impact on the accuracy of target detection.

[0008] The prior art with patent number CN202310826843.6 realizes the global self-attention function by introducing the attention mechanism, which helps to improve the detection accuracy. However, the computational complexity of the attention mechanism is quadratic with the size of the input feature map. Especially when dealing with high-resolution deep space images, it will significantly increase the number of model parameters and computational volume, resulting in a decrease in the inference speed. This may make it difficult to deploy the model in resource-constrained spacecraft embedded systems. In addition, this model only supports single-class detection, which greatly limits its generality and practicality in actual applications. Because in actual application scenarios, there are often more than one type of space target, so this model cannot meet the diverse needs in actual operation. Summary of the Invention

[0009] To solve the technical problems existing in the background art, the present invention provides an artificial intelligence detection method and system for small space targets. The technical solution adopted by the present invention is as follows:

[0010] In the first aspect of the present invention, an artificial intelligence detection method for small space targets is provided. The method includes the following steps:

[0011] Extract multi-scale features of the input image through a preset YOLO backbone network to obtain a feature map of the first scale, a feature map of the second scale, and a feature map of the third scale;

[0012] Generate deeper feature maps introduced by the first scale, deeper feature maps introduced by the second scale, and deeper feature maps introduced by the third scale by performing a convolutional embedding operation with an adaptive number of layers on the feature map of the first scale, the feature map of the second scale, and the feature map of the third scale;

[0013] Concatenate the deeper feature maps introduced by the first scale, the deeper feature maps introduced by the second scale, and the deeper feature maps introduced by the third scale;

[0014] Refine key features of the concatenated deeper feature maps through a spatial feature extraction branch and a channel feature extraction branch respectively to obtain an effective deep-scale feature map;

[0015] Perform multi-scale fusion on the effective deep-scale feature map and the original feature map extracted by the YOLO backbone network to obtain a first fusion feature map, a second fusion feature map, and a third fusion feature map;

[0016] Output target prediction results through the detection head for the first fusion feature map, the second fusion feature map, and the third fusion feature map.

[0017] As a preferred solution, the YOLO backbone network sequentially includes an initial convolutional layer, a C3K2 module, an SPPF module, and a C2PSA module. Among them, the C3K2 module adopts a CSP structure with two small convolutional kernels, and the C2PSA module is used to enhance the spatial attention of the feature map;

[0018] Output feature maps of three scales through the YOLO backbone network:

[0019]

[0020] Among them, CBS represents the CBS module, C3K2 represents the C3K2 module, x represents the input image, respectively represent the feature map of the first scale, the feature map of the second scale, and the feature map of the third scale.

[0021] As a preferred solution, the method for generating the first deeper - level scale information feature map, the second deeper - level scale information feature map, and the third deeper - level scale information feature map by embedding the feature map of the first scale, the feature map of the second scale, and the feature map of the third scale through a convolutional embedding operation with an adaptive number of layers includes:

[0022] Perform convolution operations, batch normalization processing, and non - linear activation processing on the multi - scale feature map in sequence. The specific formulas are as follows:

[0023] CBR l (x)=ReLU(BN(Conv k×k (x)))

[0024] CB k (x)=BN(Conv k×k (x))

[0025] f i conv =CB1(CBR3(CBR1(f i Bone )))+CB1(f i Bone )

[0026] Among them, Conv k×k represents a convolutional layer with a kernel size of k×k, BN represents batch normalization, ReLU represents the ReLU activation function, f i Bone represents the feature map of the i - th scale extracted by the YOLO backbone network, and f i Conv represents the enhanced feature map of the i - th scale;

[0027] Perform a convolutional embedding operation with the corresponding number of layers M on the enhanced feature map of the i - th scale. At the same time, perform a layer - by - layer convolutional embedding operation by applying small - kernel convolution. The specific formula is as follows:

[0028] Conv M (x)=[Conv 2×2 (x)] ×M

[0029] f i P =Conv M (ReLU(f i Conv ))

[0030] Among them, when i = 1, M = 3; when i = 2, M = 2; when i = 3, M = 1, f i ConvRepresents the enhanced feature map of the i-th scale, f i P Represents the deeper feature map introduced by the i-th scale.

[0031] As a preferred solution, the method of concatenating the deeper feature maps introduced by the first scale, the deeper feature maps introduced by the second scale, and the deeper feature maps introduced by the third scale includes:

[0032] Concatenate the deeper feature maps introduced by the first scale, the deeper feature maps introduced by the second scale, and the deeper feature maps introduced by the third scale through the Concat operation. The specific formula is as follows:

[0033]

[0034] where f CP is the concatenated deeper feature map, are respectively the deeper feature map introduced by the first scale, the deeper feature map introduced by the second scale, and the deeper feature map introduced by the third scale.

[0035] As a preferred solution, the method of refining key features of the concatenated deeper feature map through a spatial feature extraction branch and a channel feature extraction branch to obtain an effective deep-scale feature map includes:

[0036] Input the concatenated deeper feature map into the spatial feature extraction branch and the channel feature extraction branch respectively;

[0037] The spatial feature extraction branch is processed through preset Reshape, Linear, Mamba modules and RMSNorm to generate a spatial feature map. The specific formula is:

[0038] token LP = Reshape(f CP )

[0039] f LP = Reshape(RMSNorm(Mamba(token LP )))

[0040] where f CP is the concatenated deeper feature map, Reshape represents the reshaping operation, Mamba represents the Mamba module, and RMSNorm represents the RMS normalization layer; fL LP represents the spatial feature map

[0041] The channel feature extraction branch generates a channel feature map through Reshape, Manba module, RMSNorm, and Sigmoid weighted processing. The specific formula is as follows:

[0042] token UP = Linear(Reshape(f CP ))

[0043] f UP = Reshape(RMSNorm(Mamba(token UP )))

[0044] where f CP is the deeper feature map obtained by concatenation. Linear represents the linear layer operation, Reshape represents the reshaping operation, Mamba represents the Mamba module, RMSNorm represents the RMS normalization layer, and f UP represents the channel feature map;

[0045] The spatial feature map and the channel feature map are subjected to a dot product operation to output an effective deep-scale feature map. The specific formula is as follows:

[0046] f DPSM = f LP * σ(Linear(f UP ))

[0047] f LP represents the spatial feature map, f UP represents the channel feature map, Linear represents the linear layer operation, σ represents the Sigmoid activation function, * represents the dot product operation in the channel dimension, and f DPSM represents the effective deep-scale feature map.

[0048] As a preferred solution, the method for multi-scale fusion of the effective deep-scale feature map and the original feature map extracted by the YOLO backbone network to obtain the first fusion feature map, the second fusion feature map, and the third fusion feature map includes:

[0049] The effective deep-scale feature map is successively upsampled to the same size as the original feature map extracted by the YOLO backbone network to obtain the first effective deep-scale feature map, the second effective deep-scale feature map, and the third effective deep-scale feature map. The specific formula is as follows:

[0050]

[0051] where represents the first effective deep-scale feature map, represents the second effective deep-scale feature map, It represents the third effective deep - scale feature map, and Upsample represents the up - sampling operation;

[0052] The first effective deep - scale feature map, the second effective deep - scale feature map, and the third effective deep - scale feature map are respectively concatenated with the feature maps of the first scale, the second scale, and the third scale. Then, the information of the three scales is fused with the introduced deeper - level scale information to obtain the first fusion feature map, the second fusion feature map, and the third fusion feature map. The specific formula is:

[0053]

[0054] Among them, represents the first fusion feature map, represents the second fusion feature map, represents the third fusion feature map, respectively represent the feature maps of the first scale, the second scale, and the third scale, and C3K2 represents the C3K2 module operation of YOLO.

[0055] As a preferred solution, the specific formula for outputting the target prediction result through the detection head with the first fusion feature map, the second fusion feature map, and the third fusion feature map is:

[0056]

[0057] Among them, Detect represents the YOLO detection head, represents the first fusion feature map, represents the second fusion feature map, represents the third fusion feature map, and Result represents the prediction result.

[0058] In the second aspect of the present invention, an artificial intelligence detection system for small space targets is provided. The system includes a multi - scale feature extraction module, an adaptive feature enhancement module, a splicing module, a dual - branch feature extraction module, a feature fusion module, and a detection module;

[0059] The multi - scale feature extraction module is used to extract the multi - scale features of the input image through a preset YOLO backbone network to obtain the feature maps of the first scale, the second scale, and the third scale;

[0060] The adaptive feature enhancement module is used to generate the deeper - level feature maps introduced by the first scale, the deeper - level feature maps introduced by the second scale, and the deeper - level feature maps introduced by the third scale by embedding operations with an adaptive number of convolutional layers for the feature maps of the first scale, the feature maps of the second scale, and the feature maps of the third scale;

[0061] The splicing module is used to splice the deeper feature maps introduced by the first scale, the deeper feature maps introduced by the second scale, and the deeper feature maps introduced by the third scale;

[0062] The dual-branch feature extraction module is used to refine key features of the spliced deeper feature maps through a spatial feature extraction branch and a channel feature extraction branch respectively, to obtain effective deep-scale feature maps;

[0063] The feature fusion module is used to perform multi-scale fusion on the effective deep-scale feature maps and the original feature maps extracted by the YOLO backbone network to obtain a first fusion feature map, a second fusion feature map, and a third fusion feature map;

[0064] The detection module is used to output target prediction results through a detection head for the first fusion feature map, the second fusion feature map, and the third fusion feature map.

[0065] The third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the foregoing artificial intelligence detection method for small space targets are implemented.

[0066] The fourth aspect of the present invention provides a computer device, including a storage medium, a processor, and a computer program stored in the storage medium and executable by the processor. When the computer program is executed by the processor, the steps of the foregoing artificial intelligence detection method for small space targets are implemented.

[0067] Compared with the prior art, the beneficial effects of the present invention are:

[0068] (1) Existing space target detections are mostly optimized based on conventional-scale target datasets or single-class small target datasets, with insufficient coverage and systematicness for multi-class small target features. The present invention uses a more comprehensive and detailed dataset for training the model, and this dataset covers 9 types of typical small space targets. By deeply analyzing the features of small space targets, the present invention designs a targeted target recognition strategy, which not only enhances the sensitivity of the model to tiny targets, but also improves its recognition accuracy in complex space backgrounds.

[0069] (2) The existing methods have insufficient feature scale information for model detection and are difficult to effectively capture tiny abnormal targets in complex space environments. The present invention designs an adaptive feature enhancement module, introduces deeper scale information, and significantly improves the model's detection ability for targets of different scales by fusing multi-scale feature maps.

[0070] (3) In order to improve the effect of feature fusion, the present invention proposes a dual-branch feature extraction module, which integrates global feature information in the spatial dimension and enhances the ability to capture information of small targets in complex backgrounds. In the channel dimension, Mamba's global attention mechanism is used to allocate weights and adaptively adjust the importance of different feature channels. This dual-channel strategy enables the model to extract and utilize key information more efficiently and accurately during the feature fusion process.

[0071] (4) Due to the limited computing resource constraints of spacecraft equipment, existing algorithms often face the problem of excessive computational complexity. The present invention combines the linear complexity Mamba sequence modeling technology with the lightweight architecture of YOLO, takes advantage of the global dependency modeling of the state space model, and uses the lightweight design of the deep separable convolution detection head design and the adaptive feature enhancement strategy. While maintaining the lightweight of the model, it also enhances the detection accuracy of the model for small targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 A flow chart of an artificial intelligence detection method for small space targets provided in this embodiment;

[0073] Figure 2 A schematic diagram of the network structure of an artificial intelligence detection method for small space targets provided in this embodiment;

[0074] Figure 3 A schematic diagram of the structure of the adaptive feature enhancement module provided in this embodiment;

[0075] Figure 4 A schematic diagram of the structure of a dual-branch feature extraction module provided in this embodiment. DETAILED DESCRIPTION

[0076] The accompanying drawings are only used for illustrative purposes and are not to be construed as limiting the present invention;

[0077] It should be clear that the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the embodiments of the present application.

[0078] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the embodiments of the present application. The singular forms of "a", "said" and "the" used in the embodiments of the present application and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.

[0079] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims. In the description of the present application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects, and do not have to be used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.

[0080] In addition, in the description of the present application, unless otherwise specified, "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0081] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0082] Embodiment 1

[0083] Please refer to Figure 1 and Figure 2 , this embodiment provides an artificial intelligence detection method for small space targets, and the method includes the following steps:

[0084] S1: Extract multi-scale features of the input image through a preset YOLO backbone network to obtain a feature map of the first scale, a feature map of the second scale, and a feature map of the third scale;

[0085] In a specific embodiment, the YOLO backbone network sequentially includes an initial convolutional layer, a C3K2 module, an SPPF module, and a C2PSA module, where the C3K2 module adopts a CSP structure with double small convolutional kernels, and the C2PSA module is used to enhance the spatial attention of the feature map;

[0086] Output three-scale feature maps through the YOLO backbone network:

[0087]

[0088] Among them, CBS represents the CBS module, C3K2 represents the C3K2 module, and x represents the input image. respectively represent the feature maps of the first scale, the feature maps of the second scale, and the feature maps of the third scale.

[0089] S2: Generate the deeper feature maps introduced by the first scale, the deeper feature maps introduced by the second scale, and the deeper feature maps introduced by the third scale by performing a convolutional embedding operation with an adaptive number of layers on the feature maps of the first scale, the feature maps of the second scale, and the feature maps of the third scale;

[0090] Please refer to Figure 3 , in a specific embodiment, the method for generating the first deeper scale information feature map, the second deeper scale information feature map, and the third deeper scale information feature map by performing a convolutional embedding operation with an adaptive number of layers on the feature maps of the first scale, the feature maps of the second scale, and the feature maps of the third scale includes:

[0091] Perform convolutional operations, batch normalization processing, and non-linear activation processing on the multi-scale feature maps in sequence. The specific formula is as follows:

[0092] CBR k (x) = ReLU(BN(Conv k×k (x)))

[0093] CB k (x) = BN(Conv k×k (x))

[0094]

[0095] Among them, Conv k×k represents a convolutional layer with a kernel size of k×k, BN represents batch normalization, ReLU represents the ReLU activation function, f i Bone represents the feature map of the i-th scale extracted by the YOLO backbone network, and f i Conv represents the enhanced feature map of the i-th scale;

[0096] Perform a convolutional embedding operation with the corresponding number of layers M on the enhanced feature map of the i-th scale. At the same time, perform a layer-by-layer convolutional embedding operation by applying a small kernel convolution. The specific formula is as follows:

[0097] Conv M (x) = [Conv 2×2 (x)] ×M

[0098] f i P = Conv M (ReLU(f iConv ))

[0099] Among them, when i is 1, M is 3; when i is 2, M is 2; when i is 3, M is 1, f i Conv represents the enhanced feature map of the i-th scale, and f i P represents the deeper feature map introduced by the i-th scale.

[0100] S3: Concatenate the deeper feature maps introduced by the first scale, the deeper feature maps introduced by the second scale, and the deeper feature maps introduced by the third scale;

[0101] In a specific embodiment, the method for concatenating the deeper feature maps introduced by the first scale, the deeper feature maps introduced by the second scale, and the deeper feature maps introduced by the third scale includes:

[0102] Concatenate the deeper feature maps introduced by the first scale, the deeper feature maps introduced by the second scale, and the deeper feature maps introduced by the third scale through the Concat operation. The specific formula is as follows:

[0103]

[0104] where, f CP is the concatenated deeper feature map, are respectively the deeper feature maps introduced by the first scale, the deeper feature maps introduced by the second scale, and the deeper feature maps introduced by the third scale.

[0105] S4: Refine the key features of the concatenated deeper feature maps through the spatial feature extraction branch and the channel feature extraction branch respectively to obtain the effective deep-scale feature maps;

[0106] Please refer to Figure 4 In a specific embodiment, the method for refining the key features of the concatenated deeper feature maps through the spatial feature extraction branch and the channel feature extraction branch respectively to obtain the effective deep-scale feature maps includes:

[0107] Input the concatenated deeper feature maps into the spatial feature extraction branch and the channel feature extraction branch respectively;

[0108] The spatial feature extraction branch is processed through the preset Reshape, Linear, Manba modules and RMSNorm to generate the spatial feature map. The specific formula is:

[0109] token LP= Reshape(f CP )

[0110] f LP = Reshape(RMSNorm(Mamba(token LP )))

[0111] where f CP is a deeper feature map obtained by concatenation, Reshape represents the reshaping operation, Mamba represents the Mamba module, RMSNorm represents the RMS normalization layer; f LP represents the spatial feature map

[0112] The channel feature extraction branch generates a channel feature map through Reshape, Manba module, RMSNorm and Sigmoid weighted processing. The specific formula is:

[0113] token UP = Linear(Reshape(f CP ))

[0114] f UP = Reshape(RMSNorm(Mamba(token UP )))

[0115] where f CP is a deeper feature map obtained by concatenation, Linear represents the linear layer operation, Reshape represents the reshaping operation, Mamba represents the Mamba module, RMSNorm represents the RMS normalization layer, f uP represents the channel feature map;

[0116] Multiply the spatial feature map and the channel feature map pointwise to output an effective deep-scale feature map. The specific formula is:

[0117] f DPSM = f LP * σ(Linear(f UP ))

[0118] f LP represents the spatial feature map, f UP represents the channel feature map, Linear represents the linear layer operation, σ represents the Sigmoid activation function, * represents the pointwise multiplication operation in the channel dimension, f DPSM represents the effective deep-scale feature map.

[0119] S5: Multiscale fuse the effective deep-scale feature map and the original feature map extracted by the YOLO backbone network to obtain the first fused feature map, the second fused feature map, and the third fused feature map;

[0120] In a specific embodiment, the method for multi-scale fusion of the effective deep-scale feature map and the original feature map extracted by the YOLO backbone network to obtain the first fusion feature map, the second fusion feature map, and the third fusion feature map includes:

[0121] The effective deep-scale feature map is sequentially upsampled to the same size as the original feature map extracted by the YOLO backbone network to obtain the first effective deep-scale feature map, the second effective deep-scale feature map, and the third effective deep-scale feature map. The specific formula is:

[0122]

[0123]

[0124] Among them, represents the first effective deep-scale feature map, represents the second effective deep-scale feature map, represents the third effective deep-scale feature map, and Upsample represents the upsampling operation;

[0125] The first effective deep-scale feature map, the second effective deep-scale feature map, and the third effective deep-scale feature map are respectively corresponding concatenated with the feature map of the first scale, the feature map of the second scale, and the feature map of the third scale, and then the three-scale information is fused with the introduced deeper-scale information to obtain the first fusion feature map, the second fusion feature map, and the third fusion feature map. The specific formula is:

[0126]

[0127] Among them, represents the first fusion feature map, represents the second fusion feature map, represents the third fusion feature map, respectively represent the feature map of the first scale, the feature map of the second scale, and the feature map of the third scale, and C3K2 represents the C3K2 module operation of YOLO.

[0128] S6: Output the target prediction result through the detection head using the first fusion feature map, the second fusion feature map, and the third fusion feature map;

[0129] In a specific embodiment, the specific formula for outputting the target prediction result through the detection head using the first fusion feature map, the second fusion feature map, and the third fusion feature map is:

[0130]

[0131] Among them, Detect represents the YOLO detection head, represents the first fused feature map, represents the second fused feature map, represents the third fused feature map, and Result represents the prediction result.

[0132] Embodiment 2

[0133] This embodiment provides an artificial intelligence detection system for small space targets. The system includes a multi-scale feature extraction module, an adaptive feature enhancement module, a splicing module, a dual-branch feature extraction module, a feature fusion module, and a detection module;

[0134] The multi-scale feature extraction module is used to extract multi-scale features of the input image through a preset YOLO backbone network, and obtain a feature map of the first scale, a feature map of the second scale, and a feature map of the third scale;

[0135] Please refer to Figure 3 , the adaptive feature enhancement module is used to generate deeper feature maps introduced by the first scale, deeper feature maps introduced by the second scale, and deeper feature maps introduced by the third scale by performing convolutional embedding operations with an adaptive number of layers on the feature map of the first scale, the feature map of the second scale, and the feature map of the third scale;

[0136] The splicing module is used to splice the deeper feature maps introduced by the first scale, the deeper feature maps introduced by the second scale, and the deeper feature maps introduced by the third scale;

[0137] Please refer to Figure 4 , the dual-branch feature extraction module is used to refine key features of the spliced deeper feature maps through a spatial feature extraction branch and a channel feature extraction branch respectively, and obtain an effective deep-scale feature map;

[0138] The feature fusion module is used to perform multi-scale fusion on the effective deep-scale feature map and the original feature map extracted by the YOLO backbone network to obtain a first fused feature map, a second fused feature map, and a third fused feature map;

[0139] The detection module is used to output the target prediction result through the detection head for the first fused feature map, the second fused feature map, and the third fused feature map.

[0140] Embodiment 3

[0141] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, it implements the steps of the artificial intelligence detection method for small space targets described in Embodiment 1.

[0142] Example 4

[0143] A computer device includes a storage medium, a processor, and a computer program stored in the storage medium and executable by the processor. When the computer program is executed by the processor, it implements the steps of an artificial intelligence detection method for small space targets described in Example 1.

[0144] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. An artificial intelligence detection method for small space targets, characterized in that, The method includes the following steps: Extract multi-scale features of the input image through a preset YOLO backbone network to obtain a feature map of the first scale, a feature map of the second scale, and a feature map of the third scale; Generate deeper feature maps introduced by the first scale, deeper feature maps introduced by the second scale, and deeper feature maps introduced by the third scale by performing a convolutional embedding operation with an adaptive number of layers on the feature map of the first scale, the feature map of the second scale, and the feature map of the third scale; Concatenate the deeper feature maps introduced by the first scale, the deeper feature maps introduced by the second scale, and the deeper feature maps introduced by the third scale; Refine key features of the concatenated deeper feature maps through a spatial feature extraction branch and a channel feature extraction branch respectively to obtain an effective deep-scale feature map; Perform multi-scale fusion on the effective deep-scale feature map and the original feature map extracted by the YOLO backbone network to obtain a first fusion feature map, a second fusion feature map, and a third fusion feature map; Output the target prediction result through the detection head using the first fusion feature map, the second fusion feature map, and the third fusion feature map.

2. The artificial intelligence detection method for small space targets according to claim 1, characterized in that, The YOLO backbone network sequentially includes an initial convolutional layer, a C3K2 module, an SPPF module, and a C2PSA module, where the C3K2 module adopts a CSP structure with double small convolutional kernels, and the C2PSA module is used to enhance the spatial attention of the feature map; Output feature maps of three scales through the YOLO backbone network: Among them, CBS represents the CBS module, C3K2 represents the C3K2 module, and x represents the input image. They respectively represent the feature map of the first scale, the feature map of the second scale, and the feature map of the third scale.

3. An artificial intelligence detection method for small space targets according to claim 1, characterized in that, The method of generating a first deeper-scale information feature map, a second deeper-scale information feature map, and a third deeper-scale information feature map by performing a convolutional embedding operation with an adaptive number of layers on the feature map of the first scale, the feature map of the second scale, and the feature map of the third scale includes: Perform convolutional operations, batch normalization processing, and non-linear activation processing on the multi-scale feature maps in sequence. The specific formulas are as follows: CBR k (x) = ReLU(BN(Conv k×k (x))) CB k (x) = BN(Conv k×k (x)) Among them, Conv k×k represents a convolutional layer with a kernel size of k×k, BN represents batch normalization, and ReLU represents the ReLU activation function. represents the feature map of the i-th scale extracted by the YOLO backbone network, and f i Conv represents the enhanced feature map of the i-th scale; Perform a convolutional embedding operation with the corresponding number of layers M on the enhanced feature map of the i-th scale. At the same time, perform a layer-by-layer convolutional embedding operation by applying a small-kernel convolution. The specific formula is as follows: Conv M (x) = [Conv 2×2 (x)] ×M Among them, when i is 1, M is 3; when i is 2, M is 2; when i is 3, M is 1, f i Conv represents the enhanced feature map of the i-th scale, represents the deeper feature map introduced by the i-th scale.

4. An artificial intelligence detection method for small space targets according to claim 1, characterized in that, The method of concatenating the deeper feature maps introduced by the first scale, the deeper feature maps introduced by the second scale, and the deeper feature maps introduced by the third scale includes: Concatenate the deeper feature maps introduced by the first scale, the deeper feature maps introduced by the second scale, and the deeper feature maps introduced by the third scale through a Concat operation. The specific formula is as follows: Among them, f CP is the deeper feature map obtained by splicing, are respectively the deeper feature map introduced by the first scale, the deeper feature map introduced by the second scale, and the deeper feature map introduced by the third scale.

5. An artificial intelligence detection method for small space targets according to claim 1, characterized in that, The method of refining key features of the concatenated deeper feature maps through a spatial feature extraction branch and a channel feature extraction branch respectively to obtain an effective deep-scale feature map includes: Input the concatenated deeper feature maps into a spatial feature extraction branch and a channel feature extraction branch respectively; The spatial feature extraction branch generates a spatial feature map through preset Reshape, Linear, Manba modules and RMSNorm processing. The specific formula is: token LP = Reshape(f CP ) f LP = Reshape(RMSNorm(Mamba(token LP ))) Among them, f CP is the deeper feature map obtained by splicing, Reshape represents the reshaping operation, Mamba represents the Mamba module, and RMSNorm represents the RMS normalization layer; f LP represents the spatial feature map The channel feature extraction branch generates a channel feature map through Reshape, Manba module, RMSNorm, and Sigmoid weighting processing. The specific formula is as follows: token UP = Linear(Reshape(f CP )) f UP = Reshape(RMSNorm(Mamba(token UP ))) Among them, f CP is the deeper feature map obtained by splicing, Linear represents the linear layer operation, Reshape represents the reshaping operation, Mamba represents the Mamba module, RMSNorm represents the RMS normalization layer, and f UP represents the channel feature map; Perform a dot product operation on the spatial feature map and the channel feature map to output an effective deep-scale feature map. The specific formula is as follows: f DPSM = f LP * σ(Linear(f UP )) f LP represents the spatial feature map, and f UP represents the channel feature map, Linear represents the linear layer operation, σ represents the Sigmoid activation function, * represents the dot product operation in the channel dimension, and f DPSM represents the effective deep-scale feature map.

6. The artificial intelligence detection method for small space targets according to claim 1, characterized in that The method of performing multi-scale fusion on the effective deep-scale feature map and the original feature map extracted by the YOLO backbone network to obtain the first fusion feature map, the second fusion feature map, and the third fusion feature map includes: Upsample the effective deep-scale feature map sequentially to the same size as the original feature map extracted by the YOLO backbone network to obtain the first effective deep-scale feature map, the second effective deep-scale feature map, and the third effective deep-scale feature map. The specific formula is as follows: Among them, represents the first effective deep-scale feature map, represents the second effective deep-scale feature map, represents the third effective deep-scale feature map, and Upsample represents the upsampling operation; Perform corresponding splicing on the first effective deep-scale feature map, the second effective deep-scale feature map, and the third effective deep-scale feature map with the feature maps of the first scale, the second scale, and the third scale respectively. Then fuse the three-scale information with the introduced deeper-scale information to obtain the first fusion feature map, the second fusion feature map, and the third fusion feature map. The specific formula is as follows: Among them, represents the first fused feature map, represents the second fused feature map, represents the third fused feature map, respectively represent the feature maps of the first scale, the second scale, and the third scale. C3K2 represents the C3K2 module operation of YOLO.

7. An artificial intelligence detection method for small space targets according to claim 1, characterized in that, The specific formula for outputting the target prediction result by the detection head through the first fusion feature map, the second fusion feature map, and the third fusion feature map is as follows: Among them, Detect represents the YOLO detection head, represents the first fused feature map, represents the second fused feature map, represents the third fused feature map, and Result represents the prediction result.

8. An artificial intelligence detection system for small space targets, characterized in that, The system includes a multi-scale feature extraction module, an adaptive feature enhancement module, a splicing module, a dual-branch feature extraction module, a feature fusion module, and a detection module; The multi-scale feature extraction module is used to extract multi-scale features of the input image through a preset YOLO backbone network to obtain the feature maps of the first scale, the second scale, and the third scale; The adaptive feature enhancement module is used to generate the deeper-level features introduced by the first scale, the deeper-level features introduced by the second scale, and the deeper-level features introduced by the third scale by embedding operations of convolutions with an adaptive number of layers on the feature maps of the first scale, the feature maps of the second scale, and the feature maps of the third scale; The splicing module is used to splice the deeper-level features introduced by the first scale, the deeper-level features introduced by the second scale, and the deeper-level features introduced by the third scale; The dual-branch feature extraction module is used to extract key features from the spliced deeper-level feature map through a spatial feature extraction branch and a channel feature extraction branch respectively to obtain an effective deep-scale feature map; The feature fusion module is used to perform multi-scale fusion on the effective deep-scale feature map and the original feature map extracted by the YOLO backbone network to obtain the first fusion feature map, the second fusion feature map, and the third fusion feature map; The detection module is used to output the target prediction result through the detection head with the first fusion feature map, the second fusion feature map, and the third fusion feature map.

9. A computer-readable storage medium storing a computer program thereon, characterized in that: When the computer program is executed by a processor, it implements the steps of an artificial intelligence detection method for small space targets as described in any one of claims 1 to 7.

10. A computer device, characterized in that: It includes a storage medium, a processor, and a computer program stored in the storage medium and executable by the processor. When the computer program is executed by the processor, it implements the steps of an artificial intelligence detection method for small space targets as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Space small target detection method based on position coding

    CN117079098A

  • Space target detection method based on Atlas200DK and related equipment

    CN117830857A