Deep learning-based celestial body faint target feature extraction method

By building a target neural network, extracting and fusing multi-scale and weak target information, the problem of low accuracy of weak target detection in celestial object detection is solved, and the detection rate is improved.

CN120355933AActive Publication Date: 2025-07-22PLA PEOPLES LIBERATION ARMY OF CHINA STRATEGIC SUPPORT FORCE AEROSPACE ENG UNIV

Patent Information

Application Number
CN202510845707.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect dark and weak celestial targets in celestial target detection, resulting in a decrease in detection accuracy.

Method used

Using a deep learning-based method, a high-resolution feature map is generated by building a target neural network, including a backbone network, a feature extraction module and a feature fusion module.

Benefits of technology

The detection rate of dark and weak celestial targets is improved, the network's attention to dark and weak targets is enhanced, and information loss problem is alleviated during multi-scale feature extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355933A_ABST
    Figure CN120355933A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image processing, particularly relates to a celestial body faint target feature extraction method based on deep learning, and aims to solve the problem that in the prior art, a faint celestial body target is difficult to effectively detect when the celestial body target is detected, so that the detection accuracy is influenced. The method provided by the invention comprises the following steps: acquiring a foundation optical image to be detected; inputting the ground-based optical image into a pre-constructed target neural network to obtain a target feature map output by the target neural network, the target feature map being used for indicating detail information and semantic information of a celestial body faint target in the ground-based optical image; the target neural network comprises a backbone network, a feature extraction module and a feature fusion module. Based on the method provided by the invention, a feature extraction module (Res-MBDA) and a feature fusion module (CAFFM) are adopted to enhance the attention on the information of the faint celestial body target, and the detection rate of the faint celestial body target is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a method for extracting features of faint celestial objects based on deep learning. Background Art

[0002] The deep learning method based on data-driven can learn target features from data and is widely used in celestial object detection. According to the characteristic that the scales of celestial objects vary greatly, the existing detection modules focus on extracting the multi-scale characteristics of targets, and effectively detect celestial objects of different sizes through downsampling methods such as cascaded convolutional layers and feature pyramid structures.

[0003] However, the existing modules ignore the characteristic that the brightness distribution of celestial objects has a large dynamic range when extracting multi-scale features of targets. The downsampling (a technique commonly used in signal processing and image processing to reduce the sampling rate or resolution of data) method ignores the effective retention of faint target information, resulting in the loss of information of faint celestial objects with low contrast to the background, which affects the detection accuracy. Summary of the Invention

[0004] To solve the above problems in the prior art, that is, it is difficult to effectively detect faint celestial objects during celestial object detection in the prior art, thus affecting the detection accuracy, the present invention proposes a method for extracting features of faint celestial objects based on deep learning. The method includes: Obtain a ground-based optical image to be detected; Input the ground-based optical image into a pre-constructed target neural network to obtain a target feature map output by the target neural network. The target feature map is used to indicate the detail information and semantic information of faint celestial objects in the ground-based optical image; Wherein, the target neural network includes: A backbone network for performing multi-level convolution on the ground-based optical image to obtain input feature maps at multiple scales; A feature extraction module for extracting features from the input feature maps to obtain output feature maps. The output feature maps have the same resolution as the input feature maps. The feature extraction module is composed of multiple branches, and the dilation rates of the dilated convolutions corresponding to each branch increase in sequence, so that each branch obtains the target information of the input feature maps according to the receptive field range increasing in sequence. The target information includes the detail information, context information, and multi-scale information of the input feature maps; A feature fusion module is used to extract the bottommost feature map and the topmost feature map from the output feature maps for adaptive weighted fusion processing to obtain a fused feature map. After channel dimension stacking and convolution operations are performed on the fused feature map with a high-resolution feature map, the target feature map is generated. The high-resolution feature map is an output feature map that has not been downsampled.

[0005] In some preferred embodiments, the input feature map F of the feature extraction module includes all the shallow feature maps of the same layer as well as two adjacent layer feature maps and .

[0006] In some preferred embodiments, the input feature map F of the feature extraction module satisfies: ; In the formula, represents the feature stacking operation in the channel dimension, represents the convolution operation with a stride of 2 and a convolution kernel of 4, represents the transposed convolution operation corresponding to the convolution operation.

[0007] In some preferred embodiments, the processing process of the feature extraction module for the input feature map F includes: ; In the formula, represents the output feature map of the feature extraction module, represents the intermediate feature map used to calculate , represents the standard convolution operation, represents the CBAM attention mechanism based on residual connection, represents the ReLU activation function, represents the element-wise addition operation.

[0008] In some preferred embodiments, the satisfies: ; ; In the formula, represents the output feature map of the corresponding branch, , represents three different branches, represents the dilated convolution, and the subscript is the dilation rate, represents the standard convolution operation, represents the scSE attention mechanism.

[0009] In some preferred embodiments, the processing process of the feature fusion module satisfies: ; ; ; ; ; wherein, is the bottommost feature map, is the topmost feature map, is the processing result based on the bottommost feature map and the topmost feature map, and respectively represent the attention features in the channel dimension and the spatial dimension, is the processing result based on the attention features in the channel dimension and the spatial dimension, and respectively represent the global average pooling operations across the channel dimension and the cross-spatial dimension, represents the global max pooling operation across the spatial dimension, represents the channel shuffle operation, represents the group convolution, and the subscript is the convolution kernel size, represents the sigmoid activation function, represents the fusion weight.

[0010] In some preferred embodiments, the weighted feature maps of the topmost feature map and the bottommost feature map satisfy: ; ; wherein, is the weighted feature map of the topmost feature map, is the weighted feature map of the bottommost feature map, is the topmost feature map, is the bottommost feature map.

[0011] In some preferred embodiments, the feature fusion module includes at least two branches to retain the detailed information of the input feature map, and the finally output fused feature map satisfies: ; wherein, represents the standard convolution operation with a convolution kernel of 1, is the weighted feature map of the topmost feature map, is the weighted feature map of the bottommost feature map, is the topmost feature map, is the bottommost feature map.

[0012] In some preferred embodiments, the process of performing channel dimension stacking and convolution operations on the high-resolution feature maps in the fused feature map and the output feature map satisfies: ; In the formula, represents a standard convolution operation, represents the fused feature map finally output by the feature fusion module, represents the 0th layer, that is, the high-resolution feature map without downsampling, from the 1st high-resolution feature map to the jth high-resolution feature map, where j represents the serial number of the feature map in the penultimate position.

[0013] In some preferred embodiments, the method further includes: Inputting the target feature map into a pre-constructed inference network to obtain the target segmentation result output by the inference network; According to the target segmentation result, obtaining target position and prediction box information by performing connected component extraction and centroid localization operations, where the target position and prediction box information are used to indicate faint celestial targets in the ground-based optical image.

[0014] Advantages of the present invention: Considering conditions such as large differences in the scales of celestial targets and limited information of faint celestial targets, the method proposed by the present invention takes into account the extraction of multi-scale information and faint target information when extracting celestial target features, alleviates the problem of decreased target detection rate caused by the loss of faint target information due to downsampling when extracting multi-scale features. The feature extraction module (Res-MBDA) effectively extracts detailed information, context information, and scale information of celestial targets, improving the utilization rate of faint celestial target information; the feature fusion module (CAFFM) performs adaptive weighted fusion on the output feature map, and the fused feature map fully retains the detailed information and semantic information of faint celestial targets, improving the fusion quality. The above two modules can be plugged and played on various networks and have high efficiency. Through this method, the network's attention to faint celestial targets can be enhanced, and the network detection rate can be improved. Description of the Drawings

[0015] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objectives, and advantages of the present application will become more obvious: Figure 1 is a schematic flowchart of a method for extracting features of faint celestial targets based on deep learning proposed by an embodiment of the present invention; Figure 2 is a schematic diagram of the processing flow of the target neural network proposed by an embodiment of the present invention; Figure 3It is a schematic structural diagram of the Res-MBDA feature extraction module proposed in an embodiment of the present invention; Figure 4 It is a schematic structural diagram of the CAFFM feature fusion module proposed in an embodiment of the present invention; Figure 5 It is a schematic structural diagram of a computer system proposed in an embodiment of the present invention. Detailed implementation manners

[0016] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. Additionally, it should be noted that for the sake of description, only parts related to the relevant invention are shown in the drawings.

[0017] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.

[0018] Please refer to Figure 1 , the first embodiment of the present application provides a method for extracting features of faint celestial objects based on deep learning, and the method includes: Step S10, obtaining a ground-based optical image to be detected; Step S20, inputting the ground-based optical image into a pre-constructed target neural network to obtain a target feature map output by the target neural network, where the target feature map is used to indicate the detail information and semantic information of the faint celestial object in the ground-based optical image; Among them, the target neural network includes: A backbone network, which is used to perform multi-level convolution on the ground-based optical image to obtain input feature maps at multiple scales; A feature extraction module, which is used to extract features from the input feature maps to obtain output feature maps. The output feature maps have the same resolution as the input feature maps. The feature extraction module is composed of multiple branches, and the dilation rates of the atrous convolutions corresponding to each branch increase in sequence, so that each branch obtains the target information of the input feature maps according to the increasing receptive field range. The target information includes the detail information, context information, and multi-scale information of the input feature maps; A feature fusion module, which is used to extract the bottommost feature map and the topmost feature map in the output feature maps for adaptive weighted fusion processing to obtain a fusion feature map. After channel dimension stacking and convolution operations with a high-resolution feature map, the target feature map is generated. The high-resolution feature map is the output feature map without downsampling.

[0019] It should be noted that considering the characteristics of large differences in the scales of celestial targets and large dynamic ranges in brightness distributions, in this embodiment, a U-Net network based on dense nesting is used as the above-mentioned backbone network. Specifically, it can be a DNA-Net (i.e., Dense nested attention network) for infrared small target detection networks.

[0020] The above-mentioned backbone network is also used to perform up / down sampling processing on the output feature map obtained by the feature extraction module. Specifically, convolution / transposed convolution is used to implement sampling. The parameters of this sampling method are learnable, and it can adaptively adjust the sampling parameters according to image characteristics, task requirements, etc. It has stronger flexibility and can effectively enhance the network's ability to represent and restore the features of weak celestial targets, which is beneficial for the backbone network to maintain and enhance the weak target information in the network. After the up / down sampling processing, in order to make the most of the limited information of faint celestial targets, all the high-resolution feature maps that have not been downsampled are stacked in the channel dimension with the fusion feature map, and the final target feature map is obtained through a simple convolutional layer.

[0021] Specifically, please refer to Figure 2 , Figure 2 which shows the process of processing ground-based optical images by the target neural network based on this embodiment. Among them, the feature extraction module in this embodiment is specifically a Multi-Branch Dilated Attention block with Residual Connections (hereinafter simply referred to as Res-MBDA), and the feature fusion module in this embodiment is specifically a Cross Attention Feature Fusion Module (hereinafter simply referred to as CAFFM).

[0022] In a feasible implementation manner, Res-MBDA receives the feature map information from three directions as input to ensure that the input features fully contain the faint target information. In order to effectively utilize the limited information of faint targets during the feature extraction process, Res-MBDA abandons the traditional downsampling method and realizes the extraction of multi-scale features of the target and enhances the context information correlation ability between features through a multi-branch parallel structure of dilated convolutions with different dilation coefficients and the scSE attention mechanism; particularly, a linear mapping branch can also be designed to retain the faint target information in the input feature map. Finally, the CBAM attention mechanism based on residual connections further explores the correlation between features in different dimensions of the fusion feature map and outputs the result.

[0023] In a feasible implementation, CAFFM realizes the adaptive weighted fusion of the topmost feature map and the bottommost feature map: First, perform an element-wise addition operation on the two feature maps to achieve a rough mixing of the two feature maps. Generate corresponding attention maps for the mixed feature map in the spatial dimension and the channel dimension to highlight the importance of features at different spatial positions and channel positions. After adding the attention maps of the two dimensions element-wise, the important information of the target under different dimensions can be effectively retained, avoiding the important information being overwhelmed by interference information such as noise or background. Then, perform channel shuffle and group convolution to further ensure the full circulation of features, and generate the corresponding fusion weights of the two feature maps through the sigmoid activation function. In addition, in order to retain the original detail information and semantic information in the two feature maps, two additional branches can be designed to transmit the information of the original feature maps. Finally, perform element-wise addition and convolution operations on the information from the four branches to achieve the adaptive weighted fusion of the feature maps. Finally, in order to make full use of the limited information of faint celestial object targets, stack and perform convolution operations on all high-resolution feature maps and CAFFM fusion feature maps in the channel dimension to generate the above-mentioned target feature map.

[0024] This embodiment mainly proposes a feature extraction module Res-MBDA and a feature fusion module CAFFM. Res-MBDA is mainly used to mine and utilize the limited information from faint celestial object targets in the celestial object information containing many scales, and continuously enhance the feature representation of the information of faint celestial object targets through the way of module stacking to generate high-quality feature maps (i.e., high-resolution feature maps); CAFFM is mainly used to adaptively weight and fuse the bottommost feature map and the topmost feature map, so as to highlight the detail information and semantic information of faint celestial object targets in the fused feature map and improve the quality of feature map fusion. In addition, in order to make full use of the limited information of faint celestial object targets, stack and perform convolution on all high-resolution feature maps and CAFFM fusion feature maps in the channel dimension. Finally, the generated feature map highlights the key areas of faint celestial object targets, improving the celestial object detection rate of the network. Based on this method, the information of faint celestial object targets can be fully retained and utilized in the whole processing process, and the celestial object detection rate of ground-based optical images can be improved through the effective interaction between detail information and semantic information.

[0025] Specifically, the input feature map F of the feature extraction module includes all shallow feature maps of the same layer as well as two adjacent layer feature maps and .

[0026] For example, when applying this method based on the infrared small target detection network DNA-Net (Dense nested attention network, where DNA-Net is an instance of the dense nested U-Net network), feature extraction is performed using Res-MBDA on the same-level nodes of the backbone network. Let the number of the j-th feature map in the i-th layer be , then the information input to Res-MBDA comes from three directions: all shallow feature maps of the same layer and two adjacent-layer feature maps and .

[0027] More specifically, the input feature map F of the feature extraction module satisfies: ; In the formula, represents the feature stacking operation in the channel dimension, represents the convolution operation with a stride of 2 and a convolution kernel of 4, represents the transposed convolution operation corresponding to the convolution operation.

[0028] More specifically, the processing process of the feature extraction module for the input feature map F includes: ; In the formula, represents the output feature map of the feature extraction module, represents the intermediate feature map used to calculate , represents the standard convolution operation, represents the CBAM attention mechanism based on the residual connection, represents the ReLU activation function, represents the element-wise addition operation.

[0029] More specifically, the satisfies: ; ; In the formula, represents the output feature map of the corresponding branch, , represents three different branches, represents the dilated convolution, and the subscript is the dilation rate, represents the standard convolution operation, represents the scSE attention mechanism.

[0030] Based on this embodiment, Res-MBDA is used to fully exploit the target detail information, context information, and scale information in the feature map, and the scSE attention mechanism is used to integrate the above information in the spatial dimension and channel dimension. By suppressing useless information, an enhanced representation of useful information is achieved, and a feature map with the same resolution as the input feature map is output, effectively retaining the faint target information and achieving an enhanced representation of the faint target information, thereby improving the network's attention to faint targets.

[0031] Further, please refer to Figure 3 In the Res-MBDA feature extraction module shown in Figure 3 , is mainly used for the transformation of the number of channels. The purpose of using the 1×3 and 3×1 forms in the second and third branches is to reduce the computational amount, and its essence is equivalent to a 3×3 convolution. is used to provide different receptive field ranges to achieve the extraction of the detail information, context information, and scale information of celestial targets, which is achieved by taking different dilation rate parameters rate. Integrates the feature information through feature interaction in the spatial dimension and channel dimension, and realizes the enhanced representation of useful information by suppressing useless information; the Res-CBAM attention mechanism is mainly used to generate the attention map of the feature to ensure that the important information of the target is retained in the output feature map and not overwhelmed by the remaining information.

[0032] In this embodiment, after being processed by the Res-MBDA feature extraction module, a series of output feature maps will be obtained. After these output maps are fused by the CAFFM feature fusion module, a target feature map that can effectively indicate the faint celestial target information will be finally obtained.

[0033] In this embodiment, the processing flow of the above feature fusion module includes: ; ; ; ; ; where is the bottom-layer feature map, is the top-layer feature map, is the processing result based on the bottom-layer feature map and the top-layer feature map, and respectively represent the attention features in the channel dimension and the spatial dimension, is the processing result based on the attention features in the channel dimension and the spatial dimension, and respectively represent the global average pooling operations across the channel dimension and the spatial dimension, represents the global max pooling operation across the spatial dimension, represents the channel shuffle operation, represents the group convolution, and the subscript is the convolution kernel size, represents the sigmoid activation function, represents the fusion weight.

[0034] Specifically, the weighted feature maps of the top-level feature map and the low-level feature map satisfy: ; ; In the formula, is the weighted feature map of the topmost feature map, is the weighted feature map of the bottommost feature map, is the topmost feature map, is the bottommost feature map.

[0035] More specifically, in order to further retain the original information in the input feature map, especially the detailed information in the low-level feature map, which is prone to being submerged in the background, two additional branches are constructed to retain the information of the input feature map. Therefore, the cross-attention fusion feature finally output by CAFFM can be expressed as: ; In the formula, represents the standard convolution operation with a convolution kernel of 1, is the weighted feature map of the topmost feature map, is the weighted feature map of the bottommost feature map, is the topmost feature map, is the bottommost feature map.

[0036] Furthermore, please refer to Figure 4 In the CAFFM feature fusion module shown in Figure 4 , first, for the bottommost feature map and the topmost feature The rough mixing of features is achieved through element-wise addition operations. Then, the feature attention maps of the mixed feature map in the spatial dimension and the channel dimension are obtained through spatial attention and channel attention respectively. The attention maps in the two dimensions are added element-wise to achieve the interaction between the high-level feature map and the low-level feature map in the spatial and channel dimensions. Channel shuffle (CS) and group convolution (GConv) further promote the flow and fusion of features with less computational cost, ensuring sufficient interaction between features at different spatial positions and different channels. Finally, the sigmoid activation function is applied to generate the feature map fusion weights. To maximize the retention of information in the input feature map, two additional branches are added to retain the information in []. Finally, CAFFM performs element-wise addition on the four-branch feature maps to complete the adaptive weighted fusion of the feature maps.

[0037] In this embodiment, the fused feature map obtained by CAFFM is stacked with all high-resolution feature maps in the backbone network in the channel dimension. Since the high-resolution feature map contains the richest information of faint celestial object targets, stacking is used to make full use of the limited information of faint celestial object targets and improve the detection rate of the network for such targets.

[0038] Specifically, the process of stacking the fused feature map with the remaining output feature maps in the channel dimension satisfies: ; In the formula, represents the standard convolution operation, represents the fused feature map finally output by the feature fusion module, represents the first to the j-th high-resolution feature maps of the 0th layer (i.e., the non-downsampled high-resolution feature map layer), and j represents the serial number of the feature map in the penultimate position.

[0039] In this embodiment, the above method further includes: inputting the target feature map into a pre-constructed inference network to obtain the target segmentation result output by the inference network; according to the target segmentation result, obtaining the target position and prediction box information through performing connected component extraction and centroid localization operations, and the target position and prediction box information are used to indicate the faint celestial object targets in the ground-based optical image.

[0040] Specifically, the pre-constructed inference network in this embodiment can be a semantic segmentation network based on a fully convolutional structure, which can generate the target segmentation result through pixel-by-pixel classification; at the same time, according to the target segmentation result, combined with the well-known "connected component extraction → centroid localization → prediction box generation" process in the art, the final faint celestial object target can be obtained. The algorithms adopted in this process and the relevant parameters of the prediction box information, etc., can all be implemented by any feasible solution in the art, and this embodiment does not make any limitations.

[0041] The second embodiment of the present application provides a system for extracting features of faint celestial objects based on deep learning, including: A data acquisition module, configured to acquire ground-based optical images to be detected; A data output module, configured to input the ground-based optical images into a pre-constructed target neural network, and obtain a target feature map output by the target neural network, where the target feature map is used to indicate the detail information and semantic information of faint celestial objects in the ground-based optical images; Wherein, the target neural network includes: A backbone network, configured to perform multi-level convolution on the ground-based optical images to obtain input feature maps at multiple scales; A feature extraction module, configured to extract features from the input feature maps to obtain output feature maps, where the output feature maps have the same resolution as the input feature maps, and the feature extraction module is composed of multiple branches, and the dilation rates of the atrous convolutions corresponding to each branch increase in sequence, so that each branch obtains the target information of the input feature maps according to the increasing receptive field range, and the target information includes the detail information, context information and multi-scale information of the input feature maps; A feature fusion module, configured to extract the bottommost feature map and the topmost feature map in the output feature maps for adaptive weighted fusion processing to obtain a fusion feature map, and after performing channel dimension stacking and convolution operations on the fusion feature map with a high-resolution feature map, generate the target feature map, where the high-resolution feature map is the output feature map without downsampling.

[0042] The third embodiment of the present application further proposes an electronic device, including: At least one processor; and a memory communicatively connected to at least one of the processors; wherein, the memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the method as described in the first embodiment.

[0043] The fourth embodiment of the present application further proposes a computer-readable storage medium, where the computer-readable storage medium stores computer instructions, and the computer instructions are used to be executed by the computer to implement the method as described in the first embodiment.

[0044] Next, refer to Figure 5 , which shows a schematic structural diagram of a computer system of a server suitable for implementing the method, system, and device embodiments of the present application. Figure 5 The server shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0045] As Figure 5As shown, the computer system includes a central processing unit (CPU), 301, which can perform various appropriate actions and processes according to a program stored in a read only memory (ROM), 302, or a program loaded from a storage section, 308, into a random access memory (RAM), 303. In the RAM, 303, various programs and data required for system operations are also stored. The CPU, 301, the ROM, 302, and the RAM, 303, are connected to each other via a bus, 304. An input / output (I / O) interface, 305, is also connected to the bus, 304.

[0046] The following components are connected to the I / O interface, 305: an input section, 306, including a keyboard, a mouse, etc.; an output section, 307, including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section, 303, including a hard disk, etc.; and a communication section, 309, including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section, 309, performs communication processing via a network such as the Internet. A drive, 310, is also connected to the I / O interface, 305, as required. A removable medium, 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive, 310, as required so that a computer program read therefrom can be installed into the storage section, 308, as required.

[0047] Specifically, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product that includes a computer program carried on a computer-readable medium, the computer program including program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section, 309, and / or installed from the removable medium, 311. When the computer program is executed by a central processing unit (CPU), 301, the above-described functions defined in the method of the present application are performed. It should be noted that the computer-readable medium described above in the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above.

[0048] More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present application, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device. And in the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0049] Computer program code for performing the operations of the present application may be written in one or more programming languages or combinations thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0050] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions denoted in the blocks may occur in a different order than that denoted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0051] The terms "first", "second", etc. are used to distinguish similar objects and not to describe or indicate a particular order or sequence.

[0052] The term "comprising" or any other similar term is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus / device that comprises a series of elements includes not only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or apparatus / device.

[0053] Thus far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings.

[0054] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A method for extracting features of faint celestial objects based on deep learning, characterized in that The method includes: Obtaining a ground-based optical image to be detected; Inputting the ground-based optical image into a pre-constructed target neural network to obtain a target feature map output by the target neural network, where the target feature map is used to indicate the detail information and semantic information of faint celestial targets in the ground-based optical image; Among them, the target neural network includes: A backbone network for performing multi-level convolution on the ground-based optical image to obtain input feature maps at multiple scales; A feature extraction module for extracting features from the input feature maps to obtain output feature maps. The output feature maps have the same resolution as the input feature maps. The feature extraction module is composed of multiple branches, and the dilation rates of the atrous convolutions corresponding to each branch increase sequentially, so that each branch obtains the target information of the input feature maps according to the sequentially increasing receptive field ranges. The target information includes the detail information, context information, and multi-scale information of the input feature maps; A feature fusion module for adaptively weighted fusion processing of the bottommost feature map and the topmost feature map in the output feature maps to obtain a fused feature map. After channel dimension stacking and convolution operations with a high-resolution feature map, the target feature map is generated. The high-resolution feature map is the output feature map without downsampling.

2. The method for extracting features of faint celestial objects based on deep learning according to claim 1, wherein, The input feature map F of the feature extraction module includes all the shallow feature maps of the same layer as well as the feature maps of two adjacent layers and .

3. The method for extracting features of faint celestial objects based on deep learning according to claim 2, wherein The input feature map F of the feature extraction module satisfies: ; In the formula, represents the feature stacking operation performed along the channel dimension, represents the convolution operation with a stride of 2 and a convolution kernel of 4, represents the transposed convolution operation corresponding to the convolution operation.

4. The method for extracting features of faint celestial objects based on deep learning according to claim 1, wherein, The processing process of the feature extraction module for the input feature map F includes: ; Wherein, represents the output feature map of the feature extraction module, represents the intermediate feature map used to calculate , represents the standard convolution operation, represents the CBAM attention mechanism based on residual connection, represents the ReLU activation function, represents the element-wise addition operation.

5. The method for extracting features of faint celestial objects based on deep learning according to claim 3, wherein The said satisfies that: ; ; In the formula, represents the output feature map of the corresponding branch, , representing three different branches, represents dilated convolution, with the subscript being the dilation rate, represents a standard convolution operation, represents the scSE attention mechanism.

6. The method for extracting features of faint celestial objects based on deep learning according to claim 1, wherein The processing process of the feature fusion module satisfies: ; ; ; ; ; In the formula, is the bottommost feature map, is the topmost feature map, is the processing result based on the bottommost feature map and the topmost feature map, and respectively represent the attention features in the channel dimension and the spatial dimension, is the processing result based on the attention features in the channel dimension and the spatial dimension, and respectively represent the global average pooling operations across the channel dimension and across the spatial dimension, represents the global max pooling operation across the spatial dimension, represents the channel shuffle operation, represents group convolution, and the subscript is the convolution kernel size, represents the sigmoid activation function, represents the fusion weight.

7. The method for extracting features of faint celestial objects based on deep learning according to claim 6, wherein The weighted feature maps of the topmost feature map and the bottommost feature map satisfy: ; ; In the formula, is the weighted feature map of the topmost feature map, is the weighted feature map of the bottommost feature map, is the topmost feature map, is the bottommost feature map.

8. The method for extracting features of faint celestial objects based on deep learning according to claim 6, wherein The feature fusion module includes at least two branches to retain the detailed information of the input feature map, and the finally output fused feature map of the feature fusion module satisfies: ; In the formula, represents the standard convolution operation with a convolution kernel of 1, is the weighted feature map of the topmost feature map, is the weighted feature map of the bottommost feature map, is the topmost feature map, is the bottommost feature map.

9. The method for extracting features of faint celestial objects based on deep learning according to claim 1, wherein The process of channel dimension stacking and convolution operations between the fused feature map and the high-resolution feature map in the output feature maps satisfies: ; In the formula, represents the standard convolution operation, represents the fused feature map finally output by the feature fusion module, represents the 0th layer, that is, the high-resolution feature layer without downsampling, from the 1st high-resolution feature map to the jth high-resolution feature map, where j represents the serial number of the feature map at the penultimate position.

10. The method for extracting features of faint celestial objects based on deep learning according to claim 1, wherein The method further includes: Inputting the target feature map into a pre-constructed inference network to obtain a target segmentation result output by the inference network; According to the target segmentation result, obtaining target position and prediction box information by performing connected component extraction and centroid localization operations. The target position and prediction box information are used to indicate faint celestial targets in the ground-based optical image.

Citation Information

Patent Citations

  • Infrared weak and small target detection method and device, computing equipment and storage medium

    CN115546586A

  • Infrared weak and small target detection method and device, computing equipment and storage medium

    CN115661573A

  • Auxiliary intelligent driving target detection method based on improved YOLOV4

    CN115761697A

  • Infrared weak and small target tracking method and device, electronic equipment and storage medium

    CN116740135A

  • Display screen defect detection method and system based on enhanced feature extraction network

    CN118570212A

Cited By

  • Spatial faint target detection method based on multi-scale wavelet features

    CN121458966A