An enhanced detection method for dim space targets based on optical feature fusion
Through multi-spectral data fusion and deep learning technology, the problem of insufficient accuracy of dark object detection in complex spatial environments is solved, and more efficient target detection and analysis is achieved.
Patent Information
- Application Number
- CN202510513213.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The prior art is difficult to effectively integrate multiple optical perception features in complex spatial environments, resulting in increased difficulty in detecting and analyzing targets in dark and weak spaces, especially in the condition of unknown samples.
Multispectral data acquisition, preprocessing, adaptive weight fusion algorithm, deep learning object detection model combined with attention mechanism and self-supervised learning methods are used to detect and position weak targets, including Gaussian filtering, wavelet transformation, histogram equalization and other technical means.
It improves the detection accuracy and robustness of dark and weak space targets, makes up for the limitations of a single data source, and improves the intelligent detection and intelligence analysis capabilities of space targets.
Smart Images

Figure CN120030502B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for enhancing the detection of dim space targets based on optical feature fusion, and belongs to the field of target detection. Background Art
[0002] To improve the detection ability of dim space targets, fully explore the intrinsic correlation features of various types of perception information including multi-spectral imaging, mid-long wave infrared imaging, visible light imaging, laser imaging, etc., make up for the limitations of a single data source, so as to obtain more comprehensive and detailed target characteristic information, and enhance the space target detection ability.
[0003] In terms of feature fusion analysis, combining the data of multi-source sensors and using technologies such as computer vision, machine learning, and multi-modal analysis to establish intelligent detection and autonomous analysis of space targets is an important task. However, in the face of the continuous growth of the number of targets and unknown changes in space activities, it is difficult to obtain sufficient learnable sample sets in a timely manner, resulting in increased difficulty in target detection and analysis. Therefore, there is an urgent need to establish a feature fusion enhancement detection technology for rare sample space targets that is not limited to known samples and can integrate different optical perception features, improve the intelligent detection and intelligence analysis capabilities of space targets, and provide strong technical support for space situation awareness. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for enhancing the detection of dim space targets based on optical feature fusion, which can improve the detection ability of dim space targets, make up for the limitations of a single data source, obtain more comprehensive and detailed target characteristic information, and enhance the intelligent detection and intelligence analysis capabilities of space targets.
[0005] To achieve the above purpose, the present invention provides a method for enhancing the detection of dim space targets based on optical feature fusion, including the following steps:
[0006] Step1. Collect multi-spectral data in a dim space scene, including optical information such as mid-long wave infrared imaging, visible light imaging, and laser imaging;
[0007] Step2. Preprocess the multi-spectral data collected in Step1, including noise removal, image enhancement, and feature extraction;
[0008] Step3. Use an adaptive weight fusion algorithm to perform feature fusion on the preprocessed multi-spectral data in Step2 to generate a fused multi-scale feature map;
[0009] Step4. Use a deep learning target detection model combined with an attention mechanism to detect and locate dim targets in the fused feature map;
[0010] Step 5. Introduce a self-supervised learning method to train the deep learning object detection model, and utilize unlabeled data to enhance the generalization ability of the model;
[0011] Step 6. Post-process the detection results of Step 4, including background noise filtering and target precise positioning, and output the position information and feature description of the dim space target.
[0012] Furthermore, for the visible light imaging collected in Step 2, Gaussian filtering is used to remove noise while retaining large-scale feature information. For the infrared imaging and laser imaging collected in Step 2, wavelet transform denoising method is used. By decomposing the signal into frequency components of different scales, then performing threshold processing on the high-frequency noise components, and reconstructing the signal, histogram equalization method is used to enhance the brightness and contrast of the image, and transform the gray histogram of the input image into a uniformly distributed histogram, thereby enhancing the visual effect of the image. The steps for preprocessing are as follows:
[0013] Step 2.1-1. Let the input original image be I and the Gaussian kernel size be k, calculate the weight value of the Gaussian kernel G. The mathematical expression of the two-dimensional Gaussian function is:
[0014] ;
[0015] where x and y represent spatial coordinates; σ is the standard deviation, which determines the distribution range of the Gaussian function, and G(x,y) is the weight value of the Gaussian kernel. The Gaussian filter generates a filter kernel using the Gaussian function. According to the above formula, a Gaussian filter kernel is generated, where k represents the radius size of the kernel, that is, k = 3σ;
[0016] Step 2.1-2. Perform edge padding at the boundary of the input image I(x,y), and then perform a convolution operation on each pixel point with the Gaussian kernel to generate the filtered image , and calculate the output value of each pixel point. The formula is as follows:
[0017] ;
[0018] where I(x-i, y-i) is the pixel value of the original image at the position (x-i, y-i); G(x,y) is the weight at the position (x,y) in the Gaussian kernel;
[0019] Step 2.1-3. In order to keep the image brightness unchanged, the Gaussian kernel needs to be normalized so that the sum of all weights is equal to 1. The normalization formula is as follows:
[0020] ;
[0021] Step2.2-1. Use the wavelet transform denoising method to denoise the collected infrared imaging and laser imaging. Specifically, decompose the image into different scales and frequency ranges. The noise is mainly distributed in the high-frequency part. By performing threshold processing on the high-frequency noise coefficients, retain the important feature information of the image and remove the noise. Perform L-layer wavelet decomposition on the image f(x,y), and the result is: ;
[0022] The wavelet decomposition uses a low-pass filter h(n) and a high-pass filter g(n), and completes the decomposition through convolution operations:
[0023] ;
[0024] where A L is the low-frequency coefficient of the Lth layer, D j is the high-frequency coefficient of the jth layer, A j+1 and D j+1 represent the low-frequency and high-frequency components of the (j + 1)th layer respectively, and 2k represents downsampling.
[0025] Step2.2-2. Use the soft threshold method to denoise the high-frequency component D j . Specifically, by smoothing the high-frequency coefficients, effectively remove the noise while retaining the main features of the image. The main operation is to subtract the threshold T from the absolute value of the high-frequency coefficients, and then restore the direction of the coefficients according to the sign. The mathematical expression used is:
[0026] ;
[0027] Step2.2-3. The selection of the threshold T has a great influence on the denoising effect. Therefore, determine the optimal threshold through the global threshold. The formula is:
[0028] ;
[0029] where N is the number of pixels in the image, is the standard deviation of the noise, which can be estimated by the median of the high-frequency coefficients. The expression is .
[0030] Step2.2-4. Perform wavelet inverse transform on the processed high-frequency component in Step2.2-2 and the low-frequency component A L in Step2.2-1 to reconstruct the denoised image. The wavelet inverse transform formula is as follows:
[0031] ;
[0032] The inverse transform process is opposite to the decomposition. Combine the low-pass filter h(n) and the high-pass filter g(n) for upsampling and convolution operations.
[0033] Step2.3-1. Use the histogram equalization method to enhance the brightness and contrast of the image. Specifically, by redistributing the gray values of the image pixels, the brightness and contrast of the image are improved. The basic idea is to transform the gray histogram of the input image into a uniformly distributed histogram, thereby enhancing the visual effect of the image. Count the number of pixels at each gray level in the input image Then calculate the gray level probability density , and the calculation expression is as follows:
[0034] ;
[0035] where represents the number of pixels with a gray value of , represents the total number of pixels in the image.
[0036] Step2.3-2. For each gray level , calculate its cumulative distribution function represents the cumulative probability from gray level 0 to k, and the expression is as follows:
[0037] ;
[0038] Step2.3-3. According to the cumulative distribution function, map the input gray value to the new output gray , and replace the gray value of all pixels in the input image with the corresponding output gray value to generate the histogram equalized image g(x,y), and the expression is as follows:
[0039] ;
[0040] where L-1 is the range of gray values of the output image, and round() rounds the result to the nearest integer.
[0041] Furthermore, the specific steps of Step3 are as follows:
[0042] The adaptive weight fusion algorithm optimizes the fusion effect by dynamically calculating the weights of each input data (or feature). Its main purpose is to assign weights according to the importance of the data (such as energy, confidence, or attention, etc.) and achieve efficient information fusion. First, define the contribution weight of the i-th spectral data feature map , weight Calculations are performed for each position (h, w), and adaptive calculations are carried out according to the local contribution degree based on the attention mechanism, which is expressed by the following formula: ;
[0043] where β is the temperature parameter, used to adjust the smoothness of the weight distribution;
[0044] Secondly, according to the adaptive weight all feature maps are weighted and summed to generate a fused feature map. The fusion formula is:
[0045] ;
[0046] where, is the input multi-spectral feature map, is the locally adaptive weight calculated dynamically;
[0047] Then, convolutional layers with different receptive fields are used to fuse features at different levels, strengthening the attention ability of the deep network to dim targets. While obtaining the semantic feature expression of dim targets, the accuracy of target localization is also taken into account, improving the detection effect. Through multiple convolutions and downsamplings, feature maps with different resolutions are generated, and the expression is as follows:
[0048] ;
[0049] where is the feature map of the l-th layer, conv represents the convolution operation, k is the convolution kernel size, s is the stride, and L is the number of multi-scale layers.
[0050] Furthermore, the steps of detecting and localizing dim targets in the fused feature map by using the deep learning object detection model combined with the attention mechanism in Step4 are as follows:
[0051] First, the feature map is enhanced by using the attention mechanism. Specifically, the core idea of the attention mechanism is to enhance the performance of important regions or features in the image by focusing on them, while suppressing unimportant parts. Global pooling is performed on each channel to obtain the global description vector of the channel. The formula for global average pooling is:
[0052] ;
[0053] Then, channel weights are generated through two fully connected networks and activation functions. The weight calculation formula is as follows:
[0054] ;
[0055] where and are weight matrices, is the ReLU activation function, is the Sigmoid function.
[0056] Finally, the channel weights act on the feature map, and the formula is:
[0057] ;
[0058] Furthermore, the specific steps of the said Step5 are:
[0059] Maximize the similarity between positive samples and minimize the similarity between negative samples through the contrastive loss function. The formula is as follows:
[0060] ;
[0061] where, T is the temperature parameter, M is the total number of samples in the batch, z1 and z2 are the feature representations of two views, sim(z1, z2) is the cosine similarity between the feature representations, and z j is the feature representation of other samples.
[0062] Furthermore, the steps of background noise filtering and target precise positioning for the detection results in the said Step6 are:
[0063] Use the NMS algorithm to filter redundant detection boxes and retain the optimal boxes. Specifically, NMS is mainly used to filter out those low-confidence boxes that overlap with other boxes, thereby improving the accuracy of target detection; first, sort the candidate boxes in descending order of confidence scores, and sequentially select the detection box Bmax with the highest confidence, calculate the IoU between it and the remaining boxes B j and remove the detection boxes whose IoU with Bmax is greater than the threshold until all boxes are processed. The calculation formula is as follows:
[0064] ;
[0065] Further optimize the positioning accuracy of the target through the bounding box regression branch of the target detection model, and the optimization target is the minimum regression error.
[0066] Beneficial effects: Through multi-spectral feature fusion and deep learning detection technology, the problem of insufficient single optical imaging information is effectively solved, and the detection accuracy and robustness of dim targets in complex space environments are improved; it can improve the detection ability of dim space targets, make up for the limitations of a single data source, obtain more comprehensive and detailed target characteristic information, and enhance the ability of intelligent detection and intelligence analysis of space targets. Description of the Drawings
[0067] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0068] Figure 1 is the flow schematic diagram of the present invention;
[0069] Figure 2 is the principle diagram of denoising using wavelet transform in the present invention;
[0070] Figure 3 is the schematic diagram of the adaptive feature processing module for fusing the features of the preprocessed multi-spectral data using the adaptive weight fusion algorithm in the present invention to generate the fused multi-scale feature map;
[0071] Figure 4 is the architecture diagram of training the detection model using the self-supervised learning method in the present invention. Specific Embodiments
[0072] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0073] As Figure 1 shown, a method for enhancing the detection of dim space targets based on optical feature fusion includes the following steps:
[0074] Step1. Collect multi-spectral data in a dim space scene, including optical information such as mid-wave and long-wave infrared imaging, visible light imaging, and laser imaging;
[0075] Step2. Preprocess the multi-spectral data collected in Step1, including noise removal, image enhancement, and feature extraction;
[0076] Step3. Use the adaptive weight fusion algorithm to perform feature fusion on the preprocessed multi-spectral data in Step2 to generate a fused multi-scale feature map;
[0077] Step4. Use the deep learning object detection model combined with the attention mechanism to detect and locate the dim targets in the fused feature map;
[0078] Step 5: Introduce self-supervised learning methods to train the deep learning object detection model, and use unlabeled data to enhance the generalization ability of the model;
[0079] Step 6: Post-process the detection results of Step 4, including background noise filtering and target precise positioning, and output the position information and feature description of the dim space target.
[0080] As a preferred implementation, for the visible light imaging collected in Step 2, Gaussian filtering is used to remove noise while retaining large-scale feature information. For the infrared imaging and laser imaging collected in Step 2, wavelet transform denoising method is used. By decomposing the signal into frequency components of different scales, then performing threshold processing on the high-frequency noise components, and reconstructing the signal. The principle diagram of wavelet transform denoising is as Figure 2 shown. Use histogram equalization method to enhance the brightness and contrast of the image, transform the gray histogram of the input image into a uniformly distributed histogram, so as to enhance the visual effect of the image for preprocessing;
[0081] Step 2.1-1: Let the input original image be I and the Gaussian kernel size be k, calculate the weight value of the Gaussian kernel G. The mathematical expression of the two-dimensional Gaussian function is:
[0082] ;
[0083] where x, y represent spatial coordinates; σ is the standard deviation, which determines the distribution range of the Gaussian function, and G(x,y) is the weight value of the Gaussian kernel. The Gaussian filter uses the Gaussian function to generate a filter kernel. According to the above formula, a Gaussian filter kernel is generated, where k represents the radius size of the kernel, that is, k = 3σ;
[0084] Step 2.1-2: Perform edge padding at the boundary of the input image I(x,y), and then perform a convolution operation on each pixel point with the Gaussian kernel to generate the filtered image , calculate the output value of each pixel point , and the formula is as follows:
[0085] ;
[0086] where I(x-i, y-i) is the pixel value of the original image at the position (x-i, y-i); G(x,y) is the weight at the position (x,y) in the Gaussian kernel;
[0087] Step 2.1-3: In order to keep the image brightness unchanged, the Gaussian kernel needs to be normalized so that the sum of all weights is equal to 1. The normalization formula is as follows:
[0088] ;
[0089] Step2.2-1. Apply the wavelet transform denoising method to the collected infrared imaging and laser imaging for denoising. Specifically, decompose the image into different scales and frequency ranges. The noise is mainly distributed in the high-frequency part. By performing threshold processing on the high-frequency noise coefficients, retain the important feature information of the image and remove the noise. Perform L-layer wavelet decomposition on the image f(x,y), and the result is: ;
[0090] The wavelet decomposition uses a low-pass filter h(n) and a high-pass filter g(n), and completes the decomposition through convolution operations:
[0091] ;
[0092] Among them, A L is the low-frequency coefficient of the Lth layer, D j is the high-frequency coefficient of the jth layer, A j+1 and D j+1 respectively represent the low-frequency and high-frequency components of the (j + 1)th layer, and 2k represents downsampling.
[0093] Step2.2-2. Use the soft threshold method to denoise the high-frequency component D j . Specifically, by smoothing the high-frequency coefficients, effectively remove the noise while retaining the main features of the image. The main operation is to subtract the threshold T from the absolute value of the high-frequency coefficients, and then restore the direction of the coefficients according to the sign. The mathematical expression used is:
[0094] ;
[0095] Step2.2-3. The selection of the threshold T has a greater impact on the denoising effect. Therefore, determine the optimal threshold through the global threshold. The formula is:
[0096] ;
[0097] Among them, N is the number of pixels in the image, is the standard deviation of the noise, which can be estimated by the median of the high-frequency coefficients. The expression is ;
[0098] Step2.2-4. Perform wavelet inverse transform on the processed high-frequency component in Step2.2-2 and the low-frequency component A L in Step2.2-1 to reconstruct the denoised image. The wavelet inverse transform formula is as follows:
[0099] ;
[0100] The inverse transformation process is the opposite of the decomposition, and upsampling and convolution operations are performed by combining the low-pass filter h(n) and the high-pass filter g(n);
[0101] Step2.3-1. Use the histogram equalization method to enhance the brightness and contrast of the image. Specifically, by redistributing the gray values of the image pixels, the brightness and contrast of the image are improved; its basic idea is to transform the gray histogram of the input image into a uniformly distributed histogram, thereby enhancing the visual effect of the image, and counting the number of pixels at each gray level in the input image Then calculate the gray level probability density , and the calculation expression is as follows:
[0102] ;
[0103] where represents the number of pixels with a gray value of , represents the total number of pixels in the image;
[0104] Step2.3-2. For each gray level , calculate its cumulative distribution function represents the cumulative probability from gray level 0 to k, and the expression is as follows:
[0105] ;
[0106] Step2.3-3. According to the cumulative distribution function, map the input gray value to the new output gray , and replace the gray values of all pixels in the input image with the corresponding output gray value to generate the histogram-equalized image g(x, y), and the expression is as follows:
[0107] ;
[0108] where L-1 is the range of gray values of the output image, and round() rounds the result to the nearest integer.
[0109] As a preferred implementation, the adaptive weight fusion algorithm optimizes the fusion effect by dynamically calculating the weights of each input data. Its main purpose is to allocate weights according to the importance of the data and achieve efficient information fusion, as Figure 3 shown. The specific steps of Step3 are as follows:
[0110] First, define the contribution weight of the i-th spectral data feature map , weight Calculate for each position (h, w), and perform adaptive calculation according to the local contribution degree based on the attention mechanism, which is expressed by the following formula: ;
[0111] where β is the temperature parameter, used to adjust the smoothness of the weight distribution;
[0112] Secondly, according to the adaptive weight perform weighted summation on all feature maps to generate a fused feature map. The fusion formula is:
[0113] ;
[0114] where, is the input multi-spectral feature map, is the locally adaptive weight calculated dynamically;
[0115] Then use convolutional layers with different receptive fields to fuse features at different levels, strengthen the attention ability of the deep network to dim targets, and take into account the accuracy of target localization while obtaining the semantic feature expression of dim targets, improving the detection effect. Through multiple convolutions and downsamplings, generate feature maps with different resolutions, and the expression is as follows:
[0116] ;
[0117] where is the feature map of the l-th layer, conv represents the convolution operation, k is the convolution kernel size, s is the stride, and L is the number of multi-scale layers.
[0118] As a preferred implementation manner, the specific steps of Step4 are:
[0119] First, use the attention mechanism to enhance the feature map. Specifically, the core idea of the attention mechanism is to enhance the performance of important regions or features in the image by focusing on them, while suppressing unimportant parts. Perform global pooling on each channel to obtain the global description vector of the channel. The formula for global average pooling is:
[0120] ;
[0121] Then generate channel weights through two fully connected networks and activation functions. The weight calculation formula is as follows:
[0122] ;
[0123] where and are weight matrices, is the ReLU activation function, is the Sigmoid function;
[0124] Finally, the channel weights act on the feature map, and the formula is:
[0125] ;
[0126] As a preferred implementation manner, the specific steps of the Step5 are:
[0127] Maximize the similarity between positive samples and minimize the similarity between negative samples through the contrast loss function. The formula is as follows:
[0128] ;
[0129] where T is the temperature parameter, M is the total number of samples in the batch, z1 and z2 are the feature representations of two views, sim(z1, z2) is the cosine similarity between the feature representations, and z j is the feature representation of other samples.
[0130] As a preferred implementation manner, the specific steps of the Step6 are:
[0131] Use the NMS algorithm to filter redundant detection boxes and retain the optimal boxes. Specifically, NMS is mainly used to filter out those low-confidence boxes that overlap with other boxes, thereby improving the accuracy of object detection. First, sort the candidate boxes in descending order of confidence score, and sequentially select the detection box Bmax with the highest confidence. Calculate the IoU between the remaining boxes B j and remove the detection boxes whose IoU with Bmax is greater than the threshold until all boxes are processed. The calculation formula is as follows:
[0132] ;
[0133] Through the bounding box regression branch of the object detection model, further optimize the localization accuracy of the object, and the optimization goal is the minimum regression error.
[0134] In summary, the enhanced detection method for dim space targets based on optical feature fusion can improve the detection ability of dim space targets, make up for the limitations of a single data source, obtain more comprehensive and detailed target characteristic information, and enhance the ability of intelligent detection and intelligence analysis of space targets.
[0135] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.
Claims
1. An enhanced detection method for dim space targets based on optical feature fusion, characterized in that, It includes the following steps: Step 1: Collect multi-spectral data in a dim space scene, including mid-wave and long-wave infrared imaging, visible light imaging, and laser imaging optical information; Step 2: Preprocess the multi-spectral data collected in Step 1, including noise removal, image enhancement, and feature extraction; The specific steps are as follows: Step 2.1-1: Assume the input original image is I and the Gaussian kernel size is k, calculate the weight value of the Gaussian kernel G, and the mathematical expression of the two-dimensional Gaussian function is: ; where x and y represent spatial coordinates; σ is the standard deviation, and G(x, y) is the weight value of the Gaussian kernel; the Gaussian filter generates a filter kernel using the Gaussian function. According to the above formula, a Gaussian filter kernel of is generated, where k represents the radius size of the kernel, i.e., k = 3σ; Step 2.1-2. Perform edge padding at the boundary of the input image I(x, y), and then perform a convolution operation on each pixel with a Gaussian kernel to generate a filtered image , and calculate the output value of each pixel . The formula is as follows: ; where I(x - i, y - i) is the pixel value of the original image at position (x - i, y - j); G(x, y) is the weight at position (x, y) in the Gaussian kernel; Step 2.1-3: The Gaussian kernel needs to be normalized so that the sum of all weights is equal to 1. The normalization formula is as follows: ; Step 2.2-1: Use the wavelet transform denoising method to denoise the collected infrared imaging and laser imaging. Decompose the image f(x, y) into L layers of wavelets, and the result is: ; Wavelet decomposition uses a low-pass filter h(n) and a high-pass filter g(n) to complete the decomposition through convolution operations: ; Among which A L is the low-frequency coefficient of the L-th layer, D j is the high-frequency coefficient of the j-th layer, A j+1 and D j+1 respectively represent the low-frequency and high-frequency components of the (j + 1)-th layer, and 2k represents downsampling; Step 2.2-2: Denoise the high-frequency component Dj by the soft threshold method. Subtract the threshold T from the absolute value of the high-frequency coefficient, and then restore the direction of the coefficient according to the sign. The mathematical expression used is: ; Step 2.2-3: Determine the optimal threshold T through the global threshold. The formula is: ; where N is the number of pixels in the image, is the standard deviation of the noise, which can be estimated by the median of the high-frequency coefficients, and the expression is: ; Step2.2-4. Perform wavelet inverse transform on the high-frequency components processed in Step2.2-2 and the low-frequency component A in Step2.2-1 L to reconstruct the denoised image. The wavelet inverse transform formula is as follows: ; The inverse transformation process is the opposite of the decomposition. Combine the low-pass filter h(n) and the high-pass filter g(n) for upsampling and convolution operations; Step 2.3-1. Use the histogram equalization method to enhance the brightness and contrast of the image, and count the number of pixels at each gray level in the input image Then calculate the gray level probability density , and the calculation expression is as follows: ; Among them represents the number of pixels with a grayscale value of , represents the total number of pixels in the image; Step2.3-2. For each gray level , calculate its cumulative distribution function represents the cumulative probability from gray level 0 to k, and the expression is as follows: ; Step 2.3-3. Map the input grayscale value according to the cumulative distribution function to the new output grayscale . Replace the grayscale values of all pixels in the input image with the corresponding output grayscale values to generate the histogram-equalized image g(x, y). The expression is as follows: ; Step 3: Use the adaptive weight fusion algorithm to perform feature fusion on the preprocessed multi-spectral data in Step 2 to generate a fused multi-scale feature map; Step 4: Use a deep learning object detection model combined with an attention mechanism to detect and locate dim targets in the fused feature map; Step 5: Introduce a self-supervised learning method to train the deep learning object detection model and use unlabeled data to enhance the generalization ability of the model; Step 6: Post-process the detection results in Step 4, including background noise filtering and target precise positioning, and output the position information and feature description of the dim space target.
2. The enhanced detection method for dim space targets based on optical feature fusion according to claim 1, characterized in that The specific steps of Step 3 are as follows: First, define the contribution weight of the i-th spectral data feature map . The weight is adaptively calculated based on the local contribution degree of the attention mechanism and is expressed by the following formula: ; where β is the temperature parameter, used to adjust the smoothness of the weight distribution; Then, according to the adaptive weights for all feature maps perform weighted summation to generate the fused feature map. The fusion formula is as follows: ; Among them, is the input multi-spectral feature map, is the locally adaptive weight calculated dynamically.
3. The enhanced detection method for dim space targets based on optical feature fusion according to claim 1, wherein The specific steps of Step 4 are as follows: First, use the attention mechanism to enhance the feature map. Perform global pooling on each channel to obtain the global description vector of the channel. The formula for global average pooling is: ; Then generate channel weights through a two-layer fully connected network and an activation function. The weight calculation formula is as follows: ; Among them and are weight matrices, is the ReLU activation function, is the Sigmoid function; Finally, the channel weights are applied to the feature map, and the formula is: 。 4. The enhanced detection method for dim space targets based on optical feature fusion according to claim 1, characterized in that The specific steps of Step 5 are as follows: Maximize the similarity between positive samples and minimize the similarity between negative samples through the contrast loss function. The formula is as follows: ; where T is the temperature parameter, M is the total number of samples in the batch, z1 and z2 are the feature representations of two views, sim(z1, z2) is the cosine similarity between the feature representations, and z j is the feature representation of other samples.
5. The enhanced detection method for dim space targets based on optical feature fusion according to claim 1, wherein The specific steps of Step 6 are as follows: Use the NMS algorithm to filter redundant detection boxes and retain the optimal boxes. Sort the candidate boxes in descending order according to the confidence score, and sequentially select the detection box B with the highest confidence max , calculate the IoU between it and the remaining boxes Bj, and remove the detection boxes with IoU greater than the threshold with B max . The calculation formula is as follows until all boxes are processed: ; Through the bounding box regression branch of the object detection model, further optimize the positioning accuracy of the target, and the optimization target is the minimum regression error.
Citation Information
Patent Citations
Weak-luminescence and fluorescence optical imaging device and imaging method thereof
CN101634748A
Joint topographic mapping method based on space-borne SAR image and optical image
CN109100719A