Infrared long-distance spatial adjacent target super-resolution method

The infrared long-distance spatial adjacent target super-resolution method is constructed through the method of dynamic iterative contraction threshold, which solves the detection problem of closely adjacent targets in infrared imaging, and achieves adaptive efficient detection and robustness improvement.

CN120543378AActive Publication Date: 2025-08-26NANJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510644028.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-08-26
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

In the prior art, it is difficult to effectively demix and super-resolution identify closely adjacent small infrared targets in infrared imaging, especially in long-distance detection and monitoring tasks. Traditional methods are sensitive to hyperparameter adjustment, and deep learning methods are difficult to adapt to target changes after feature extraction and parameter fixation, resulting in poor detection results.

Method used

The infrared long-distance spatial adjacent target super-resolution method based on dynamic iterative contraction threshold is adopted. By analyzing the target imaging characteristics, a deep expansion network framework is built, combined with sparse reconstruction algorithm and dynamic near-end mapping module, the network parameters are dynamically adjusted to adapt to the input data, and the joint estimation of target quantity, position and radiation intensity is achieved.

Benefits of technology

It improves the detection accuracy and robustness of adjacent targets in infrared space, can adaptively adjust network parameters, improves detection effect, and is suitable for various practical scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120543378A_ABST
    Figure CN120543378A_ABST
Patent Text Reader

Abstract

The invention discloses an infrared long-distance spatial adjacent target super-resolution method, and the method comprises the steps: S1, analyzing the imaging characteristics of an optical detection system of a long-distance target and the imaging characteristics of multiple spatial adjacent targets, carrying out the imaging modeling of a point target, and generating a simulation image data set; s2, initializing an input simulation image based on linear mapping; s3, adopting a sparse reconstruction algorithm to realize infrared space adjacent target super-resolution modeling, and constructing a deep expansion network framework; s4, enabling the initialized image to sequentially pass through each stage of the deep expansion network; and S5, constructing a GPU training environment, setting a configuration file training model of a data loader, a model and an evaluator, applying the trained model to a test set, generating a high-resolution image, extracting target coordinate information through post-processing, inputting the target coordinate information into the evaluator, obtaining an evaluation index based on an infrared long-distance space adjacent target, and calculating an average detection rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of infrared imaging technology, in particular to an infrared long-distance spatial proximity target super-resolution method. Background Art

[0002] Infrared imaging plays a key role in various long-range detection and surveillance missions due to its extreme sensitivity to thermal radiation and its indifference to lighting conditions. However, the radiation intensity captured from distant targets is actually weak due to their large distance from the imaging system. This challenge is exacerbated when targets appear in dense, spatially close clusters, such as those in close proximity, resulting in overlapping spherical light spots. This phenomenon obscures the perception of the number, precise location, and radiation intensity of individual targets, significantly hindering the subsequent detection, tracking, and identification stages of infrared search and tracking (IRST) systems. Therefore, it is of great significance to explore effective techniques for demixing and super-resolution reconstruction of such closely-spaced infrared small targets (CSIST) to accurately discern their exact location and radiation intensity.

[0003] Although super-resolution of small infrared targets at close range is crucial in various applications, research on this specific task is still very scarce. In existing technologies, the main detection techniques include traditional optimization methods, infrared small target detection methods, and super-resolution methods based on depth expansion frameworks. These methods still have some shortcomings in practical applications:

[0004] Traditional optimization methods analyze the optical properties of nearby targets in infrared space and the imaging mechanism, formulate the problem as a parameter estimation task, and employ optimization algorithms to solve it. Leveraging the sparsity of targets on the imaging plane, a sparse reconstruction method based on discrete sampling was designed. This method uses an overcomplete dictionary and solves a second-order cone programming problem under l1-norm regularization. However, the performance of these optimization-based models is highly dependent on careful hyperparameter tuning, which poses significant challenges in real-world scenarios.

[0005] Traditional infrared small target detection techniques are primarily categorized into sparse plus low-rank decomposition and deep learning methods. Sparse low-rank decomposition methods aim to separate targets from background by decomposing the original image into a sparse target representation and a low-rank background representation. However, real-world distractors, which appear as outliers in the background, pose a significant challenge. Traditional methods rely solely on grayscale intensity or basic hand-crafted features, making it difficult to semantically distinguish true targets from distractors. Furthermore, these methods are highly sensitive to hyperparameter configuration and require extensive manual tuning, hindering their practical application.

[0006] Deep learning research has primarily focused on developing multi-scale feature fusion models to offset the scarcity of intrinsic target features. These methods primarily focus on detecting overlapping objects as a whole, but due to a lack of datasets, subsequent tasks such as sub-pixel localization and radiation intensity estimation are difficult to implement. Super-resolution methods based on deep unwrapping frameworks, such as ISTA-Net (based on a soft-thresholding iterative network) and ADMM-Net (based on the alternating direction multiplier method), reinterpret traditional optimization algorithms as fully connected feedforward neural networks. This approach effectively generalizes to new samples and, through end-to-end training, achieves high image reconstruction performance with fewer iterations and faster inference speed. However, current methods still exhibit low performance, and even slight changes in the number or position of objects can lead to significant artifacts in the image. This is primarily due to the limited feature extraction capabilities of current networks and their inability to effectively utilize multi-scale feature information. Furthermore, once the network is trained, key parameters such as the soft threshold and step size are fixed, making it difficult to adjust them to the specific characteristics of the target. Summary of the Invention

[0007] The present invention provides an infrared long-distance spatial neighboring target super-resolution method based on dynamic iterative shrinkage threshold, which has stronger robustness to hyperparameter changes and can be effectively applied to various practical scenarios, and can solve at least one of the above technical problems.

[0008] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0009] A method for super-resolution of infrared long-distance spatially adjacent targets comprises the following steps:

[0010] S1. Analyze the imaging characteristics of the optical detection system for long-range targets and the imaging characteristics of spatially adjacent multi-targets, model point target imaging, and generate a simulation image dataset;

[0011] S2. Initializing the input simulation image based on linear mapping to form an initialized image;

[0012] S3, using sparse reconstruction algorithm to achieve super-resolution modeling of adjacent targets in infrared space and build a deep unfolding network framework;

[0013] S4, passing the image initialized by S2 through each stage of the deep expansion network in S3 in sequence, each stage is a gradient descent module and a dynamic proximal mapping module in sequence;

[0014] S5. Build a GPU training environment, set up the data loader, infrared long-distance spatial proximity super-resolution model, and evaluator configuration files to train the model. Apply the trained model to the test set to generate high-resolution images. Extract target coordinate information through post-processing and input it into the evaluator to obtain evaluation indicators based on infrared long-distance spatial proximity targets and calculate the average detection rate.

[0015] Furthermore, the S1 further includes:

[0016] S1.1. Analysis of target imaging characteristics: Consider the distant target as a point source target. The point source target is imaged on the detector using a Gaussian function as the point spread function (PSF). Multi-target imaging satisfies the linear superposition of Gaussian energies. The point spread function expression is:

[0017]

[0018] Among them, σ PSF represents the diffusion variance, (x t ,y t ) is the coordinate of the target on the focal plane;

[0019] S1.2. Simulation target: Assume that the number of adjacent targets in infrared long-distance space is a natural number between 1 and 5, set the distance between targets to be less than one pixel, and the centers of multiple targets are concentrated in a single pixel on the detector image plane, thus forming an aliased cluster and indistinguishable target features. The target features include at least the number of targets, their positions, and their radiation intensity.

[0020] Furthermore, the S2 further includes:

[0021] S2.1. Given a dataset Among them, z i Represents the image of the infrared space adjacent target imaging, s i It represents the image after super-resolution c times. According to the point target information (x i ,y i ,g i ) and c calculate the high-resolution image The grayscale of the pixel at is the peak intensity g i , which generates s i ;

[0022] S2.2, let Z = [z1,…,z M ],S=[s1,…,s M ],initialization The matrix Q init = It is expressed as:

[0023]

[0024] Where T represents the matrix transpose and F represents the Frobenius norm.

[0025] Furthermore, the S3 further includes:

[0026] S3.1. Model and analyze infrared spatial neighboring targets in a pure Gaussian framework. The image area contains N spatial neighboring targets, and the target is represented by t i =(x i ,y i ,g i ,σ PSF ), where i∈1,..,N, (x i ,y i ), represents the spatial sub-pixel coordinate, g i is the peak intensity, σ PSF represents the diffusion variance, the center point set and peak intensity of all targets are represented by and express;

[0027] S3.2. Predicting target sets using the trained infrared long-range spatial proximity super-resolution model The corresponding peak intensity set Among them, M is the number of predicted targets, Forecast point The peak intensity is obtained by minimizing the predicted intensity With the true intensity g i The difference between them is used to achieve joint estimation of target quantity, sub-pixel position and radiation intensity;

[0028] S3.3. Establish an infrared long-distance spatial proximity target imaging model for the infrared focal plane. Define the target response at pixel (i, j) as the integral value of the point spread function (PSF) within the pixel boundary. The specific expression is:

[0029]

[0030] Among them, (x i,j ,y i,j ) represents the pixel center position and indicates the uniform width of pixel D;

[0031] Assume that an infrared focal plane consists of U×V pixels and there are K targets with target coordinates (x k ,y k ), k=1,…,K, by rearranging the columns of the focal plane matrix into a vector of UV×1, the focal plane measurement model is constructed as:

[0032] z=[g c (x1,y1)gc (x2,y2)…g c (x K ,y K )]s+n=G(x,y)s+n

[0033] Among them, G(x,y) is a UV×K steering matrix, which represents the influence on the response of each pixel. c (x k ,y k ) represents the contribution of target k to the pixel response, and the peak intensity of the target is represented by the vector s=[s1,s2,…,s K ] T Indicates that n represents the variance Gaussian white noise, n is independent between different pixels;

[0034] S3.4, the sparse reconstruction algorithm is used to model the infrared space neighboring target image, and the sub-pixel center is used as the possible location of the target. By exhausting all possible sub-pixel location sets Ω = {(x l ,y l )} l=1,…,L , construct a spatial neighboring target location set, where L = UVn 2 , n is the size of the sub-pixel grid, each sub-pixel contains at most one target, and the maximum deviation of the target position from the nearest sub-pixel center does not exceed Dividing each pixel into an n×n sub-pixel grid, we can extend the overcomplete representation of the focal plane measurement model:

[0035]

[0036] Where G(Ω) is a matrix with the steering vectors in the position set Ω as columns, represents the sparse signal vector, w is the noise;

[0037] Through sparse signal reconstruction, the target signal strength is optimized using l1 norm regularization to best fit the observation value z. The formula is:

[0038]

[0039] Among them, λ is the preset regularization parameter;

[0040] S3.5, the reconstruction method based on the iterative soft threshold algorithm ISTA, alternately performs the following two update steps to solve the sparsity problem:

[0041] Gradient Descent:

[0042] Near-end mapping:

[0043] The above two steps are respectively designed as a gradient descent module and a dynamic proximal mapping module, and a deep unfolding network framework is formed by alternately connecting the gradient descent module and the dynamic proximal mapping module multiple times.

[0044] Furthermore, the proximal mapping module adopts a trainable nonlinear transformation module Dynamic soft threshold module soft(·,θ) and inverse transform module Simulate the iterative process of the iterative soft threshold algorithm ISTA.

[0045] Furthermore, in said S4, said dynamic proximal mapping module is composed of a dynamic transformation module Dynamic soft threshold module soft(·,θ) and inverse transform module It consists of three parts. In the process of constructing the dynamic iterative shrinkage threshold network DISTA-Net, each iteration is expressed as:

[0046]

[0047] Furthermore, the dynamic transformation module It consists of a two-branch network, where the main branch is a Cov-ReLU-Conv structure, which is used to extract input image features, and the secondary branch is used to fine-tune the features extracted by the main branch. The dynamic transformation structure process is:

[0048] S4.a.1. The output of the k-1th stage is scaled by a multi-layer perceptron to obtain the weight vector W:

[0049]

[0050] Where f(·) represents a fully connected layer;

[0051] S4.a.2. Use the features as weights of the convolution kernel r (k) Perform convolution operation:

[0052] wr (k) =C(W,r (k) )

[0053] Where C(·) represents the convolutional layer;

[0054] S4.a.3, the nonlinearity of the secondary branch is ensured by the sigmoid function, and it is added to the main branch according to different weights to obtain

[0055]

[0056] Among them, A and B represent two convolution operations respectively, and α is a predefined hyperparameter used to adjust the weights of the two branches.

[0057] Furthermore, the dynamic soft threshold module soft(·,θ) introduces a selectivity mechanism, and the implementation process is as follows:

[0058] S4.b.1. Extract features through two convolutions to obtain different receptive field information and Channel-based average pooling and maximum pooling extract spatial relationships:

[0059]

[0060] S4.b.2. Use convolution to jointly extract features to better interact with the information after average pooling and maximum pooling:

[0061]

[0062] S4.b.3. Use the sigmoid function to obtain the mask for each spatial selection:

[0063]

[0064] By weighting different receptive field feature maps with corresponding spatial selection masks, a dynamic threshold is obtained:

[0065]

[0066] S4.b.4, set the threshold value of the threshold function to θ, and set the dynamic transformation module The output features are input into the dynamic soft threshold network to make the features sparse.

[0067] Furthermore, the inverse transformation module It is used to remap the sparse features output by the dynamic soft threshold module soft(·,θ) back to the single-channel image space. It consists of a simple Cov-ReLU-Conv structure and the implementation process is as follows:

[0068] S4.c.1, each convolution size is 3×3, and the loss function L constraint In the inverse transform module To constrain and ensure Among them, I is the identity matrix, and the loss function L constraint for:

[0069]

[0070] Among them, M is the size of the training set, N is the number of iterative network modules, and S is the size of N s ;

[0071] S4.c.2. Loss function Ldiscrepancy Constrained super-resolution image Similar to s, the loss function L discrepancy for:

[0072]

[0073] S4.c.3, the total loss function Loss is:

[0074] Loss = L discrepancy +γL constraint

[0075] Among them, γ is the parameter for balancing loss weight.

[0076] Furthermore, the S5 further includes:

[0077] S5.1. Evaluator Setup: Based on the calculation method of the average precision of the target detection algorithm evaluation metric, a specific evaluation metric is designed for coordinate detection of point target non-bounding box detection, namely the average precision of spatially adjacent targets, to effectively determine whether the predicted point target is a true positive TP or a false positive FP.

[0078] S5.2. When a prediction point target is classified as TP or FP, a binary list is constructed, where TP predictions are represented by 1 and FP predictions are represented by 0. Based on this binary list, a range of precision and recall values ​​are obtained by varying the confidence threshold for positive predictions. A precision-recall curve (PR) is generated. The area under the PR curve is calculated to obtain the average precision, thereby comprehensively evaluating the model performance.

[0079] S5.3. Other configurations: Build the GPU training environment, set the network training epochs, batch_size, load the image PNG file with the data loader, convert the XML file into a super-resolution image, count the number of targets, and configure the optimizer.

[0080] S5.4. The model validation evaluator sets an intensity threshold. The super-resolution image predicted by the model uses this threshold to suppress the data below this value to 0. The remaining elements need to be projected back to the original coordinate space and used as the predicted points and the actual points to calculate CSO_mAP. The projection rule is:

[0081] (x i ,y i ,g i ) corresponds to the predicted point Among them, (x i ,y i ) represents the i-th non-zero prediction point in the predicted image, g i represents the predicted radiation intensity;

[0082] S5.5. The prediction results are passed through the evaluator to obtain the final indicators. The hyperparameters are saved in the configuration file. The configuration file is loaded during training. The model training, verification and testing are completed under the open source deep learning MMengine framework, and the model weights are saved.

[0083] The beneficial effects of the present invention are embodied in:

[0084] This paper provides a comprehensive approach to spatial proximity object detection, bridging the entire process from dataset development to model construction and evaluation metrics. Specifically, it presents a baseline spatial proximity object super-resolution model that adaptively generates convolution weights and threshold parameters to adjust the reconstruction process in real time. Unlike previous methods, the parameters associated with proximity mapping, including nonlinear transformations, shrinkage thresholds, and step sizes, are dynamically adjusted based on the input data, rather than being manually designed or fixed after training. This approach effectively improves infrared spatial proximity object detection.

[0085] This paper reconstructs the traditional sparse reconstruction problem based on a dynamic network framework, integrates the deep expansion method and attention mechanism, effectively improves the network feature extraction capability, and thus improves the accuracy and robustness of detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0086] The drawings described herein are used to provide further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute improper limitations on the present application.

[0087] Figure 1 It is a schematic diagram of the overall flow of the method according to an embodiment of the present invention.

[0088] Figure 2 It is a schematic diagram of a specific flow chart of the method according to an embodiment of the present invention.

[0089] Figure 3 3 is a schematic diagram of the model structure of the dynamic iterative shrinkage threshold network according to an embodiment of the present invention.

[0090] Figure 4 It is a schematic diagram of the structure of the dynamic transformation module according to an embodiment of the present invention.

[0091] Figure 5 4 is a schematic structural diagram of a dynamic threshold module according to an embodiment of the present invention.

[0092] Figure 6 It is a structural block diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0093] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. In the absence of conflict, the embodiments in this application and the features in the embodiments can be combined with each other. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0094] It should be noted that the meaning of "and / or" appearing throughout the text includes three parallel solutions. Taking "A and / or B" as an example, it includes solution A, solution B, or solutions in which both A and B are satisfied. In addition, "multiple" refers to more than two. In addition, the technical solutions between the various embodiments can be combined with each other, but this must be based on the fact that ordinary technicians in this field can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0095] See also Figure 1-Figure 2 The embodiment of the present invention provides an infrared long-distance spatial proximity target super-resolution method, comprising the following steps:

[0096] S1. Analyze the imaging characteristics of the optical detection system for long-range targets and the imaging characteristics of spatially adjacent multi-targets, model point target imaging, and generate a simulation image dataset;

[0097] S2. Initializing the input simulation image based on linear mapping to form an initialized image;

[0098] S3, using sparse reconstruction algorithm to achieve super-resolution modeling of adjacent targets in infrared space and build a deep unfolding network framework;

[0099] S4, passing the image initialized by S2 through each stage of the deep expansion network in S3 in sequence, each stage is a gradient descent module and a dynamic proximal mapping module in sequence;

[0100] S5. Build a GPU training environment, set up the data loader, infrared long-distance spatial proximity super-resolution model, and evaluator configuration files to train the model. Apply the trained model to the test set to generate high-resolution images. Extract target coordinate information through post-processing and input it into the evaluator to obtain evaluation indicators based on infrared long-distance spatial proximity targets and calculate the average detection rate.

[0101] In this embodiment, the S1 further includes:

[0102] S1.1. Analyze target imaging characteristics: Consider distant targets as point source targets, analyze the imaging characteristics of the optical detection system of distant targets, model the imaging of power source targets, analyze the image characteristics of multiple spatially adjacent targets, simulate 100,000 image data, and record the position and radiation intensity of the target in each image as labels. The dataset is divided into training set, validation set, and test set for each stage of model training.

[0103] S1.2. When a monochromatic point light source is imaged through a lens, an "Airy disk" appears at the focal point. This is characterized by a central bright spot that accounts for 84% of the total energy, surrounded by multiple concentric diffraction rings. This diffraction effect is usually approximated by a point spread function (PSF), whose standard deviation depends on the focal ratio (f-number) and detection band of the sensor and is used to quantify the energy spread.

[0104] The point source target is imaged on the detector with a Gaussian function as the point spread function (PSF). Multi-target imaging satisfies the linear superposition of Gaussian energy. The point spread function expression is:

[0105]

[0106] Among them, σ PSF represents the diffusion variance, (x t ,y t ) are the coordinates of the target on the focal plane.

[0107] In multi-target imaging, the intensity captured by each pixel is the cumulative response of superimposed point sources. Especially for distant objects, each object is treated as a point source, and the resulting "Airy disk" radius is determined by 1.22λ / D, where λ is the wavelength and D is the lens diameter. This radius is equal to 1.9σ of the Gaussian PSF, marking the physical resolution limit of the sensor and conforming to the definition of Rayleigh units. In this example, the 84% energy concentration definition is used as the resolution criterion, and σ is used as the resolution criterion. PSF Set it to 0.5 pixels.

[0108] S1.3. Simulation target: Assume that the number of adjacent targets in infrared long-distance space is a natural number between 1 and 5. Set the distance between targets to be small, and the centers of multiple targets are concentrated in a single pixel on the detector image plane, thus forming an aliased cluster and indistinguishable target features. The target features include at least the number of targets, position, and radiation intensity.

[0109] The simulation experiment of this embodiment covers images containing 1 to 5 targets, each target is represented by sub-pixel coordinates and peak intensity t i =(x i ,y i ,g i ,σ PSF), where i∈1,…,N, is the labeled learning target, using an 11×11 pixel grid as each target set To test the super-resolution performance of this infrared long-range spatial proximity super-resolution model, this example simulates multiple targets within a single pixel, ensuring that the distance between them exceeds 0.52 Rayleigh units. In this configuration, a dataset containing 100,000 samples is simulated, of which 80,000 are allocated to the training set, and the remaining 20,000 samples are evenly distributed between the validation set and the test set.

[0110] In this embodiment, the S2 further includes:

[0111] S2.1, as with traditional iterative algorithms, such as ISTA and DISTA-Net, also requires initialization, i.e. Initialize and use linear mapping to calculate the initial value. Specifically, given a data set Among them, z i Represents the image of the infrared space adjacent target imaging, s i Indicates the image after super-resolution by a factor of c. For example, when the pixel division factor is c, the probability of aliasing of point targets still meets the preset requirement.

[0112] S2.2 According to the point target information (x i ,y i ,g i ) and c calculate the high-resolution image The grayscale of the pixel at is the peak intensity g i , which generates s i ;

[0113] Let Z = [z1,…,z M ],S=[s1,…,s M ],initialization The matrix Q init =argmin Q ||QZ-SF2=SZTZZT-1, s0 is expressed as:

[0114]

[0115] Where T represents the matrix transpose and F represents the Frobenius norm.

[0116] See also Figure 3 In this embodiment, S3 further includes:

[0117] S3.1. Model and analyze infrared spatial neighboring targets in a pure Gaussian framework. The image area contains N spatial neighboring targets, and the target is represented by ti =(x i ,y i ,g i ,σ PSF ), where i∈1,..,N, (x i ,y i ), represents the spatial sub-pixel coordinate, g i is the peak intensity, σ PSF represents the diffusion variance, the center point set and peak intensity of all targets are represented by and express.

[0118] S3.2. Predicting target sets using the trained infrared long-range spatial proximity super-resolution model The corresponding peak intensity set Among them, M is the number of predicted targets, Forecast point The peak intensity is obtained by minimizing the predicted intensity With the true intensity g i The difference between them is used to achieve joint estimation of target quantity, sub-pixel position and radiation intensity.

[0119] S3.3. Establish an infrared focal plane imaging model for long-distance spatially adjacent targets. In an infrared imaging system, due to the influence of the point spread function (PSF) of the optical system, the infrared radiation emitted by adjacent targets in infrared space is dispersed between pixels on the focal plane. Each pixel integrates the incident radiation of the overlapping targets and generates a response. The pixel responses to the targets are linearly superimposed, thus forming a linear focal plane imaging model. Specifically, the target response of pixel (i, j) is defined as the integral value of the point spread function PSF within the pixel boundary, expressed as:

[0120]

[0121] Among them, (x i,j ,y i,j ) represents the pixel center position and indicates the uniform width of pixel D;

[0122] Assume that an infrared focal plane consists of U×V pixels and there are K targets with target coordinates (x k ,y k ), k=1,…,K, by rearranging the columns of the focal plane matrix into a vector of UV×1, the focal plane measurement model is constructed as:

[0123] z=[g c (x1,y1)g c (x2,y2)…g c (x K ,yK )]s+n=G(x,y)s+n

[0124] #(3)

[0125] Among them, G(x,y) is a UV×K steering matrix, which represents the influence on the response of each pixel. c (x k ,y k ) represents the contribution of target k to the pixel response, and the peak intensity of the target is represented by the vector s=[s1,s2,…,s K ] T Indicates that n represents the variance Gaussian white noise, n is independent between different pixels;

[0126] Through the above modeling, this embodiment aims to more accurately capture and analyze the signals of small infrared targets to improve the performance of target detection and image super-resolution.

[0127] S3.4. Use sparse reconstruction algorithm to model the infrared space neighboring target image. Under the premise of allowing a certain quantization error, the possible positions of the closely distributed infrared weak targets on the image plane are limited. Take the sub-pixel center as the possible position of the target, and exhaustively enumerate all possible sub-pixel position sets Ω = {(x l ,y l )} l=1,…,L , construct a spatial neighboring target location set, where L = UVn 2 , n is the size of the sub-pixel grid, each sub-pixel contains at most one target, and the maximum deviation of the target position from the nearest sub-pixel center does not exceed like Figure 5 As shown in Figure 1, it reflects the process of dividing pixels into 3×3 sub-pixel grids. Each pixel is divided into n×n sub-pixel grids, and the over-complete representation of the focal plane measurement model is obtained by extension:

[0128]

[0129] in:

[0130] G(Ω) is a matrix with the steering vectors in the position set Ω as columns. When n>1, the number of columns of G(Ω) is significantly greater than the number of rows (L>UV). The sparse signal vector after expansion Zero padding is performed, taking non-zero values ​​only at the target positions within the grid;

[0131] represents a sparse signal vector;

[0132] w is noise;

[0133] By reformulating the measurement model as the sparse representation in equation (4), the super-resolution task of closely distributed infrared dim targets is transformed into a sparse signal reconstruction problem. This requires selecting K basis functions from G(Ω) and optimizing the signal strength to best fit the observation value z. This sparse recovery problem can be formalized using the l1 norm regularization method as follows:

[0134]

[0135] Among them, λ is the preset regularization parameter.

[0136] Solved Afterwards, various properties of closely distributed infrared weak targets can be extracted through post-processing procedures, including target quantity, radiation intensity and sub-pixel positioning. The non-zero values ​​of are thresholded to obtain the number of targets in the densely distributed infrared dim targets. These non-zero values ​​represent the peak intensity of the target, and the related basis function positions give the coordinate estimates of the densely distributed infrared dim targets.

[0137] S3.5, the reconstruction method based on the iterative soft thresholding algorithm (ISTA) alternately performs the following two update steps to solve the sparsity problem:

[0138]

[0139] in, Indicates that under a given transformation Ψ The transformation coefficients of , k and ρ represent the index and step size of ISTA iteration respectively.

[0140] Although formula (6) is relatively intuitive, formula (7) is actually a special proximal mapping problem, namely prox λφ (r (k) ),when When , this proximal mapping associated with the regularizer φ is defined as:

[0141]

[0142] In pursuit of better reconstruction performance, ISTA-Net adopts a trainable nonlinear transformation function F(·) to replace the traditional transformation Ψ to induce the sparsity of natural images.

[0143] By introducing a linear relationship Formula (5) can be restated as:

[0144]

[0145] Here, λ and α are combined into a learnable parameter θ.

[0146] By solving formula (9), it is easy to prove that:

[0147]

[0148]

[0149] Due to the reversible property of the nonlinear transformation function F(·), the left inverse of F(·) is introduced satisfy where Ι represents the unit operator. Specifically, The function of the network structure is symmetric with F(·). and F(x) are both learnable, then we can Designed as a linear convolution operator structure separated by ReLU operators, just make sure to constrain According to formula (10), It can be calculated in closed form as:

[0150]

[0151] S3.6, the ISTA iteration process is represented as a module of the network. By using the powerful fitting ability of the deep neural network, only a limited number of such modules can be reused to simulate the ISTA solution process. In order to increase the capacity of the network, it is possible to not limit F and θ must be the same, that is, The network structures and weights involved in F and θ do not need to be exactly the same. Then formula (11) can also be expressed as:

[0152]

[0153] See also Figure 4 In this embodiment, a trainable nonlinear transformation function in a neural network is used. This method of replacing the traditional transformation Ψ with the sparsity of natural images cleverly solves a series of drawbacks of traditional algorithms using Ψ to solve problems, allowing the network to spontaneously learn a better transformation during training;

[0154] However, experiments have shown that using only the simple design of Conv-Relu-Conv cannot fully achieve the above-mentioned sparsity-induced requirements.

[0155] In many tasks, the attention mechanism is a simple and effective way to enhance the representation ability of neural networks, because each stage is further sparse. From formula (6), we can see that It has guiding significance for r. In view of this, another branch is designed, in which a method that can be used is used. The convolution kernel, which serves as auxiliary information to guide r, adaptively promotes the expressiveness of image features, thereby achieving more effective nonlinear transformation.

[0156] Therefore, the dynamic transformation module It consists of a two-branch network, where the main branch is a Cov-ReLU-Conv structure, which is used to extract input image features, and the secondary branch is used to fine-tune the features extracted by the main branch. The dynamic transformation structure process is:

[0157] S4.a.1. The output of the k-1th stage is reduced to a feature size of 9 by a multi-layer perceptron to obtain the weight vector W:

[0158]

[0159] Where f(·) represents a fully connected layer;

[0160] S4.a.2. Use the features as weights of the convolution kernel r (k) Perform convolution operation:

[0161] wr (k) =C(W,r (k) )

[0162] Where C(·) represents the convolutional layer;

[0163] S4.a.3, the nonlinearity of the secondary branch is ensured by the sigmoid function, and it is added to the main branch according to different weights to obtain

[0164]

[0165] Among them, A and B represent two convolution operations respectively, and α is a predefined hyperparameter used to adjust the weights of the two branches.

[0166] In this embodiment, considering that the threshold θ of ISTA-Net does not change after model training, and the performance of the traditional ISTA algorithm is extremely sensitive to the threshold, it is recommended that the network be able to dynamically adjust θ based on the information contained in the input image itself, rather than forcing the network to learn a fixed θ. This allows the network to focus more on the distribution information of nearby infrared targets in the image.

[0167] In order to enhance the ability of the network to focus on the most relevant spatial context area to detect aliased targets, a threshold dynamic module is designed. The sparsity features of the induced natural image extracted by convolution are passed through two layers of convolution to obtain different receptive fields, and then information of different scales is extracted. Different weights are obtained according to the information at different locations. Different weights are assigned to the spatial context information using a selective mechanism, and finally a threshold θ related to the input image is obtained. The result obtained by the module is output through the soft threshold function.

[0168] See also Figure 5 In this embodiment, the dynamic soft threshold module soft(·,θ) introduces a selectivity mechanism, and the implementation process is as follows:

[0169] S4.b.1. Extract features through two convolutions to obtain different receptive field information and Channel-based average pooling and maximum pooling extract spatial relationships:

[0170]

[0171] Max pooling can effectively retain the most significant features and effectively highlight significant information such as edges in the image;

[0172] Average pooling does not simply select local significant features, but tends to retain the information of all pixels in the area. The calculation method of average pooling is more sensitive to the details of the image.

[0173] S4.b.2. Use convolution to jointly extract features to better interact with the information after average pooling and maximum pooling:

[0174]

[0175] Among them, Conv represents the convolution operation;

[0176] S4.b.3. Use the sigmoid function to obtain the mask for each spatial selection:

[0177]

[0178] By weighting different receptive field feature maps with corresponding spatial selection masks, a dynamic threshold is obtained:

[0179]

[0180] S4.b.4, set the threshold value of the threshold function to θ, and set the dynamic transformation module The output features are input into the dynamic soft threshold network to make the features sparse.

[0181] In this embodiment, the inverse transformation module It is used to remap the sparse features output by the dynamic soft threshold module soft(·,θ) back to the single-channel image space. It consists of a simple Cov-ReLU-Conv structure and the implementation process is as follows:

[0182] S4.c.1, each convolution size is 3×3, and the loss function L constraint In the inverse transform module To constrain and ensure Among them, I is the identity matrix, and the loss function L constraint for:

[0183]

[0184] Among them, M is the size of the training set, N is the number of iterative network modules, and S is the size of N s ;

[0185] S4.c.2. Loss function L discrepancy Constrained super-resolution image Similar to s, the loss function L discrepancy for:

[0186]

[0187] S4.c.3, the total loss function Loss is:

[0188] Loss = L discrepancy +γL constraint

[0189] Here, γ is a parameter for balancing loss weights, which is set to 0.01 in this embodiment.

[0190] In this embodiment, the S5 further includes:

[0191] S5.1. Evaluator settings: Based on the calculation method of the average precision of the target detection algorithm evaluation index, a specific evaluation index is designed for the coordinate detection of point target non-bounding box detection, namely the average precision of spatially adjacent targets, to effectively judge whether the predicted point target is a true positive TP or a false positive FP.

[0192] Specifically, during the inference process, the predicted point target closest to the true value point is usually considered the best prediction. In order to avoid the situation where a true value point matches multiple predicted points, this embodiment introduces a CSO matching criterion, in which only when the predicted point Corresponding to the true value t i When it is designated as TP, it is ensured that t iThere is no previous match with another predicted point target with higher intensity. Formally, the criterion is defined as:

[0193]

[0194] Among them, δ k ∈{0.05, 0.1, 0.15, 0.2, 0.25} is a series of distance thresholds, where k = 1, 2, 3, 4, 5, used to control the required localization accuracy.

[0195] When the predicted point target is classified as TP or FP, a binary list can be constructed according to the common practice in COCO, where TP prediction is represented by 1 and FP prediction is represented by 0. Based on this binary list, by changing the confidence threshold of the positive prediction, a series of precision and recall values ​​are obtained to generate the precision-recall curve PR. The average precision AP is obtained by calculating the area under the PR curve. AP provides a comprehensive evaluation of the performance of the model under various confidence thresholds, effectively summarizing the trade-off between precision and recall.

[0196] S5.2. Other configurations: Build the GPU training environment, network training epochs, batch_size, load the image PNG file with the data loader, convert the XML file into a super-resolution image, count the number of targets, and configure the optimizer.

[0197] In this example, the following settings are used: a 3090 GPU with 24GB of video memory; 150 network training epochs; a data loader that loads image PNG files, converts XML files into super-resolution images, and counts information such as the number of objects; a batch_size of 64 images per batch, discarding any images less than 64; an optimizer type of Adam, and a fixed learning rate of 0.0001.

[0198] S5.3. Set the intensity threshold of the model validation evaluator. The intensity threshold can be set to 50. The super-resolution image predicted by the model will suppress the data below this value to 0. The remaining elements need to be projected back to the original 11x11 coordinate space and used as the predicted points and the actual points to calculate CSO_mAP. The projection rule is:

[0199] (x i ,y i ,g i ) corresponds to the predicted point Among them, (x i ,y i ) represents the i-th non-zero prediction point in the predicted image, g i Represents the predicted radiation intensity.

[0200] S5.4. The prediction results are passed through the evaluator to obtain the final indicators. The hyperparameters are saved in the configuration file. The configuration file is loaded during training. The model training, verification and testing are completed under the open source deep learning MMengine framework, and the model weights are saved.

[0201] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of the above-mentioned infrared long-distance spatially adjacent target super-resolution method.

[0202] See also Figure 6 An embodiment of the present invention further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the above-mentioned infrared long-distance spatial proximity target super-resolution method.

[0203] An embodiment of the present invention further provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the steps of the above-mentioned infrared long-distance spatially adjacent target super-resolution method.

[0204] It is understandable that the system, device and storage medium provided in the embodiments of the present invention correspond to the method provided in the embodiments of the present invention. The explanation, examples and beneficial effects of the relevant contents can refer to the corresponding parts of the above-mentioned infrared long-distance spatial proximity target super-resolution method.

[0205] It should be noted that those skilled in the art will understand that all or part of the steps implemented in the embodiments of the present invention can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using hardware, it can be implemented in whole or in part in the form of purchased standard parts or modified parts. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiments of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium (eg, a solid state disk (SSD)).

[0206] In summary, this paper proposes a dynamic iterative shrinkage threshold network (DISTA-Net) for detecting neighboring targets in infrared space, which reconceptualizes the traditional coefficient reconstruction into a dynamic framework. It can dynamically generate convolution weights and threshold parameters to adjust the reconstruction process in real time, greatly improving the detection capability of neighboring targets in infrared space.

[0207] It should be understood that the examples and implementation methods described herein are for illustrative purposes only and are not intended to limit the present invention. Those skilled in the art may make various modifications or changes based on them. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for super-resolution of infrared long-distance spatial proximity targets, characterized in that: The following steps are involved: S1. Analyze the imaging characteristics of the optical detection system for long-range targets and the imaging characteristics of spatially adjacent multi-targets, model point target imaging, and generate a simulation image dataset; S2. Initializing the input simulation image based on linear mapping to form an initialized image; S3, using sparse reconstruction algorithm to achieve super-resolution modeling of adjacent targets in infrared space and build a deep unfolding network framework; S4, passing the image initialized by S2 through each stage of the deep expansion network in S3 in sequence, each stage is a gradient descent module and a dynamic proximal mapping module in sequence; S5. Build a GPU training environment, set up the data loader, infrared long-distance spatial proximity super-resolution model, and evaluator configuration files to train the model. Apply the trained model to the test set to generate high-resolution images. Extract target coordinate information through post-processing and input it into the evaluator to obtain evaluation indicators based on infrared long-distance spatial proximity targets and calculate the average detection rate.

2. The infrared long-distance spatial proximity target super-resolution method according to claim 1, characterized in that: Said S1 further comprises: S1.

1. Analysis of target imaging characteristics: Consider the distant target as a point source target. The point source target is imaged on the detector using a Gaussian function as the point spread function (PSF). Multi-target imaging satisfies the linear superposition of Gaussian energies. The point spread function expression is: Among them, σ PSF represents the diffusion variance, (x t ,y t ) is the coordinate of the target on the focal plane; S1.

2. Simulation target: Assume that the number of adjacent targets in infrared long-distance space is a natural number between 1 and 5, set the distance between targets to be less than one pixel, and the centers of multiple targets are concentrated in a single pixel on the detector image plane, thus forming an aliased cluster and indistinguishable target features. The target features include at least the number of targets, their positions, and their radiation intensity.

3. The infrared long-distance spatial proximity target super-resolution method according to claim 1, characterized in that: Said S2 further comprises: S2.

1. Given a dataset Among them, z i Represents the image of the infrared space adjacent target imaging, s i It represents the image after super-resolution c times. According to the point target information (x i ,y i ,g i ) and c calculate the high-resolution image The grayscale of the pixel at is the peak intensity g i , which generates s i ; S2.2, let Z = [z1,…,z M ],S=[s1,…,s M ],initialization Matrix It is expressed as: Where T represents the matrix transpose and F represents the Frobenius norm.

4. The infrared long-distance spatial proximity target super-resolution method according to claim 1, characterized in that: Said S3 further comprises: S3.

1. Model and analyze infrared spatial neighboring targets in a pure Gaussian framework. The image area contains N spatial neighboring targets, and the target is represented by t i =(x i ,y i ,g i ,σ PSF ), where i∈1,..,N, (x i ,y i ), represents the spatial sub-pixel coordinate, g i is the peak intensity, σ PSF represents the diffusion variance, the center point set and peak intensity of all targets are represented by and express; S3.

2. Predicting target sets using the trained infrared long-range spatial proximity super-resolution model The corresponding peak intensity set Among them, M is the number of predicted targets, Forecast point The peak intensity is obtained by minimizing the predicted intensity With the true intensity g i The difference between them is used to achieve joint estimation of target quantity, sub-pixel position and radiation intensity; S3.

3. Establish an infrared long-distance spatial proximity target imaging model for the infrared focal plane. Define the target response at pixel (i, j) as the integral value of the point spread function (PSF) within the pixel boundary. The specific expression is: Among them, (x i,j ,y i,j ) represents the pixel center position and indicates the uniform width of pixel D; Assume that an infrared focal plane consists of U×V pixels and there are K targets with target coordinates (x k ,y k ), k=1,…,K, by rearranging the columns of the focal plane matrix into a vector of UV×1, the focal plane measurement model is constructed as: z=[g c (x1,y1)g c (x2,y2)…g c (x K ,y K )]s+n=G(x,y)s+n Among them, G(x,y) is a UV×K steering matrix, which represents the influence on the response of each pixel. c (x k ,y k ) represents the contribution of target k to the pixel response, and the peak intensity of the target is represented by the vector Indicates that n represents the variance Gaussian white noise, n is independent between different pixels; S3.4, the sparse reconstruction algorithm is used to model the infrared space neighboring target image, and the sub-pixel center is used as the possible location of the target. By exhausting all possible sub-pixel location sets Ω = {(x l ,y l )} l=1,…,L , construct a spatial neighboring target location set, where L = UVn 2 , n is the size of the sub-pixel grid, each sub-pixel contains at most one target, and the maximum deviation of the target position from the nearest sub-pixel center does not exceed Dividing each pixel into an n×n sub-pixel grid, we can extend the overcomplete representation of the focal plane measurement model: Where G(Ω) is a matrix with the steering vectors in the position set Ω as columns, represents the sparse signal vector, w is the noise; Through sparse signal reconstruction, using Norm regularization optimizes the target signal strength to best fit the observed value z, and the formula is: Among them, λ is the preset regularization parameter; S3.5, the reconstruction method based on the iterative soft threshold algorithm ISTA, alternately performs the following two update steps to solve the sparsity problem: Gradient Descent: Near-end mapping: The above two steps are respectively designed as a gradient descent module and a dynamic proximal mapping module, and a deep unfolding network framework is formed by alternately connecting the gradient descent module and the dynamic proximal mapping module multiple times.

5. The infrared long-distance spatial proximity target super-resolution method according to claim 4, characterized in that: The proximal mapping module uses a trainable nonlinear transformation module Dynamic soft threshold module soft(·,θ) and inverse transform module Simulate the iterative process of the iterative soft threshold algorithm ISTA.

6. The infrared long-distance spatial proximity target super-resolution method according to claim 1, characterized in that: In S4, the dynamic proximal mapping module is composed of a dynamic transformation module Dynamic soft threshold module soft(·,θ) and inverse transform module It consists of three parts. In the process of constructing the dynamic iterative shrinkage threshold network DISTA-Net, each iteration is expressed as:

7. The infrared long-distance spatial proximity target super-resolution method according to claim 6, characterized in that: The dynamic transformation module It consists of a two-branch network, where the main branch is a Cov-ReLU-Conv structure, which is used to extract input image features, and the secondary branch is used to fine-tune the features extracted by the main branch. The dynamic transformation structure process is: S4.a.

1. The output of the k-1th stage is scaled by a multi-layer perceptron to obtain the weight vector W: Where f(·) represents a fully connected layer; S4.a.

2. Use the features as weights of the convolution kernel r (k) Perform convolution operation: wr (k) =C(W,r (k) ) Where C(·) represents the convolutional layer; S4.a.3, the nonlinearity of the secondary branch is ensured by the sigmoid function, and it is added to the main branch according to different weights to obtain Among them, A and B represent two convolution operations respectively, and α is a predefined hyperparameter used to adjust the weights of the two branches.

8. The infrared long-distance spatial proximity target super-resolution method according to claim 6, characterized in that: The dynamic soft threshold module soft(·,θ) introduces a selectivity mechanism, and the implementation process is as follows: S4.b.

1. Extract features through two convolutions to obtain different receptive field information and Channel-based average pooling and maximum pooling extract spatial relationships: S4.b.

2. Use convolution to jointly extract features to better interact with the information after average pooling and maximum pooling: S4.b.

3. Use the sigmoid function to obtain the mask for each spatial selection: By weighting different receptive field feature maps with corresponding spatial selection masks, a dynamic threshold is obtained: S4.b.4, set the threshold value of the threshold function to θ, and set the dynamic transformation module The output features are input into the dynamic soft threshold network to make the features sparse.

9. The infrared long-distance spatial proximity target super-resolution method according to claim 6, characterized in that: The inverse transform module It is used to remap the sparse features output by the dynamic soft threshold module soft(·,θ) back to the single-channel image space. It consists of a simple Cov-ReLU-Conv structure and the implementation process is as follows: S4.c.1, each convolution size is 3×3, and the loss function L constraint In the inverse transform module To constrain and ensure Among them, I is the identity matrix, and the loss function L constraint for: Among them, M is the size of the training set, N is the number of iterative network modules, and S is the size of N s ; S4.c.

2. Loss function L discrepancv Constrained super-resolution image Similar to s, the loss function L discrepancv for: S4.c.3, the total loss function Loss is: Loss=L discrepancy +γL constraint Among them, γ is the parameter for balancing loss weight.

10. The infrared long-distance spatial proximity target super-resolution method according to claim 1, characterized in that: Said S5 further comprises: S5.

1. Evaluator Setup: Based on the calculation method of the average precision of the target detection algorithm evaluation metric, a specific evaluation metric is designed for coordinate detection of point target non-bounding box detection, namely the average precision of spatially adjacent targets, to effectively determine whether the predicted point target is a true positive TP or a false positive FP. S5.

2. When a prediction point target is classified as TP or FP, a binary list is constructed, where TP predictions are represented by 1 and FP predictions are represented by 0. Based on this binary list, a range of precision and recall values ​​are obtained by varying the confidence threshold for positive predictions. A precision-recall curve (PR) is generated. The area under the PR curve is calculated to obtain the average precision, thereby comprehensively evaluating the model performance. S5.

3. Other configurations: Build the GPU training environment, set the network training epochs, batch_size, load the image PNG file with the data loader, convert the XML file into a super-resolution image, count the number of targets, and configure the optimizer. S5.

4. The model validation evaluator sets an intensity threshold. The super-resolution image predicted by the model uses this threshold to suppress the data below this value to 0. The remaining elements need to be projected back to the original coordinate space and used as the predicted points and the actual points to calculate CSO_mAP. The projection rule is: (x i ,y i ,g i ) corresponds to the predicted point Among them, (x i ,y i ) represents the i-th non-zero prediction point in the predicted image, g i represents the predicted radiation intensity; S5.

5. The prediction results are passed through the evaluator to obtain the final indicators. The hyperparameters are saved in the configuration file. The configuration file is loaded during training. The model training, verification and testing are completed under the open source deep learning MMengine framework, and the model weights are saved.

Citation Information

Patent Citations

  • Image super-resolution reconstruction method based on genetic algorithm and regular prior model

    CN104408697A

  • Image super-resolution reconstruction method based on deep convolution sparse coding

    CN112907449A

  • Infrared image super-resolution and small target detection method

    CN113222824A

  • SAR moving target clutter suppression method based on deep neural network

    CN117710512A

  • Infrared space adjacent target super-resolution sub-pixel positioning method

    CN119152237A

Cited By

  • Infrared small target cluster sub-pixel positioning method based on physical structure prior and deep expansion network

    CN122243934A

  • An infrared small target cluster sub-pixel positioning method based on physical structure prior and deep unfolding network

    CN122243934B