Infrared long-distance space adjacent target super-resolution method

By constructing a super-resolution method for infrared long-distance spatially nearby targets with a dynamically iteratively shrinking threshold, the problem of overlapping light spots blurring in infrared small target detection is solved, and adaptive detection and high-precision estimation of closely adjacent targets are achieved.

CN120543378BActive Publication Date: 2026-06-02NANJING UNIV OF POSTS & TELECOMM

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING UNIV OF POSTS & TELECOMM
Filing Date
2025-05-19
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing infrared small target detection technologies suffer from overlapping spherical light spots that blur the identification of the number, location, and radiation intensity of individual targets when dealing with closely adjacent targets. Traditional methods rely on handcrafted features and are sensitive to hyperparameter adjustment, while deep learning methods are insufficient in feature extraction and parameter adjustment, resulting in poor detection performance.

Method used

An infrared long-range spatial proximity target super-resolution method based on dynamic iterative shrinkage threshold is adopted. By analyzing the optical imaging characteristics, a deep unfolding network framework is constructed. Combined with sparse reconstruction algorithm and dynamic near-end mapping module, the network parameters are dynamically adjusted to improve the detection effect.

Benefits of technology

It achieves adaptive detection of nearby targets in infrared space, improving detection accuracy and robustness. It can adjust parameters in the reconstruction process in real time, improving the estimation accuracy of target number, location and radiation intensity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120543378B_ABST
    Figure CN120543378B_ABST
Patent Text Reader

Abstract

The application discloses an infrared long-distance space adjacent target super-resolution method, wherein the method comprises the following steps: S1, analyzing the imaging characteristics of a long-distance target optical detection system and the imaging characteristics of space adjacent multiple targets, modeling point target imaging, and generating a simulation image dataset; S2, initializing the input simulation image based on linear mapping; S3, realizing infrared space adjacent target super-resolution modeling by adopting a sparse reconstruction algorithm, and constructing a deep unfolding network framework; S4, sequentially passing the initialized image through each stage of the deep unfolding network; S5, constructing a GPU training environment, setting a data loader, a model and an evaluator configuration file, training the model, applying the trained model to a test set, generating a high-resolution image, extracting target coordinate information by post-processing and inputting the target coordinate information into the evaluator, obtaining evaluation indexes based on infrared long-distance space adjacent targets, and calculating an average detection rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of infrared imaging technology, and more specifically to a super-resolution method for infrared long-distance spatial proximity targets. Background Technology

[0002] Infrared imaging plays a crucial role in various long-range detection and surveillance missions because it is extremely sensitive to thermal radiation and unaffected by illumination conditions. However, due to the considerable distance between distant targets and the imaging system, the radiation intensity captured from remote targets is actually quite weak. This challenge is exacerbated when targets appear in densely packed clusters in space, such as spatially adjacent targets, resulting in overlapping spherical spots. This phenomenon obscures the perception of the number, precise location, and radiation intensity of individual targets, significantly hindering the subsequent detection, tracking, and identification phases of infrared search and track (IRST) systems. Therefore, exploring effective techniques for unmixing and super-resolution reconstruction of such closely-spaced infrared small targets (CSIST) to accurately identify their precise location and radiation intensity is of great significance.

[0003] While near-range infrared small target super-resolution is crucial in various applications, research specifically targeting this task remains scarce. Existing techniques primarily include traditional optimization methods, infrared small target detection methods, and depth unfolding framework-based super-resolution methods. However, these methods still have some limitations in practical applications.

[0004] Traditional optimization methods analyze the optical characteristics and imaging mechanisms of nearby targets in infrared space, formulating the problem as a parameter estimation task and employing optimization algorithms to solve it. This paper utilizes the sparsity of targets on the imaging plane to design a sparse reconstruction method based on discrete sampling. This method uses an overcomplete dictionary and solves a second-order cone programming problem under l1-norm regularization. However, the performance of these optimization-based models is highly dependent on meticulous hyperparameter tuning, which poses a significant challenge in real-world scenarios.

[0005] Traditional infrared small target detection techniques are mainly divided into sparse low-rank decomposition and deep learning methods. Sparse low-rank decomposition aims to separate targets from the background by decomposing the original image into a sparse target representation and a low-rank background representation. However, real-world interference factors that manifest as outliers in the background pose a significant challenge. Traditional methods rely solely on grayscale intensity or basic handcrafted features, making it difficult to semantically distinguish between genuine targets and interference. Furthermore, these methods are highly sensitive to hyperparameter configuration, requiring extensive manual tuning, which hinders their practical application.

[0006] Deep learning research primarily focuses on developing multi-scale feature fusion models to compensate for the scarcity of inherent target features. These methods mainly concentrate on detecting overlapping targets as a whole, but due to a lack of datasets, they struggle with subsequent tasks such as sub-pixel localization and radiance estimation. Super-resolution methods within deep unfolding frameworks, such as ISTA-Net based on a soft-thresholding iterative network and ADMM-Net based on the alternating direction multiplier method, reinterpret traditional optimization algorithms as fully connected feedforward neural networks. This approach effectively extends to new samples, achieving high image reconstruction results with fewer iterations and faster inference speed through end-to-end training. However, current methods still exhibit relatively low performance metrics; even slight changes in the number or location of targets can lead to severe artifacts. This is mainly due to two factors: firstly, the current networks still lack sufficient feature extraction capabilities and fail to effectively utilize multi-scale feature information; secondly, once the network is trained, some key parameters, such as soft threshold and stride, are fixed and difficult to adjust based on the specific characteristics of the target. Summary of the Invention

[0007] The present invention provides a super-resolution method for infrared long-distance spatial proximity targets based on dynamic iterative shrinkage threshold, which is more robust to hyperparameter changes and can be effectively applied to various practical scenarios, and can at least solve one of the above-mentioned technical problems.

[0008] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:

[0009] A super-resolution method for infrared long-range spatial proximity targets includes the following steps:

[0010] S1. Analyze the imaging characteristics of the optical detection system for distant targets and the imaging characteristics of multiple targets in spatial proximity, model the imaging of point targets, and generate a simulation image dataset.

[0011] S2. Initialize the input simulation image based on linear mapping to form an initialized image;

[0012] S3. Employ a sparse reconstruction algorithm to achieve super-resolution modeling of infrared spatial proximity targets and construct a deep unfolding network framework;

[0013] S4. The image initialized by S2 is passed sequentially through each stage of the depth unrolling network in S3, with each stage being a gradient descent module and a dynamic proximal mapping module.

[0014] S5. Build a GPU training environment, set up the configuration files for the data loader, the infrared long-range spatial proximity super-resolution model, and the evaluator, train the model, apply the trained model to the test set, generate high-resolution images, extract target coordinate information through post-processing and input it into the evaluator, obtain the evaluation index based on infrared long-range spatial proximity targets, and calculate the average detection rate.

[0015] Furthermore, S1 further includes:

[0016] S1.1 Analyzing Target Imaging Characteristics: Considering distant targets as point source targets, the image of these point source targets on the detector uses a Gaussian function as the point spread function (PSF). Multi-target imaging satisfies the linear superposition of Gaussian energies. The expression for the point spread function is:

[0017]

[0018] Where, σ PSF Represents the diffusion variance, (x t ,y t () represents the coordinates of the target on the focal plane;

[0019] S1.2 Simulated Target: Assuming that the number of nearby targets in the infrared long-distance space is a natural number from 1 to 5, the distance between targets is set to be less than one pixel, and the centers of multiple targets are concentrated in a single pixel of the detector image plane, thus forming an aliased cluster with indistinguishable target features. These target features include at least the number of targets, their positions, and radiation intensity.

[0020] Furthermore, S2 further includes:

[0021] S2.1, Given a dataset Among them, z i An image representing an infrared spatially proximal target, s i This represents the image after super-resolution by c times, based on the point target information (x). i ,y i ,g i ) and c calculate its high-resolution image The grayscale value of the pixel at that location is the peak intensity g. i s is thus generated i ;

[0022] S2.2, Let Z = [z1, ..., z M ],S=[s1,…,s M ],initialization matrix Q init = This is represented as:

[0023]

[0024] Where T represents the matrix transpose and F represents the Frobenes norm.

[0025] Furthermore, S3 further includes:

[0026] S3.1. Modeling and analyzing infrared spatially neighboring targets within a pure Gaussian framework. The image region contains N spatially neighboring targets, denoted as t. i =(x i ,y i ,g i ,σ PSF ), where i∈1,..,N, (x i ,y i ), representing spatial sub-pixel coordinates, g i It is the peak intensity, σ PSF The diffusion variance is represented by the set of center points of all targets and the peak intensity, respectively. and express;

[0027] S3.2 Predict the target set using the trained infrared long-range spatial proximity super-resolution model. With the corresponding peak intensity set Where M is the predicted number of targets. For prediction points The peak intensity is obtained by minimizing the predicted intensity. Compared with the true strength g i The differences between them enable joint estimation of target quantity, sub-pixel position, and radiation intensity;

[0028] S3.3. Establish an infrared long-range spatial proximity target imaging model based on the infrared focal plane. Define the response of the target at pixel (i,j) as the integral value of the point spread function (PSF) within the pixel boundary. The specific expression is as follows:

[0029]

[0030] Among them, (x i,j ,y i,j ) represents the center position of the pixel, and indicates the uniform width of pixel D;

[0031] Assume an infrared focal plane consists of U×V pixels, and there are K targets with coordinates (x, y, y). k ,y k (k = 1, ..., K), by rearranging the columns of the focal plane matrix into a UV×1 vector, the focal plane measurement model is constructed as follows:

[0032] z = [g c (x1,y1)gc (x2,y2)…g c (x K ,y K )]s+n=G(x,y)s+n

[0033] Where G(x,y) is a UV×K steering matrix, representing the effect on the response of each pixel, with columns g c (x k ,y k The peak intensity of the target is represented by the vector s = [s1, s2, ..., s]. K ] T This indicates that n represents the variance. Gaussian white noise, where n is independent between different pixels;

[0034] S3.4. A sparse reconstruction algorithm is used to model the infrared spatially nearby target image, taking the sub-pixel center as the possible location of the target. This is achieved by exhaustively enumerating all possible sub-pixel location sets Ω={(x l ,y l )} l=1,…,L Construct a set of spatially nearest target locations, where L = UVn 2 Where n is the size of the subpixel grid, each subpixel contains at most one target, and the maximum deviation between the target position and the center of the nearest subpixel does not exceed [a certain value]. Dividing each pixel into an n×n sub-pixel grid, we obtain an overcomplete representation of the focal plane measurement model:

[0035]

[0036] Where G(Ω) is a matrix with the guiding vectors in the position set Ω as its columns. This represents a sparse signal vector, where w is noise;

[0037] By reconstructing the sparse signal and optimizing the target signal intensity using L1 norm regularization, the best-fit observation value z is obtained, as shown in the formula:

[0038]

[0039] Where λ is a preset regularization parameter;

[0040] S3.5. A reconstruction method based on the iterative soft thresholding algorithm ISTA, which alternately performs the following two update steps to address the sparsity problem:

[0041] Gradient descent:

[0042] Proximal mapping:

[0043] The two steps described above are designed as gradient descent modules and dynamic proximal mapping modules, respectively. The deep unfolding network framework is formed by repeatedly connecting the gradient descent modules and dynamic proximal mapping modules in an alternating manner.

[0044] Furthermore, the proximal mapping module employs a trainable nonlinear transformation module. Dynamic soft thresholding module soft(·,θ) and inverse transform module Simulate the iterative process of the Iterative Soft Thresholding Algorithm (ISTA).

[0045] Furthermore, in step S4, the dynamic proximal mapping module is composed of a dynamic transformation module. Dynamic soft thresholding module soft(·,θ) and inverse transform module Composed of three parts, each iteration in the construction of the dynamic iterative shrinking threshold network DISTA-Net is represented as follows:

[0046]

[0047] Furthermore, the dynamic transformation module It consists of a two-branch network, where the main branch is a Cov-ReLU-Conv structure used to extract features from the input image, and the secondary branches are used to fine-tune the features extracted by the main branch. The dynamic transformation process is as follows:

[0048] S4.a.1, The output of the (k-1)th stage is processed by a multilayer perceptron to scale the features and obtain the weight vector W:

[0049]

[0050] Where f(·) represents a fully connected layer;

[0051] S4.a.2, Using features as weights for the convolution kernel r (k) Perform convolution operations:

[0052] wr (k) =C(W,r) (k) )

[0053] Where C(·) represents a convolutional layer;

[0054] S4.a.3. The non-linearity of the secondary branch is ensured by using the sigmoid function, and it is added to the main branch with different weights to obtain...

[0055]

[0056] Here, A and B represent two convolution operations, and α is a predefined hyperparameter used to adjust the weights of the two branches.

[0057] Furthermore, the dynamic soft threshold module soft(·,θ) introduces a selection mechanism, the implementation process of which is as follows:

[0058] S4.b.1. By extracting features through two convolutions, different receptive field information is obtained. and Extracting spatial relationships based on channel-based average pooling and max pooling:

[0059]

[0060] S4.b.2. Use convolution to jointly extract features in order to better interact the information after average pooling and max pooling:

[0061]

[0062] S4.b.3, Use the sigmoid function to obtain the mask for each spatial selection:

[0063]

[0064] A dynamic threshold is obtained by weighting different receptive field feature maps with their corresponding spatial selection masks:

[0065]

[0066] S4.b.4, Set the threshold of the threshold function to θ, and then set the dynamic transformation module... The output features are input into a dynamic soft thresholding network to sparsify the features.

[0067] Furthermore, the inverse transformation module The function used to remap the sparsified features output by the dynamic soft thresholding module soft(·,θ) back to the single-channel image space consists of a simple Cov-ReLU-Conv structure, and the implementation process is as follows:

[0068] S4.c.1, Each convolution is 3×3 in size, and the loss function L... constraint The inverse transformation module Set constraints to ensure Where I is the identity matrix, and L is the loss function. constraint for:

[0069]

[0070] Where M is the size of the training set, N is the number of iterative network modules, and S is of size N. s ;

[0071] S4.c.2, Loss Function Ldiscrepancy Constrained super-resolution image Similar to s, the loss function L discrepancy for:

[0072]

[0073] S4.c.3, The total loss function Loss is:

[0074] Loss = L discrepancy +γL constraint

[0075] Where γ is the parameter for balancing the loss weight.

[0076] Furthermore, S5 further includes:

[0077] S5.1 Evaluator settings: Based on the calculation method of the average accuracy of the target detection algorithm evaluation index, a specific evaluation index is designed for the coordinate detection of point targets without bounding boxes, namely the average accuracy of spatially adjacent targets, to effectively judge whether the predicted point target is a true positive (TP) or a false positive (FP).

[0078] S5.2 When the predicted point target is classified as TP or FP, a binary list is constructed, where TP prediction is represented by 1 and FP prediction is represented by 0. Based on this binary list, a series of precision and recall values ​​are obtained by changing the confidence threshold of positive prediction, and a precision-recall curve PR is generated. The average precision is obtained by calculating the area under the PR curve, thereby comprehensively evaluating the model performance.

[0079] S5.3 Other configurations: Build the GPU training environment, network training epochs, batch size, load PNG files of images with the data loader, convert XML files to super-resolution images, count the number of targets, and configure the optimizer;

[0080] S5.4 The model validation evaluator sets an intensity threshold. Data below this threshold in the model's predicted super-resolution image is suppressed to 0. The remaining elements need to be projected back into the original coordinate space and used as predicted points to calculate CSO_mAP with the ground truth points. The projection rules are as follows:

[0081] (x i ,y i ,g i The corresponding prediction point is Among them, (x i ,y i ) represents the i-th non-zero prediction point in the prediction image, g i Indicates the predicted radiation intensity;

[0082] S5.5. The prediction results are passed through the evaluator to obtain the final index. The hyperparameters are saved in the configuration file. The configuration file is loaded during training. The model training, validation and testing, as well as the model weights are saved, are completed in the open-source deep learning MMengine framework.

[0083] The beneficial effects of this invention are reflected in:

[0084] This invention provides a complete solution for spatial proximity target detection, encompassing the entire process from dataset preparation to model building and evaluation metrics. Specifically, it offers a benchmark spatial proximity target super-resolution model that adaptively generates convolutional weights and threshold parameters to adjust the reconstruction process in real time. Unlike previous methods, the parameters related to proximity mapping, including nonlinear transformations, shrinkage thresholds, and stride, are dynamically adjusted based on the input data, rather than being manually designed or fixed after training. This effectively improves the detection performance of infrared spatial proximity targets.

[0085] This invention reconstructs the traditional sparse reconstruction problem based on a dynamic network framework, integrating deep unfolding methods and attention mechanisms to effectively improve the network's feature extraction capabilities, thereby enhancing the accuracy and robustness of detection. Attached Figure Description

[0086] The accompanying drawings, which are provided to further illustrate this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application.

[0087] Figure 1 This is a schematic diagram of the overall process of the method according to an embodiment of the present invention.

[0088] Figure 2 This is a schematic diagram illustrating the specific process of the method according to an embodiment of the present invention.

[0089] Figure 3 This is a schematic diagram of the model structure of the dynamic iterative shrinking threshold network according to an embodiment of the present invention.

[0090] Figure 4 This is a schematic diagram of the dynamic transformation module structure according to an embodiment of the present invention.

[0091] Figure 5 This is a schematic diagram of the structure of the dynamic threshold module in an embodiment of the present invention.

[0092] Figure 6 This is a structural block diagram of a computer device according to an embodiment of the present invention. Detailed Implementation

[0093] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Unless otherwise specified, the embodiments and features in the embodiments of this application can be combined with each other. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0094] It should be noted that the meaning of "and / or" throughout the text includes three parallel solutions. Taking "A and / or B" as an example, it includes solution A, solution B, or a solution that simultaneously satisfies A and B. Furthermore, "multiple" refers to two or more. Additionally, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.

[0095] See Figures 1-2 This invention provides a super-resolution method for infrared long-range spatial proximity targets, comprising the following steps:

[0096] S1. Analyze the imaging characteristics of the optical detection system for distant targets and the imaging characteristics of multiple targets in spatial proximity, model the imaging of point targets, and generate a simulation image dataset.

[0097] S2. Initialize the input simulation image based on linear mapping to form an initialized image;

[0098] S3. Employ a sparse reconstruction algorithm to achieve super-resolution modeling of infrared spatial proximity targets and construct a deep unfolding network framework;

[0099] S4. The image initialized by S2 is passed sequentially through each stage of the depth unrolling network in S3, with each stage being a gradient descent module and a dynamic proximal mapping module.

[0100] S5. Build a GPU training environment, set up the configuration files for the data loader, the infrared long-range spatial proximity super-resolution model, and the evaluator, train the model, apply the trained model to the test set, generate high-resolution images, extract target coordinate information through post-processing and input it into the evaluator, obtain the evaluation index based on infrared long-range spatial proximity targets, and calculate the average detection rate.

[0101] In this embodiment, S1 further includes:

[0102] S1.1 Analyze the imaging characteristics of the target: Treat the distant target as a point source target, analyze the imaging characteristics of the optical detection system of the distant target, model the imaging of the power source target, analyze the image characteristics of multiple targets in spatial proximity, simulate 100,000 image data, and record the position and radiation intensity of the target in each image as a label. Divide the dataset into training set, validation set and test set for each stage of model training.

[0103] S1.2 When a monochromatic point light source is imaged through a lens, an "Airy disk" appears at the focal point. Its characteristic is that the central bright spot accounts for 84% of the total energy, surrounded by multiple concentric diffraction rings. This diffraction effect is usually approximated by the point spread function (PSF), whose standard deviation depends on the focal ratio (f-value) of the sensor and the detection band, and is used to quantify energy diffusion.

[0104] Imaging point source targets on the detector using a Gaussian function as the point spread function (PSF), multi-target imaging satisfies the linear superposition of Gaussian energies. The expression for the point spread function is:

[0105]

[0106] Where, σ PSF Represents the diffusion variance, (x t ,y t ) is the coordinate of the target on the focal plane.

[0107] In multi-object imaging, the intensity captured by each pixel is the cumulative response of superimposed point sources. Especially for distant objects, each object is considered a point source, and the resulting "Airy disk" radius is determined by 1.22λ / D, where λ is the wavelength and D is the lens diameter. This radius equals 1.9σ of the Gaussian PSF, marking the physical resolution limit of the sensor and conforming to the definition of Rayleigh units. In this embodiment, 84% energy concentration is used as the resolution criterion, and σ... PSF Set to 0.5 pixels.

[0108] S1.3 Simulated Target: Assuming that the number of nearby targets in the infrared long-distance space is a natural number from 1 to 5, the distance between the targets is set to be small, and the centers of multiple targets are concentrated in a single pixel of the detector image plane, thus forming an aliased cluster with indistinguishable target features. These target features include at least the number of targets, their positions, and their radiation intensity.

[0109] The simulation experiment in this embodiment covers images containing 1 to 5 targets, each target represented by sub-pixel coordinates and peak intensity t. i =(x i ,y i ,g i ,σ PSF), where i∈1,…,N, are the learning objectives for annotation, and an 11×11 pixel grid is used for each objective set. The image space. To test the super-resolution performance of this infrared long-range spatial proximity super-resolution model, this embodiment simulates multiple targets within a single pixel range, while ensuring that the distance between them exceeds 0.52 Rayleigh units. In this configuration, a dataset containing 100,000 samples is simulated, with 80,000 samples allocated to the training set and the remaining 20,000 samples evenly distributed between the validation and test sets.

[0110] In this embodiment, S2 further includes:

[0111] S2.1 Similar to traditional iterative algorithms, such as ISTA and DISTA-Net, initialization is also required, i.e., initialization is performed on... Initialization is performed using linear mappings to compute initial values; specifically, given a dataset... Among them, z i An image representing an infrared spatially proximal target, s i This indicates the image after super-resolution by c times. For example, when the pixel division factor is c times, the probability of aliasing of point targets still meets a preset requirement.

[0112] S2.2 Based on the point target information (x) of the original image i ,y i ,g i ) and c calculate its high-resolution image The grayscale value of the pixel at that location is the peak intensity g. i s is thus generated i ;

[0113] Let Z = [z1, ..., z M ],S=[s1,…,s M ],initialization matrix Q init =argmin Q ||QZ-SF2=SZTZZT-1, s0 is represented as:

[0114]

[0115] Where T represents the matrix transpose and F represents the Frobenes norm.

[0116] See Figure 3 In this embodiment, S3 further includes:

[0117] S3.1. Modeling and analyzing infrared spatially neighboring targets within a pure Gaussian framework. The image region contains N spatially neighboring targets, denoted as t.i =(x i ,y i ,g i ,σ PSF ), where i∈1,..,N, (x i ,y i ), representing spatial sub-pixel coordinates, g i It is the peak intensity, σ PSF The diffusion variance is represented by the set of center points of all targets and the peak intensity, respectively. and express.

[0118] S3.2 Predict the target set using the trained infrared long-range spatial proximity super-resolution model. With the corresponding peak intensity set Where M is the predicted number of targets. For prediction points The peak intensity is obtained by minimizing the predicted intensity. Compared with the true strength g i The differences between them enable joint estimation of target quantity, sub-pixel position, and radiation intensity.

[0119] S3.3. Establish an infrared long-range spatial proximity target imaging model on the infrared focal plane. In an infrared imaging system, due to the influence of the point spread function (PSF) of the optical system, the infrared radiation emitted by the infrared spatial proximity target is dispersed among the pixels on the focal plane. Each pixel integrates the incident radiation of the overlapping target to generate a response. The responses of the pixels to the target are linearly superimposed, thus forming a linear focal plane imaging model. Specifically, the response of pixel (i,j) to the target is defined as the integral value of the point spread function PSF within the boundary of that pixel, expressed as:

[0120]

[0121] Among them, (x i,j ,y i,j ) represents the center position of the pixel, and indicates the uniform width of pixel D;

[0122] Assume an infrared focal plane consists of U×V pixels, and there are K targets with coordinates (x, y, y). k ,y k (k = 1, ..., K), by rearranging the columns of the focal plane matrix into a UV×1 vector, the focal plane measurement model is constructed as follows:

[0123] z = [g c (x1,y1)g c (x2,y2)…g c (x K ,yK )]s+n=G(x,y)s+n

[0124] #(3)

[0125] Where G(x,y) is a UV×K steering matrix, representing the effect on the response of each pixel, with columns g c (x k ,y k The peak intensity of the target is represented by the vector s = [s1, s2, ..., s]. K ] T This indicates that n represents the variance. Gaussian white noise, where n is independent between different pixels;

[0126] Through the modeling described above, this embodiment aims to more accurately capture and analyze the signals of weak infrared targets in order to improve the performance of target detection and image super-resolution.

[0127] S3.4. A sparse reconstruction algorithm is used to model infrared spatially neighboring target images. Under the premise of allowing a certain quantization error, the possible locations of densely distributed infrared weak targets on the image plane are finite. The sub-pixel center is taken as the possible target location. This is achieved by exhaustively enumerating the set of all possible sub-pixel locations Ω={(x... l ,y l )} l=1,…,L Construct a set of spatially nearest target locations, where L = UVn 2 Where n is the size of the subpixel grid, each subpixel contains at most one target, and the maximum deviation between the target position and the center of the nearest subpixel does not exceed [a certain value]. like Figure 5 As shown, this illustrates the process of dividing a pixel into a 3×3 sub-pixel grid. Dividing each pixel into an n×n sub-pixel grid expands to obtain an overcomplete representation of the focal plane measurement model.

[0128]

[0129] in:

[0130] G(Ω) is a matrix with the steering vectors in the position set Ω as its columns. When n>1, the number of columns in G(Ω) is significantly greater than the number of rows (L>UV). This results in a sparse signal vector after expansion. Zero-filling was applied, taking non-zero values ​​only at the target location within the grid;

[0131] Represents a sparse signal vector;

[0132] w represents noise;

[0133] By reformulating the measurement model into the sparse representation in Equation (4) above, the super-resolution task of densely distributed infrared weak targets is transformed into a sparse signal reconstruction problem. This requires selecting K basis functions from G(Ω) and optimizing the signal strength to best fit the observed value z. This sparse recovery problem can be formalized using l1 norm regularization, as shown in the following formula:

[0134]

[0135] Where λ is a preset regularization parameter.

[0136] Solving for the results Subsequently, post-processing procedures can extract various attributes of closely distributed, low-resolution infrared targets, including target quantity, radiation intensity, and sub-pixel localization. By thresholding the non-zero terms, the number of targets among closely distributed weak infrared targets can be obtained. These non-zero values ​​represent the peak intensity of the targets, and the associated basis function positions provide coordinate estimates for the closely distributed weak infrared targets.

[0137] S3.5. A reconstruction method based on the Iterative Soft Thresholding Algorithm (ISTA) alternately performs the following two update steps to address the sparsity problem:

[0138]

[0139] in, Indicates that under a given transformation Ψ The transformation coefficients, k and ρ represent the index and step size of the ISTA iteration, respectively.

[0140] Although formula (6) is relatively intuitive, formula (7) is actually a special problem of prox mapping, i.e., prox λφ (r (k) ),when When, the proximal mapping associated with the regularizer φ is defined as:

[0141]

[0142] In pursuit of better reconstruction performance, ISTA-Net replaces the traditional transform Ψ with a trainable nonlinear transform function F(·) to induce sparsity in natural images.

[0143] By introducing a linear relationship Formula (5) can be restated as:

[0144]

[0145] Here, λ and α are combined into a single learnable parameter θ.

[0146] By solving formula (9), it is easy to prove that:

[0147]

[0148]

[0149] Due to the invertibility of the nonlinear transformation function F(·), the left inverse of F(·) is introduced. satisfy Where Ι represents the unit operator. Specifically, The functionality of the network structure is symmetric with F(·). Because Since both F(x) and F(x) are learnable, then we can... Designed as a linear convolution operator structure separated by ReLU operators, as long as constraints are ensured during training... According to formula (10), It can then be calculated in closed-form as follows:

[0150]

[0151] S3.6. Represent one iteration of ISTA as a module of the network. Leveraging the powerful fitting ability of deep neural networks, only a finite number of such modules need to be reused to simulate the ISTA solution process. To increase the network capacity, there is no limit to the number of modules. F and θ must be the same, that is The network structures and weights involved in F and θ do not need to be exactly the same. Therefore, formula (11) can also be expressed as:

[0152]

[0153] See Figure 4 In this embodiment, a trainable nonlinear transformation function in a neural network is used. This method, which replaces the traditional transformation Ψ to induce sparsity in natural images, cleverly solves a series of drawbacks of traditional algorithms that use Ψ to solve problems, allowing the network to spontaneously learn a better transformation during training.

[0154] However, experiments have shown that simply using the Conv-Relu-Conv design is insufficient to fully meet the aforementioned requirements for sparsity-induced geometries.

[0155] In many tasks, attention mechanisms are a simple and effective way to enhance the representational power of neural networks, as each stage further sparsifies the representation. As can be seen from formula (6), This is instructive for r; therefore, another branch was designed, in which a method that can utilize... As an auxiliary information guide, the convolution kernel r adaptively enhances the expressive power of image features, thereby achieving more efficient nonlinear transformations.

[0156] Therefore, the dynamic transformation module It consists of a two-branch network, where the main branch is a Cov-ReLU-Conv structure used to extract features from the input image, and the secondary branches are used to fine-tune the features extracted by the main branch. The dynamic transformation process is as follows:

[0157] S4.a.1, The output of the (k-1)th stage is processed by a multilayer perceptron to reduce the feature size from 1089 to a scaled size of 9, resulting in the weight vector W:

[0158]

[0159] Where f(·) represents a fully connected layer;

[0160] S4.a.2, Using features as weights for the convolution kernel r (k) Perform convolution operations:

[0161] wr (k) =C(W,r) (k) )

[0162] Where C(·) represents a convolutional layer;

[0163] S4.a.3. The non-linearity of the secondary branch is ensured by using the sigmoid function, and it is added to the main branch with different weights to obtain...

[0164]

[0165] Here, A and B represent two convolution operations, and α is a predefined hyperparameter used to adjust the weights of the two branches.

[0166] In this embodiment, considering that the threshold θ of ISTA-Net does not change after the model training is completed, and the performance of the traditional algorithm ISTA is extremely sensitive to the threshold, it is recommended that the network be able to dynamically adjust θ according to the information contained in the input image itself, rather than having the network learn a fixed θ, so as to make the network focus more on the distribution information of infrared spatial neighboring targets in the image.

[0167] To enhance the network's ability to focus on the most relevant spatial context regions for detecting aliased targets, a dynamic thresholding module was designed. This module extracts sparse features from the induced natural image through convolution, passing them through two convolutional layers to obtain different receptive fields, thereby extracting information at different scales. Different weights are assigned based on the information at different locations, and a selective mechanism is used to assign these weights to the spatial context information, ultimately resulting in a threshold θ related to the input image. The data then undergoes... The results obtained by the module are output after passing through a soft threshold function.

[0168] See Figure 5 In this embodiment, the dynamic soft threshold module soft(·,θ) introduces a selection mechanism, the implementation process of which is as follows:

[0169] S4.b.1. By extracting features through two convolutions, different receptive field information is obtained. and Extracting spatial relationships based on channel-based average pooling and max pooling:

[0170]

[0171] Max pooling can effectively preserve the most salient features and effectively highlight salient information such as edges in an image;

[0172] Average pooling does not simply select local salient features, but tends to retain information from all pixels within a region. The calculation method of average pooling is more sensitive to the details of the image.

[0173] S4.b.2. Use convolution to jointly extract features in order to better interact the information after average pooling and max pooling:

[0174]

[0175] Where Conv represents the convolution operation;

[0176] S4.b.3, Use the sigmoid function to obtain the mask for each spatial selection:

[0177]

[0178] A dynamic threshold is obtained by weighting different receptive field feature maps with their corresponding spatial selection masks:

[0179]

[0180] S4.b.4, Set the threshold of the threshold function to θ, and then set the dynamic transformation module... The output features are input into a dynamic soft thresholding network to sparsify the features.

[0181] In this embodiment, the inverse transformation module The function used to remap the sparsified features output by the dynamic soft thresholding module soft(·,θ) back to the single-channel image space consists of a simple Cov-ReLU-Conv structure, and the implementation process is as follows:

[0182] S4.c.1, Each convolution is 3×3 in size, and the loss function L... constraint The inverse transformation module Set constraints to ensure Where I is the identity matrix, and L is the loss function. constraint for:

[0183]

[0184] Where M is the size of the training set, N is the number of iterative network modules, and S is of size N. s ;

[0185] S4.c.2, Loss Function L discrepancy Constrained super-resolution image Similar to s, the loss function L discrepancy for:

[0186]

[0187] S4.c.3, The total loss function Loss is:

[0188] Loss = L discrepancy +γL constraint

[0189] Wherein, γ is the parameter for balancing the loss weight, which is set to 0.01 in this embodiment.

[0190] In this embodiment, S5 further includes:

[0191] S5.1 Evaluator Settings: Based on the calculation method of the average accuracy of the target detection algorithm evaluation index, a specific evaluation index is designed for the coordinate detection of point targets without bounding boxes, namely the average accuracy of spatially nearby targets, to effectively judge whether the predicted point target is a true positive (TP) or a false positive (FP).

[0192] Specifically, during the inference process, the predicted point closest to the true value is typically considered the best prediction. To avoid a situation where one true value matches multiple predicted points, this embodiment introduces a CSO matching criterion, where a prediction is only considered best if the predicted point... Corresponding to the true value t i Only when it is specified as TP, and at the same time ensure t iThere was no previous match with another predicted point target of higher intensity. Formally, this criterion is defined as:

[0193]

[0194] Where, δ k ∈{0.05,0.1,0.15,0.2,0.25} is a series of distance thresholds, where k=1,2,3,4,5, used to control the required positioning accuracy.

[0195] When the predicted point target is classified as TP or FP, a binary list can be constructed according to the general practice in COCO, where TP predictions are represented by 1 and FP predictions by 0. Based on this binary list, a series of precision and recall values ​​are obtained by changing the confidence threshold of positive predictions, generating a precision-recall curve PR. The average precision AP is obtained by calculating the area under the PR curve. AP provides a comprehensive evaluation of the model's performance at various confidence thresholds, effectively summarizing the trade-off between precision and recall.

[0196] S5.2 Other configurations: Build a GPU training environment, network training epochs, batch size, load PNG files of images using a data loader, convert XML files to super-resolution images, count the number of targets, and configure the optimizer.

[0197] In this embodiment, the following settings can be configured: a 3090 GPU with 24GB of VRAM; 150 training epochs; a data loader that loads PNG images, converts XML files to super-resolution images, and counts the number of targets; a batch size of 64; and discarding any batches with fewer than 64 images. The optimizer type is set to Adam, and the learning rate is fixed at 0.0001.

[0198] S5.3 The model validation evaluator sets an intensity threshold, which can be set to 50. The super-resolution image predicted by the model will suppress data below this threshold to 0. The remaining elements need to be projected back into the original 11x11 coordinate space as predicted points and used to calculate CSO_mAP with the ground truth points. The projection rules are as follows:

[0199] (x i ,y i ,g i The corresponding prediction point is Among them, (x i ,y i ) represents the i-th non-zero prediction point in the prediction image, g i This indicates the predicted radiation intensity.

[0200] S5.4 The prediction results are passed through the evaluator to obtain the final index. The hyperparameters are saved in the configuration file. The configuration file is loaded during training. The model training, validation and testing, as well as the model weights are saved, are completed in the open-source deep learning MMengine framework.

[0201] This invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the infrared long-range spatial proximity target super-resolution method described above.

[0202] See Figure 6 The present invention also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the infrared long-range spatial proximity target super-resolution method described above.

[0203] This invention also provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the steps of the infrared long-range spatial proximity target super-resolution method described above.

[0204] It is understood that the systems, devices and storage media provided in the embodiments of the present invention correspond to the methods provided in the embodiments of the present invention, and the explanations, examples and beneficial effects of the relevant contents can be referred to the corresponding parts of the above-described infrared long-range spatial proximity target super-resolution method.

[0205] It should be noted that those skilled in the art will understand that all or part of the steps implemented in the embodiments of the present invention can be implemented entirely or partially by software, hardware, firmware, or any combination thereof. When implemented in hardware, it can be implemented entirely or partially by purchasing standard parts or modifications. When implemented in software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid state disks (SSDs)).

[0206] In summary, this invention proposes a dynamic iterative shrinking threshold network (DISTA-Net) for infrared spatial proximity targets. It reconceptualizes traditional coefficient reconstruction into a dynamic framework that can dynamically generate convolution weights and threshold parameters to adjust the reconstruction process in real time, greatly improving the detection capability of infrared spatial proximity targets.

[0207] It should be understood that the examples and embodiments described herein are for illustrative purposes only and are not intended to limit the invention. Those skilled in the art can make various modifications or changes based on them. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the invention should be included within the protection scope of the invention.

Claims

1. A super-resolution method for infrared long-range spatial proximity targets, characterized in that, Includes the following steps: S1. Analyze the imaging characteristics of the optical detection system for distant targets and the imaging characteristics of multiple targets in spatial proximity, model the imaging of point targets, and generate a simulation image dataset. S2. Initialize the input simulation image based on linear mapping to form an initialized image; S3. Employ a sparse reconstruction algorithm to achieve super-resolution modeling of infrared spatial proximity targets and construct a deep unfolding network framework; S4. The image initialized by S2 is passed sequentially through each stage of the depth unrolling network in S3, with each stage being a gradient descent module and a dynamic proximal mapping module. S5. Build a GPU training environment, set up the configuration files for the data loader, infrared long-range spatial proximity super-resolution model and evaluator, train the model, apply the trained model to the test set, generate high-resolution images, extract target coordinate information through post-processing and input it into the evaluator, obtain the evaluation index based on infrared long-range spatial proximity targets, and calculate the average detection rate. S3 further includes: S3.

1. Modeling and analyzing infrared spatially neighboring targets within a pure Gaussian framework. The image region contains N spatially neighboring targets, denoted as t. i =(x i ,y i ,g i ,σ PSF ), where i∈1,.., N, (x i ,y i ) represents the spatial sub-pixel coordinates, g i It is the peak intensity, σ PSF The diffusion variance is represented by the set of center points of all targets and the peak intensity, respectively. and express; S3.2 Predict the target set using the trained infrared long-range spatial proximity super-resolution model. With the corresponding peak intensity set ,in, To predict the target quantity, For prediction points The peak intensity is obtained by minimizing the predicted intensity. Compared with the true strength g i The differences between them enable joint estimation of target quantity, sub-pixel position, and radiation intensity; S3.

3. Establish an infrared long-range spatial proximity target imaging model based on the infrared focal plane. Define the response of the target at pixel (i,j) as the integral value of the point spread function (PSF) within the pixel boundary. The specific expression is as follows: Among them, (x i,j ,y i,j ) represents the center position of the pixel, and D represents the uniform width of the pixel; Assume an infrared focal plane consists of U×V pixels, and there are N targets with coordinates (x, y, y). k ,y k (k = 1, ..., N), by rearranging the columns of the focal plane matrix into a UV×1 vector, the focal plane measurement model is constructed as follows: z=[g c (x1,y1)g c (x2,y2)…g c (x N ,y N )]s+ε=G(x,y)s+ε Where G(x,y) is a UV×N steering matrix, representing the effect on the response of each pixel, with columns g c (x k ,y k () represents the contribution of target k to the pixel response. The representative has variance Gaussian white noise, It is independent between different pixels; S3.

4. A sparse reconstruction algorithm is used to model the infrared spatially nearby target image, taking the sub-pixel center as the possible location of the target. This is achieved by exhaustively enumerating all possible sub-pixel location sets Ω={(x l ,y l )} l=1,…,L Construct a set of spatially nearest target locations, where L = UVn 2 Where n is the size of the subpixel grid, each subpixel contains at most one target, and the maximum deviation between the target position and the center of the nearest subpixel does not exceed [a certain value]. Dividing each pixel into an n×n sub-pixel grid, we obtain an overcomplete representation of the focal plane measurement model: Where G(Ω) is a matrix with the guiding vectors in the position set Ω as its columns. This represents a sparse signal vector, where w is noise; By reconstructing sparse signals, utilizing Norm regularization optimizes the target signal intensity, thereby best fitting the observed value z, as shown in the formula: Where λ is a preset regularization parameter; S3.

5. A reconstruction method based on the iterative soft thresholding algorithm ISTA, which alternately performs the following two update steps to address the sparsity problem: Gradient descent: Proximal mapping: T represents matrix transpose; gradient descent module and dynamic proximal mapping module are designed for the above two steps respectively, and a deep unfolding network framework is formed by repeatedly connecting gradient descent module and dynamic proximal mapping module in alternation; Input the evaluator to obtain evaluation metrics based on infrared long-distance spatial proximity targets and calculate the average detection rate.

2. The infrared long-range spatial proximity target super-resolution method as described in claim 1, characterized in that, S1 further includes: S1.1 Analyzing Target Imaging Characteristics: Considering distant targets as point source targets, the image of these point source targets on the detector uses a Gaussian function as the point spread function (PSF). Multi-target imaging satisfies the linear superposition of Gaussian energies. The expression for the point spread function is: S1.2 Simulated Target: Assuming that the number of nearby targets in the infrared long-distance space is a natural number from 1 to 5, the distance between targets is set to be less than one pixel, and the centers of multiple targets are concentrated in a single pixel of the detector image plane, thus forming an aliased cluster with indistinguishable target features. These target features include at least the number of targets, their positions, and radiation intensity.

3. The infrared long-range spatial proximity target super-resolution method as described in claim 1, characterized in that, S2 further includes: S2.1, Given a dataset , where z n An image representing an infrared spatially proximal target, s n This represents the image after super-resolution by c times, based on the point target information (x). i ,y i ,g i ) and c calculate its high-resolution image The grayscale value of the pixel at that location is the peak intensity g. i s is thus generated n ; S2.2, Let Z = [z1, ..., z M ],S=[s1,…,s M ],initialization matrix This is represented as: Where F represents the Frobenz norm.

4. The infrared long-range spatial proximity target super-resolution method as described in claim 1, characterized in that, The proximal mapping module employs a trainable nonlinear transformation module. Dynamic soft thresholding module soft(·,θ) and inverse transform module Simulate the iterative process of the Iterative Soft Thresholding Algorithm (ISTA).

5. The infrared long-range spatial proximity target super-resolution method as described in claim 1, characterized in that, In step S4, the dynamic proximal mapping module consists of a dynamic transformation module. Dynamic soft thresholding module soft(·,θ) and inverse transform module Composed of three parts, each iteration in the construction of the dynamic iterative shrinking threshold network DISTA-Net is represented as follows:

6. The infrared long-range spatial proximity target super-resolution method as described in claim 5, characterized in that, The dynamic transformation module It consists of a two-branch network, where the main branch is a Cov-ReLU-Conv structure used to extract features from the input image, and the secondary branches are used to fine-tune the features extracted by the main branch. The dynamic transformation process is as follows: S4.a.1, The output of the (k-1)th stage is processed by a multilayer perceptron to scale the features and obtain the weight vector W: Where f(·) represents a fully connected layer; S4.a.2, Using features as weights for the convolution kernel r (k) Perform convolution operations: Where C(·) represents a convolutional layer; S4.a.

3. The non-linearity of the secondary branch is ensured by using the sigmoid function, and it is added to the main branch with different weights to obtain... Here, A and B represent two convolution operations, and α is a predefined hyperparameter used to adjust the weights of the two branches.

7. The infrared long-range spatial proximity target super-resolution method as described in claim 5, characterized in that, The dynamic soft threshold module soft(·,θ) introduces a selection mechanism, the implementation process of which is as follows: S4.b.

1. By extracting features through two convolutions, different receptive field information is obtained. and Extracting spatial relationships based on channel-based average pooling and max pooling: S4.b.

2. Use convolution to jointly extract features in order to better interact the information after average pooling and max pooling: S4.b.

3. Use the sigmoid function to obtain the mask for each spatial selection; A dynamic threshold is obtained by weighting different receptive field feature maps with corresponding spatial selection masks; S4.b.4, Set the threshold of the threshold function to θ, and then set the dynamic transformation module... The output features are input into a dynamic soft thresholding network to sparsify the features.

8. The infrared long-range spatial proximity target super-resolution method as described in claim 1, characterized in that, S5 further includes: S5.1 Evaluator settings: Based on the calculation method of the average accuracy of the target detection algorithm evaluation index, a specific evaluation index is designed for the coordinate detection of point targets without bounding boxes, namely the average accuracy of spatially adjacent targets, to effectively judge whether the predicted point target is a true positive (TP) or a false positive (FP). S5.2 When the predicted point target is classified as TP or FP, a binary list is constructed, where TP prediction is represented by 1 and FP prediction is represented by 0. Based on this binary list, a series of precision and recall values ​​are obtained by changing the confidence threshold of positive prediction, and a precision-recall curve PR is generated. The average precision is obtained by calculating the area under the PR curve, thereby comprehensively evaluating the model performance. S5.3 Other configurations: Build the GPU training environment, network training epochs, batch size, load PNG files of images with the data loader, convert XML files to super-resolution images, count the number of targets, and configure the optimizer; S5.4 The model validation evaluator sets an intensity threshold. Data below this threshold in the model's predicted super-resolution image is suppressed to 0. The remaining elements need to be projected back into the original coordinate space and used as predicted points to calculate CSO_mAP with the ground truth points. The projection rules are as follows: (x i ,y i ,g i The corresponding prediction point is ; S5.

5. The prediction results are passed through the evaluator to obtain the final index. The hyperparameters are saved in the configuration file. The configuration file is loaded during training. The model training, validation and testing, as well as the model weights are saved, are completed in the open-source deep learning MMengine framework.