Multi-frame space adjacent infrared small target unmixing method and network based on model and data dual drive

Through the multi-frame space adjacent infrared small target demix method based on the model and data dual-driven multi-frame space, the detection problem of infrared small target detection is solved when it is less than the Rayleigh radius, effective demix and positioning of infrared point targets is achieved, and detection application scenarios are expanded.

CN120495627APending Publication Date: 2025-08-15XIAN INST OF OPTICS & PRECISION MECHANICS CHINESE ACAD OF SCI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510564960.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing infrared small-object detection method has degraded detection performance when the target space distance is less than the Rayleigh radius, and lacks spatially adjacent infrared small-objective demixing methods, so it is impossible to effectively identify and locate infrared point targets.

Method used

The multi-frame space adjacent infrared small target demix method based on dual-driven model and data is adopted. Through initialization, feature extraction, position encoding, convolution and time deformable alignment, demixing of infrared point targets whose target space distance is less than the Rayleigh radius is achieved.

Benefits of technology

The feature extraction capability and inter-frame alignment accuracy are improved, and the detection and sub-pixel positioning of infrared point targets with target spatial distances smaller than the Rayleigh radius can be achieved, which broadens the application scenarios of infrared small target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495627A_ABST
    Figure CN120495627A_ABST
Patent Text Reader

Abstract

The invention relates to a spatial adjacent infrared small target unmixing method and network, in particular to a multi-frame spatial adjacent infrared small target unmixing method and network based on model and data dual drive, and solves the problem that the prior art cannot meet the detection requirement of an infrared point target of which the target spatial distance is smaller than the Rayleigh radius. The method comprises the following steps: step 1, inputting slice images of S frames; step 2, initialization is carried out; 3, feature extraction is carried out through W times of iteration; step 4, adding the time sequence information into the mapping image after the Wth iteration; 5, sequentially and independently executing the steps a to b on each sequence image according to the sequence that the first frame to the (2N + 1)-th frame, the second frame to the (2N + 2)-th frame,..., the (S-2N)-th frame and the sequence of the sequences of the images after the time sequence information is added; the steps a to b are as follows: step a: convolution; and b, time deformable alignment is carried out, and then tail convolution is carried out.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and a network for unmixing small infrared targets in close proximity to each other in space, and in particular to a method and a network for unmixing small infrared targets in close proximity to each other in multiple frames based on dual-drive of a model and data. Background Art

[0002] Infrared imaging technology, unaffected by changes in illumination, offers irreplaceable advantages in military reconnaissance, security monitoring, search and rescue, and other fields. However, due to limitations in infrared sensor resolution and the detection requirements of early warning systems, distant targets often appear as tiny, pixel-sized objects in imaging. Furthermore, interference from atmospheric scattering, sensor noise, and complex backgrounds results in an extremely low signal-to-noise ratio. These targets lack distinct features such as shape and texture, making them difficult to effectively identify using commonly used detection methods (such as contour- or contrast-based algorithms).

[0003] The existing mainstream infrared small target detection (IRSTD) methods for infrared small targets can be divided into two categories: traditional infrared small target detection methods and deep learning infrared small target detection methods.

[0004] Traditional infrared small target detection methods rely on artificially designed mathematical models or prior knowledge, mainly including single-frame methods that separate targets from backgrounds through spatial filtering (such as Top-Hat transform) or low-rank sparse decomposition (such as IPI model), and rely on the spatial distribution characteristics of targets (such as local contrast) for detection, as well as multi-frame methods that utilize the motion trajectory characteristics of targets in the time domain (such as dynamic programming and optical flow method) to suppress random noise through inter-frame correlation.

[0005] Among existing deep learning methods for infrared small target detection, early methods, such as densely nested attention networks, achieve multi-scale feature fusion through dense cross-layer semantic interactions, and combine attention mechanisms to enhance the saliency of small targets. The core of these methods is to retain target details through nested residual structures, significantly reducing false alarm rates in complex backgrounds (such as clouds and sea surfaces). Subsequent methods, whether through multi-scale progressive fusion in semantic segmentation (taking the U-Net in the U-Net (UIU-Net) framework as an example) or regression-based detection paradigms (taking the local to global fusion network (LoGoNet) as an example), all enhance detection effects through multi-scale fusion and task-specific structural optimization.

[0006] Although the above-mentioned infrared small target detection method has achieved certain results, there are still two problems:

[0007] First, in terms of the application scenarios of the detection method, there are the following problems:

[0008] (1) Its application is mainly limited to scenes with large field of view and sparse objects.

[0009] (2) Its implicit assumption is that "there is a strict one-to-one correspondence between the detected targets in the image and the actual targets in the real world."

[0010] Therefore, due to the limitations of sensor resolution, when the target's spatial distance is less than the Rayleigh spot radius, its radiation energy will be severely aliased during the imaging process, ultimately forming a blurred light spot on the image plane. In such scenarios, the detection performance of existing methods is significantly degraded. Common contour detection strategies are ineffective for point targets, and even if dense target spots are detected, the precise number and location cannot be determined.

[0011] Second, there are the following problems in the design of detection methods:

[0012] (1) Existing detection methods rely on a general residual block stacking backbone network for feature extraction, and their feature extraction capabilities are limited.

[0013] (2) The innovation of existing detection methods is mostly to complicate the model to achieve the effect of improving the detection rate, but it is still unable to detect small infrared targets whose target space distance is less than the Rayleigh spot radius, and it is not conducive to practical application deployment.

[0014] In addition, according to the survey, the existing methods for detecting dense small targets are mostly limited to visible light, and the scenes are limited to dense vehicles in traffic backgrounds, crowds in stations, etc. There is no relevant proposal for unmixing of spatially adjacent infrared small targets, and there is a lack of spatially adjacent infrared small target unmixing methods and networks with sequence benchmarks.

[0015] In summary, there is an urgent need to develop a demixing method and network for spatially adjacent small infrared targets to meet the detection requirements of infrared point targets whose target space distance is less than the Rayleigh spot radius.

[0016] The spatially adjacent infrared small target involved in the present invention refers to an infrared point target whose target spatial distance is smaller than the Rayleigh spot radius. Summary of the Invention

[0017] The purpose of the present invention is to solve the technical problems that the application scenarios of existing infrared small target detection methods are limited to scenes with large fields of view and sparse targets, the feature extraction capabilities are limited, and the existing dense small target detection methods are mostly limited to visible light, and there is no spatially adjacent infrared small target unmixing method, so that the existing technology cannot meet the detection requirements of infrared point targets with a target spatial distance less than the Rayleigh spot radius. Instead, a multi-frame spatially adjacent infrared small target unmixing method and network based on dual-drive of model and data is provided.

[0018] In order to solve the above technical problems, the technical solution adopted by the present invention is:

[0019] A multi-frame spatially adjacent infrared small target unmixing method based on dual-drive model and data, wherein the spatially adjacent infrared small target refers to an infrared point target whose target spatial distance is less than the Rayleigh radius; the method is special in that it includes the following steps:

[0020] Step 1: Input the slice images containing the light spots of S frames to be demixed on the same trajectory in order from the earliest to the last frame, where S is an integer greater than or equal to 3;

[0021] Step 2: Initializing the slice images of each frame input in step 1 respectively to obtain initialized images of each frame with an initialization super-resolution ratio of c, where c is an integer greater than 1;

[0022] Step 3: For each frame of the initialized image obtained in step 2, feature extraction is performed through W iterations, each iteration including one gradient descent module iteration and one proximal mapping module iteration, where 5≤W≤15, and W is an integer, to obtain a mapped image after the Wth iteration of feature extraction for each frame;

[0023] Step 4: Position encoding is performed on each of the S frames in step 1 to obtain the encoded time sequence information of each frame in the S frames; then, the encoded time sequence information of each frame in the S frames is added one by one to the mapped image after the Wth iteration of the feature extraction of each frame obtained in step 3 to obtain an image of each frame after adding the time sequence information;

[0024] Step 5: For the images after adding timing information of each frame obtained in step 4, according to the 1st frame to the (2N+1)th frame, the 2nd frame to the (2N+2)th frame, ..., the (S-2N)th frame to the Sth frame, each as a sequence and its sequence order, starting from the 1st sequence, in sequence order, independently perform steps a to b on the images after adding timing information of the consecutive (2N+1) frames corresponding to each sequence to obtain the feature map after unmixing of the intermediate frames of the corresponding sequence, until steps a to b are performed on all sequences, and the feature maps after unmixing of the intermediate frames of all sequences constitute the unmixed image of the slice image containing the light spot of the (S-2N) frame after removing the first and last N frames from the S frames in step 1, thereby completing the unmixing; N is an integer, and 3≤2N+1≤S; steps a to b are:

[0025] Step a: Perform convolution to obtain the intermediate frame feature map of the corresponding sequence and 2N reference frame feature maps;

[0026] Step b: Perform time-deformable alignment on the reference frame feature maps obtained in step a so that the reference frame feature maps are aligned with the intermediate frame feature map obtained in step a in the spatiotemporal dimension, and obtain a mapping image after the alignment of the feature maps of each frame of the corresponding sequence; then perform tail convolution on the mapping image after the alignment of the feature maps of each frame of the corresponding sequence, and obtain a feature map after the unmixing of the intermediate frame of the corresponding sequence.

[0027] Furthermore, the step b is specifically as follows:

[0028] Step b.1: The reference frame feature maps obtained in step a are compared with the intermediate frame feature maps obtained in step a one by one, and the features useful for alignment in adjacent frames are selected through a selective attention mechanism to create aggregated mapping images of the reference frame feature maps and the intermediate frame feature maps.

[0029] Step b.2: performing convolution on the aggregated mapping images of each reference frame feature map and the intermediate frame feature map created in step b.1, respectively, to predict sampling parameters corresponding to each reference frame feature map;

[0030] Step b.3: According to the sampling parameters corresponding to each reference frame feature map predicted in step b.2, features are adaptively extracted from each reference frame feature map obtained in step a in a one-to-one correspondence through bilinear interpolation, so as to align and integrate the complementary information in each reference frame feature map with the overall feature representation, so that each reference frame feature map is aligned with the intermediate frame feature map obtained in step a in the spatiotemporal dimension, and each reference frame feature map is modulated into a mapping image after the alignment of each reference frame feature map. The mapping image after the alignment of each reference frame feature map and the image after the timing information is added to the intermediate frame of the corresponding sequence constitute the mapping image after the alignment of each frame feature map of the corresponding sequence as a whole;

[0031] Step b.4: Perform convolution on the mapped image after alignment of the feature maps of each frame of the corresponding sequence obtained in step b.3, plus V residual blocks plus the tail convolution of the convolution structure, where V is a natural number greater than zero, and its specific value is determined according to the design requirements, to obtain the feature map of the intermediate frame of the corresponding sequence after demixing.

[0032] Furthermore, in step 2, the initialization is performed using the following formula:

[0033]

[0034] Where: Q init Represents the preset initialization matrix; L i represents the i-th frame slice image input in step 1, where i is a natural number, 1≤i≤S; Represents the image after initialization of the i-th frame.

[0035] Furthermore, in step 3, the gradient descent module is as follows:

[0036]

[0037] Where: represents the mapping image after the (k-1)th iteration in the feature extraction of the i-th frame, where k is 1, 2, ..., W, and k is an integer. When k = 1, is the image of the i-th frame after initialization obtained in step 2; ρ represents the step size of gradient descent, which is obtained through learning; Φ represents the preset sampling matrix; Φ T represents the transposed matrix of Φ; express Direct reconstruction results at the kth iteration;

[0038] The proximal mapping module is as follows:

[0039]

[0040] Where: H GTi represents the Ground Truth corresponding to the i-th frame image; G(·) represents a trainable nonlinear transformation function that satisfies in: represents the left inverse of G(·), I represents the identity matrix; λ satisfies θ=λ·α, where λ represents the predefined l1 regularization parameter and α represents the predefined l2 regularization parameter; represents the mapped image after the kth iteration in the feature extraction of the i-th frame.

[0041] In this way, the deep unfolding paradigm replaces the hand-crafted matrix Ψ with a trainable nonlinear transformation function G(·), which can reduce the amount of computation and save computational costs.

[0042] Furthermore, in step 4, position coding is performed on each frame in the S frames in step 1 to obtain the time sequence information of each frame in the S frames after coding. The following formula is used to obtain the time sequence information of each frame in the S frames after coding:

[0043] B i =Sigmoid(MLP(Encoder(b i )));

[0044] Where: b i represents the position code of the i-th frame in the S frame described in step 1, b i Take the number (i-1); Encoder (·) means expanding the frame position to the batch size and normalizing it; MLP (·) means multi-layer perceptron; Sigmoid (·) means activation function; Bi represents the timing information of the i-th frame in the S frames after encoding in step 1;

[0045] When the encoded time sequence information of each frame in the S frames is added one by one to the mapped image after the Wth iteration in the feature extraction of each frame obtained in step 3, the following formula is used for addition:

[0046]

[0047] Where: represents the mapped image after the Wth iteration in the feature extraction of the i-th frame; ⊙ represents the Hadamard product of two matrices; H i ′ Indicates the image of the i-th frame after adding timing information.

[0048] Furthermore, in step b.1, the attention selection mechanism adopts the following formula:

[0049]

[0050] Where:

[0051] F i When the image after adding the timing information of the i-th frame is the reference frame of the sequence in step 5, the reference frame feature map obtained after the convolution in step a is obtained, F i ∈{F t-N ,…,F t+N}, and F i ≠F t , said t represents the middle frame of the sequence in step 5;

[0052] F t An image representing the intermediate frame of the sequence in step 5 after adding timing information, and the intermediate frame feature map obtained after performing convolution in step a;

[0053] SelectiveAttention(F i ,F t ) means reducing F by convolution first i and F t The channel dimension of the two feature maps is then connected, and the maximum pooling feature and the average pooling feature are extracted through the average maximum pooling module, and activated by the Sigmoid function to obtain the maximum pooling weight and the average pooling weight. The maximum pooling weight is then applied to the reference frame feature map, and the average pooling weight is applied to the intermediate frame feature map. Finally, the weighted features obtained from the two are connected along the channel dimension to create the reference frame feature map F i Aggregate map image with the intermediate frame feature map;

[0054] Represents the reference frame feature map F i Aggregate map image with the intermediate frame feature map;

[0055] In step b.2, the aggregated mapping images of each reference frame feature map created in step b.1 and the intermediate frame feature map are convolved to predict the sampling parameters corresponding to each reference frame feature map, using the following formula for convolution:

[0056]

[0057] Where: η i Represents the reference frame feature map F i The corresponding sampling parameters are the content-related offset matrices;

[0058] In step b.3, according to the sampling parameters corresponding to each reference frame feature map predicted in step b.2, features are adaptively extracted from each reference frame feature map obtained in step a one-to-one by bilinear interpolation, so as to align and integrate the complementary information in each reference frame feature map with the overall feature representation, so that each reference frame feature map is aligned with the intermediate frame feature map obtained in step a in the spatiotemporal dimension. When each reference frame feature map is modulated into a mapped image after the reference frame feature map is aligned, its expression is:

[0059]

[0060] Where: represents bilinear interpolation; H i ″ represents the reference frame feature map F i Aligned mapped image.

[0061] Furthermore, in step 3, when performing feature extraction, the constraint loss should satisfy the following formula:

[0062]

[0063] Where: X represents the number of frames in the entire training set; Y represents the size of the initialized image obtained in step 2; m represents the mth frame of the entire training set; W represents the number of iterations during feature extraction in step 3; s m Represents the Ground Truth corresponding to the m-th frame image of the entire training set; G (k) (·) represents the k-th iteration result of G(·); express The k-th iteration result of ; represents the constraint loss;

[0064] In step b, when performing the time-deformable alignment, the alignment loss should satisfy the following formula:

[0065]

[0066] Where: H t ′ The image representing the intermediate frame corresponding to each sequence with the timing information added; represents the alignment loss;

[0067] The regression loss should satisfy the following formula:

[0068]

[0069] Where: X / S represents the number of trajectories; H ji and s ji Represent the reconstruction result of the i-th frame of the j-th trajectory and its corresponding Ground Truth respectively; represents the regression loss.

[0070] Furthermore, in order to load a complete trajectory each time, in step 1, S is equal to the number of frames of a complete trajectory.

[0071] At the same time, the present invention also provides a multi-frame spatial proximity infrared small target unmixing network based on dual-drive of model and data, which is used to implement the above-mentioned multi-frame spatial proximity infrared small target unmixing method based on dual-drive of model and data. The special features of the network are:

[0072] It includes initialization module, feature extraction module, position encoding module, convolution module and time deformable alignment module;

[0073] The initialization module is used to initialize the slice images containing the light spots of the S frames to be demixed on the same input trajectory, respectively, to obtain the initialized images of each frame with an initialization super-resolution ratio of c, where c is an integer greater than 1; and S is an integer greater than or equal to 3;

[0074] The feature extraction module is used to perform feature extraction on the initialized image of each frame output by the initialization module through W iterations, each iteration including one gradient descent module iteration and one proximal mapping module iteration, where 5≤W≤15, and W is an integer, to obtain the mapping image after the Wth iteration in the feature extraction of each frame;

[0075] The position encoding module is used to perform position encoding on each frame of the S frames to be demixed on the same input trajectory to obtain time sequence information of each frame in the S frames after encoding; then, the time sequence information of each frame in the S frames after encoding is added one by one to the mapped image after the Wth iteration of the feature extraction of each frame output by the feature extraction module, to obtain an image of each frame after adding the time sequence information;

[0076] The convolution module is used to perform convolution on the images after adding timing information of each frame output by the position encoding module, starting from the first sequence and in sequence order, independently and sequentially on the images after adding timing information of consecutive (2N+1) frames corresponding to each sequence, according to the sequence order, according to which the 1st frame to the (2N+1)th frame, the 2nd frame to the (2N+2)th frame, ..., the (S-2N)th frame to the Sth frame are each a sequence and their sequence order, to obtain an intermediate frame feature map and 2N reference frame feature maps of the corresponding sequence; N is an integer, and 3≤2N+1≤S;

[0077] The temporal deformable alignment module includes a temporal deformable alignment block and a tail convolution block;

[0078] The temporal deformable alignment block is used to perform temporal deformable alignment on the reference frame feature maps of the corresponding sequence output by the convolution module, so that the reference frame feature maps of the corresponding sequence are aligned with the intermediate frame feature maps of the corresponding sequence output by the convolution module in the spatiotemporal dimension, thereby obtaining a mapped image after the feature maps of each frame of the corresponding sequence are aligned;

[0079] The tail convolution block is used to perform tail convolution on the mapped image after the feature maps of each frame of the corresponding sequence output by the time deformable alignment block are aligned to obtain the feature map of the intermediate frame of the corresponding sequence after unmixing; the feature maps of the intermediate frames of all sequences after unmixing as a whole constitute an unmixed image of the slice image containing the light spot after removing the first and last N frames from the S frames to be unmixed on the same input trajectory (S-2N).

[0080] Furthermore, for optimization, the Adam optimizer and the MMEngine framework were used, and the learning rate was always kept at 10 -4 .

[0081] The beneficial effects of the present invention are:

[0082] (1) The present invention is based on a multi-frame spatially adjacent infrared small target unmixing method driven by both model and data. When performing feature extraction, a sparse-driven deep expansion architecture is used to replace the residual block stacking architecture, thereby improving the feature extraction capability of the network. When aligning the feature maps of each reference frame to the feature map of the intermediate frame in the spatiotemporal dimension, a lightweight multi-frame temporal deformable alignment method based on implicit learning is used to replace the traditional optical flow method based on explicit learning, thereby improving the inter-frame alignment accuracy. Therefore, it can realize the unmixing of multi-frame spatially adjacent infrared point targets whose target spatial distance is less than the Rayleigh (Rayleigh spot) radius, and thus can meet the detection requirements of infrared point targets whose target spatial distance is less than the Rayleigh (Rayleigh spot) radius. Therefore, the present invention solves the technical problems that the application scenarios of existing infrared small target detection methods are limited to scenes with large fields of view and sparse targets, the feature extraction capability is limited, and the existing dense small target detection methods are mostly limited to visible light, and there is no spatially adjacent infrared small target unmixing method, so that the existing technology cannot meet the detection requirements of infrared point targets whose target spatial distance is less than the Rayleigh (Rayleigh spot) radius.

[0083] (2) The present invention proposes the unmixing problem of multi-frame spatially adjacent infrared small targets for the first time, broadening the definition of infrared small target detection (IRSTD). It can realize the unmixing of multi-frame spatially adjacent infrared point targets whose target space distance is less than the Rayleigh (Rayleigh spot) radius, and can meet the detection requirements of infrared point targets whose target space distance is less than the Rayleigh (Rayleigh spot) radius, thereby making the detection of dense small targets no longer limited to visible light.

[0084] (3) The present invention is based on a multi-frame spatial adjacent infrared small target unmixing method driven by both model and data, which can perform sub-pixel positioning and unmixing of sub-targets in light spots with energy aliasing.

[0085] (4) The present invention is based on a multi-frame spatially adjacent infrared small target unmixing network driven by both models and data, namely a multi-frame deformable refinement network (DeRefNet), which includes an initialization module, a feature extraction module, a position encoding module, a convolution module, and a temporal deformable alignment module. It creatively combines the idea of deep unfolding with multi-frame deformable alignment, and is applied to the field of multi-frame spatially adjacent infrared small target unmixing for the first time.

[0086] (5) In the multi-frame spatially adjacent infrared small target unmixing network driven by both the model and the data, when the convolution module performs convolution and the subsequent time deformable alignment module performs alignment, the first frame to the (2N+1)th frame, the second frame to the (2N+2)th frame, ..., the (S-2N)th frame to the Sth frame are each a sequence and their sequence order. Starting from the first sequence, the continuous (2N+1) frames corresponding to each sequence are processed independently in sequence order, and a total of (S-2N) times are processed instead of S / (2N+1) times. Therefore, the data utilization rate is improved and the model performance is better. BRIEF DESCRIPTION OF THE DRAWINGS

[0087] Figure 1 It is a technical roadmap of an embodiment of the multi-frame spatial proximity infrared small target unmixing method based on dual-drive of model and data of the present invention;

[0088] Figure 2 This is a system overall architecture diagram of a multi-frame spatial proximity infrared small target unmixing method and network embodiment based on a dual-drive model and data of the present invention;

[0089] Figure 3 This is an example of a partial data set corresponding to each of the three trajectories in the data set used when inputting slice images containing light spots of S frames to be unmixed on the same trajectory in step 1 of an embodiment of the multi-frame spatial adjacent infrared small target unmixing method based on dual-drive of model and data of the present invention;

[0090] Figure 4 This is an implementation example of the multi-frame spatial proximity infrared small target unmixing method (i.e., DeRefNet) based on model and data dual-drive using ISTA, ISTANet, and the present invention. The three unmixing methods are used to unmix the slice images on the four input trajectories A, B, C, and D, respectively, and the visualization comparison chart is compared with GT (the ground truth corresponding to the input slice image).

[0091] The descriptions of the numbers in the figure are as follows:

[0092] 1- Initialization module, 2- Feature extraction module, 3- Position encoding module, 4- Convolution module, 5- Time deformable alignment module, 51- Time deformable alignment block, 52- Tail convolution block, 01- Slice image, 02- Initialized image, 03- Image after adding timing information, 04- Aligned mapping image, 05- Feature map after unmixing of the intermediate frame. DETAILED DESCRIPTION

[0093] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0094] See also Figure 1 and Figure 2 The present invention provides a multi-frame spatially adjacent infrared small target unmixing method based on dual-drive of model and data. The spatially adjacent infrared small target refers to an infrared point target whose target space distance is less than the Rayleigh radius. All GTs involved in the present invention are abbreviations of Ground Truth, which means true value in Chinese. The multi-frame spatially adjacent infrared small target unmixing method based on dual-drive of model and data includes the following steps:

[0095] Step 1: Input the slice images 01 containing the spot of the S frames to be unmixed on the same trajectory in order from the earliest to the last frame, where S is an integer greater than or equal to 3. In order to load the complete trajectory each time, the batch size is usually set to the number of frames of a complete trajectory in the experiment, that is, let S equal the number of frames of a complete trajectory. In this embodiment, S is equal to 20.

[0096] Step 2: Initialize the slice images 01 of each frame input in step 1 using the following formula:

[0097]

[0098] Where: Q init Represents the preset initialization matrix; L i represents the i-th frame slice image input in step 1, i is a natural number, 1≤i≤S; Represents the image after initialization of the i-th frame;

[0099] After initialization, an initialized image O2 of each frame with an initial super-resolution ratio of c is obtained, where c is an integer greater than 1. In this embodiment, the initial super-resolution ratio c is set to 3. The size of the slice image O1 containing the light spot input in step 1 is 11×11, so the size of its corresponding GT (Ground Truth) image and the size of the final output unmixed image are both 33×33. Figure 3 This is an example of a partial data set corresponding to each of the three trajectories in the data set used when inputting slice images containing light spots of S frames to be unmixed on the same trajectory in step 1 of the embodiment of the multi-frame spatial adjacent infrared small target unmixing method based on dual drive of model and data of the present invention. In this figure, GT represents the ground truth corresponding to the input slice image;

[0100] Step 3: For the initialized image O2 of each frame obtained in step 2, feature extraction is performed through W iterations, each iteration includes one gradient descent module iteration and one proximal mapping module iteration, 5≤W≤15, and W is an integer, and the mapping image after the Wth iteration in the feature extraction of each frame is obtained;

[0101] The formulas of the gradient descent module and the proximal mapping module are as follows:

[0102]

[0103] Where: represents the mapping image after the (k-1)th iteration in the feature extraction of the i-th frame, where k is 1, 2, ..., W, and k is an integer. When k = 1, is the image of the i-th frame after initialization obtained in step 2; ρ represents the step size of gradient descent, which is obtained through learning; Φ represents the preset sampling matrix; Φ T represents the transposed matrix of Φ; express The direct reconstruction result in the kth iteration; H GTi represents the Ground Truth corresponding to the i-th frame image; λ represents the predefined l1 regularization parameter; Ψ represents a hand-crafted matrix; represents the mapped image after the kth iteration in the feature extraction of the i-th frame;

[0104] Since solving the more complex non-orthogonal (or even nonlinear) transformation Ψ It is difficult and usually requires a large number of iterations to reach the theoretical optimal value, which usually incurs a large amount of computational cost. In this embodiment, the deep expansion paradigm replaces the hand-crafted matrix Ψ with a trainable nonlinear transformation function G(·); given the reversibility G(·), its left inverse Defined as I represents the identity matrix; the formula of the above proximal mapping module becomes:

[0105]

[0106] Where: G(·) represents a trainable nonlinear transformation function that satisfies in: represents the left inverse of G(·), and I represents the identity matrix;

[0107] We expect that in the formula of the proximal mapping module, and H GTi The difference between will be minimized, i.e. Satisfy the following formula:

[0108]

[0109] Where α represents the predefined l2 regularization parameter; at this time, the formula of the proximal mapping module is:

[0110]

[0111] Where: θ satisfies θ=λ·α;

[0112] Combining λ and α into θ, we can get the following formula:

[0113]

[0114] Where: soft represents soft threshold;

[0115] Then through the inverse operation, we can get the following formula:

[0116]

[0117] Since G(·) and is learnable, see Figure 2 ,so It can also be expressed as the following soft threshold:

[0118]

[0119] Figure 2 Medium R (k) represents the direct reconstruction result in the kth iteration of feature extraction, H (k) represents the mapping image after the kth iteration in feature extraction;

[0120] When performing the above feature extraction, the constraint loss should satisfy the following formula:

[0121]

[0122] Where: X represents the number of frames in the entire training set; Y represents the size of the initialized image 02 obtained in step 2; m represents the mth frame of the entire training set; W represents the number of iterations during feature extraction in step 3; s m Represents the Ground Truth corresponding to the m-th frame image of the entire training set; G (k) (·) represents the k-th iteration result of G(·); express The k-th iteration result of ; represents the constraint loss;

[0123] Step 4: Position-code each frame in the S-frames in step 1, and use the following formula to obtain the time sequence information of each frame in the S-frame after coding:

[0124] B i =Sigmoid(MLP(Encoder(b i )));

[0125] Where: b i represents the position code of the i-th frame in the above S frames in step 1, b iTake the number (i-1); Encoder (·) means expanding the frame position to the batch size and normalizing it; MLP (·) means multi-layer perceptron; Sigmoid (·) means activation function; B i represents the timing information of the i-th frame in the S frames after encoding in step 1;

[0126] Then, the following formula is used to add the encoded time sequence information of each frame in the S frames to the mapped image after the Wth iteration of the feature extraction of each frame obtained in step 3, and the image 03 after adding the time sequence information of each frame is obtained:

[0127]

[0128] Where: represents the mapped image after the Wth iteration in the feature extraction of the i-th frame; ⊙ represents the Hadamard product of two matrices; H i ′ represents the image of the i-th frame after adding timing information;

[0129] Step 5: For the image 03 after adding timing information of each frame obtained in step 4, according to the 1st frame to the (2N+1)th frame, the 2nd frame to the (2N+2)th frame, ..., the (S-2N)th frame to the Sth frame, each as a sequence and its sequence order, starting from the 1st sequence, in sequence order, independently perform steps a to b on the image 03 after adding timing information of the consecutive (2N+1) frames corresponding to each sequence to obtain the feature map 05 after unmixing of the intermediate frame of the corresponding sequence; the above N is an integer, and 3≤2N+1≤S; in this embodiment, N is 2, that is, according to the 1st frame to the 5th frame, the 2nd frame to the 6th frame, ..., the 16th frame to the 20th frame, each as a sequence and its sequence order, starting from the 1st sequence, in sequence order, independently process the 5 consecutive frames corresponding to each sequence, for a total of 16 times, instead of 4 times, thereby improving data utilization and improving model performance; the above steps a to b are:

[0130] Step a: Perform convolution to obtain the intermediate frame feature map of the corresponding sequence and 2N reference frame feature maps; in this embodiment, each sequence has 4 reference frame feature maps;

[0131] Step b: Perform time-deformable alignment on the feature maps of each reference frame obtained in step a, so that the feature maps of each reference frame are aligned with the feature map of the intermediate frame obtained in step a in the spatiotemporal dimension, and obtain the mapped image 04 after the feature maps of each frame of the corresponding sequence are aligned; then perform tail convolution on the mapped image 04 after the feature maps of each frame of the corresponding sequence are aligned, and obtain the unmixed feature map 05 of the intermediate frame of the corresponding sequence, specifically:

[0132] Step b.1: The feature maps of each reference frame obtained in step a are compared with the feature map of the intermediate frame obtained in step a one by one. Through the selective attention mechanism, features useful for alignment in adjacent frames are selected to create aggregated mapping images of each reference frame feature map and the intermediate frame feature map. The selective attention mechanism adopts the following formula:

[0133]

[0134] Where:

[0135] F i When the image after adding the timing information of the i-th frame is the reference frame of the sequence in step 5, the reference frame feature map obtained after the convolution in step a is obtained, F i ∈{F t-N ,…,F t+N}, and F i ≠F t , t represents the middle frame of the sequence in step 5;

[0136] F t The image representing the middle frame of the sequence in step 5 after adding the timing information, and the feature map of the middle frame obtained after the convolution in step a;

[0137] SelectiveAttention(F i ,F t ) means reducing F by convolution first i and F t The channel dimension of the two feature maps is then connected, and the maximum pooling feature and the average pooling feature are extracted through the average maximum pooling module, and activated by the Sigmoid function to obtain the maximum pooling weight and the average pooling weight. The maximum pooling weight is then applied to the reference frame feature map, and the average pooling weight is applied to the intermediate frame feature map. Finally, the weighted features obtained from the two are connected along the channel dimension to create the reference frame feature map F i Aggregate map image with the intermediate frame feature map;

[0138] Represents the reference frame feature map F i Aggregate map image with the intermediate frame feature map;

[0139] Step b.2: For the aggregated mapping images of each reference frame feature map and the intermediate frame feature map created in step b.1, convolution is performed using the following formula to predict the sampling parameters corresponding to each reference frame feature map:

[0140]

[0141] Where: η i Represents the reference frame feature map Fi The corresponding sampling parameters are the content-related offset matrices;

[0142] Step b.3: According to the sampling parameters corresponding to each reference frame feature map predicted in step b.2, features are adaptively extracted from each reference frame feature map obtained in step a one-to-one correspondence through bilinear interpolation to align and integrate the complementary information in each reference frame feature map with the overall feature representation, so that each reference frame feature map is aligned with the intermediate frame feature map obtained in step a in the spatiotemporal dimension, and each reference frame feature map is modulated into a mapped image after the reference frame feature map is aligned. The expression is:

[0143]

[0144] Where: represents bilinear interpolation; H i ″ represents the reference frame feature map F i Aligned mapped image;

[0145] The obtained mapping image after the feature maps of each reference frame are aligned, together with the image of the intermediate frame of the corresponding sequence after adding the timing information, constitute the mapping image 04 after the feature maps of each frame of the corresponding sequence are aligned;

[0146] Step b.4: Convolution is performed on the mapped image 04 obtained in step b.3 after the alignment of the feature maps of each frame of the corresponding sequence, and V residual blocks are added to the tail convolution of the convolution structure. The above V is a natural number greater than zero. Its specific value is determined according to the design requirements, usually to make the evaluation indicator mAP (mean Average Precision) as high as possible, to obtain the feature map 05 of the middle frame of the corresponding sequence after demixing. In this embodiment, the number of residual blocks V is set to 5;

[0147] When performing the above time-deformable alignment, the alignment loss should satisfy the following formula:

[0148]

[0149] Where: H t ′ The image representing the intermediate frame corresponding to each sequence with the timing information added; represents the alignment loss;

[0150] After executing steps a to b for all sequences, the feature maps 05 after unmixing of the intermediate frames of all sequences constitute the unmixed image of the slice image 01 containing the light spot of the (S-2N) frames after removing the first and last N frames from the above S frames in step 1, and the unmixing is completed.

[0151] The present invention is based on a multi-frame spatial proximity infrared small target unmixing method driven by both model and data. Its regression loss should satisfy the following formula:

[0152]

[0153] Where: X / S represents the number of trajectories; H ji and s ji Represent the reconstruction result of the i-th frame of the j-th trajectory and its corresponding Ground Truth respectively; represents the regression loss.

[0154] See also Figure 2 At the same time, the present invention also provides a multi-frame spatially adjacent infrared small target unmixing network based on dual-drive of model and data, which is used to implement the above-mentioned multi-frame spatially adjacent infrared small target unmixing method based on dual-drive of model and data. The multi-frame spatially adjacent infrared small target unmixing network based on dual-drive of model and data comprises an initialization module 1, a feature extraction module 2, a position encoding module 3, a convolution module 4 and a time-deformable alignment module 5;

[0155] The initialization module 1 is used to initialize the slice images 01 containing the light spots of the S frames to be demixed on the same input trajectory, respectively, to obtain the initialized images 02 of each frame with an initialization super-resolution ratio of c, where c is an integer greater than 1; and S is an integer greater than or equal to 3;

[0156] The feature extraction module 2 is used to perform feature extraction on the initialized image O2 of each frame output by the initialization module 1 through W iterations, each iteration including one gradient descent module iteration and one proximal mapping module iteration, where 5≤W≤15, and W is an integer, to obtain the mapped image after the Wth iteration in the feature extraction of each frame;

[0157] The position encoding module 3 is used to perform position encoding on each frame of the S frames to be demixed on the same input trajectory, and obtain the time sequence information of each frame in the S frames after the encoding process; then the time sequence information of each frame in the S frames after the encoding process is added one by one to the mapping image after the Wth iteration of the feature extraction of each frame output by the feature extraction module 2, and obtain the image 03 of each frame after the time sequence information is added;

[0158] The convolution module 4 is used to convolve the image 03 after adding the timing information of each frame output by the position encoding module 3, according to the 1st frame to the (2N+1)th frame, the 2nd frame to the (2N+2)th frame, ..., the (S-2N)th frame to the Sth frame, starting from the 1st sequence, and independently in sequence order on the image 03 after adding the timing information of the consecutive (2N+1) frames corresponding to each sequence, to obtain the intermediate frame feature map of the corresponding sequence and 2N reference frame feature maps; the above N is an integer, and 3≤2N+1≤S;

[0159] The time deformable alignment module 5 includes a time deformable alignment block 51 and a tail convolution block 52;

[0160] The time deformable alignment block 51 is used to perform time deformable alignment on the reference frame feature maps of the corresponding sequence output by the convolution module 4, so that the reference frame feature maps of the corresponding sequence are aligned with the intermediate frame feature maps of the corresponding sequence output by the convolution module 4 in the spatiotemporal dimension, thereby obtaining a mapped image 04 after the alignment of the frame feature maps of the corresponding sequence;

[0161] The tail convolution block 52 is used to perform tail convolution on the mapped image 04 after the feature maps of each frame of the corresponding sequence output by the time deformable alignment block 51 are aligned, and obtain the feature map 05 after the intermediate frame of the corresponding sequence is unmixed; the feature map 05 after the intermediate frame of all sequences is unmixed as a whole and constitutes the unmixed image of the slice image 01 containing the light spot after removing the first and last N frames (S-2N) from the S frames to be unmixed on the same input trajectory.

[0162] In the embodiment of the multi-frame spatially adjacent infrared small target unmixing network based on dual-drive of model and data of the present invention, all convolutions involved use 32-channel convolutions with a convolution kernel of 3×3.

[0163] In order to optimize, it is preferred to use the Adam optimizer and the MMEngine framework to execute the multi-frame spatial proximity infrared small target unmixing network based on the dual-driven model and data of the present invention, and the learning rate is always kept at 10 -4 .

[0164] Figure 4 This is an example of the multi-frame spatial proximity infrared small target unmixing method (i.e., DeRefNet) based on dual-driven model and data using ISTA, ISTANet, and the present invention. The three unmixing methods unmix the slice images on the four input tracks A, B, C, and D, respectively, and the visual comparison chart with GT (the ground truth corresponding to the input slice image). Figure 4It can be seen that the demixed image obtained by the multi-frame spatial neighboring infrared small target demixing method based on dual-drive of model and data, namely DeRefNet, is closer to GT (Ground Truth corresponding to the input slice image) than the demixed image obtained by ISTA and ISTANet, and the demixing effect is better.

[0165] In summary, the present invention is based on a multi-frame spatially adjacent infrared small target unmixing method and network driven by both models and data, creatively combines the idea of depth unfolding with multi-frame deformable alignment, and applies it for the first time to the field of multi-frame spatially adjacent infrared small target unmixing. It can realize the unmixing of multi-frame spatially adjacent infrared point targets whose target space distance is less than the Rayleigh (Rayleigh spot) radius, and can meet the detection requirements of infrared point targets whose target space distance is less than the Rayleigh (Rayleigh spot) radius.

Claims

1. A multi-frame spatially adjacent infrared small target unmixing method based on dual-driven model and data, wherein the spatially adjacent infrared small target refers to an infrared point target whose target spatial distance is less than the Rayleigh radius; characterized in that: The following steps are involved: Step 1: S frames of slice images (01) containing light spots to be unmixed on the same trajectory are inputted sequentially from the earliest to the last frame in the order of the frames, where S is an integer greater than or equal to 3; Step 2: Initializing the slice images (01) of each frame input in step 1, respectively, to obtain initialized images (02) of each frame with an initialization super-resolution ratio of c, where c is an integer greater than 1; Step 3: For the initialized image (02) of each frame obtained in step 2, feature extraction is performed through W iterations, each iteration including one gradient descent module iteration and one proximal mapping module iteration, where 5≤W≤15, and W is an integer, to obtain a mapping image after the Wth iteration in the feature extraction of each frame; Step 4: Position encoding is performed on each of the S frames in step 1 to obtain the time sequence information of each frame in the S frames after the encoding process; then the time sequence information of each frame in the S frames after the encoding process is added one by one to the mapping image after the Wth iteration of the feature extraction of each frame obtained in step 3 to obtain the image of each frame after the addition of the time sequence information (03); Step 5: for the image (03) after adding timing information of each frame obtained in step 4, according to the 1st frame to the (2N+1)th frame, the 2nd frame to the (2N+2)th frame, ..., the (S-2N)th frame to the Sth frame, each as a sequence and its sequence order, starting from the 1st sequence, in sequence order, independently perform steps a to b on the image (03) after adding timing information of the continuous (2N+1) frames corresponding to each sequence, so as to obtain the feature map (05) after unmixing of the intermediate frames of the corresponding sequence, until all the sequences are performed. The feature maps (05) after unmixing of the intermediate frames of all the sequences constitute the unmixed image of the slice image (01) containing the light spot of the (S-2N)th frame after removing the first and last N frames in the S frames in step 1, thereby completing the unmixing; N is an integer, and 3≤2N+1≤S; the steps a to b are: Step a: Perform convolution to obtain the intermediate frame feature map of the corresponding sequence and 2N reference frame feature maps; Step b: performing time-deformable alignment on the reference frame feature maps obtained in step a, so that the reference frame feature maps are aligned with the intermediate frame feature maps obtained in step a in the spatiotemporal dimension, and obtaining a mapping image (04) after the alignment of the feature maps of each frame of the corresponding sequence; and then performing tail convolution on the mapping image (04) after the alignment of the feature maps of each frame of the corresponding sequence, and obtaining a feature map (05) after the unmixing of the intermediate frame of the corresponding sequence.

2. The multi-frame spatial proximity infrared small target unmixing method based on dual-drive model and data according to claim 1 is characterized by: The step b is specifically as follows: Step b.1: The reference frame feature maps obtained in step a are compared with the intermediate frame feature maps obtained in step a one by one, and the features useful for alignment in adjacent frames are selected through a selective attention mechanism to create aggregated mapping images of the reference frame feature maps and the intermediate frame feature maps. Step b.2: performing convolution on the aggregated mapping images of each reference frame feature map and the intermediate frame feature map created in step b.1, respectively, to predict sampling parameters corresponding to each reference frame feature map; Step b.3: According to the sampling parameters corresponding to the reference frame feature maps predicted in step b.2, features are adaptively extracted from the reference frame feature maps obtained in step a one-to-one by bilinear interpolation, so as to align and integrate the complementary information in each reference frame feature map with the overall feature representation, so that each reference frame feature map is aligned with the intermediate frame feature map obtained in step a in the spatiotemporal dimension, and each reference frame feature map is modulated into a mapping image after the reference frame feature map is aligned. The mapping image after the reference frame feature map is aligned and the image of the intermediate frame of the corresponding sequence after adding the timing information constitute the mapping image after the frame feature map of the corresponding sequence is aligned (04). Step b.4: The mapped image (04) obtained in step b.3 after alignment of the feature maps of each frame of the corresponding sequence is convolved with V residual blocks and a tail convolution of the convolution structure, where V is a natural number greater than zero and its specific value is determined according to the design requirements, to obtain the unmixed feature map (05) of the intermediate frame of the corresponding sequence.

3. The multi-frame spatial proximity infrared small target unmixing method based on dual-drive model and data according to claim 2 is characterized by: In step 2, the initialization is performed using the following formula: Where: Q init Represents the preset initialization matrix; L i represents the i-th frame slice image input in step 1, where i is a natural number, 1≤i≤S; Represents the image after initialization of the i-th frame.

4. The multi-frame spatial proximity infrared small target unmixing method based on dual-drive model and data according to claim 3 is characterized by: In step 3, the gradient descent module is as follows: Where: represents the mapping image after the (k-1)th iteration in the feature extraction of the i-th frame, where k is 1, 2, ..., W, and k is an integer. When k = 1, is the image of the i-th frame after initialization obtained in step 2; ρ represents the step size of gradient descent, which is obtained through learning; Φ represents the preset sampling matrix; Φ T represents the transposed matrix of Φ; express Direct reconstruction results at the kth iteration; The proximal mapping module is as follows: Where: H GTi represents the Ground Truth corresponding to the i-th frame image; G(·) represents a trainable nonlinear transformation function that satisfies in: represents the left inverse of G(·), I represents the identity matrix; θ satisfies θ=λ·α, where λ represents the predefined l1 regularization parameter and α represents the predefined l2 regularization parameter; represents the mapped image after the kth iteration in the feature extraction of the i-th frame.

5. The multi-frame spatial proximity infrared small target unmixing method based on dual-drive model and data according to claim 4 is characterized by: In step 4, position coding is performed on each frame in the S frames in step 1 to obtain the time sequence information of each frame in the S frames after coding. The following formula is used to obtain the time sequence information of each frame in the S frames after coding: B i =Sigmoid(MLP(Encoder(b i ))); Where: b i represents the position code of the i-th frame in the S frame described in step 1, b i Take the number (i-1); Encoder (·) means expanding the frame position to the batch size and normalizing it; MLP (·) means multi-layer perceptron; Sigmoid (·) means activation function; B i represents the timing information of the i-th frame in the S frames after encoding in step 1; When the encoded time sequence information of each frame in the S frames is added one by one to the mapped image after the Wth iteration in the feature extraction of each frame obtained in step 3, the following formula is used for addition: Where: represents the mapped image after the Wth iteration in the feature extraction of the i-th frame; ⊙ represents the Hadamard product of two matrices; H i ′ Indicates the image of the i-th frame after adding timing information.

6. The multi-frame spatial proximity infrared small target unmixing method based on dual-drive model and data according to claim 5 is characterized by: In step b.1, the attention selection mechanism adopts the following formula: Where: F i When the image after adding the timing information of the i-th frame is the reference frame of the sequence in step 5, the reference frame feature map obtained after the convolution in step a is obtained, F i ∈{F t-N ,…,F t+N }, and F i ≠F t , said t represents the middle frame of the sequence in step 5; F t An image representing the intermediate frame of the sequence in step 5 after adding timing information, and the intermediate frame feature map obtained after performing convolution in step a; SelectiveAttention(F i ,F t ) means reducing F by convolution first i and F t The channel dimension of the two feature maps is then connected, and the maximum pooling feature and the average pooling feature are extracted through the average maximum pooling module, and activated by the Sigmoid function to obtain the maximum pooling weight and the average pooling weight. The maximum pooling weight is then applied to the reference frame feature map, and the average pooling weight is applied to the intermediate frame feature map. Finally, the weighted features obtained from the two are connected along the channel dimension to create the reference frame feature map F i Aggregate map image with the intermediate frame feature map; Represents the reference frame feature map F i Aggregate map image with the intermediate frame feature map; In step b.2, the aggregated mapping images of each reference frame feature map created in step b.1 and the intermediate frame feature map are convolved to predict the sampling parameters corresponding to each reference frame feature map, using the following formula for convolution: Where: η i Represents the reference frame feature map F i The corresponding sampling parameters are the content-related offset matrices; In step b.3, according to the sampling parameters corresponding to each reference frame feature map predicted in step b.2, features are adaptively extracted from each reference frame feature map obtained in step a one-to-one by bilinear interpolation, so as to align and integrate the complementary information in each reference frame feature map with the overall feature representation, so that each reference frame feature map is aligned with the intermediate frame feature map obtained in step a in the spatiotemporal dimension. When each reference frame feature map is modulated into a mapped image after the reference frame feature map is aligned, its expression is: Where: represents bilinear interpolation; H i ″ represents the reference frame feature map F i Aligned mapped image.

7. The multi-frame spatial proximity infrared small target unmixing method based on dual-drive model and data according to claim 6 is characterized by: In step 3, when performing feature extraction, the constraint loss should satisfy the following formula: Where: X represents the number of frames in the entire training set; Y represents the size of the initialized image (02) obtained in step 2; m represents the mth frame of the entire training set; W represents the number of iterations during feature extraction in step 3; s m Represents the Ground Truth corresponding to the m-th frame image of the entire training set; G (k) (·) represents the k-th iteration result of G(·); express The k-th iteration result of ; represents the constraint loss; In step b, when performing the time-deformable alignment, the alignment loss should satisfy the following formula: Where: H t ′ The image representing the intermediate frame corresponding to each sequence with the timing information added; represents the alignment loss; The regression loss should satisfy the following formula: Where: X / S represents the number of trajectories; H ji and s ji Represent the reconstruction result of the i-th frame of the j-th trajectory and its corresponding Ground Truth respectively; represents the regression loss.

8. The multi-frame spatial proximity infrared small target unmixing method based on dual-drive model and data according to any one of claims 1 to 7, characterized in that: In step 1, S is equal to the number of frames of a complete trajectory.

9. A multi-frame spatially adjacent infrared small target unmixing network based on dual-model and data driving, used to implement the multi-frame spatially adjacent infrared small target unmixing method based on dual-model and data driving according to any one of claims 1 to 8, characterized in that: It includes an initialization module (1), a feature extraction module (2), a position encoding module (3), a convolution module (4) and a time-deformable alignment module (5); The initialization module (1) is used to initialize the slice images (01) containing light spots of S frames to be demixed on the same input trajectory, respectively, to obtain initialized images (02) of each frame with an initialization super-resolution ratio of c, where c is an integer greater than 1; and S is an integer greater than or equal to 3; The feature extraction module (2) is used to perform feature extraction on the initialized image (02) of each frame output by the initialization module (1) through W iterations, each iteration including one gradient descent module iteration and one proximal mapping module iteration, wherein 5≤W≤15, and W is an integer, to obtain a mapping image after the Wth iteration in the feature extraction of each frame; The position encoding module (3) is used to perform position encoding on each frame of the S frames to be demixed on the same input trajectory, thereby obtaining time sequence information of each frame in the S frames after encoding; then, the time sequence information of each frame in the S frames after encoding is added one by one to the mapped image after the Wth iteration of the feature extraction of each frame output by the feature extraction module (2), thereby obtaining an image (03) of each frame after adding the time sequence information; The convolution module (4) is used to convolve the images (03) after adding the time sequence information of each frame output by the position encoding module (3), starting from the first sequence, and independently performing convolution on the images (03) after adding the time sequence information of the consecutive (2N+1) frames corresponding to each sequence in sequence order, according to the sequence order, so as to obtain the intermediate frame feature map and 2N reference frame feature maps of the corresponding sequence; N is an integer, and 3≤2N+1≤S; The time deformable alignment module (5) includes a time deformable alignment block (51) and a tail convolution block (52); The time deformable alignment block (51) is used to perform time deformable alignment on each reference frame feature map of the corresponding sequence output by the convolution module (4), so that each reference frame feature map of the corresponding sequence is aligned with the intermediate frame feature map of the corresponding sequence output by the convolution module (4) in the spatiotemporal dimension, thereby obtaining a mapping image (04) after the alignment of each frame feature map of the corresponding sequence; The tail convolution block (52) is used to perform tail convolution on the mapped image (04) after alignment of the feature maps of each frame of the corresponding sequence output by the time deformable alignment block (51), so as to obtain the feature map (05) after unmixing of the intermediate frames of the corresponding sequence; the feature maps (05) after unmixing of the intermediate frames of all sequences constitute as a whole an unmixed image of the slice image (01) containing the light spot after removing the first and last N frames from the S frames to be unmixed on the same input trajectory.

10. The multi-frame spatial proximity infrared small target unmixing network based on dual-drive of model and data according to claim 9, characterized in that: The Adam optimizer and MMEngine framework are used for execution, and the learning rate is always kept at 10 -4 .

Citation Information

Cited By

  • Infrared small target cluster sub-pixel positioning method based on physical structure prior and deep expansion network

    CN122243934A