Infrared image super-resolution method and device based on state space model and medium
Through the state-space model method, the long-range spatiotemporal dependencies of infrared images are captured and features are enhanced, which solves the edge blur and texture distortion problems of reconstructed images in traditional methods, and realizes efficient super-resolution reconstruction of infrared images, which is suitable for real-time processing.
Patent Information
- Application Number
- CN202511263037.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-05
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-09-05
AI Technical Summary
Traditional infrared image super-resolution methods have difficulty in effectively modeling the long-range spatiotemporal continuity in infrared images, resulting in edge blur and texture distortion in the reconstructed images. In addition, the ViT-based method has high computational complexity in high-resolution image processing and is difficult to meet real-time processing requirements.
A state-space model-based approach is adopted to capture long-range dependencies through a spatial attention module, combine multi-scale convolution with state-space equations to model heat diffusion granularity features, enhance features through nonlinear gating units, and finally reconstruct through residual local feature blocks and sub-pixel convolution layers.
It effectively overcomes the problems of edge blur and texture distortion, reduces computational complexity and memory consumption, and achieves high-fidelity super-resolution reconstruction of infrared images, making it suitable for real-time application scenarios.
Smart Images

Figure CN120746839A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computers, and in particular to a method, device, and medium for infrared image super-resolution based on a state-space model. Background Art
[0002] Infrared image super-resolution technology is a research area in computer vision and is highly valuable in applications such as military reconnaissance, medical diagnosis, and security monitoring. The core goal of this technology is to improve the quality of low-resolution infrared images and achieve high-precision reconstruction, including enhancing image detail and suppressing noise.
[0003] In traditional solutions, mainstream infrared image super-resolution methods are mainly based on Convolutional Neural Network (CNN) and Vision Transformer (ViT) architectures.
[0004] However, in traditional solutions, CNN-based methods are limited by their local receptive field characteristics and are difficult to effectively model the long-range spatiotemporal continuity caused by the thermal diffusion effect in infrared images, which often leads to problems such as edge blur and texture distortion in reconstructed images. On the other hand, although the ViT-based method can capture global context information, the self-attention mechanism it relies on has The computational complexity is of the order of magnitude. In high-resolution image processing scenarios, this mechanism will lead to increased memory consumption and it is difficult to meet the application requirements of real-time processing.
[0005] At the same time, due to the significant thermal diffusion effect that is prevalent in infrared images, thermal signals exhibit continuity characteristics in the spatial dimension. Traditional schemes generally fail to model this key characteristic in a targeted manner, and therefore are prone to defects such as artifacts and loss of details in their reconstruction results. Summary of the Invention
[0006] To solve the above problems, this application proposes an infrared image super-resolution method based on a state-space model, including: The low-resolution infrared image is input into the shallow convolution layer, and the initial image features are output through feature extraction; Input the initial image features into the spatial attention module, and output fused features that incorporate long-range dependencies through key information fusion; The fused features are input into a gated state space module, the heat diffusion granularity features are captured through multi-scale convolution, and the state space equation is used to model the model and output the state space features; Inputting the state space feature into a nonlinear gating unit for feature enhancement, and outputting the state space feature after feature enhancement; The state space features after feature enhancement are used as input, and long-range context information is aggregated through multi-level efficient state space group iterative processing to output enhanced global features; The global features are input into the residual local feature block for local detail refinement, and up-sampled through a sub-pixel convolution layer to output a reconstructed high-resolution infrared image.
[0007] In one example, the initial image features are input into a spatial attention module, a query vector, a key vector, and a value vector are generated in parallel through depthwise separable convolution, and the cosine similarity between each pixel corresponding to the initial image features is calculated; Selecting key information corresponding to the highest multiple key areas according to the cosine similarity to obtain an attention weight matrix; The attention weight matrix and the value vector corresponding to the initial image features are taken as input, and through weighted aggregation, a fused feature that incorporates long-range dependencies is output.
[0008] In one example, the fused features are input into a gated state space module, and convolution operations are performed on the fused features at multiple scales to obtain corresponding intermediate features. The intermediate features corresponding to multiple scales are spliced together to obtain multi-scale fusion features as heat diffusion granularity features.
[0009] In one example, the heat diffusion granularity feature is reshaped into a feature sequence according to the size of the feature; Modeling is performed through state space equations to obtain a state space model; The feature sequence is input into the state space model, each element in the feature sequence is processed to obtain an output sequence, and the output sequence is reshaped to obtain a state space feature.
[0010] In one example, the state space feature is input into a nonlinear gating unit, and the state space feature is segmented by channel segmentation to obtain a first feature branch and a second feature branch; For the first feature branch, channel weights are learned through layer normalization and activated through Gaussian error linear units to obtain a channel-adaptive weight vector; The second feature branch and the channel-adaptive weight vector are taken as input, and channel-adaptive weighted fusion is performed through element-by-element multiplication to output feature-enhanced state space features.
[0011] In one example, the feature-enhanced state space feature is used as input, and the data processing processes corresponding to the spatial attention module, the gated state space module, and the nonlinear gating unit are sequentially performed through each level of the multi-level efficient state space group; The output of the previous level efficient state space group is used as the input of the next level efficient state space group, and the enhanced global features are output by the last level efficient state space group.
[0012] In one example, the enhanced global features are used as input residual local feature blocks, and high-frequency detail refinement and local information compensation are performed through multi-layer local convolution operations and non-linear activation functions to obtain detail-enhanced local features; The detail-enhanced local features are used as input, and 4 times upsampling resolution is reconstructed through a sub-pixel convolution layer to output a reconstructed high-resolution infrared image.
[0013] In one example, through joint optimization training, a real high-resolution infrared image is used as a supervision target, and the model parameter training and optimization of at least some modules in the shallow convolution layer, the spatial attention module, the gated state space module, the nonlinear gating unit, the multi-level efficient state space group, and the sub-pixel convolution layer are performed through the L1 loss function.
[0014] On the other hand, the present application also proposes an infrared image super-resolution device based on a state-space model, comprising: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the infrared image super-resolution method based on the state-space model as described in any of the above examples.
[0015] On the other hand, the present application also proposes a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured as the infrared image super-resolution method based on the state-space model described in any of the above examples.
[0016] The infrared image super-resolution method based on the state-space model proposed in this application can bring the following beneficial effects: By introducing the spatial attention mechanism and gated state space modeling, the long-range spatiotemporal dependencies caused by heat diffusion in infrared images are effectively captured, overcoming the edge blur and texture distortion problems caused by the local receptive field limitation of the CNN method.
[0017] By taking advantage of the linear computational complexity of the state-space model, memory consumption and computational overhead are significantly reduced, solving the bottleneck of the ViT method in real-time application in high-resolution image processing.
[0018] By collaboratively modeling the granularity characteristics of heat diffusion through multi-scale convolution and state-space equations, and combining nonlinear gating enhancement with local detail refinement, the detail reconstruction capability is significantly improved, and artifacts and detail loss are reduced.
[0019] While maintaining low computational complexity, high-fidelity super-resolution reconstruction of infrared images is achieved, which is suitable for application scenarios with high real-time requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 Schematic diagram of the process of the infrared image super-resolution method based on the state space model in an embodiment of the present application; Figure 2 Schematic diagram of an infrared image super-resolution device based on a state-space model in an embodiment of the present application. DETAILED DESCRIPTION
[0021] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0022] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.
[0023] like Figure 1 As shown, the embodiment of the present application provides an infrared image super-resolution method based on a state-space model, including: S101: Input the low-resolution infrared image into the shallow convolution layer and output the initial image features through feature extraction.
[0024] A low-resolution infrared image refers to an infrared image with a resolution lower than a preset level, which is lower than the resolution of a “high-resolution infrared image” mentioned below in this article.
[0025] Through shallow feature extraction, an initial convolution operation is performed on the input low-resolution infrared image to extract the initial image features. The initial image features can also be called the basic feature map, which can be shown as Formula 1: Formula 1; in, is a low-resolution infrared image. For convolution operation, the size of the corresponding convolution kernel is 3×3. is the initial image feature.
[0026] S102: Input the initial image features into the spatial attention module, and output fused features that incorporate long-range dependencies through key information fusion.
[0027] Specifically, the initial image features are fed into the spatial attention module, which uses depthwise separable convolution to generate query, key, and value vectors in parallel. The module then calculates the cosine similarity between the pixels corresponding to the initial image features. The spatial attention module, also known as the spatial Top-k attention module, performs dynamic attention fusion.
[0028] Based on the cosine similarity, the key information corresponding to the highest number of key areas is selected to obtain the attention weight matrix. The top k key areas (key heat signal areas) are dynamically selected as the corresponding key information to reduce redundant calculations, and the corresponding feature map is output through weighted aggregation, as shown in Formula 2 and Formula 3: Formula 2; Among them, M is the attention weight matrix, It is a depth-separable convolution operation, and the size of the corresponding convolution kernel is 3×3. is the initial image feature, T is the matrix transpose operation, is the key vector.
[0029] Formula 3; in, is the attention weight matrix after sparseness, is the element in row i and column j of the sparse attention weight matrix, is the element in row i and column j in the attention weight matrix, Sort all elements in the jth column of the attention weight matrix M and select the set of top k values with the largest values.
[0030] The attention weight matrix and the value vector corresponding to the initial image features are taken as input. Through weighted aggregation, the output is a fusion feature that integrates long-range dependencies. The dynamic nature of the spatial attention module can focus on key information interactions and improve computational efficiency, as shown in Formula 4: Formula 4; in, To fusion features, Is the normalization function, the sparse attention weight matrix Normalize each column of so that the sum of the weights of all non-zero elements in each column is 1; It is a convolution operation, and the size of the corresponding convolution kernel is 1×1. Perform convolution operation to generate a value vector.
[0031] S103: Input the fusion features into a gated state space module, capture the heat diffusion granularity features through multi-scale convolution, model them through state space equations, and output state space features.
[0032] Specifically, when capturing the heat diffusion granularity feature, the fused feature is input into the gated state space module. Through convolution operations at multiple scales, the fused feature is convolved separately to obtain the corresponding intermediate features. The intermediate features corresponding to multiple scales are spliced to obtain the multi-scale fused feature as the heat diffusion granularity feature, which can be shown as Formula 5: Formula 5; in, To fusion features, 、 、 All are convolution operations, and the corresponding convolution kernel sizes are 1×1, 3×3, and 5×5; For splicing operation, It is a multi-scale fusion feature.
[0033] During modeling, the heat diffusion granularity features are reshaped into feature sequences according to their size; the heat diffusion granularity features are reshaped into feature sequences. (The size is [H, W, C], i.e., height, width, and number of channels) reshaped into a feature sequence, flattening the spatial dimension to obtain a sequence of length L = H * W, and the feature dimension of each element in the sequence is C.
[0034] By modeling the state space equation, the state space model is obtained, which can be shown as Formula 6 and Formula 7: Formula 6; Formula 7; Where t is the discrete time step, which corresponds to the sequence position after flattening the image feature map in image processing. is the input at time t, corresponding to the multi-scale fusion feature The feature vector at sequence position t; is the hidden state of the system at time t, which contains the memory of all historical information from the beginning of the sequence to time t; is the hidden state of the system at time t+1; is the output at time t, corresponding to the state space feature; A, B, C, and D are the state transfer matrix, input matrix, output matrix, and feedforward matrix, respectively, and are all learnable parameter matrices.
[0035] The feature sequence is input into the state space model, each element in the feature sequence is processed to obtain the output sequence, and the output sequence is reshaped to obtain the state space feature.
[0036] When all elements in the sequence have been processed, an output sequence of the same length L is obtained. The output sequence is then reshaped back to the original spatial size [H, W, C'], and the reshaped feature map is the state space feature.
[0037] S104: Inputting the state space feature into a nonlinear gating unit for feature enhancement, and outputting the state space feature after feature enhancement.
[0038] Specifically, the state space feature is input into the nonlinear gating unit, and the state space feature is segmented by channel segmentation to obtain the first feature branch and the second feature branch, which can be shown as Formula 8: Formula 8; in, 、 They are the first feature branch and the second feature branch respectively. These two branches have the same spatial size (height, width), but the number of channels is half of the state space feature y(t); is the channel splitting operation, is the state space feature.
[0039] For the first feature branch, the channel weight is learned by layer normalization and activated by the Gaussian error linear unit to obtain the channel-adaptive weight vector. The second feature branch and the channel-adaptive weight vector are used as input, and the channel-adaptive weighted fusion is performed by element-by-element multiplication. The state space features after feature enhancement are output. The channel-adaptive weighted fusion is achieved by element-by-element multiplication, highlighting important channel features, suppressing noise, further enhancing the model's ability to characterize infrared thermal dynamic characteristics, and enhancing thermal dynamic adaptability. It can be shown as Formula 9: Formula 9; in, is the state space feature after feature enhancement, For the layer normalization operation, the first feature branch Each eigenvector in is standardized so that its mean is 0 and its variance is 1; is the Gaussian error linear unit activation function, which introduces a smooth nonlinear transformation for the normalized features; It is element-by-element multiplication, which means multiplying the elements of corresponding positions of two matrices or tensors of the same dimension; The whole constitutes an adaptive gating signal. This signal contains the nonlinear transformation from Contextual information about the branch.
[0040] S105: taking the state space features after feature enhancement as input, aggregating long-range context information through multi-level efficient state space group iterative processing, and outputting enhanced global features.
[0041] Specifically, the state space features after feature enhancement are used as input, and the data processing processes corresponding to the spatial attention module, the gated state space module, and the nonlinear gated unit are executed in sequence through each level of the multi-level efficient state space group. The output of the previous level of efficient state space group is used as the input of the next level of efficient state space group, and the enhanced global features are output by the last level of efficient state space group.
[0042] Furthermore, taking the four-level efficient state space group as an example, it can be expressed as Formula 10: Formula 10; in, To enhance the global features, is the state space feature after feature enhancement, The system is a four-dimensional efficient state space group. Within each level of the efficient state space group, the corresponding data processing steps of the spatial attention module, the gated state space module, and the nonlinear gating unit are repeatedly executed. Through skip connections or feature fusion mechanisms between layers, it iteratively aggregates and deepens long-range contextual information. Lower layers capture basic dependencies, while higher layers integrate more abstract and profound contextual information, gradually improving understanding of complex heat diffusion patterns.
[0043] S106: Input the global features into the residual local feature block to refine the local details, and perform upsampling through a sub-pixel convolution layer to output a reconstructed high-resolution infrared image.
[0044] Specifically, the enhanced global features are used as the input residual local feature blocks, and high-frequency details are refined and local information (including high-frequency details such as image edges and textures) is compensated through multi-layer local convolution operations and non-linear activation functions to obtain detail-enhanced local features; the detail-enhanced local features are used as input and reconstructed with a 4x upsampling resolution through a sub-pixel convolution layer, and the reconstructed high-resolution infrared image with enhanced details and suppressed artifacts is output.
[0045] High-frequency detail refinement and local information compensation can be expressed as Formula 11: Formula 11; in, For local features of detail enhancement, For convolution operation, the size of the corresponding convolution kernel is 3×3. is the rectified linear unit activation function, For enhanced global features.
[0046] The up-sampled resolution reconstruction can be expressed as Formula 12: Formula 12; in, It is a high-resolution infrared image; For pixel reorganization operation, amplification is achieved through convolution + periodic screening. Assuming the target is 4 times upsampling, the number of channels is first expanded to times, and then rearrange the feature maps according to specific rules, so that the spatial size (H, W) is expanded by 4 times, while the number of channels is reduced to 1 / 16 of the original; For convolution operation, the size of the corresponding convolution kernel is 3×3.
[0047] By introducing the spatial attention mechanism and gated state space modeling, the long-range spatiotemporal dependencies caused by heat diffusion in infrared images are effectively captured, overcoming the edge blur and texture distortion problems caused by the local receptive field limitation of the CNN method.
[0048] By taking advantage of the linear computational complexity of the state-space model, memory consumption and computational overhead are significantly reduced, solving the bottleneck of the ViT method in real-time application in high-resolution image processing.
[0049] By collaboratively modeling the granularity characteristics of heat diffusion through multi-scale convolution and state-space equations, and combining nonlinear gating enhancement with local detail refinement, the detail reconstruction capability is significantly improved, and artifacts and detail loss are reduced.
[0050] While maintaining low computational complexity, high-fidelity super-resolution reconstruction of infrared images is achieved, which is suitable for application scenarios with high real-time requirements.
[0051] In one embodiment, during model training, a real high-resolution infrared image is used as a supervision target through joint optimization training, and the model parameters of at least some modules in the shallow convolution layer, spatial attention module, gated state space module, nonlinear gated unit, multi-level efficient state space group, and sub-pixel convolution layer are optimized through L1 loss function to generate high-resolution infrared image output with enhanced details and suppressed artifacts. The loss function can be shown as Formula 13: Formula XIII; in, To reconstruct a high-resolution infrared image, For real high-resolution infrared images, is the L1 norm.
[0052] like Figure 2 As shown, the embodiment of the present application further proposes an infrared image super-resolution device based on a state-space model, comprising: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the infrared image super-resolution method based on the state-space model as described in any of the above embodiments.
[0053] The embodiment of the present application further proposes a non-volatile computer storage medium storing computer executable instructions, wherein the computer executable instructions are configured as the infrared image super-resolution method based on the state-space model described in any of the above embodiments.
[0054] The various embodiments in this application are described in a progressive manner. Similar portions between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the device and medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simple. For relevant portions, refer to the descriptions of the method embodiments.
[0055] The devices and media provided in the embodiments of the present application correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects to their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.
[0056] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0057] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0058] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0059] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0060] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0061] Memory may include non-permanent storage in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0062] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0063] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0064] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A super-resolution method for infrared images based on a state-space model, characterized in that: include: The low-resolution infrared image is input into the shallow convolution layer, and the initial image features are output through feature extraction; Input the initial image features into the spatial attention module, and output fused features that incorporate long-range dependencies through key information fusion; The fused features are input into a gated state space module, the heat diffusion granularity features are captured through multi-scale convolution, and the state space equation is used to model the model and output the state space features; Inputting the state space feature into a nonlinear gating unit for feature enhancement, and outputting the state space feature after feature enhancement; The state space features after feature enhancement are used as input, and long-range context information is aggregated through multi-level efficient state space group iterative processing to output enhanced global features; The global features are input into the residual local feature block for local detail refinement, and up-sampled through a sub-pixel convolution layer to output a reconstructed high-resolution infrared image.
2. The method according to claim 1, characterized in that The initial image features are input into the spatial attention module, and through key information fusion, the fused features that integrate long-range dependencies are output, specifically including: Input the initial image features into the spatial attention module, generate the query vector, key vector and value vector in parallel through depthwise separable convolution, and calculate the cosine similarity between the pixels corresponding to the initial image features; Selecting key information corresponding to the highest multiple key areas according to the cosine similarity to obtain an attention weight matrix; The attention weight matrix and the value vector corresponding to the initial image features are taken as input, and through weighted aggregation, a fused feature that incorporates long-range dependencies is output.
3. The method according to claim 1, characterized in that The fused features are input into the gated state space module to capture the heat diffusion granularity features through multi-scale convolution, specifically including: Input the fused features into the gated state space module, and perform convolution operations on the fused features at multiple scales to obtain corresponding intermediate features; The intermediate features corresponding to multiple scales are spliced together to obtain multi-scale fusion features as heat diffusion granularity features.
4. The method according to claim 3, characterized in that Modeling is performed through state space equations to output state space features, including: reshape into a feature sequence according to the size of the thermal diffusion granularity feature; Modeling is performed through state space equations to obtain a state space model; The feature sequence is input into the state space model, each element in the feature sequence is processed to obtain an output sequence, and the output sequence is reshaped to obtain a state space feature.
5. The method according to claim 1, wherein Inputting the state space feature into a nonlinear gating unit for feature enhancement, and outputting the state space feature after feature enhancement, specifically comprising: Inputting the state space feature into a nonlinear gating unit, and segmenting the state space feature by channel segmentation to obtain a first feature branch and a second feature branch; For the first feature branch, channel weights are learned through layer normalization and activated through Gaussian error linear units to obtain a channel-adaptive weight vector; The second feature branch and the channel-adaptive weight vector are taken as input, and channel-adaptive weighted fusion is performed through element-by-element multiplication to output feature-enhanced state space features.
6. The method according to claim 5, characterized in that The state space features after feature enhancement are used as input, and multi-level efficient state space group iterative processing is performed to aggregate long-range context information and output enhanced global features, specifically including: The state space feature after feature enhancement is used as input, and the data processing processes corresponding to the spatial attention module, the gated state space module, and the nonlinear gating unit are sequentially executed through each level of the multi-level efficient state space group; The output of the previous level efficient state space group is used as the input of the next level efficient state space group, and the enhanced global features are output by the last level efficient state space group.
7. The method according to claim 1, characterized in that The global features are input into the residual local feature block for local detail refinement, and up-sampled through the sub-pixel convolution layer to output a reconstructed high-resolution infrared image, specifically including: The enhanced global features are used as input residual local feature blocks, and high-frequency detail refinement and local information compensation are performed through multi-layer local convolution operations and non-linear activation functions to obtain detail-enhanced local features; The detail-enhanced local features are used as input, and 4 times upsampling resolution is reconstructed through a sub-pixel convolution layer to output a reconstructed high-resolution infrared image.
8. The method according to claim 1, characterized in that The method further comprises: Through joint optimization training, real high-resolution infrared images are used as supervision targets, and the model parameter training and optimization of at least some modules in the shallow convolution layer, the spatial attention module, the gated state space module, the nonlinear gating unit, the multi-level efficient state space group, and the sub-pixel convolution layer are performed through the L1 loss function.
9. An infrared image super-resolution device based on a state-space model, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the infrared image super-resolution method based on the state-space model as described in any one of claims 1 to 8.
10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are configured as the infrared image super-resolution method based on a state-space model as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Super-resolution reconstruction model, method and device for efficient multi-attention feature fusion and storage medium
CN115660955A
Infrared image super-resolution reconstruction method based on edge enhancement
CN116071243A
Diffusion model, multi-scale and attention module medical ultrasonic image segmentation method
CN119180826A
Remote sensing image super-resolution reconstruction method and system based on prior diffusion model
CN119251054A
Remote sensing image building extraction method fusing double-space attention features
CN120198800A
Cited By
Extra-high pressure valve hall-oriented multispectral collaborative perception and adaptive diagnosis method and equipment
CN121679247A
Multi-spectral cooperative sensing and adaptive diagnosis method and device for extra-high voltage valve hall
CN121679247B
Image feature enhancement method and system, medium, equipment and program product
CN122199290A