An infrared image super-resolution reconstruction method and system

By combining the progressive residual with the edge method, the infrared image features are gradually extracted and fused, which solves the problem of low resolution of infrared images and achieves higher-precision image reconstruction.

CN120318075BActive Publication Date: 2025-10-10STATE GRID JIANGXI ELECTRIC POWER CO LTD RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510772875.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-10-10
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

The resolution of infrared images is generally low, and existing super-resolution reconstruction technology is difficult to effectively integrate edge features, resulting in insufficient reconstruction accuracy.

Method used

A progressive residual and edge combination method is adopted to gradually reduce feature redundancy, fuse edge features and improve reconstruction accuracy through shallow and deep feature extraction modules, recursive structure modules and upsampling networks.

Benefits of technology

The accuracy of super-resolution reconstruction of infrared images is improved and the reconstruction effect of edge details is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318075B_ABST
    Figure CN120318075B_ABST
Patent Text Reader

Abstract

The application discloses an infrared image super-resolution reconstruction method and system. The method comprises the following steps: performing first feature extraction on the low-resolution image according to a preset shallow feature extraction module to obtain shallow features; performing second feature extraction on the low-resolution image according to a preset deep feature extraction module to obtain deep features, wherein the deep feature extraction module comprises a plurality of progressive residual edge blocks connected in sequence and a fusion module connected with each progressive residual edge block; superimposing the shallow features into the deep features processed by convolution according to a preset recursive structure module, and performing up-sampling on the superimposed result based on a preset up-sampling network to obtain enlarged features; and performing convolution operation on the enlarged features to obtain a reconstructed high-resolution image. The method can improve the accuracy of infrared image super-resolution reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of image processing, and particularly relates to an infrared image super-resolution reconstruction method and system. BACKGROUND

[0002] The object surface temperature reflected by an infrared image is a single-layer image, and the infrared image resolution is generally low due to the hardware limitations such as imaging components.

[0003] The infrared image lacks the description of object texture details, but contains more edges formed by temperature differences between regions, so the edges are important detail features of the infrared image and should be focused on in the super-resolution reconstruction. SUMMARY

[0004] The application provides an infrared image super-resolution reconstruction system to solve the technical problem of low accuracy of the existing super-resolution reconstruction.

[0005] In a first aspect, an embodiment of the application provides an infrared image super-resolution reconstruction method combining progressive residual and edges, and the method comprises the following steps.

[0006] A low-resolution image is obtained, a first feature extraction is performed on the low-resolution image according to a preset shallow feature extraction module, and a shallow feature is obtained , and the expression is as follows:

[0007] ,

[0008] In the expression, is a feature extraction operation, is the input low-resolution image, is the shallow feature.

[0009] A second feature extraction is performed on the low-resolution image according to a preset deep feature extraction module, and a deep feature is obtained , wherein the deep feature extraction module comprises a plurality of progressive residual edge blocks connected in sequence and a fusion module connected with each progressive residual edge block.

[0010] The expression of the deep feature is as follows:

[0011] ,

[0012] ,

[0013] In the expression, is the feature extraction operation of the nth progressive residual edge block, is the feature extraction operation of the n-1th progressive residual edge block, is the feature extraction operation of the first progressive residual edge block, is the feature extraction result of the nth progressive residual edge block, is the number of progressive residual edge blocks, is the fusion function, For deep features;

[0014] According to the preset recursive structure module, the shallow features Superimposed on the deep features after convolution processing The superimposed result is upsampled based on the preset upsampling network to obtain the amplified features. , the expression is:

[0015] ,

[0016] Where, It is an upsampling network, including convolutional layers and pixel shuffling layers. Extract features for the third layer of the recursive block;

[0017] The magnified features Perform convolution operation to obtain the reconstructed high-resolution image , the expression is:

[0018] ,

[0019] Where, The convolution operation with kernel size 3×3.

[0020] In a second aspect, an embodiment of the present invention provides an infrared image super-resolution reconstruction system combining progressive residual and edge, the system comprising:

[0021] The first extraction unit is configured to obtain a low-resolution image, perform first feature extraction on the low-resolution image according to a preset shallow feature extraction module, and obtain shallow features. , the expression is:

[0022] ,

[0023] Where, is the feature extraction operation, is the input low-resolution image, It is a shallow feature;

[0024] The second extraction unit is configured to perform a second feature extraction on the low-resolution image according to a preset deep feature extraction module to obtain a deep feature. , wherein the deep feature extraction module includes a plurality of sequentially connected progressive residual edge blocks and a fusion module connected to each progressive residual edge block;

[0025] The deep features The expression is:

[0026] ,

[0027] ,

[0028] Where, is the feature extraction operation of the nth progressive residual edge block, is the feature extraction operation of the n-1th progressive residual edge block, is the feature extraction operation of the first progressive residual edge block, is the feature extraction result of the nth progressive residual edge block, is the number of progressive residual edge blocks, is the fusion function, For deep features;

[0029] The superposition unit is configured to superimpose shallow features according to a preset recursive structure module Superimposed on the deep features after convolution processing The superimposed result is upsampled based on the preset upsampling network to obtain the amplified features. , the expression is:

[0030] ,

[0031] Where, It is an upsampling network, including convolutional layers and pixel shuffling layers. Extract features for the third layer of the recursive block;

[0032] The reconstruction unit is configured to reconstruct the magnified features Perform convolution operation to obtain the reconstructed high-resolution image , the expression is:

[0033] ,

[0034] Where, The convolution operation with kernel size 3×3.

[0035] In a third aspect, an electronic device is provided, which includes at least one processor, and a memory connected with the at least one processor in communication, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the multi-intent recognition training or using method of any of the embodiments of the present application.

[0036] In a fourth aspect, the embodiments of the present application further provide a computer readable storage medium, which stores a computer program, and the computer program includes program instructions, and the program instructions, when executed by a computer, enable the computer to perform the steps of the multi-intent recognition training or using method of any of the embodiments of the present application.

[0037] The infrared image super-resolution reconstruction method and system of the present application designs a progressive residual edge block, reduces feature redundancy through multi-module down-sampling and up-sampling operations, and at the same time, in order to reduce the loss of edge details, the progressive edge features are integrated, which can improve the accuracy of infrared image super-resolution reconstruction. BRIEF DESCRIPTION OF DRAWINGS

[0038] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0039] Figure 1 A flowchart of an infrared image super-resolution reconstruction method provided by an embodiment of the present application is provided.

[0040] Figure 2 A structural block diagram of an infrared image super-resolution reconstruction method provided by an embodiment of the present application is provided.

[0041] Figure 3 A flowchart of a group enhancement convolution module provided by an embodiment of the present application is provided.

[0042] Figure 4 A flowchart of a token dictionary cross-attention block provided by an embodiment of the present application is provided.

[0043] Figure 5 A flowchart of adaptive class attention provided by an embodiment of the present application is provided.

[0044] Figure 6 A flowchart of a progressive residual edge module provided by an embodiment of the present application is provided.

[0045] Figure 7 A flowchart of a residual learning structure provided by an embodiment of the present application is provided.

[0046] Figure 8 A structural block diagram of an edge feature extraction submodule and an intermediate processing submodule provided in one embodiment of the present invention;

[0047] Figure 9 A flowchart of a recursive structure module provided by an embodiment of the present invention

[0048] Figure 10 A structural block diagram of an infrared image super-resolution reconstruction system provided by one embodiment of the present invention;

[0049] Figure 11 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0051] See also Figure 1 , which shows a flowchart of an infrared image super-resolution reconstruction method of the present application.

[0052] like Figure 1 As shown, the infrared image super-resolution reconstruction method specifically includes the following steps:

[0053] Step S101: obtain a low-resolution image, perform a first feature extraction on the low-resolution image according to a preset shallow feature extraction module, and obtain shallow features. .

[0054] In this step, shallow features The expression is:

[0055] ,

[0056] Where, is the feature extraction operation, is the input low-resolution image, It is a shallow feature.

[0057] Specifically, the shallow feature extraction module includes a group enhancement convolution module, which includes a convolution layer, layer, Layers, depth-wise separable convolutional layers, and channel shuffle operations;

[0058] According to the convolution layer, the low-resolution image Perform convolution operation to obtain features with 64 channels, and divide the obtained features into two groups according to 1 / 4 and 3 / 4 of the number of feature channels;

[0059] The convolution kernel is 3×3 Layer, processing the 1 / 4 group of features, the number of input channels is 16, the number of output channels is 16, and the convolution kernel is 3×3 Layer, processing the 3rd / 4th group of features, the number of input channels is 48, and the number of output channels is 48;

[0060] Will The output of the layer is used as the next Layer and The input of the layer is further split into two groups, 1 / 4 and 3 / 4 of the number of feature channels. A channel shuffle and a depthwise separable convolution layer are applied to the feature branches in the 1 / 4 group. The purpose of channel shuffling is to enable more information interaction between different channels, thereby improving feature representation. The depthwise separable convolution layer preserves spatial information and ignores information interaction between channels. This achieves clear spatial and channel-wise feature representation.

[0061] The features extracted by the group-enhanced convolution module can provide complementary wide information for the deep feature extraction part, thereby improving the performance of super-resolution reconstruction. The expression of the group-enhanced convolution module is:

[0062] ,

[0063] ,

[0064] ,

[0065] ,

[0066] ,

[0067] ,

[0068] ,

[0069] ,

[0070] ,

[0071] ,

[0072] ,

[0073] Where, is a convolution operation, is a depthwise separable convolution layer, is a channel shuffle operation, is a Relu activation function operation, is a token dictionary cross attention block, is an adaptive class attention, is the first time the 1 / 4 group feature map processing result, is the n-th time the 3 / 4 group feature map processing result, is a concatenation operation, represents taking the current input feature / 4 group channel number, Figure 1 represents taking the current input feature / 4 group channel number, Figure 3 is a shallow feature.

[0074] Among them, in order to promote the fusion of features at different stages, a token dictionary cross attention block and an adaptive class attention are designed.

[0075] In the process of specific embodiments, as shown in Figure 4 , compared with the existing multi-head self-attention (MSA) algorithm which generates query, key and value tokens from the input feature itself, an additional dictionary D is introduced, which summarizes external prior knowledge in the training stage, and uses the learned token dictionary D to generate Key dictionary and Value dictionary , and uses the input feature X to generate query token

[0076] ,

[0077] ,

[0078] ,

[0079] The three transformations can be regarded as a kind of re-encoding of the original input, the purpose of which is to extract information from different angles so as to capture the correlation between elements more effectively when calculating the attention score subsequently.

[0080] Then, the query token is enhanced by cross attention calculation using the key dictionary and the value dictionary, and the expression is:

[0081] ,

[0082] ,

[0083] ​Normalized cosine distance is used instead of dot product operation because it is hoped that each token in the dictionary has a similarity map S that is converted into an attention map A using the SoftMax function for subsequent calculations.

[0084] The above TDCA operation first selects similar tokens in the keyword dictionary and obtains the attention map. Then TDCA uses the similarity values ​​to jointly combine the corresponding tokens in the value dictionary, which is equivalent to reconstructing the HR block using HR dictionary atoms and representation coefficients. TDCA can embed external priors into the learned dictionary to enhance the input image features.

[0085] Since the convolution operation projects each layer of image features into a different feature space, it is necessary to learn a different token dictionary to provide each layer with an external prior for each specific feature space. This will result in a large number of additional parameters. Therefore, an adaptive refinement strategy is designed, which refines the token dictionary of the previous layer based on the similarity graph and the features updated in the form of opposite attention. In order to introduce the proposed adaptive refinement strategy, the layer index is set for the input features and token dictionary. ,Right now and Respectively represent The input features and token dictionaries of the layer. In the following layers, each dictionary Based on the For the first Layer dictionary { For each marker in}(i=1,...,M), select the enhanced feature More specifically, Represented as an attention map The i-th column of with all N query tags Therefore, based on each , you can choose the corresponding enhancement token To reconstruct the new token dictionary element , and combine them to form :

[0086] ,

[0087] ,

[0088] Where, The attention map generated in the cross-attention block for the token dictionary, Cross attention blocks for token dictionary, It means calculating the cosine similarity between two tokens. SoftMax is a SoftMax function for adjusting the range of similarity values, , and are linear transformations of query tokens, key dictionary tokens and value dictionary tokens respectively, X is the input feature, is the query token, is the dictionary, is the Key dictionary, is the Value dictionary, is the l+1th index layer of the dictionary, is a learnable parameter, is the re-probability distribution of the lth layer dictionary, is a normalization layer for adjusting the range of attention maps, is the enhanced feature of the l+1th layer, is the transpose matrix of the lth layer attention map.

[0089] The above adaptive refinement strategy starts from the initial feature dictionary introduced from the external prior information , and gradually selects relevant features from the whole image to refine the dictionary. The refined dictionary can span the boundary of the self-attention window and summarize the typical local structure of the whole image, thereby improving the image features using global information.

[0090] Due to the quadratic computational complexity of self-attention, most existing methods must divide the input features into rectangular windows before performing attention. This window-based attention calculation severely limits the range of the receptive field. In addition, this content-independent division strategy may cause irrelevant tokens to be grouped into the same window, potentially affecting the accuracy of the attention map. In order to better utilize self-attention, adaptive feature segmentation is a more suitable choice. In the specific embodiment process, as shown in Figure 5 , since the attention mapping between the input features obtained by TDCA and the token dictionary implicitly contains the class information of each pixel, the input features can be classified. According to the maximum case of the dictionary token, each pixel is classified into various categories , ,… .

[0091] The expression of adaptive class attention is:

[0092] ,

[0093] In the formula, is the class of pixel classification, is the pixel value, is a function in Numpy, is the jth row and kth column of the attention map, For the i-th category.

[0094] if exist , ,…, The highest pixel Will be classified as , which shows that More likely to be associated with the i-th token in the dictionary belong to the same class. Therefore, each class can be perceived as an irregularly shaped window containing labels of the same class.

[0095] Step S102: Perform a second feature extraction on the low-resolution image according to a preset deep feature extraction module to obtain a deep feature , wherein the deep feature extraction module includes a plurality of sequentially connected progressive residual edge blocks and a fusion module connected to each progressive residual edge block.

[0096] In this step, deep features The expression is:

[0097] ,

[0098] ,

[0099] Where, is the feature extraction operation of the nth progressive residual edge block, is the feature extraction operation of the n-1th progressive residual edge block, is the feature extraction operation of the first progressive residual edge block, is the feature extraction result of the nth progressive residual edge block, is the number of progressive residual edge blocks, is the fusion function, For deep features;

[0100] The progressive residual edge module in the deep feature extraction part adopts the U-Net structure with fused edge. The U-Net structure includes 6 residual learning structures (RLA), 3 edge feature extraction submodules and 2 intermediate processing submodules. is the input of the progressive residual edge module, is the output of the progressive residual edge module. Figure 7 The residual learning architecture (RLA) is recorded from left to right as ,Will and The result is downsampled as and The input, and 、 and The process of integrating edge features into the feature map can be described as:

[0101] ,

[0102] Where, It is the fusion result of edge features and feature maps of residual learning structure. It is The output features of the edge feature extraction submodule. 、 and As 、 and Input, 、 The result is upsampled as and input.

[0103] Progressive residual edge modules (such as Figure 6 ) includes a first edge feature extraction submodule, a first intermediate processing submodule, a second edge feature extraction submodule, a second intermediate processing submodule, a third edge feature extraction submodule and a residual learning structure;

[0104] The first edge feature extraction submodule, the second edge feature extraction submodule, and the third edge feature extraction submodule are all connected to a 2×2 maximum pooling layer to form a downsampling mode;

[0105] The residual learning structure (such as Figure 7 As shown in Figure 2, it contains three residual units and a fusion layer. Each residual unit contains convolution, activation function, and convolution in sequence.

[0106] Among them, the first edge feature extraction submodule includes two convolutions with a kernel size of 3×3 and a channel number c of 64. Each convolution is followed by a convolution with a kernel size of 1×1 and a channel number of 21 for dimensionality increase. The two convolution results are added in parallel and then reduced in dimension by a convolution with a kernel size of 1×1 and a channel number of 1. The result after dimensionality reduction is used as the first output result of the edge feature extraction submodule. ;

[0107] First output result The data is input to the first intermediate processing submodule, which includes two convolutions with a kernel size of 3×3 and a channel number c of 64, which are serialized and then subjected to a 2×2 maximum pooling. The output of the first intermediate processing submodule serves as the input of the second edge feature extraction submodule.

[0108] The second edge feature extraction submodule includes two convolutions with a kernel size of 3×3 and a channel number c of 128. Each convolution is followed by a convolution with a kernel size of 1×1 and a channel number of 21 for dimensionality increase. The two convolution results are added in parallel and then reduced by a convolution with a kernel size of 1×1 and a channel number of 1. The result of dimensionality reduction is used as the second output result of the second edge feature extraction submodule. ;

[0109] Second output result The data is input to the second intermediate processing submodule, which includes two convolutions with a kernel size of 3×3 and a channel number c of 64, followed by a 2×2 maximum pooling. The output of the second intermediate processing submodule serves as the input of the third edge feature extraction submodule.

[0110] The third edge feature extraction submodule includes three convolutions with a kernel size of 3×3 and a channel number c of 256. Each convolution is followed by a convolution with a kernel size of 1×1 and a channel number of 21 for dimensionality increase. The two convolution results are added in parallel and then reduced by a convolution with a kernel size of 1×1 and a channel number of 1. The result of dimensionality reduction is used as the third output result of the third edge feature extraction submodule. , the expression is:

[0111] ,

[0112] ,

[0113] ,

[0114] Where, Represents low-resolution images Perform bicubic interpolation. is the result of bicubic interpolation processing, is the first edge feature extraction submodule, is the first output result of the first edge feature extraction submodule, For the Intermediate processing submodule, For the edge feature extraction submodule, It is the first output result of the j-th edge feature extraction submodule.

[0115] Among them, the three edge feature extraction submodules and the two intermediate processing submodules have the same structure (such as Figure 8 shown).

[0116] Step S103: The shallow features are transformed into Superimposed on the deep features after convolution processing The superimposed result is upsampled based on the preset upsampling network to obtain the amplified features. .

[0117] In this step, the enlarged features The expression is:

[0118] ,

[0119] Where, It is an upsampling network, including convolutional layers and pixel shuffling layers. Extract features for the 3rd layer of the recursive block.

[0120] Specifically, the recursive structure modules of the reconstruction part (such as Figure 9 (as shown in Figure 2) superimposes shallow features onto deep features after convolution. The recursive structure module uses a recursive learning approach to stack multiple residual units in a residual branch to achieve a global residual learning effect. The recursive block in the recursive structure module is a residual unit, which includes convolution, normalization, and activation functions. During the recursive block recursion, a multi-path structure is used, and all residual units share the same branch input, and the weight set is shared among these residual units. The expression of the recursive structure module is:

[0121] ,

[0122] ,

[0123] ,

[0124] ,

[0125] Where, is a shallow auxiliary feature, Extract features for the first layer of the recursive block, Extract features for the second layer of the recursive block, Extract features for the 3rd layer of the recursive block, is the convolution operation, for Activation function operation, is the batch normalization layer, For the dropout operation, It is a linear transformation operation that converts the output into a linear structure so that the model can better fit the training data.

[0126] Step S104: amplify the features Perform convolution operation to obtain the reconstructed high-resolution image , the expression is:

[0127] ,

[0128] Where, The convolution operation with kernel size 3×3.

[0129] In summary, the method of the present application designs a progressive residual edge block, and reduces feature redundancy in multiple modules by gradually utilizing downsampling and upsampling operations. At the same time, in order to reduce the loss of edge details, the gradient edge feature is incorporated, which can improve the accuracy of super-resolution reconstruction of infrared images.

[0130] See also Figure 10 , which shows a structural block diagram of an infrared image super-resolution reconstruction system of the present application.

[0131] like Figure 10 As shown, the infrared image super-resolution reconstruction system 200 includes a first extraction unit 210 , a second extraction unit 220 , a superposition unit 230 and a reconstruction unit 240 .

[0132] The first extraction unit 210 is configured to obtain a low-resolution image, perform first feature extraction on the low-resolution image according to a preset shallow feature extraction module, and obtain shallow features. , the expression is:

[0133] ,

[0134] Where, is the feature extraction operation, is the input low-resolution image, It is a shallow feature;

[0135] The second extraction unit 220 is configured to perform a second feature extraction on the low-resolution image according to a preset deep feature extraction module to obtain a deep feature. , wherein the deep feature extraction module includes a plurality of sequentially connected progressive residual edge blocks and a fusion module connected to each progressive residual edge block;

[0136] The deep features The expression is:

[0137] ,

[0138] ,

[0139] Where, is the feature extraction operation of the nth progressive residual edge block, is the feature extraction operation of the n-1th progressive residual edge block, is the feature extraction operation of the first progressive residual edge block, is the feature extraction result of the nth progressive residual edge block, is the number of progressive residual edge blocks, is the fusion function, For deep features;

[0140] The superposition unit 230 is configured to superimpose shallow features according to a preset recursive structure module. Superimposed on the deep features after convolution processing The superimposed result is upsampled based on the preset upsampling network to obtain the amplified features. , the expression is:

[0141] ,

[0142] Where, It is an upsampling network, including convolutional layers and pixel shuffling layers. Extract features for the third layer of the recursive block;

[0143] The reconstruction unit 240 is configured to reconstruct the magnified features Perform convolution operation to obtain the reconstructed high-resolution image , the expression is:

[0144] ,

[0145] Where, The convolution operation with kernel size 3×3.

[0146] It should be understood that Figure 10 Modules and references documented in Figure 1 Therefore, the operations and features described above for the method and the corresponding technical effects also apply to Figure 10 The modules in it will not be described in detail here.

[0147] In other embodiments, embodiments of the present invention further provide a computer-readable storage medium having a computer program stored thereon, wherein when the program instructions are executed by a processor, the processor is caused to execute the infrared image super-resolution reconstruction method in any of the above method embodiments;

[0148] As an embodiment, the computer-readable storage medium of the present invention stores computer-executable instructions, and the computer-executable instructions are configured as follows:

[0149] Acquire a low-resolution image, perform a first feature extraction on the low-resolution image according to a preset shallow feature extraction module, and obtain shallow features. , the expression is:

[0150] ,

[0151] Where, is the feature extraction operation, is the input low-resolution image, It is a shallow feature;

[0152] The second feature extraction is performed on the low-resolution image according to the preset deep feature extraction module to obtain the deep feature , wherein the deep feature extraction module includes a plurality of sequentially connected progressive residual edge blocks and a fusion module connected to each progressive residual edge block;

[0153] The deep features The expression is:

[0154] ,

[0155] ,

[0156] Where, is the feature extraction operation of the nth progressive residual edge block, is the feature extraction operation of the n-1th progressive residual edge block, is the feature extraction operation of the first progressive residual edge block, is the feature extraction result of the nth progressive residual edge block, is the number of progressive residual edge blocks, is the fusion function, For deep features;

[0157] According to the preset recursive structure module, the shallow features Superimposed on the deep features after convolution processing The superimposed result is upsampled based on the preset upsampling network to obtain the amplified features. , the expression is:

[0158] ,

[0159] Where, It is an upsampling network, including convolutional layers and pixel shuffling layers. Extract features for the third layer of the recursive block;

[0160] The magnified features Perform convolution operation to obtain the reconstructed high-resolution image , the expression is:

[0161] ,

[0162] Where, The convolution operation with kernel size 3×3.

[0163] The computer-readable storage medium may include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for the function; the data storage area may store data generated by the infrared image super-resolution reconstruction system. Furthermore, the computer-readable storage medium may include high-speed random access memory and may also include storage, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the computer-readable storage medium may optionally include storage remote from the processor. Such remote storage may be connected to the infrared image super-resolution reconstruction system via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0164] Figure 11 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention, such as Figure 11 As shown, the device includes: a processor 310 and a memory 320. The electronic device may also include: an input device 330 and an output device 340. The processor 310, the memory 320, the input device 330 and the output device 340 may be connected via a bus or other means. Figure 11 The example of a bus connection is shown. Memory 320 is the aforementioned computer-readable storage medium. Processor 310 executes the various server functional applications and data processing by running the non-volatile software programs, instructions, and modules stored in memory 320, thereby implementing the infrared image super-resolution reconstruction method of the aforementioned method embodiment. Input device 330 can receive input digital or character information and generate key signal input related to user settings and function control of the infrared image super-resolution reconstruction system. Output device 340 may include a display device such as a display screen.

[0165] The electronic device can execute the method provided by the embodiment of the present invention, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method provided by the embodiment of the present invention.

[0166] As an embodiment, the electronic device is applied to an infrared image super-resolution reconstruction system and is used for a client, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to:

[0167] Acquire a low-resolution image, perform a first feature extraction on the low-resolution image according to a preset shallow feature extraction module, and obtain shallow features. , the expression is:

[0168] ,

[0169] Where, is the feature extraction operation, is the input low-resolution image, It is a shallow feature;

[0170] The second feature extraction is performed on the low-resolution image according to the preset deep feature extraction module to obtain the deep feature , wherein the deep feature extraction module includes a plurality of sequentially connected progressive residual edge blocks and a fusion module connected to each progressive residual edge block;

[0171] The deep features The expression is:

[0172] ,

[0173] ,

[0174] Where, is the feature extraction operation of the nth progressive residual edge block, is the feature extraction operation of the n-1th progressive residual edge block, is the feature extraction operation of the first progressive residual edge block, is the feature extraction result of the nth progressive residual edge block, is the number of progressive residual edge blocks, is the fusion function, For deep features;

[0175] According to the preset recursive structure module, the shallow features Superimposed on the deep features after convolution processing The superimposed result is upsampled based on the preset upsampling network to obtain the amplified features. , the expression is:

[0176] ,

[0177] Where, It is an upsampling network, including convolutional layers and pixel shuffling layers. Extract features for the third layer of the recursive block;

[0178] The magnified features Perform convolution operation to obtain the reconstructed high-resolution image , the expression is:

[0179] ,

[0180] Where, The convolution operation with kernel size 3×3.

[0181] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods of each embodiment or certain portions of the embodiments.

[0182] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A super-resolution reconstruction method for infrared images, characterized in that: The method comprises: Acquire a low-resolution image, perform a first feature extraction on the low-resolution image according to a preset shallow feature extraction module, and obtain shallow features. , the expression is: , Where, is the feature extraction operation, is the input low-resolution image, It is a shallow feature; The shallow feature extraction module includes a group enhancement convolution module, and the group enhancement convolution module includes a convolution layer, layer, Layers, depth-wise separable convolutional layers, and channel shuffle operations; According to the convolution layer, the low-resolution image Perform convolution operation to obtain features with 64 channels, and divide the obtained features into two groups according to 1 / 4 and 3 / 4 of the number of feature channels; The convolution kernel is 3×3 Layer, processing the 1 / 4 group of features, the number of input channels is 16, the number of output channels is 16, and the convolution kernel is 3×3 Layer, processing the 3rd / 4th group of features, the number of input channels is 48, and the number of output channels is 48; Will The output of the layer is used as the next Layer and The input of the layer is further divided into two groups according to 1 / 4 and 3 / 4 of the number of feature channels. Among them, the channel shuffling operation and the depth-wise separable convolution layer are added to the 1 / 4 group of feature branches. The expression is: , , , , , , , , , , , Where, is the convolution operation, is a depth-wise separable convolutional layer, is the channel shuffle operation, It is the Relu activation function operation, Cross attention blocks for token dictionary, is the adaptive category attention, For the The processing results of 1 / 4 group of feature maps are For the The processing results of 3 / 4 groups of feature maps are For splicing operations, Indicates taking 1 / 4 of the number of channels of the current input feature map. Indicates taking 3 / 4 groups of channels of the current input feature map, It is a shallow feature; The second feature extraction is performed on the low-resolution image according to the preset deep feature extraction module to obtain the deep feature , wherein the deep feature extraction module includes a plurality of sequentially connected progressive residual edge blocks and a fusion module connected to each progressive residual edge block; The deep features The expression is: , , Where, is the feature extraction operation of the nth progressive residual edge block, is the feature extraction operation of the n-1th progressive residual edge block, is the feature extraction operation of the first progressive residual edge block, is the feature extraction result of the nth progressive residual edge block, is the number of progressive residual edge blocks, is the fusion function, For deep features; According to the preset recursive structure module, the shallow features Superimposed on the deep features after convolution processing The superimposed result is upsampled based on the preset upsampling network to obtain the amplified features. , the expression is: , Where, It is an upsampling network, including convolutional layers and pixel shuffling layers. Extract features for the third layer of the recursive block; The magnified features Perform convolution operation to obtain the reconstructed high-resolution image , the expression is: , Where, The convolution operation with kernel size 3×3; The progressive residual edge block includes a first edge feature extraction submodule, a first intermediate processing submodule, a second edge feature extraction submodule, a second intermediate processing submodule, a third edge feature extraction submodule and a residual learning structure.

2. The infrared image super-resolution reconstruction method according to claim 1, characterized in that: in, The expression of the token dictionary cross attention block is: , , , , , , , Where, The attention map generated in the cross-attention block for the token dictionary, Cross attention blocks for token dictionary, It means calculating the cosine similarity between two tokens. To adjust the range of similarity values, SoftMax is the SoftMax function. , and are the linear transformations of query tokens, key dictionary tokens, and value dictionary tokens, respectively. X is the input feature. To query the token, For the dictionary, For the Key dictionary, For the Value dictionary, is the l+1th index level of the dictionary, is a learnable parameter. is the re-probability distribution of the l-th layer dictionary, is a normalization layer used to adjust the range of the attention map, is the enhanced feature of the l+1th layer, is the transposed matrix of the attention map of the lth layer.

3. The infrared image super-resolution reconstruction method according to claim 1, characterized in that: in, The expression of adaptive category attention is: , Where, is the category of pixel classification, is the pixel value, is a function in Numpy, is the jth row and kth column of the attention map, For the i-th category.

4. The infrared image super-resolution reconstruction method according to claim 1, characterized in that: The first edge feature extraction submodule, the second edge feature extraction submodule, and the third edge feature extraction submodule are all connected to a 2×2 maximum pooling layer to form a downsampling mode; The residual learning structure includes three residual units and a fusion layer, and each residual unit includes convolution, activation function and convolution in sequence; Among them, the first edge feature extraction submodule includes two convolutions with a kernel size of 3×3 and a channel number c of 64. Each convolution is followed by a convolution with a kernel size of 1×1 and a channel number of 21 for dimensionality increase. The two convolution results are added in parallel and then reduced in dimension by a convolution with a kernel size of 1×1 and a channel number of 1. The result after dimensionality reduction is used as the first output result of the edge feature extraction submodule. ; First output result The data is input into the first intermediate processing submodule, which includes two convolutions with a kernel size of 3×3 and a channel number c of 64, which are serialized and then subjected to a 2×2 maximum pooling. The output of the first intermediate processing submodule serves as the input of the second edge feature extraction submodule. The second edge feature extraction submodule includes two convolutions with a kernel size of 3×3 and a channel number c of 128. Each convolution is followed by a convolution with a kernel size of 1×1 and a channel number of 21 for dimensionality increase. The two convolution results are added in parallel and then reduced by a convolution with a kernel size of 1×1 and a channel number of 1. The result of dimensionality reduction is used as the second output result of the second edge feature extraction submodule. ; Second output result The data is input to the second intermediate processing submodule, which includes two convolutions with a kernel size of 3×3 and a channel number c of 64, followed by a 2×2 maximum pooling. The output of the second intermediate processing submodule serves as the input of the third edge feature extraction submodule. The third edge feature extraction submodule includes three convolutions with a kernel size of 3×3 and a channel number c of 256. Each convolution is followed by a convolution with a kernel size of 1×1 and a channel number of 21 for dimensionality increase. The two convolution results are added in parallel and then reduced by a convolution with a kernel size of 1×1 and a channel number of 1. The result of dimensionality reduction is used as the third output result of the third edge feature extraction submodule. , the expression is: , , , Where, Represents low-resolution images Perform bicubic interpolation. is the result of bicubic interpolation processing, is the first edge feature extraction submodule, is the first output result of the first edge feature extraction submodule, For the Intermediate processing submodule, For the edge feature extraction submodule, For the The first output result of the edge feature extraction submodule.

5. The infrared image super-resolution reconstruction method according to claim 1, characterized in that: in, The recursive structure module uses a multi-branch structure in the recursive process, and all recursive blocks in the recursive structure module share the same branch input. The expression of the recursive structure module is: , , , , Where, is a shallow auxiliary feature, Extract features for the first layer of the recursive block, Extract features for the second layer of the recursive block, Extract features for the 3rd layer of the recursive block, is the convolution operation, for Activation function operation, is the batch normalization layer, For the dropout operation, It is a linear transformation operation that converts the output into a linear structure so that the model can better fit the training data.

6. An infrared image super-resolution reconstruction system, characterized in that: The system comprises: The first extraction unit is configured to obtain a low-resolution image, perform first feature extraction on the low-resolution image according to a preset shallow feature extraction module, and obtain shallow features. , the expression is: , Where, is the feature extraction operation, is the input low-resolution image, It is a shallow feature; The shallow feature extraction module includes a group enhancement convolution module, and the group enhancement convolution module includes a convolution layer, layer, Layers, depth-wise separable convolutional layers, and channel shuffle operations; According to the convolution layer, the low-resolution image Perform convolution operation to obtain features with 64 channels, and divide the obtained features into two groups according to 1 / 4 and 3 / 4 of the number of feature channels; The convolution kernel is 3×3 Layer, processing the 1 / 4 group of features, the number of input channels is 16, the number of output channels is 16, and the convolution kernel is 3×3 Layer, processing the 3rd / 4th group of features, the number of input channels is 48, and the number of output channels is 48; Will The output of the layer is used as the next Layer and The input of the layer is further divided into two groups according to 1 / 4 and 3 / 4 of the number of feature channels. Among them, the channel shuffling operation and the depth-wise separable convolution layer are added to the 1 / 4 group of feature branches. The expression is: , , , , , , , , , , , Where, is the convolution operation, is a depth-wise separable convolutional layer, is the channel shuffle operation, It is the Relu activation function operation, Cross attention blocks for token dictionary, is the adaptive category attention, For the The result of processing 1 / 4 of the feature maps is The processing results of 3 / 4 groups of feature maps are For splicing operations, Indicates taking 1 / 4 of the number of channels of the current input feature map. Indicates taking 3 / 4 groups of channels of the current input feature map, It is a shallow feature; The second extraction unit is configured to perform a second feature extraction on the low-resolution image according to a preset deep feature extraction module to obtain a deep feature. , wherein the deep feature extraction module includes a plurality of sequentially connected progressive residual edge blocks and a fusion module connected to each progressive residual edge block; The deep features The expression is: , , Where, is the feature extraction operation of the nth progressive residual edge block, is the feature extraction operation of the n-1th progressive residual edge block, is the feature extraction operation of the first progressive residual edge block, is the feature extraction result of the nth progressive residual edge block, is the number of progressive residual edge blocks, is the fusion function, For deep features; The superposition unit is configured to superimpose shallow features according to a preset recursive structure module Superimposed on the deep features after convolution processing The superimposed result is upsampled based on the preset upsampling network to obtain the amplified features. , the expression is: , Where, It is an upsampling network, including convolutional layers and pixel shuffling layers. Extract features for the third layer of the recursive block; The reconstruction unit is configured to reconstruct the magnified features Perform convolution operation to obtain the reconstructed high-resolution image , the expression is: , Where, The convolution operation with kernel size 3×3; The progressive residual edge block includes a first edge feature extraction submodule, a first intermediate processing submodule, a second edge feature extraction submodule, a second intermediate processing submodule, a third edge feature extraction submodule and a residual learning structure.

7. An electronic device, characterized in that: include: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Infrared image super-resolution reconstruction method based on edge enhancement

    CN116071243A

  • Real scene remote sensing image super-resolution reconstruction method and system based on progressive feature aggregation network, and storage medium

    CN120013762A