Image super-resolution reconstruction method and device

By building a multi-layer network for image feature extraction and reconstruction, the problems of feature loss and blurred details in image reconstruction are solved, and clearer image detail texture is achieved.

CN120219164APending Publication Date: 2025-06-27SHENZHEN XUMI YUNTU SPACE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510188524.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The problem of missing key features and blurred details in image reconstruction in the prior art.

Method used

Contextual interactive information network, multi-dimensional attention enhancement network, feature aggregation propagation network, shallow feature extraction network and reconstruction network are constructed through these networks, deep feature extraction network and image super-segment models are used to train models using training images, and super-resolution reconstruction of the target image.

Benefits of technology

Effectively retain key features of the image, improve the clarity of details after reconstruction, and solve the problems of feature loss and blurred details in image reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219164A_ABST
    Figure CN120219164A_ABST
Patent Text Reader

Abstract

The invention provides an image super-resolution reconstruction method and device. The method comprises the following steps: constructing a context interaction information network, a multi-dimensional attention enhancement network, a feature aggregation propagation network, a shallow feature extraction network and a reconstruction network; constructing stage networks by using a context interaction information network, a multi-dimensional attention enhancement network and a feature aggregation propagation network, and constructing a deep feature extraction network by using a plurality of stage networks; constructing an image super-division model by using the shallow feature extraction network, the deep feature extraction network and the reconstruction network; training the image super-division model by using the training image; and performing super-resolution reconstruction on the target image by using the trained image super-resolution model. By adopting the technical means, the problems that key features are lost and detail textures are blurred in image reconstruction in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technologies, and in particular, to an image super-resolution reconstruction method and apparatus. Background Art

[0002] Image super-resolution reconstruction based on deep learning improves image reconstruction performance by deepening the network, but complex networks will cause a sharp increase in the number of parameters, restricting its application on resource-constrained devices. At the same time, traditional network architectures also lack the ability to discriminate key features of images, which will result in the loss of key features after reconstruction; traditional network architectures cannot avoid adding redundant information during reconstruction, which will result in blurred detail textures after reconstruction. Summary of the Invention

[0003] In view of this, embodiments of the present disclosure provide an image super-resolution reconstruction method, apparatus, electronic device, and computer-readable storage medium to solve the problems of loss of key features and blurred detail textures in image reconstruction in the prior art.

[0004] In a first aspect of embodiments of the present disclosure, an image super-resolution reconstruction method is provided, including: constructing a context interaction information network, a multi-dimensional attention enhancement network, a feature aggregation propagation network, a shallow feature extraction network, and a reconstruction network; constructing a stage network using the context interaction information network, the multi-dimensional attention enhancement network, and the feature aggregation propagation network, and constructing a deep feature extraction network using multiple stage networks; constructing an image super-resolution model using the shallow feature extraction network, the deep feature extraction network, and the reconstruction network; training the image super-resolution model using training images; and performing super-resolution reconstruction on a target image using the trained image super-resolution model.

[0005] In a second aspect of embodiments of the present disclosure, an image super-resolution reconstruction apparatus is provided, including: a first construction module configured to construct a context interaction information network, a multi-dimensional attention enhancement network, a feature aggregation propagation network, a shallow feature extraction network, and a reconstruction network; a second construction module configured to construct a stage network using the context interaction information network, the multi-dimensional attention enhancement network, and the feature aggregation propagation network, and construct a deep feature extraction network using multiple stage networks; a third construction module configured to construct an image super-resolution model using the shallow feature extraction network, the deep feature extraction network, and the reconstruction network; a training module configured to train the image super-resolution model using training images; and a reconstruction module configured to perform super-resolution reconstruction on a target image using the trained image super-resolution model.

[0006] In a third aspect of embodiments of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the steps of the above method are implemented.

[0007] In the fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0008] The beneficial effects of the embodiments of the present disclosure compared with the prior art are as follows: constructing a context interaction information network, a multi-dimensional attention enhancement network, a feature aggregation propagation network, a shallow feature extraction network, and a reconstruction network; using the context interaction information network, the multi-dimensional attention enhancement network, and the feature aggregation propagation network to construct a stage network, and using multiple stage networks to construct a deep feature extraction network; using the shallow feature extraction network, the deep feature extraction network, and the reconstruction network to construct an image super-resolution model; training the image super-resolution model with training images; and performing super-resolution reconstruction on a target image using the trained image super-resolution model. By adopting the above technical means, the problems of missing key features and blurred detail textures in image reconstruction in the prior art can be solved, thereby retaining key features and making the detail textures clear. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0010] Figure 1 is a schematic flowchart of an image super-resolution reconstruction method provided by an embodiment of the present disclosure;

[0011] Figure 2 is a schematic flowchart of another image super-resolution reconstruction method provided by an embodiment of the present disclosure;

[0012] Figure 3 is a schematic structural diagram of an image super-resolution reconstruction device provided by an embodiment of the present disclosure;

[0013] Figure 4 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0014] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, the detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.

[0015] The following will describe in detail an image super-resolution reconstruction method and apparatus according to an embodiment of the present disclosure with reference to the accompanying drawings.

[0016] Figure 1 It is a schematic flowchart of an image super-resolution reconstruction method provided by an embodiment of the present disclosure. Figure 1 The image super-resolution reconstruction method can be executed by a computer or a server, or software on a computer or a server. As Figure 1 shown, the image super-resolution reconstruction method includes:

[0017] S101, constructing a contextual interaction information network, a multi-dimensional attention enhancement network, a feature aggregation and propagation network, a shallow feature extraction network, and a reconstruction network;

[0018] S102, constructing a stage network by using the contextual interaction information network, the multi-dimensional attention enhancement network, and the feature aggregation and propagation network, and constructing a deep feature extraction network by using multiple stage networks;

[0019] S103, constructing an image super-resolution model by using the shallow feature extraction network, the deep feature extraction network, and the reconstruction network;

[0020] S104, training the image super-resolution model by using training images;

[0021] S105, performing super-resolution reconstruction on a target image by using the trained image super-resolution model.

[0022] Contextual Interaction Information Block (CIIB), Multi-dimensional Attention Enhancement Block (MAEB), Feature Aggregation and Propagation Block (FAPB).

[0023] Design a multi-dimensional attention enhancement network to improve the network's discrimination ability for key features, and extract high-frequency information in both the channel and spatial dimensions. Finally, a feature aggregation and propagation network is proposed to effectively aggregate deep detail information, remove redundant information, and promote the propagation of effective information in the network, making the detail texture of the reconstructed image clearer.

[0024] According to the technical solution provided by the embodiments of the present application, a context interaction information network, a multi-dimensional attention enhancement network, a feature aggregation and propagation network, a shallow feature extraction network, and a reconstruction network are constructed; a stage network is constructed by using the context interaction information network, the multi-dimensional attention enhancement network, and the feature aggregation and propagation network, and a deep feature extraction network is constructed by using multiple stage networks; an image super-resolution model is constructed by using the shallow feature extraction network, the deep feature extraction network, and the reconstruction network; the image super-resolution model is trained by using training images; and the target image is super-resolution reconstructed by using the trained image super-resolution model. By adopting the above technical means, the problems of missing key features and blurred detail textures in image reconstruction in the prior art can be solved, and thus the key features can be retained and the detail textures can be made clear.

[0025] Further, training the image super-resolution model by using training images includes: inputting the training images into the image super-resolution model: processing the training images through the shallow feature extraction network to obtain image shallow features; processing the image shallow features through the deep feature extraction network to obtain image deep features; processing the training images and the image deep features through the reconstruction network to obtain a reconstruction result; calculating the loss between the labels of the training images and the reconstruction result, and optimizing the model parameters of the image super-resolution model according to the loss.

[0026] Further, constructing the context interaction information network includes: using a combination of multiple convolutional layers and activation layers to form two feature processing branches, wherein the dilation rates of the convolutional layers in the two feature processing branches are different; and constructing the context interaction information network by using a segmentation layer, the two feature processing branches, and a splicing layer.

[0027] For example, the dilation rate of the convolutional layer in one feature processing branch is 2, and the dilation rate of the convolutional layer in the other feature processing branch is 3. The feature processing branch includes a combination of three convolutional layers and activation layers. That is, inside the feature processing branch, there are a convolutional layer, an activation layer, a convolutional layer, an activation layer, a convolutional layer, and an activation layer in sequence.

[0028] Denote the input of the context interaction information network as the first feature. Inside the context interaction information network: the segmentation layer divides the first feature into two segmentation features in a preset ratio; the two feature processing branches respectively process one segmentation feature; and the splicing layer splices the results processed by the two feature processing branches to obtain the second feature. The second feature is the output of the context interaction information network.

[0029] Furthermore, a multi-dimensional attention enhancement network is constructed, including: constructing a channel attention network using an average pooling layer, a max pooling layer, a concatenation layer, a convolutional layer, an activation layer, and a multiplication layer; constructing a spatial attention network using an average pooling layer, a max pooling layer, a multi-layer perceptron, an addition layer, an activation layer, and a multiplication layer; constructing two feature subdivision branches using a convolutional layer, an activation layer, and a multiplication layer, where the receptive fields of the convolutional layers in the two feature processing branches are different; constructing a feature refinement network using the two feature subdivision branches and an addition layer; constructing a multi-dimensional attention enhancement network using the channel attention network, the spatial attention network, and the feature refinement network.

[0030] The inputs of the multi-dimensional attention enhancement network and the context interaction information network in the same stage network are the same. That is, the input of the multi-dimensional attention enhancement network is denoted as the first feature. Then, inside the multi-dimensional attention enhancement network:

[0031] The channel attention network processes the first feature. Inside the channel attention network: the average pooling layer and the max pooling layer process the first feature respectively to obtain an average pooling feature and a max pooling feature; the concatenation layer concatenates the average pooling feature and the max pooling feature together to obtain a concatenated feature; the convolutional layer and the activation layer process the concatenated feature in sequence to obtain a first activation feature; the multiplication layer multiplies the first activation feature and the first feature to obtain a first multiplication feature.

[0032] The spatial attention network and the channel attention network are parallel, and the inputs of both the spatial attention network and the channel attention network are the input of the multi-dimensional attention enhancement network. The spatial attention network processes the first feature. Inside the channel attention network: the average pooling layer and the max pooling layer process the first feature respectively to obtain an average pooling feature and a max pooling feature; the multi-layer perceptron processes the average pooling feature and the max pooling feature to obtain two perceptron features; the activation layer processes the two perceptron features to obtain a second activation feature; the multiplication layer multiplies the second activation feature and the first feature to obtain a second multiplication feature.

[0033] The first addition feature obtained after adding the first multiplication feature and the second multiplication feature is used as the input of the feature refinement network. Inside the feature refinement network: the two feature subdivision branches process the feature after adding the first multiplication feature and the second multiplication feature respectively to obtain two subdivision features; the addition layer uses the second addition feature obtained by adding the two subdivision features as the output of the multi-dimensional attention enhancement network. Inside the feature subdivision branch: the convolutional layer and the activation layer process the first addition feature to obtain a third activation feature; the multiplication layer multiplies the first addition feature and the third activation feature to obtain a subdivision feature.

[0034] The receptive field can be changed by adjusting the convolution kernel size, stride, and the number of convolutional layers.

[0035] Furthermore, a feature aggregation and propagation network is constructed using convolutional layers, activation layers, multiplication layers, addition layers, and concatenation layers; a shallow feature extraction network is constructed using convolutional layers; and a reconstruction network is constructed using sub-pixel convolutional upsampling layers and bicubic interpolation upsampling layers.

[0036] The feature aggregation and propagation network is a method for enhancing feature representations in convolutional neural networks (CNNs). It improves model performance by effectively integrating multi-scale information and cross-layer connections. It internally includes convolutional layers, activation layers, multiplication layers, addition layers, and concatenation layers.

[0037] Inside the reconstruction network: The sub-pixel convolutional upsampling layer processes the target deep features to obtain the first upsampled feature; the bicubic interpolation upsampling layer processes the target image to obtain the second upsampled feature; the first upsampled feature and the second upsampled feature are added together as the output of the reconstruction network, which is the reconstruction result.

[0038] Sub-pixel Convolutional Upsampling, also known as Pixel Shuffle or ESPCN (Efficient Sub-Pixel Convolutional Neural Network), is a technique used in image super-resolution (SR) and other tasks that require upsampling. It generates high-resolution images by rearranging low-resolution pixels in the feature map instead of directly increasing the size of the feature map.

[0039] Bicubic Interpolation upsampling is a commonly used method for enlarging images in the fields of image processing and computer vision. It creates a higher-resolution image by estimating the new value of each pixel, which is based on the weighted average of the surrounding 16 pixels (4x4 neighborhood), and the weights are calculated according to the distance and a specific cubic function.

[0040] Furthermore, a deep feature extraction network is constructed using multiple stage networks, including: serially connecting five stage networks to obtain the deep feature extraction network.

[0041] Figure 2 It is a schematic flowchart of another image super-resolution reconstruction method provided by an embodiment of the present disclosure, as Figure 2 shown, the method includes:

[0042] S201, input the target image into the image super-resolution model:

[0043] S202, process the target image through the shallow feature extraction network to obtain the target shallow features;

[0044] S203. Process the target shallow features through a deep feature extraction network to obtain target deep features;

[0045] S204. Process the target image and the target deep features through a reconstruction network to obtain a target reconstruction result.

[0046] The deep feature extraction network processes the target shallow features to obtain target deep features, including: the first-stage network processes the target shallow features to obtain first-stage features; the second-stage network processes the target shallow features to obtain second-stage features; the third-stage network processes the target shallow features to obtain third-stage features; the fourth-stage network processes the target shallow features to obtain fourth-stage features; the fifth-stage network processes the target shallow features to obtain target deep features.

[0047] The processing inside the networks of the five stages is the same. Taking the first-stage network as an example. Inside the first-stage network:

[0048] Inside the context interaction information network: The segmentation layer divides the target deep features into two segmentation features with a preset ratio; two feature processing branches respectively process one segmentation feature; the splicing layer splices the results processed by the two feature processing branches to obtain interaction features.

[0049] Inside the multi-dimensional attention enhancement network: The channel attention network processes the target deep features to obtain channel attention features; the spatial attention network processes the target deep features to obtain spatial attention features; the feature refinement network processes the spatial attention features and the channel attention features to obtain refined features.

[0050] The feature aggregation and propagation network processes the refined features to obtain first-stage features.

[0051] All the above optional technical solutions can be combined arbitrarily to form optional embodiments of the present application, which will not be elaborated one by one here.

[0052] The following is an embodiment of the device of the present disclosure, which can be used to execute the embodiment of the method of the present disclosure. For the details not disclosed in the embodiment of the device of the present disclosure, please refer to the embodiment of the method of the present disclosure.

[0053] Figure 3 It is a schematic diagram of an image super-resolution reconstruction device provided by an embodiment of the present disclosure. As Figure 3 shown, the image super-resolution reconstruction device includes:

[0054] The first construction module 301 is configured to construct a context interaction information network, a multi-dimensional attention enhancement network, a feature aggregation and propagation network, a shallow feature extraction network, and a reconstruction network;

[0055] The second construction module 302 is configured to construct a stage network by using a context interaction information network, a multi-dimensional attention enhancement network, and a feature aggregation propagation network, and construct a deep feature extraction network by using multiple stage networks;

[0056] The third construction module 303 is configured to construct an image super-resolution model by using a shallow feature extraction network, a deep feature extraction network, and a reconstruction network;

[0057] The training module 304 is configured to train the image super-resolution model by using training images;

[0058] The reconstruction module 305 is configured to perform super-resolution reconstruction on a target image by using the trained image super-resolution model.

[0059] According to the technical solution provided by the embodiments of the present application, a context interaction information network, a multi-dimensional attention enhancement network, a feature aggregation propagation network, a shallow feature extraction network, and a reconstruction network are constructed; a stage network is constructed by using the context interaction information network, the multi-dimensional attention enhancement network, and the feature aggregation propagation network, and a deep feature extraction network is constructed by using multiple stage networks; an image super-resolution model is constructed by using the shallow feature extraction network, the deep feature extraction network, and the reconstruction network; the image super-resolution model is trained by using training images; and super-resolution reconstruction is performed on a target image by using the trained image super-resolution model. By adopting the above technical means, the problems of missing key features and blurred detail textures in image reconstruction in the prior art can be solved, and thus the key features can be retained and the detail textures can be made clear.

[0060] In some embodiments, the training module 304 is further configured to input the training images into the image super-resolution model: process the training images through the shallow feature extraction network to obtain image shallow features; process the image shallow features through the deep feature extraction network to obtain image deep features; process the training images and the image deep features through the reconstruction network to obtain a reconstruction result; calculate the loss between the labels of the training images and the reconstruction result, and optimize the model parameters of the image super-resolution model according to the loss.

[0061] In some embodiments, the first construction module 301 is further configured to form two feature processing branches by using a combination of multiple convolutional layers and activation layers, wherein the dilation rates of the convolutional layers in the two feature processing branches are different; and construct a context interaction information network by using a segmentation layer, the two feature processing branches, and a splicing layer.

[0062] For example, the dilation rate of the convolutional layer in one feature processing branch is 2, and the dilation rate of the convolutional layer in the other feature processing branch is 3. The feature processing branch includes a combination of three convolutional layers and activation layers. That is, inside the feature processing branch, there are a convolutional layer, an activation layer, a convolutional layer, an activation layer, a convolutional layer, and an activation layer in sequence.

[0063] Denote the input of the context interaction information network as the first feature. Inside the context interaction information network: The segmentation layer divides the first feature into two segmented features in a preset ratio; two feature processing branches process one segmented feature respectively; the splicing layer splices the results processed by the two feature processing branches to obtain the second feature. The second feature is the output of the context interaction information network.

[0064] In some embodiments, the first construction module 301 is further configured to construct a channel attention network using an average pooling layer, a max pooling layer, a splicing layer, a convolutional layer, an activation layer, and a multiplication layer; construct a spatial attention network using an average pooling layer, a max pooling layer, a multi-layer perceptron, an addition layer, an activation layer, and a multiplication layer; construct two feature subdivision branches using a convolutional layer, an activation layer, and a multiplication layer, wherein the receptive fields of the convolutional layers in the two feature processing branches are different; construct a feature refinement network using the two feature subdivision branches and an addition layer; construct a multi-dimensional attention enhancement network using the channel attention network, the spatial attention network, and the feature refinement network.

[0065] The input of the multi-dimensional attention enhancement network and the context interaction information network in the same stage network is the same. That is, denote the input of the multi-dimensional attention enhancement network as the first feature. Inside the multi-dimensional attention enhancement network:

[0066] The channel attention network processes the first feature. Inside the channel attention network: The average pooling layer and the max pooling layer process the first feature respectively to obtain an average pooling feature and a max pooling feature; the splicing layer splices the average pooling feature and the max pooling feature together to obtain a spliced feature; the convolutional layer and the activation layer process the spliced feature in sequence to obtain a first activation feature; the multiplication layer multiplies the first activation feature and the first feature to obtain a first multiplication feature.

[0067] The spatial attention network and the channel attention network are parallel, and the inputs of both the spatial attention network and the channel attention network are the inputs of the multi-dimensional attention enhancement network. The spatial attention network processes the first feature. Inside the channel attention network: The average pooling layer and the max pooling layer process the first feature respectively to obtain an average pooling feature and a max pooling feature; the multi-layer perceptron processes the average pooling feature and the max pooling feature to obtain two perceptron features; the activation layer processes the two perceptron features to obtain a second activation feature; the multiplication layer multiplies the second activation feature and the first feature to obtain a second multiplication feature.

[0068] After the first multiplication feature and the second multiplication feature are added, the first addition feature is used as the input of the feature refinement network. Inside the feature refinement network: Two feature subdivision branches respectively process the feature after the first multiplication feature and the second multiplication feature are added to obtain two subdivided features; The addition layer uses the second addition feature obtained by adding the two subdivided features as the output of the multi-dimensional attention enhancement network. Inside the feature subdivision branch: The convolutional layer and the activation layer process the first addition feature to obtain the third activation feature; The multiplication layer multiplies the first addition feature and the third activation feature to obtain the subdivided feature.

[0069] The receptive field can be changed by adjusting the convolutional kernel size, stride, and the number of convolutional layers.

[0070] In some embodiments, the first construction module 301 is further configured to construct a feature aggregation propagation network using a convolutional layer, an activation layer, a multiplication layer, an addition layer, and a splicing layer; construct a shallow feature extraction network using a convolutional layer; construct a reconstruction network using a sub-pixel convolutional upsampling layer and a bicubic interpolation upsampling layer.

[0071] The feature aggregation propagation network is a method for enhancing feature representation in convolutional neural networks (CNNs). It improves model performance by effectively integrating multi-scale information and cross-layer connections. It includes a convolutional layer, an activation layer, a multiplication layer, an addition layer, and a splicing layer inside.

[0072] Inside the reconstruction network: The sub-pixel convolutional upsampling layer processes the target deep feature to obtain the first upsampled feature; The bicubic interpolation upsampling layer processes the target image to obtain the second upsampled feature; The first upsampled feature and the second upsampled feature are added as the output of the reconstruction network, that is, the reconstruction result.

[0073] Sub-pixel Convolutional Upsampling, also known as Pixel Shuffle or ESPCN (Efficient Sub-Pixel Convolutional Neural Network), is a technique used in image super-resolution (SR) and other tasks that require upsampling. It generates a high-resolution image by rearranging the low-resolution pixels in the feature map instead of directly increasing the size of the feature map.

[0074] Bicubic Interpolation upsampling is a commonly used method for enlarging images in the fields of image processing and computer vision. It creates a higher-resolution image by estimating the new value of each pixel, which is based on the weighted average of the surrounding 16 pixels (4x4 neighborhood), and the weights are calculated according to the distance and a specific cubic function.

[0075] In some embodiments, the second construction module 302 is further configured to construct a deep feature extraction network by using a multi-stage network, including: serially connecting five stage networks to obtain a deep feature extraction network.

[0076] In some embodiments, the reconstruction module 305 is further configured to input a target image into an image super-resolution model: process the target image through a shallow feature extraction network to obtain target shallow features; process the target shallow features through a deep feature extraction network to obtain target deep features; process the target image and the target deep features through a reconstruction network to obtain a target reconstruction result.

[0077] The deep feature extraction network processes the target shallow features to obtain target deep features, including: the first stage network processes the target shallow features to obtain first stage features; the second stage network processes the target shallow features to obtain second stage features; the third stage network processes the target shallow features to obtain third stage features; the fourth stage network processes the target shallow features to obtain fourth stage features; the fifth stage network processes the target shallow features to obtain target deep features.

[0078] The processing inside the five stage networks is the same. Taking the first stage network as an example. Inside the first stage network:

[0079] Inside the context interaction information network: the segmentation layer segments the target deep features into two segmented features with a preset ratio; two feature processing branches respectively process one segmented feature; the splicing layer splices the results processed by the two feature processing branches to obtain interaction features.

[0080] Inside the multi-dimensional attention enhancement network: the channel attention network processes the target deep features to obtain channel attention features; the spatial attention network processes the target deep features to obtain spatial attention features; the feature refinement network processes the spatial attention features and the channel attention features to obtain refined features.

[0081] The feature aggregation and propagation network processes the refined features to obtain first stage features.

[0082] It should be understood that the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present disclosure.

[0083] Figure 4 It is a schematic diagram of the electronic device 4 provided by the embodiments of the present disclosure. As Figure 4As shown, the electronic device 4 in this embodiment includes: a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable on the processor 401. When the processor 401 executes the computer program 403, the steps in the above method embodiments are implemented. Alternatively, when the processor 401 executes the computer program 403, the functions of each module / unit in the above device embodiments are implemented.

[0084] The electronic device 4 may be a desktop computer, a notebook, a palm computer, a cloud server, or other electronic devices. The electronic device 4 may include, but is not limited to, the processor 401 and the memory 402. Those skilled in the art can understand that Figure 4 This is only an example of the electronic device 4 and does not constitute a limitation on the electronic device 4. It may include more or fewer components than shown in the figure, or different components.

[0085] The processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0086] The memory 402 may be an internal storage unit of the electronic device 4. For example, the hard disk or memory of the electronic device 4. The memory 402 may also be an external storage device of the electronic device 4. For example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 4. The memory 402 may also include both an internal storage unit and an external storage device of the electronic device 4. The memory 402 is used to store computer programs and other programs and data required by the electronic device.

[0087] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. In the embodiments, each functional unit and module can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0088] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present disclosure, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned method embodiments can be implemented. The computer program can include computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device that can carry computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0089] The above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them; although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be included in the protection scope of the present disclosure.

Claims

1. A method for super-resolution image reconstruction, characterized in that: include: Construct contextual interaction information network, multi-dimensional attention enhancement network, feature aggregation propagation network, shallow feature extraction network and reconstruction network; Using the context interaction information network, the multi-dimensional attention enhancement network and the feature aggregation propagation network to construct a stage network, and using multiple stage networks to construct a deep feature extraction network; Constructing an image super-resolution model using the shallow feature extraction network, the deep feature extraction network and the reconstruction network; Using training images to train the image super-resolution model; The trained image super-resolution model is used to reconstruct the target image in super-resolution.

2. The method according to claim 1, characterized in that Constructing a contextual interaction information network, including: A combination of multiple convolutional layers and activation layers is used to form two feature processing branches, wherein the expansion rates of the convolutional layers in the two feature processing branches are different; The contextual interaction information network is constructed using a segmentation layer, two feature processing branches and a concatenation layer.

3. The method according to claim 1, characterized in that Construct a multi-dimensional attention enhancement network, including: Construct a channel attention network using average pooling layer, max pooling layer, concatenation layer, convolution layer, activation layer and multiplication layer; Construct a spatial attention network using average pooling layers, max pooling layers, multilayer perceptrons, addition layers, activation layers, and multiplication layers; Convolutional layers, activation layers, and multiplication layers are used to construct two feature segmentation branches, where the receptive fields of the convolutional layers in the two feature processing branches are different; Construct a feature refinement network using two feature segmentation branches and an addition layer; The multi-dimensional attention enhancement network is constructed using the channel attention network, the spatial attention network and the feature refinement network.

4. The method according to claim 1, characterized in that: Constructing the feature aggregation propagation network using convolutional layers, activation layers, multiplication layers, addition layers, and concatenation layers; Constructing the shallow feature extraction network using a convolutional layer; The reconstruction network is constructed using a sub-pixel convolution upsampling layer and a bicubic interpolation upsampling layer.

5. The method according to claim 1, characterized in that A deep feature extraction network is constructed using multiple stage networks, including: The five-stage networks are connected in series to obtain the deep feature extraction network.

6. The method according to claim 1, characterized in that The image super-resolution model is trained using the training image, including: Input the training image into the image super-resolution model: Processing the training image through the shallow feature extraction network to obtain shallow features of the image; Processing the shallow features of the image through the deep feature extraction network to obtain deep features of the image; Processing the training image and the deep features of the image through the reconstruction network to obtain a reconstruction result; The loss between the label of the training image and the reconstruction result is calculated, and the model parameters of the image super-resolution model are optimized according to the loss.

7. The method according to claim 1, characterized in that Use the trained image super-resolution model to reconstruct the target image in super-resolution, including: Input the target image into the image super-resolution model: Processing the target image through the shallow feature extraction network to obtain target shallow features; Processing the shallow features of the target through the deep feature extraction network to obtain the deep features of the target; The target image and the target deep features are processed by the reconstruction network to obtain a target reconstruction result.

8. An image super-resolution reconstruction device, characterized in that: include: The first building module is configured to build a contextual interaction information network, a multi-dimensional attention enhancement network, a feature aggregation propagation network, a shallow feature extraction network, and a reconstruction network; A second construction module is configured to construct a stage network using the context interaction information network, the multi-dimensional attention enhancement network and the feature aggregation propagation network, and to construct a deep feature extraction network using multiple stage networks; A third construction module is configured to construct an image super-resolution model using the shallow feature extraction network, the deep feature extraction network and the reconstruction network; A training module, configured to train the image super-resolution model using training images; The reconstruction module is configured to perform super-resolution reconstruction on the target image using the trained image super-resolution model.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.