An Image Super-Resolution Method Based on Blueprint Separable Residual Network

By using blueprints to separate residual networks and attention mechanisms in super-resolution models, the problem of high computing costs in the prior art is solved, and efficient image super-resolution reconstruction is achieved.

CN115082306BActive Publication Date: 2025-05-27SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210504706.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-10
Publication Date
2025-05-27
Estimated Expiration
2042-05-10

AI Technical Summary

Technical Problem

While improving the quality of image reconstruction, the existing super-resolution model has too high computational cost, limiting its application in practical scenarios with high efficiency or real-time requirements.

Method used

The image super-resolution method based on blueprint separable residual network is adopted. By decomposing the standard convolution layer into point-by-point convolution and depth convolution, the number of parameters and calculations of the network are reduced, and combined with the enhanced spatial attention and contrast-perceptual channel attention mechanism, the reconstruction performance is improved.

Benefits of technology

It significantly compresses the network parameter quantity and calculation quantity, while achieving better image reconstruction performance, suitable for practical scenarios with high efficiency and real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115082306B_ABST
    Figure CN115082306B_ABST
Patent Text Reader

Abstract

The present invention discloses an image super-resolution method based on a blueprint separable residual network. The method includes: obtaining a target image; inputting the target image into a trained super-resolution model to obtain a reconstructed image, the resolution of the reconstructed image being higher than that of the target image, wherein the super-resolution model includes a shallow feature extraction module, a deep feature extraction module, a multi-layer feature fusion module, and a reconstruction module, the deep feature extraction module includes a plurality of distillation modules connected in sequence, each distillation module includes multi-level feature extraction, each level includes a standard convolutional layer and a corresponding blueprint shallow residual block, with the input end as a reference, the outputs of the blueprint shallow residual blocks of the previous level are respectively transmitted to the standard convolutional layer and the blueprint shallow residual block of the next level, and the blueprint shallow residual block of the last level is connected to a blueprint separation convolutional layer. The present invention reduces the number of parameters and the amount of computation, and can achieve a strong super-resolution reconstruction effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and more specifically, to an image super-resolution method based on a blueprint separable residual network. Background Art

[0002] Single-image super-resolution (SR) is a fundamental task in the field of computer vision. It aims to reconstruct visually pleasing high-resolution (HR) images from their corresponding low-resolution (LR) observations. In recent years, the general paradigm has gradually shifted from model-based solutions to deep learning methods. These SR networks have significantly improved the quality of the restored images, and their success is partly attributed to large model capacity and dense computational properties. However, these properties largely limit the application of the model in practical scenarios with high requirements for efficiency or real-time performance. Currently, many lightweight SR networks have been proposed to address the inefficiency problem. These methods use different strategies to achieve high efficiency, including parameter sharing strategies, cascaded networks with grouped convolutions, information or feature distillation mechanisms, and attention mechanisms, etc. Although these solutions apply compact architectures and improve the mapping efficiency, there are still redundancies in the convolution operations.

[0003] In recent years, great progress has been made in the field of model compression and acceleration. These techniques can generally be divided into four categories: parameter pruning and quantization, low-rank decomposition, knowledge distillation, and transfer / compact convolution filters. For parameter pruning and quantization methods, they aim to explore the redundancy of the model architecture and attempt to remove or reduce redundant parameters. Low-rank decomposition methods use matrix or tensor decomposition to estimate more information representations of the network. Knowledge distillation methods aim to generate more compact student models from larger networks by learning the distribution of the teacher model. Methods based on the design of transfer / compact convolution filters aim to design special structured convolution filters to reduce model parameters and save storage / computation.

[0004] However, most current SR models often introduce a large amount of computational cost when bringing performance improvements, thus limiting the practical applications of these methods. For example, efficient super-resolution methods based on convolutional neural networks rely on convolutional neural network mapping to reconstruct high-resolution images and minimize the difference between the reconstructed image and the original image to achieve supervised training. However, existing network structures are often built based on standard convolutional layers, that is, one convolution needs to process all channels. In this way, there are still redundancies in the convolution operations, resulting in too high computational costs to be applied to edge devices. Summary of the Invention

[0005] The object of the present invention is to overcome the defects of the above-mentioned prior art and provide an image super-resolution method based on a blueprint separable residual network. The method includes:

[0006] Obtain a target image;

[0007] Input the target image into the trained super-resolution model to obtain a reconstructed image, where the resolution of the reconstructed image is higher than that of the target image.

[0008] Among them, the super-resolution model includes a shallow feature extraction module, a deep feature extraction module, a multi-layer feature fusion module, and a reconstruction module. The shallow feature extraction module is used to map the target image to a high-dimensional space to extract high-dimensional shallow features. The deep feature extraction module is used to extract multiple levels of deep features from the high-dimensional shallow features. The feature fusion module is used to refine and aggregate the multiple levels of deep features to obtain aggregated features. The reconstruction module is used to perform image reconstruction based on the aggregated features to obtain a reconstructed image.

[0009] Among them, the deep feature extraction module contains a plurality of distillation modules connected in sequence. Each distillation module contains multi-level feature extraction. Each level contains a standard convolutional layer and a corresponding blueprint shallow residual block. Taking the input end as a reference, the output of the blueprint shallow residual block of the previous level is respectively transmitted to the standard convolutional layer and the blueprint shallow residual block of the next level. And the blueprint shallow residual block of the last level is connected to the first blueprint separable convolutional layer, and the output of this first blueprint separable convolutional layer is fused with the features extracted by the standard convolutional layers of each level.

[0010] Compared with the prior art, the advantages of the present invention are that, aiming at the fact that a bottleneck in the efficiency of the current network is the parameter quantity and computational quantity of the standard convolutional layer, the use of blueprint separable convolution is provided. By decomposing the standard convolutional layer, a standard convolution is decomposed into a pointwise convolution and a depthwise convolution, significantly compressing the parameter quantity and computational quantity of the network. Further, a blueprint separable residual network is constructed for super-resolution image creation, achieving better reconstruction performance while reducing the parameter quantity and computational quantity.

[0011] Through the following detailed description of the exemplary embodiments of the present invention with reference to the accompanying drawings, other features and advantages of the present invention will become clear. Description of the Drawings

[0012] The drawings incorporated in the specification and constituting a part of the specification illustrate embodiments of the present invention and, together with the description, are used to explain the principles of the present invention.

[0013] Figure 1 is a flowchart of an image super-resolution method based on a blueprint separable residual network according to an embodiment of the present invention;

[0014] Figure 2 is an overall architecture diagram of an image super-resolution model based on a blueprint separable residual network according to an embodiment of the present invention;

[0015] Figure 3 is a schematic structural diagram of a distillation module according to an embodiment of the present invention;

[0016] Figure 4 is a schematic structural diagram of an enhanced spatial attention module according to an embodiment of the present invention;

[0017] Figure 5 is a schematic structural diagram of a contrast perception channel attention module according to an embodiment of the present invention;

[0018] Figure 6 is a schematic structural diagram of a blueprint shallow residual block according to an embodiment of the present invention;

[0019] Figure 7 is a schematic structural diagram of a blueprint separable convolutional layer according to an embodiment of the present invention;

[0020] Figure 8 is a schematic diagram of the processing process of a distillation module according to an embodiment of the present invention;

[0021] Figure 9 is a schematic diagram of the application process according to an embodiment of the present invention;

[0022] Figure 10 is a schematic diagram for comparing the image reconstruction effect according to an embodiment of the present invention with the prior art. Detailed Embodiments

[0023] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that: Unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present invention.

[0024] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way a limitation on the present invention or its application or use.

[0025] Techniques, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such techniques, methods, and devices should be considered as part of the specification.

[0026] In all the examples shown and discussed herein, any specific value should be construed as merely exemplary and not as a limitation. Thus, other examples of the exemplary embodiments may have different values.

[0027] It should be noted that: Similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it need not be further discussed in subsequent drawings.

[0028] See Figure 1 As shown, the provided image super-resolution method based on the blueprint separable residual network includes the following steps:

[0029] Step S110: Design a blueprint separable residual network to construct a super-resolution model.

[0030] Combined with Figure 2 As shown, this embodiment proposes a new lightweight SR network, namely the blueprint separable residual network (BSRN) as the super-resolution model. The blueprint separable residual network (BSRN) as a whole includes a shallow feature extraction module, a deep feature extraction module, a multi-layer feature fusion module, and a reconstruction module.

[0031] The shallow feature extraction module maps the input image to a high-dimensional space. The deep feature extraction maps and generates features in the high-dimensional space. As the network depth increases, the high-frequency features of the input image are gradually restored and reconstructed. The multi-layer feature fusion module integrates and processes the image features in the early and late stages of the network, and fuses them into a high-dimensional tensor with rich details. The reconstruction module performs sub-pixel reconstruction on the high-dimensional tensor and outputs a high-resolution image. In this article, the shallow features and deep features are relative. Shallow features refer to coarse-grained features, while deep features refer to finer-grained features.

[0032] In the embodiment of the present invention, the deep feature extraction module is a key module, which includes multiple efficient separable distillation modules (Efficient Separable Distillation Block, ESDB) or simply distillation modules, and is used to extract multiple levels of deep features from high-dimensional shallow features. Specifically, see Figure 3 As shown, each ESDB module uses a blueprint shallow residual block (Blueprint Shallow Residual Block, BSRB) step by step to extract and compress the input features step by step, gradually restoring the image details. After fusion, it passes through a comprehensive attention mechanism group, that is Figure 4 the enhanced spatial attention block (Enhanced Spatial Attention Block, ESA) of Figure 5 and the contrast-aware channel attention block (Contrast-Aware Attention Block, CCA) of

[0033] Specifically, combined with Figure 3As shown, each ESDB includes multi-level feature extraction. Each level includes a standard convolutional layer (conv-1) and a corresponding blueprint shallow residual block (BSRB). The output of the BSRB in the previous level is respectively passed to the standard convolutional layer and the BSRB in the next level. And the BSRB in the last level is connected to the blueprint separable convolutional layer (BSconv, or called blueprint separation convolutional layer). The output of this blueprint separable convolutional layer is fused with the features extracted by the standard convolutional layers at all levels.

[0034] It should be noted that in Figure 3 , conv-1 represents performing a standard 1x1 convolution, and the blueprint separable convolutional layer (BSconv) decomposes a standard convolution into a pointwise convolution and a depthwise convolution. See Figure 7 As shown, by decomposing the standard convolutional layer, the number of parameters and the amount of computation of the network are significantly compressed. In addition, the number of feature extraction levels included in each ESDB, the number of BSRBs, the kernel sizes of the standard convolutional layer and the depthwise convolutional layer, the number of channels, etc. can all be set according to actual needs. Preferably, when the kernel of the depthwise convolution is set to 3 and the number of fused feature branches is set to 4 (i.e., corresponding to including three levels of BSRBs), the network performance is the best.

[0035] Figure 4 is a schematic structural diagram of the enhanced spatial attention module (ESA). This module sequentially includes a standard convolutional layer (Conv-1), a strided convolutional layer (Strided Conv), a pooling layer (Pooling), grouped convolutions (Conv Groups), an upsampling layer (Upsampling), a standard convolutional layer, and activation processing (such as using Sigmoid). The spatial attention module considers the importance degree of the spatial position information in the input image, and through processes such as strided convolution and grouped convolution processing, enhances the accuracy of the spatial attention mechanism.

[0036] Figure 5 is a schematic structural diagram of the contrast-aware channel attention module (CCA). This module sequentially includes a contrast layer (Contrast), two standard convolutional layers, and Sigmoid activation processing. Compared with the traditional channel attention mechanism, the provided contrast-aware channel attention module considers the contrast between channels to obtain the weights corresponding to each channel, which helps to extract clearer texture features from the input image. And by cascading with the enhanced spatial attention module, the image reconstruction ability is further enhanced.

[0037] Figure 6It is a schematic structural diagram of a blueprint shallow residual block (BSRB), which includes a blueprint separable convolution layer (BSConv) and a Gaussian error linear unit (GELU). Through experimental verification, compared with commonly used activation processing functions such as ReLU or LeakyReLU, using GELU to construct the blueprint shallow residual block is beneficial to improving the model performance.

[0038] For clarity, Figure 8 The processing process of the ESDB is further illustrated, where H represents the height, W represents the width, and C represents the number of channels.

[0039] It should be understood that Figure 2 The shallow feature extraction module, multi-layer feature fusion module, reconstruction module, etc. in [reference] can be implemented by existing technologies, and these modules will not be elaborated here.

[0040] In summary, the present invention improves the network performance from multiple aspects, such as optimizing convolution operations and introducing effective attention modules. The blueprint separable residual network uses blueprint separable convolution to construct basic building blocks, thereby reducing redundant convolution operations. BSConv utilizes the in-kernel correlation for effective separation, which is beneficial to efficient SR. Enhanced spatial attention (ESA) and contrast-aware channel attention are introduced to enhance the model capabilities. It should be noted that in the constructed super-resolution model, using blueprint separable convolution alone or the above two attention mechanisms can also improve the model performance to a certain extent compared with the prior art.

[0041] Step S120, training the super-resolution model with the goal of minimizing the set loss function.

[0042] In one embodiment, the loss function adopts the mean absolute error, which is expressed as:

[0043]

[0044] where N is the number of image pixels, represents the super-resolution network (i.e., the super-resolution model), θ represents the network hyperparameters, X represents the input image, and Y represents the ground-truth image.

[0045] The model training process can use the gradient descent algorithm to obtain the network parameters through iterative learning. For example, the initial learning rate is set to 2e-3, the learning rate change strategy uses cosine annealing, and the period is 1.5e6. Using the collected sample data, the constructed super-resolution model is trained until convergence.

[0046] Step S130, enhancing the input target image using the trained super-resolution model to obtain a high-resolution reconstructed image.

[0047] After the model training is completed, the acquired target images can be enhanced. For example, the enhancement process includes: preprocessing the input target images, performing multi-channel replication and arranging them side by side; performing high-dimensional mapping on the preprocessed images to extract high-dimensional shallow features; refining the high-dimensional shallow features through an efficient separable distillation module to extract deep features; refining and aggregating the deep features of different network depths through a feature fusion mechanism; and performing sub-pixel operations to upsample the low-resolution deep features to generate high-resolution outputs. The present invention can be applied not only to image super-resolution but also to other low-level vision tasks, such as image denoising, image de-raining, and image de-fogging.

[0048] The above model training process can be performed offline on a server or in the cloud. Embedding the trained model into an electronic device can achieve real-time image super-resolution reconstruction. The electronic device can be a terminal device or a server. Terminal devices include any terminal devices such as mobile phones, tablets, personal digital assistants (PDAs), point-of-sale (POS) terminals, in-vehicle computers, and smart wearable devices (smart watches, virtual reality glasses, virtual reality headsets, etc.). Servers include, but are not limited to, application servers or web servers, and can be independent servers, cluster servers, or cloud servers, etc. For example, as shown in Figure 9 In the actual model application, as shown, the terminal device can acquire target images or receive uploaded target images, obtain and identify the resolution of the images, and use the images as the first images. The terminal device processes the first images using the trained model to obtain second images with a resolution higher than that of the first images, and then displays the second images to the user, or the second images can also be uploaded to the server or the cloud.

[0049] The present invention can be applied to image reconstruction in various scenarios, including but not limited to the following aspects:

[0050] 1) In intelligent sports training or video-assisted refereeing. Since the present invention is not sensitive to the fast or slow time of video actions, it can be widely applied to various sports scenarios, such as yoga with slow movements and figure skating or gymnastics with rapid movement changes.

[0051] 2) Intelligent video review. The present invention can complete abnormal action recognition and judgment on a mobile device, directly send the abnormalities to the cloud server or display them to the user, further improving the judgment speed and efficiency.

[0052] 3) Intelligent video montage. Facing a large video database, automatically extract and edit and summarize videos of unified actions.

[0053] 4) Intelligent security. Action recognition can be directly performed on intelligent terminals with limited computing resources, such as smart glasses, drones, smart cameras, etc., and abnormal behaviors can be directly fed back, improving the immediacy and accuracy of patrols and the like.

[0054] To further verify the effect of the present invention, a simulation experiment was conducted. The experimental results are shown in Table 1. Bicubic is a traditional bicubic interpolation super-resolution algorithm, and SRCNN, FSRCNN, VDSR, DRRN, MemNet, IDN, CARN, IMDN, PAN, LAPAR-A, and RFDN are existing models based on convolutional neural networks. BSRN corresponds to the present invention, and Set5, Set14, BSD100, Urban100, and Manga109 are datasets. It can be observed that the present invention outperforms existing methods on various test datasets and significantly reduces the number of model parameters and computational complexity.

[0055] Table 1 Comparison between existing super-resolution models and the model of the present invention

[0056]

[0057] In Table 1, the number of parameters is in units of K, which is an abbreviation symbol for thousand, and the computational complexity is in units of G, which is an abbreviation for gigaflops; PSNR and SSIM respectively represent peak signal-to-noise ratio and structural similarity, reflecting the reconstruction effect of the super-resolution algorithm. The larger the value, the better the restoration effect.

[0058] Table 2 is the result comparison of the convolutional decomposition method. It can be seen from Table 2 that under the same number of parameters, the present invention can significantly improve the reconstruction performance of the network.

[0059] Table 2 Result comparison of the convolutional decomposition method

[0060]

[0061] It can be seen from Table 2 that the present invention decomposes a standard convolutional layer into a pointwise convolution and a depthwise convolution through the convolutional decomposition method. Although the direct decomposition performance will decline, by increasing the network width and depth, the performance can be restored and exceed the existing technology. The current method with the optimal X3 super-resolution reconstruction effect uses 541K parameters and 42.2G of computational complexity, where the computational complexity uses Mult-Adds as a measurement index.

[0062] Table 3 is the ablation experiment of the attention mechanism. It can be observed that by comprehensively using the attention mechanism, the present invention can achieve a strong super-resolution reconstruction effect while keeping the number of parameters and computational complexity small.

[0063] Table 3 Ablation experiment of the attention mechanism.

[0064]

[0065] Table 4 shows the experimental results for different activation functions. It can be observed that the present invention uses the Gaussian Error Linear Unit (GELU) to construct the blueprint shallow residual module, which further improves the super-resolution reconstruction effect while keeping the number of parameters and the amount of computation unchanged.

[0066] Table 4 Activation Function Experiment

[0067]

[0068] Figure 10 is a comparison diagram of the image reconstruction effect after adopting the blueprint shallow residual module and the above two attention mechanisms and the prior art. It can be seen that the texture features of the output image (left figure) of the present invention are clearer.

[0069] In summary, compared with the prior art, the present invention has the following technical advantages:

[0070] 1) The provided blueprint shallow residual module, attention module, etc. can be inserted into any 2D convolutional network to achieve the effect exceeding that of 3D networks without the need for special hardware or optimization of deep learning platforms, so it has strong versatility. Currently, the mainstream video network designs all include 3D convolutions and require specific network designs to achieve the required accuracy. For example, on large datasets such as kinetics, since 3D convolutions cannot be pre-trained on ImagNet, training from scratch often requires very large computing resources. However, the model of the present invention can utilize any network pre-trained on ImagNet and converge faster on video datasets and achieve better results through the SmallBIg design.

[0071] 2) The present invention achieves a balance between the amount of computation and accuracy. Although the two currently well-known plug-and-play video recognition modules, TSM and nonlocal, can also be embedded into mainstream 2D networks, the effect of the module of the present invention is far higher than that of TSM, and the amount of computation of Nonlocal is greater than that of our SmallBIg. The results of the present invention on 2D networks are also significantly higher than those of Nonlocal + 3D networks. This shows that the design of the present invention has advantages in both the amount of computation and accuracy, and can build a more efficient SR network by reducing redundant computations and using more effective modules.

[0072] 3) The present invention is applicable to certain special application scenarios. For example, in the security field, abnormal behaviors or actions often have a short duration and rapid changes. The present invention is not sensitive to the speed of change on action frames and can model actions with different durations well. Since the kinetics dataset consists of videos of about 10s, for example, for the action of shooting a basketball, the duration from dribbling to preparing to shoot to finally scoring is long and the change is slow, while on the something-something dataset, the videos are 2 - 3s long. For example, for the action of giving a thumbs up, the change of one action does not exceed 3s. The present invention has achieved good results on both of these two datasets, which proves that the module provided by the present invention can model actions with different durations well.

[0073] The present invention can be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present invention.

[0074] The computer-readable storage medium can be a tangible device that can retain and store instructions used by an instruction execution device. The computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0075] The computer-readable program instructions described herein can be downloaded from the computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. The network adapter or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.

[0076] The computer program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, Python, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present invention.

[0077] Aspects of the present invention are described herein with reference to the flowchart and / or block diagram of a method, apparatus (system), and computer program product according to embodiments of the present invention. It should be understood that each block of the flowchart and / or block diagram, and the combinations of blocks in the flowchart and / or block diagram, can be implemented by computer-readable program instructions.

[0078] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is produced that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other devices to operate in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured article that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0079] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0080] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions. As will be apparent to those of ordinary skill in the art, implementation via hardware, implementation via software, and implementation via a combination of software and hardware are equivalent.

[0081] The embodiments of the present invention have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein. The scope of the present invention is defined by the appended claims.

Claims

1. An image super-resolution method based on a blueprint separable residual network, comprising the following steps: Obtain a target image; Input the target image into a trained super-resolution model to obtain a reconstructed image, the resolution of the reconstructed image being higher than that of the target image; Wherein, the super-resolution model includes a shallow feature extraction module, a deep feature extraction module, a multi-level feature fusion module, and a reconstruction module. The shallow feature extraction module is used to map the target image into a high-dimensional space to extract high-dimensional shallow features. The deep feature extraction module is used to extract multiple levels of deep features from the high-dimensional shallow features. The feature fusion module is used to refine and aggregate the multiple levels of deep features to obtain aggregated features. The reconstruction module is used to perform image reconstruction based on the aggregated features to obtain a reconstructed image; Wherein, the deep feature extraction module includes a plurality of distillation modules connected in sequence. Each distillation module includes multi-level feature extraction. Each level includes a standard convolutional layer and a corresponding blueprint shallow residual block. Taking the input end as a reference, the output of the blueprint shallow residual block of the previous level is respectively transmitted to the standard convolutional layer and the blueprint shallow residual block of the next level. And the blueprint shallow residual block of the last level is connected to a first blueprint separable convolutional layer. The output of the first blueprint separable convolutional layer is fused with the features extracted by the standard convolutional layers of each level. Wherein, for each of the distillation modules, after feature fusion, a cascaded spatial attention module and a contrast perception channel attention module are further provided. After the output features of the contrast perception channel attention module are fused with the input features of the corresponding distillation module, they are used as the output of the distillation module.

2. The method according to claim 1, characterized in that, the blueprint shallow residual block includes a second blueprint separable convolutional layer and a Gaussian error linear unit. The Gaussian error linear unit performs activation processing on the fused features of the input features of the blueprint shallow residual block and the features extracted by the second blueprint separable convolutional layer.

3. The method according to claim 2, characterized in that, the first blueprint separable convolutional layer and the second blueprint separable convolutional layer each include a standard convolutional layer and a depth convolutional layer.

4. The method according to claim 1, characterized in that, the spatial attention module sequentially includes a first standard convolutional layer, a strided convolutional layer, a pooling layer, a grouped convolutional layer, an upsampling layer, a second standard convolutional layer, and an activation layer. Wherein the features after the activation layer are fused with the input features of the spatial attention module and used as the output of the spatial attention module. The second standard convolutional layer extracts features from the output of the upsampling layer and the output of the first standard convolutional layer.

5. The method according to claim 1, characterized in that, the contrast perception channel attention module includes a contrast layer, two standard convolutional layers, and activation processing. Wherein the features after the activation processing are fused with the input features of the contrast perception channel attention module and used as the output. The contrast layer is used to obtain the contrast information between channels.

6. The method according to claim 3, It is characterized in that For each of the distillation modules, it includes three levels of feature extraction. The standard convolutional layers included in each level perform 1x1 convolutional operations, and the depth convolutional layers in the first blueprint separation convolutional layer and the second blueprint separation convolutional layer perform 3x3 convolutional operations.

7. The method according to claim 1, It is characterized in that The super-resolution model is trained using the mean absolute error loss function.

8. A computer-readable storage medium, on which a computer program is stored, wherein When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

9. A computer device, including a memory and a processor, and a computer program capable of running on the processor is stored on the memory, It is characterized in that When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Mine image super-resolution reconstruction method and system based on multi-scale residual network

    CN113592718A

  • Super-resolution reconstruction method based on multi-scale residual attention

    CN114331830A