An acceleration method and system for a single-image super-resolution model
By decomposing the single image super-resolution model into a basic network and a refinement network, and adding a mask prediction module, dynamic convolution adapts to the super-resolution situation of feature, solving the problems of large amount of computation and complex structure of the super-resolution model in the prior art when deploying the mobile terminal, achieving efficient acceleration effect.
Patent Information
- Application Number
- CN202210867107.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-22
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-07-22
AI Technical Summary
When deploying the existing lightweight super-resolution network model on mobile devices, the calculation is large, the structure is complex, and the optimization of efficient operators is difficult to achieve, resulting in cumbersome model deployment and the inability to universally accelerate the single-image super-resolution model.
By acquiring a single image super-resolution model and decomposing it into a basic network and a refined network, adding a mask prediction module, dynamic convolution adapts to the super-resolution situation of features, intelligently selecting feature regions that need further processing to reduce the amount of calculation and achieve acceleration.
The acceleration of the single image super-resolution model is achieved, reducing the calculation amount by 10% to 48%, and only increasing the number of parameters is relatively small, which is highly universal, helping the super-resolution method to be implemented on the mobile terminal.
Smart Images

Figure CN115131212B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image super - resolution, and particularly to an acceleration method and system for a single - image super - resolution model. Background Art
[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] Single - image super - resolution refers to super - resolving a corresponding high - resolution image from a low - resolution image. The improvement of the performance of the image super - resolution network model mainly relies on designing a larger network model. These network models with huge computational amounts can hardly be applied to mobile devices. Therefore, for better application, recent super - resolution methods are all lightweight methods. The lightweight super - resolution network model aims to achieve efficient super - resolution performance while reducing the model's computational amount.
[0004] Currently, the construction of existing lightweight super - resolution network models is achieved by relying on artificially designed high - efficiency modules. Therefore, the following defects exist.
[0005] 1. These high - efficiency modules need to be obtained through complex designs and a large amount of exploration. Therefore, the cost of obtaining the relevant lightweight super - resolution network model is high, and the structure is usually complex.
[0006] 2. The generation of high - efficiency modules is usually due to the generation of a high - efficiency complex operator, and the optimization of these operators has not been realized, resulting in poor optimization of the model when deployed to the underlying layer.
[0007] 3. A certain lightweight super - resolution network structure is designed specifically for a certain high - efficiency operator. These high - efficiency operators are complex and have poor portability, and cannot be universally plug - and - play to make the existing super - resolution methods lighter and maintain similar super - resolution performance.
[0008] 4. Currently, good large models cannot be directly deployed to mobile devices, and problems occur in actual applications.
[0009] 5. Lightweight models usually come with the generation of new lightweight operators, and the optimization of these new complex lightweight operators at the underlying layer is very difficult. Therefore, the deployment of lightweight models is also relatively cumbersome.
[0010] Therefore, although lightweight network models can maintain low computational amounts and have a certain super - resolution ability, the lightweight operators among them usually cannot be well used in other existing models to achieve the purpose of acceleration. How to universally accelerate the single - image super - resolution model has become an urgent problem to be solved at the present stage. Summary of the Invention
[0011] In view of the deficiencies in the prior art, the purpose of the present invention is to provide an acceleration method and system for a single-image super-resolution model. Based on data-driven dynamic convolution and deep learning technologies, it acts on the existing single-image super-resolution model, reducing the computational complexity of the current single-image super-resolution model and achieving the purpose of acceleration. It helps the single-image super-resolution method to be implemented and deployed on mobile devices, and has high universality.
[0012] To achieve the above object, the present invention is implemented through the following technical solutions:
[0013] The first aspect of the present disclosure provides an acceleration method for a single-image super-resolution model, including the following steps:
[0014] Obtain a single-image super-resolution model of a single refined network, decompose the model to obtain a basic network and a refined network;
[0015] Obtain the intermediate feature map and the rough super-resolution image obtained by the basic network for the image to be processed;
[0016] Perform mask prediction on the intermediate feature map, output a mask, cut out the area to be further processed according to the mask as the input of the refined network, and the output is a super-resolution image patch;
[0017] Replace the corresponding patches in the rough super-resolution image with the super-resolution image patches to obtain the final super-resolution image, and process the remaining unprocessed areas of the refined network.
[0018] Further, the steps for obtaining the mask are as follows: input the intermediate feature map into two convolutional layers and a global pooling layer respectively to obtain a spatial intermediate feature map; then apply the Softmax operation in the channel dimension of the corresponding feature blocks in the spatial intermediate feature map to obtain a spatial mask; the intermediate feature maps in the two channel dimensions of the spatial mask are respectively recorded as the processed mask and the unprocessed mask according to the well-processed feature blocks and the under-processed feature blocks, and the unprocessed mask is fed into the dynamic convolution to generate the final mask.
[0019] Furthermore, in the spatial mask, 0 represents the area that has been processed well, and 1 represents the area that needs to be further processed.
[0020] Furthermore, the specific steps for determining the area that needs to be further processed are as follows: select the K elements with the largest values from the mask, and crop the corresponding feature blocks from the intermediate feature map according to the coordinates of the K elements to obtain a set of feature blocks.
[0021] Further, according to the position information of the mask, replace the corresponding patches in the rough super-resolution image with the super-resolution image patches, and output the final super-resolution image.
[0022] The second aspect of the present disclosure provides an acceleration system for a single-image super-resolution model, including:
[0023] A data acquisition module, configured to acquire a single-image super-resolution model of a single refinement network, decompose the model to obtain a basic network and a refinement network;
[0024] A basic network module, configured to acquire an intermediate feature map and a rough super-resolved image obtained by the basic network for the image to be processed;
[0025] A mask prediction module, configured to perform mask prediction on the intermediate feature map and output a mask;
[0026] A refinement network module, configured to cut out the area to be further processed according to the mask as the input of the refinement network, and output a super-resolved image patch;
[0027] A filling-back module, configured to replace the corresponding patch in the rough super-resolved image with the super-resolved image patch to obtain the final super-resolved image, and process the remaining unprocessed area of the refinement network.
[0028] The third aspect of the present disclosure provides an acceleration method for a single-image super-resolution model, including the following steps:
[0029] Acquire a single-image super-resolution model of a multi-refinement network, decompose the single-image super-resolution model to obtain a basic network and multiple refinement networks;
[0030] Acquire an intermediate feature map and a rough super-resolved image obtained by the basic network for the image to be processed;
[0031] Perform mask prediction on the previous intermediate feature map, output a mask, and cut out the area to be further processed according to the mask as the input of the refinement network;
[0032] Replace the corresponding patch in the rough super-resolved image with the super-resolved image patch to obtain the final super-resolved image, and process the remaining unprocessed area of the refinement network.
[0033] Further, cut out the area to be further processed according to the mask as the input of the refinement network. If the refinement network is the last one, output a super-resolved image patch. If the refinement network is not the last one, output a further processed feature patch and fill it back to the position corresponding to the original feature to obtain another intermediate feature map.
[0034] The fourth aspect of the present disclosure provides an acceleration system for a single-image super-resolution model, including:
[0035] A data acquisition module, configured to acquire a single-image super-resolution model of a multi-refinement network, decompose the single-image super-resolution model to obtain a basic network and multiple refinement networks;
[0036] A basic network module, configured to acquire an intermediate feature map and a rough super-resolution image obtained by the basic network for the image to be processed;
[0037] A mask prediction module, configured to perform mask prediction on the previous intermediate feature map and output a mask;
[0038] A refinement network module, configured to cut out the area to be further processed according to the mask as the input of the refinement network;
[0039] A filling-back module, configured to replace the corresponding small block in the rough super-resolution image with the super-resolved image small block to obtain the final super-resolution image and process the remaining unprocessed area of the refinement network.
[0040] Furthermore, there will be a mask prediction module between the basic network and the refinement network and between the refinement networks.
[0041] The beneficial effects of the above embodiments of the present invention are as follows:
[0042] The present invention obtains a single-image super-resolution method, decomposes the single-image super-resolution method to obtain a basic network and a refinement network, adds a mask prediction module between the basic network and the refinement network or between the refinement networks to obtain an accelerated image super-resolution model, realizes the acceleration of a single-refinement network and a multi-refinement network, and has high universality.
[0043] The present invention can accelerate a given super-resolution method while maintaining similar method performance, reduce the computational amount of the method to 10% - 48% of the original, and only bring a small increase in the number of parameters (about 40K).
[0044] The present invention uses dynamic convolution to adaptively enable the model to intelligently select the feature area to be further processed according to the super-resolution situation of the features themselves. Description of the Drawings
[0045] The specification drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.
[0046] Figure 1 It is a flowchart of a single-refinement network for the single-image super-resolution model acceleration method in Embodiment 1 of the present invention;
[0047] Figure 2It is a flowchart of the multi-refinement network for the single-image super-resolution model acceleration method in the third embodiment of the present invention. Detailed implementation manners
[0048] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0049] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of intermediate feature maps, steps, operations, devices, components, and / or combinations thereof;
[0050] Embodiment 1:
[0051] Embodiment 1 of the present disclosure provides an acceleration method for a single-image super-resolution model, which is applicable to a single-refinement network. As Figure 1 shown, the existing single-image super-resolution model is decomposed into two parts, a base network and a refinement network. The low-quality image is initially fed into the base network to obtain an intermediate feature map F c , and the intermediate feature map F c is upsampled to obtain a rough super-resolution image I c . Then, the mask prediction process generates a mask M through the intermediate feature map F c . Under the guidance of this mask M, the mechanism will pick out the regions in the intermediate feature map F c that are not well super-resolved and crop out the corresponding feature blocks. These feature blocks are fed into the refinement network for further processing. Finally, the super-resolution image patches obtained from these feature blocks through the refinement network are used to replace the corresponding image patches in I c , so as to obtain the final super-resolution image I SR .
[0052] The specific process is as follows:
[0053] Obtain a single-image super-resolution model of a single-refinement network, decompose the model to obtain a base network and a refinement network.
[0054] Obtain the intermediate feature map and the rough super-resolution image obtained by the base network for the image to be processed.
[0055] Preferably, the decomposition of the basic network and the refinement network is not fixed, and can be weighed according to the super-resolution ability of the model and the running speed of the model (or the degree of lightweight). The basic network processes the entire image, while the refinement network processes only a part of the image. The original network is equal to the decomposed network with the basic network being all and the refinement network being zero. If significant acceleration of the original network is needed, then most of the original network can be divided into the refinement network. At this time, the super-resolution ability will decrease to a certain extent, but overall it is comparable to the original network.
[0056] Preferably, the image to be processed is the image that needs super-resolution.
[0057] Preferably, the input of the basic network is the low-quality image I LR , and the purpose of the basic network is to extract the intermediate feature map F c , F c is upsampled to the rough super-resolution image I c , and the outputs are the intermediate feature map F c and the rough super-resolution image I c .
[0058] Perform mask prediction on the intermediate feature map, output a mask, cut out the area that needs further processing according to the mask as the input of the refinement network, and the output is the super-resolved image patch.
[0059] Preferably, during the mask prediction process, the input is F c ∈R H×W×C , where R represents the set of real numbers, H represents the height of the feature, W represents the width of the feature, and C represents the number of channels. The intermediate feature map F c is input into two 3×3 convolutions and a global pooling layer to obtain the spatial intermediate feature map where p≥1 is the spatial size of the feature block F k . Note that each element in F s corresponds to a feature block F k ∈R p×p×C . Then, a Softmax operation is applied to the channel dimension of F s to obtain the spatial mask The spatial mask M s The intermediate feature maps in the two channel dimensions in the spatial mask M are respectively denoted as the processed mask M s 1 and the under-processed mask M s 2 , representing the processed feature blocks and the under-processed feature blocks respectively. The under-processed mask is fed into the dynamic convolution to generate the final mask This mechanism selects the K elements with the largest values from the mask M. According to the coordinates of these K elements, from Fc corresponding feature blocks are cropped from it to obtain a set of feature blocks
[0060] Preferably, in the spatial mask, 0 represents the processed area, and 1 represents the area that needs further processing. After the learning and training of the mask, the processed area and the area that needs further processing are divided.
[0061] The training process is as follows: First, the original network is decomposed into a basic network and a refinement network, and this whole (basic network and refinement network) is called the decomposition network. The decomposition network is trained on a specific dataset, with a batch size of 32 on each GPU, and optimized using the Adam optimizer. The cosine annealing strategy is used to optimize the learning rate. The decomposition network is trained for a total of 600,000 iterations. At the same time, a full-cut network is defined. The full-cut network (including the basic network, the cutting part, and the refinement network) sends the intermediate feature map obtained by the basic network to the cutting part for cutting, and all the feature blocks are sent into the refinement network. Then, the trained decomposition network is used as a pre-training model to train the full-cut network. The total number of training iterations for this part is also 600,000. Finally, the trained full-cut network is used as a pre-training model to train the final acceleration network. Keeping the parameters of the basic network and the refinement network unchanged, only the mask prediction module is trained. Under such learning, the final mask prediction module can accurately predict the processed area and the unprocessed area. The above training process is specifically designed for this acceleration method.
[0062] Preferably, the input of the refinement network is the set of feature blocks The output is where P represents the super-resolved image patch. p represents the side length of the image patch, and the image patch is square. K represents K elements, and the K elements correspond to K image patches. Therefore, here K represents the number of image patches selected to be processed by the refinement network. k is used as a symbol in the set, representing a certain element, and k ranges from 1 to K. S represents the super-resolution magnification factor, which is the multiple of super-resolution required.
[0063] Because that is, the data processed by the refinement network is much smaller than the corresponding part in the original original network, so this method can greatly reduce the computational amount of the original model.
[0064] The super-resolved image patches are used to replace the corresponding patches in the roughly super-resolved image to obtain the final super-resolved image, and the remaining unprocessed areas of the refinement network are processed.
[0065] Preferably, the input and the roughly super-resolved image I c, according to the position information of the mask, these image blocks are then replaced with the corresponding image blocks in the coarse super-resolution image, and the output is the final super-resolution image I SR .
[0066] Embodiment 2:
[0067] Embodiment 2 of the present disclosure provides an acceleration system for a single image super-resolution model, which is applicable to a single refinement network and includes:
[0068] The data acquisition module is used to obtain the image features that currently need super-resolution, and is specifically configured to obtain a single image super-resolution model of a single refined network, decompose the model, and obtain a basic network and a refined network.
[0069] The basic network module is configured to obtain an intermediate feature map and a coarse super-resolution image obtained by the image to be processed through the basic network.
[0070] Specifically, the image features are further processed to obtain an intermediate feature map, and the intermediate feature map is upsampled to obtain a coarse super-resolution image. The overall output of the basic network module is an intermediate feature map and a coarse super-resolution image.
[0071] The mask prediction module is configured to perform mask prediction on the intermediate feature map and output a mask.
[0072] Specifically, a mask is predicted to locate the area that is not well super-resolved. The input is an intermediate feature map, which is combined with its own information through conventional convolution blocks and dynamic convolution blocks to generate a mask. The output is a mask, the corresponding area coordinates and the corresponding area feature block.
[0073] The refinement network module is configured to cut out the area that needs further processing according to the mask as the input of the refinement network, and output the small image block with good super resolution.
[0074] Preferably, the feature area that has not been super-resolved before is further processed, the input is a feature block of the area that needs to be further processed, and the output is an image block after further super-resolvation.
[0075] The filling-back module is configured to replace the corresponding small blocks in the coarse super-resolution image with the small blocks of the super-resolution image to obtain the final super-resolution image, and process the remaining unprocessed areas of the refinement network.
[0076] Specifically, it is used to correctly fill in the small image blocks to be further processed, and then replace the corresponding small image blocks in the coarse super-resolution image with these small image blocks according to the position information of the mask, so as to obtain the final super-resolution image.
[0077] Embodiment three:
[0078] Embodiment 3 of the present disclosure provides an acceleration method for a single image super-resolution model, which is applicable to a multi-refinement network, such as Figure 2 As shown in Figure 1, the existing single image super-resolution model is decomposed into multiple parts, a base network and t refinement networks. The low-quality image is first fed into the base network to obtain an intermediate feature map F c , the intermediate feature map F c After upsampling, a rough super-resolution image I is obtained c Then the nth (1≤n≤t) mask prediction module passes through the n-1th intermediate feature map F n-1 (F0 is F c ) to generate a mask M n In this M n Under the guidance of n-1 There is no super-resolution area and the corresponding feature blocks are cropped. These feature blocks are sent to the refinement network for further processing. The nth (1≤n≤t-1) refinement network further processes these feature blocks and fills the feature blocks back into F n-1 Get F n , the tth refinement network processes these feature blocks and then passes them through the upsampling network to obtain super-resolution image blocks. Finally, the super-resolution image blocks obtained by refining these feature blocks are used to replace I c The corresponding image block in the image is obtained to obtain the final super-resolution image I SR .
[0079] The specific process is:
[0080] A single image super-resolution model of a multi-refined network is obtained, and the single image super-resolution model is decomposed to obtain a basic network and multiple refined networks.
[0081] Preferably, the decomposition of the base network and the refined network is not fixed, and a trade-off can be made based on the super-resolution capability of the model and the running speed (or the degree of lightweightness) of the model. The base network processes the entire image, while the refined network processes a part of the image. The original network is equal to the decomposed network with the base network as the whole and the refined network as 0. If the original network needs to be significantly accelerated, then most of the original network can be divided into the refined network. At this time, the super-resolution capability will be reduced to a certain extent, but it is equivalent to the original network as a whole.
[0082] Preferably, the image to be processed is an image requiring super-resolution.
[0083] Preferably, the input of the base network is a low-quality image I LR , the purpose of the basic network is to extract the intermediate feature map F c , F cIs upsampled into a rough super-resolution image I c , and the output is the intermediate feature map F c and the rough super-resolution image I c .
[0084] Perform mask prediction on the previous intermediate feature map, output a mask, and cut out the area to be further processed according to the mask as the input of the refinement network.
[0085] Preferably, in the spatial mask, 0 represents the processed area, 1 represents the area to be further processed, and the processed area and the area to be further processed are divided after the learning and training of the mask.
[0086] Preferably, the input of the nth mask prediction module is F n-1 ∈R H×W×C , and the operation is the same as the mask prediction process of the single refinement network in Embodiment 1, and the output is This mechanism selects the K elements with the largest values from the mask M n . According to the coordinates of these K elements, the corresponding feature blocks are cropped from F n-1 to obtain the feature block set
[0087] The input of the nth (1≤n≤t - 1) refinement network is the feature block set and the feature F n-1 , and these feature blocks are further processed to obtain Then the feature blocks in this set are filled back into the corresponding positions in F n-1 to obtain F n . The input of the tth refinement network is the feature block set
[0088] The output is Here S represents the magnification factor of super-resolution.
[0089] Preferably, cut out the area to be further processed according to the mask as the input of the refinement network. If the refinement network is the last one, output the super-resolved image patches. If the refinement network is not the last one, output the further processed feature patches and fill them back to the corresponding positions of the original features to obtain another intermediate feature map.
[0090] Replace the corresponding patches in the rough super-resolution image with the super-resolved image patches to obtain the final super-resolution image, and process the remaining unprocessed areas of the refinement network.
[0091] Preferably, the input and the rough super-resolution image I c, according to the position information of the mask, these image patches are then used to replace the corresponding image patches in the rough super-resolution image, and the output is the final super-resolution image I SR .
[0092] Example 4:
[0093] Embodiment 4 of the present disclosure provides an acceleration system for a single-image super-resolution model, which is applicable to multiple refinement networks
[0094] including:
[0095] The data acquisition module is used to acquire the image features that need to be super-resolved currently, and is specifically configured to acquire a single-image super-resolution model of a single refinement network, decompose the model to obtain a basic network and a refinement network
[0096] Preferably, there is a mask prediction module between the basic network and the refinement network, and between the refinement networks
[0097] Preferably, the basic network, t mask prediction modules, t refinement networks, and the filling-back module together constitute the accelerated network
[0098] The basic network module is configured to obtain the intermediate feature map and the rough super-resolution image obtained by the basic network for the image to be processed
[0099] Specifically, the image features are further processed to obtain an intermediate feature map, the intermediate feature map is upsampled to obtain a rough super-resolution image, and the overall output of the basic network module is the intermediate feature map and a rough super-resolution image
[0100] The mask prediction module is configured to perform mask prediction on the intermediate feature map and output a mask
[0101] Specifically, a mask is predicted to locate the regions that are not super-resolved well. The input is the intermediate feature map, and the intermediate feature map passes through a conventional convolution block and a dynamic convolution block to combine its own information to generate a mask. The output is a mask, the corresponding region coordinates, and the corresponding region feature blocks
[0102] The refinement network module is configured to cut out the regions that need to be further processed according to the mask as the input of the refinement network, and the output is the super-resolved image patches
[0103] Preferably, the feature regions that were not super-resolved well before are further processed. The input is the feature patches of the regions that need to be further processed, and the output is the further super-resolved image patches
[0104] Preferably, if it is the last refinement network, the output is the further super-resolved image patches. If it is not the last refinement network, the output is the further processed feature patches
[0105] The filling-back module is configured to replace the corresponding small blocks in the roughly super-resolved image with the super-resolved small image blocks to obtain the final super-resolved image, and process the remaining unprocessed areas of the refinement network.
[0106] Specifically, to correctly fill back the further processed small image blocks, according to the position information of the mask, these small image blocks are then replaced with the corresponding small image blocks in the roughly super-resolved image to obtain the final super-resolved image.
[0107] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. An acceleration method for a single-image super-resolution model, characterized in that Including the following steps: Obtain a single-image super-resolution model of a single refinement network, decompose the model to obtain a base network and a refinement network; Obtain the intermediate feature map and the rough super-resolved image obtained by the base network for the image to be processed; Perform mask prediction on the intermediate feature map, output a mask, cut out the area that needs further processing according to the mask as the input of the refinement network, and the output is a super-resolved image patch; The steps for obtaining the mask are as follows: input the intermediate feature map into two convolutional layers and a global pooling layer respectively to obtain a spatial intermediate feature map; then apply the Softmax operation in the channel dimension of the corresponding feature patches of the spatial intermediate feature map to obtain a spatial mask; the intermediate feature maps in the two channel dimensions of the spatial mask are respectively denoted as the processed mask and the unprocessed mask according to the processed feature patches and the unprocessed feature patches, and the unprocessed mask is fed into the dynamic convolution to generate the final mask; Replace the corresponding patches in the rough super-resolved image with the super-resolved image patches to obtain the final super-resolved image, and process the remaining unprocessed areas of the refinement network.
2. The acceleration method of the single-image super-resolution model according to claim 1, characterized in that, In the spatial mask, 0 represents the area that has been processed well, and 1 represents the area that needs further processing.
3. The acceleration method of the single-image super-resolution model according to claim 2, wherein The specific steps for determining the area that needs further processing are as follows: select the K elements with the largest values from the mask, and crop the corresponding feature patches from the intermediate feature map according to the coordinates of the K elements to obtain a set of feature patches.
4. The acceleration method of the single-image super-resolution model according to claim 1, wherein According to the position information of the mask, replace the corresponding patches in the rough super-resolved image with the super-resolved image patches, and output the final super-resolved image.
5. An acceleration system for a single-image super-resolution model, characterized in that, Including: A data acquisition module configured to obtain a single-image super-resolution model of a single refinement network, decompose the model to obtain a base network and a refinement network; A base network module configured to obtain the intermediate feature map and the rough super-resolved image obtained by the base network for the image to be processed; A mask prediction module configured to perform mask prediction on the intermediate feature map and output a mask; The steps for obtaining the mask are as follows: input the intermediate feature map into two convolutional layers and a global pooling layer respectively to obtain a spatial intermediate feature map; then apply the Softmax operation in the channel dimension of the corresponding feature patches of the spatial intermediate feature map to obtain a spatial mask; the intermediate feature maps in the two channel dimensions of the spatial mask are respectively denoted as the processed mask and the unprocessed mask according to the processed feature patches and the unprocessed feature patches, and the unprocessed mask is fed into the dynamic convolution to generate the final mask; A refinement network module configured to cut out the area that needs further processing according to the mask as the input of the refinement network, and the output is a super-resolved image patch; A filling-back module configured to replace the corresponding patches in the rough super-resolved image with the super-resolved image patches to obtain the final super-resolved image, and process the remaining unprocessed areas of the refinement network.
6. An acceleration method for a single-image super-resolution model, characterized in that, Including the following steps: Obtain a single-image super-resolution model of a multi-refinement network, decompose the single-image super-resolution model to obtain a base network and multiple refinement networks; Obtain the intermediate feature map and the rough super-resolved image obtained by the base network for the image to be processed; Perform mask prediction on the previous intermediate feature map, output a mask, and cut out the area to be further processed according to the mask as the input of the refinement network; The steps for obtaining the mask are as follows: Input the intermediate feature map into two convolutional layers and a global pooling layer respectively to obtain a spatial intermediate feature map; then apply the Softmax operation to the channel dimension of the corresponding feature blocks in the spatial intermediate feature map to obtain a spatial mask; the intermediate feature maps in the two channel dimensions of the spatial mask are respectively recorded as the processed mask and the unprocessed mask according to the well-processed feature blocks and the unprocessed feature blocks, and the unprocessed mask is fed into the dynamic convolution to generate the final mask; Cut out the area to be further processed according to the mask as the input of the refinement network. If the refinement network is the last one, output the super-resolved image patches. If the refinement network is not the last one, output the further processed feature patches and fill them back to the corresponding positions of the original features to obtain another intermediate feature map; Replace the corresponding patches in the roughly super-resolved image with the super-resolved image patches to obtain the final super-resolved image, and process the remaining unprocessed areas of the refinement network.
7. An acceleration system for a single-image super-resolution model, characterized in that, Include: A data acquisition module, configured to obtain a single-image super-resolution model of a multi-refinement network, decompose the single-image super-resolution model to obtain a basic network and multiple refinement networks; A basic network module, configured to obtain the intermediate feature map and the roughly super-resolved image obtained by the basic network for the image to be processed; A mask prediction module, configured to perform mask prediction on the previous intermediate feature map and output a mask; The steps for obtaining the mask are as follows: Input the intermediate feature map into two convolutional layers and a global pooling layer respectively to obtain a spatial intermediate feature map; then apply the Softmax operation to the channel dimension of the corresponding feature blocks in the spatial intermediate feature map to obtain a spatial mask; the intermediate feature maps in the two channel dimensions of the spatial mask are respectively recorded as the processed mask and the unprocessed mask according to the well-processed feature blocks and the unprocessed feature blocks, and the unprocessed mask is fed into the dynamic convolution to generate the final mask; A refinement network module, configured to cut out the area to be further processed according to the mask as the input of the refinement network; Cut out the area to be further processed according to the mask as the input of the refinement network. If the refinement network is the last one, output the super-resolved image patches. If the refinement network is not the last one, output the further processed feature patches and fill them back to the corresponding positions of the original features to obtain another intermediate feature map; A filling-back module, configured to replace the corresponding patches in the roughly super-resolved image with the super-resolved image patches to obtain the final super-resolved image, and process the remaining unprocessed areas of the refinement network.
8. The acceleration system of the single-image super-resolution model according to claim 7, characterized in that, There will be a mask prediction module between the basic network and the refinement network and between the refinement networks.
Citation Information
Patent Citations
Image super-resolution reconstruction method based on progressive perception and ultra-lightweight network
CN113096015A
Lightweight network model based on single image super-resolution, and processing method
CN113781304A