A super-resolution network structure, super-resolution module, device, and reconstruction method
By accessing the spatial scaling module and the lightweight module in the super-resolution network structure, the problem of large amount of computing and memory reading and writing in the prior art is solved, and real-time image super-resolution processing on electronic devices is realized.
Patent Information
- Application Number
- CN202110483391.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-30
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-04-30
AI Technical Summary
The existing convolutional neural network super-segment technology has large calculation amount, parameter amount and memory read and write amount, making it difficult to realize real-time image super-resolution processing on electronic devices with weak computing power such as smartphones.
A super-resolution network structure is designed, and by accessing the spatial scaling module at the beginning and end of the network, and adopting a lightweight super-segment module such as FastIMDB, the convolutional calculation amount and memory read and write amount are reduced, including the first spatial scaling module, the feature extraction module and the second spatial scaling module, and the sub-pixel operation and the lightweight convolution layer structure are used.
The calculation amount of network structure and memory read and write amount are reduced, the convergence speed of model training is improved, and real-time image super-resolution processing is realized on electronic devices.
Smart Images

Figure CN113222816B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of neural network technologies, and in particular, to a super-resolution network structure, a super-resolution module, a device, and a reconstruction method. Background Art
[0002] Image resolution is an important indicator for measuring image quality. The higher the resolution, the finer the details, the better the quality, and the richer the information content of the image. Among them, image super-resolution technology can be used to improve image resolution and enhance image performance.
[0003] Most current super-resolution technologies are based on convolutional neural network technology. A data pair is formed by using a high-definition image and a low-definition image, and then a large number of data pairs are used to train a neural network model. For example, in a typical neural network super-resolution technology, a dense connection convolutional block structure similar to the Residual in Residual Dense Block (RRDB) is used. However, the computational amount, the number of parameters, and the memory read / write amount of the convolutional neural network used in existing super-resolution technologies are all large, making it difficult to deploy on electronic devices with relatively weak computing power (such as smartphones), and it is difficult for the model inference time to meet the real-time requirement. Summary of the Invention
[0004] This application provides a super-resolution network structure, a super-resolution module, a device, and a reconstruction method, which can reduce the computational amount and memory read / write amount of the network structure, thereby reducing the model inference time of the network structure and achieving the purpose of real-time inference.
[0005] To achieve the above object, the technical solution of this application is implemented as follows:
[0006] In a first aspect, an embodiment of this application provides a super-resolution network structure, which may include a first spatial scaling module, a feature extraction module, and a second spatial scaling module; wherein,
[0007] The first spatial scaling module is configured to perform transformation processing on an input low-resolution image according to sub-pixel operation to generate a first processed image;
[0008] The feature extraction module is configured to perform feature extraction on the first processed image, and generate a second processed image according to the obtained target feature map and the first processed image;
[0009] The second spatial scaling module is configured to perform transformation processing on the second processed image according to sub-pixel operation to generate a high-resolution image.
[0010] In a second aspect, an embodiment of the present application provides a super-resolution module, which may include a third convolutional module, a second connection layer, and a fourth convolutional module; wherein,
[0011] The third convolutional module is configured to perform hierarchical convolution operations on the first feature map using a plurality of convolutional layers to generate multiple sub-feature maps;
[0012] The second connection layer is configured to fuse the multiple sub-feature maps to generate a fused sub-feature map;
[0013] The fourth convolutional module is configured to perform a convolution operation on the fused feature map and output a target sub-feature map.
[0014] In a third aspect, an embodiment of the present application provides an electronic device, which includes the super-resolution network structure as described in the first aspect.
[0015] In a fourth aspect, an embodiment of the present application provides a method for super-resolution reconstruction, which is applied to the electronic device as described in the third aspect. The method may include:
[0016] Determine the super-resolution network structure;
[0017] Input a low-resolution image into the super-resolution network structure for operation to obtain a high-resolution image.
[0018] A super-resolution network structure, a super-resolution module, a device, and a reconstruction method provided by an embodiment of the present application. The super-resolution network structure may include a first spatial scaling module, a feature extraction module, and a second spatial scaling module; wherein, the first spatial scaling module is configured to perform transformation processing on the input low-resolution image according to sub-pixel operations to generate a first processed image; the feature extraction module is configured to extract features from the first processed image and generate a second processed image based on the extracted target feature map and the first processed image; the second spatial scaling module is configured to perform transformation processing on the second processed image according to sub-pixel operations to generate a high-resolution image. In this way, since spatial scaling modules are connected at the beginning and end of the network structure, it can ensure that all convolution calculations in the feature extraction module are performed in a small spatial domain, reducing the computational amount and memory read / write amount of the network structure; moreover, a lightweight super-resolution module is designed, which can further reduce the computational amount and memory read / write amount of the network structure, thereby improving the convergence speed during model training and reducing the model inference time of the network structure, and achieving the purpose of real-time inference. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a schematic diagram of the composition of a super-resolution network structure provided by an embodiment of the present application;
[0020] Figure 2Schematic diagram of the composition of another super-resolution network structure provided by an embodiment of the present application;
[0021] Figure 3 Schematic diagram of the composition of a super-resolution module provided by an embodiment of the present application;
[0022] Figure 4 Detailed schematic diagram of the composition of a super-resolution network structure provided by an embodiment of the present application;
[0023] Figure 5 Detailed schematic diagram of the composition of a super-resolution module provided by an embodiment of the present application;
[0024] Figure 6 Schematic diagram of the composition structure of an electronic device provided by an embodiment of the present application;
[0025] Figure 7 Schematic diagram of the process of a super-resolution reconstruction method provided by an embodiment of the present application;
[0026] Figure 8 Schematic diagram of the composition structure of another electronic device provided by an embodiment of the present application. Detailed implementation manners
[0027] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. It can be understood that the specific embodiments described herein are only used to explain the related application, rather than limiting the application. In addition, it should be noted that for the sake of description, only parts related to the related application are shown in the drawings.
[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0029] In the following description, "some embodiments" are involved, which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict. It should also be pointed out that the terms "first / second / third" involved in the embodiments of the present application are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0030] It should be understood that image resolution is an important indicator for measuring image quality. The higher the resolution, the finer the details, the better the quality, and the more abundant the information contained in the image. Therefore, images with higher resolution have very important application values and research prospects in various computer vision tasks, such as military security, satellite monitoring, traffic supervision, etc.
[0031] Image super-resolution technology can improve image resolution by using signal processing methods, which is an effective way to improve image resolution and image performance. Most of the current super-resolution technologies are based on convolutional neural network technology. A data pair is composed of a high-definition image and a low-definition image, and then a large number of data pairs are used to train the neural network model. For example, in a typical neural network super-resolution technology solution, a dense connection convolutional block structure similar to the Residual in Residual Dense Block (RRDB) is used. Here, in the RRDB, more layers and connections are used, resulting in a very large number of connection layers (Concat), making the network structure more complex.
[0032] That is to say, the convolutional neural networks used in existing super-resolution technologies have large amounts of computational complexity, parameter quantity, and memory read / write volume, which are difficult to deploy for electronic devices with relatively weak computing power (such as smartphones), and the model inference time is difficult to meet the real-time requirement.
[0033] The embodiment of the present application provides a super-resolution network structure, which may include a first spatial scaling module, a feature extraction module, and a second spatial scaling module; wherein, the first spatial scaling module is used to perform transformation processing on the input low-resolution image according to sub-pixel operation to generate a first processed image; the feature extraction module is used to extract features from the first processed image, and generate a second processed image according to the extracted target feature map and the first processed image; the second spatial scaling module is used to perform transformation processing on the second processed image according to sub-pixel operation to generate a high-resolution image. In this way, since spatial scaling modules are connected at the beginning and end of the network structure, it can ensure that all convolutional calculations in the feature extraction module are performed in a small spatial domain, reducing the computational complexity and memory read / write volume of the network structure; moreover, a lightweight super-resolution module is designed, which can further reduce the computational complexity and memory read / write volume of the network structure, thereby improving the convergence speed during model training, reducing the model inference time of the network structure, and achieving the purpose of real-time inference.
[0034] The following will describe each embodiment of the present application in detail with reference to the accompanying drawings.
[0035] In an embodiment of the present application, refer to Figure 1, which shows a schematic diagram of the composition of a super-resolution network structure provided by an embodiment of the present application. As Figure 1 shown, the super-resolution network structure 10 may include a first spatial scaling module 11, a feature extraction module 12, and a second spatial scaling module 13; among them,
[0036] The first spatial scaling module 11 is used to perform transformation processing on the input low-resolution image according to sub-pixel operation to generate a first processed image;
[0037] The feature extraction module 12 is used to extract features from the first processed image, and generate a second processed image according to the obtained target feature map and the first processed image;
[0038] The second spatial scaling module 13 is used to perform transformation processing on the second processed image according to sub-pixel operation to generate a high-resolution image.
[0039] It should be noted that the input of the super-resolution network structure 10 is a low-resolution (LR) image, and the output is a high-resolution (HR) image; the function of this network structure is to convert the input low-resolution image into a high-resolution image.
[0040] In the embodiment of the present application, in order to achieve the purpose of light weight, spatial scaling modules may be connected to the start end and the end end of the super-resolution network structure 10 respectively, so as to ensure that all convolution calculations are performed in a small spatial domain. Specifically, in some embodiments, the first spatial scaling module 11 is located before the feature extraction module 12, and the second spatial scaling module 13 is located after the feature extraction module 12; among them,
[0041] The first spatial scaling module 11 is specifically used to perform spatial reduction processing on the input low-resolution image according to sub-pixel operation to obtain a first processed image;
[0042] The second spatial scaling module 13 is specifically used to perform spatial expansion processing on the second processed image according to sub-pixel operation to obtain a high-resolution image.
[0043] In the embodiment of the present application, the first spatial scaling module 11 and the second spatial scaling module 13 are PixShuffle modules with a preset multiple.
[0044] In a specific example, the value of the preset multiple is equal to 8.
[0045] Here, for the preset multiple, the value of the preset multiple may also be other values, and the embodiment of the present application does not make any limitation.
[0046] It should be noted that PixShuffle can also be called Pixelshuffle, and its implemented function is as follows: a low-resolution input image of H×W is transformed into a high-resolution input image of rH×rW through Sub-pixel operation. However, the implementation process does not directly generate this high-resolution image through interpolation or other means, but first obtains a feature map with r 2 channels (the size of the feature map is the same as that of the input low-resolution image) through convolution, and then obtains this high-resolution image through the method of periodic shuffling; among them, r can be the magnification factor of the image. Specifically, for the feature map, the number of channels is r 2 , and the r 2 channels of each pixel are rearranged into an r×r area, corresponding to an r×r-sized sub-block in the high-resolution image, so that the feature image of size r 2 ×H×W is rearranged into a high-resolution image of size 1×rH×rW.
[0047] It should also be noted that at the starting end of the super-resolution network structure 10, that is, the first spatial scaling module 11, it is mainly used to reduce the spatial domain to ensure that all convolution calculations in the subsequent feature extraction module 12 are performed in a small spatial domain. Exemplarily, if the input image is 1×8×8, it can be transformed into 64×1×1. At this time, r is a value less than 1, so that the height and width of the image become smaller to reduce the spatial domain. After such operations in the feature extraction module 12, at the end of the super-resolution network structure 10, that is, the second spatial scaling module 13, it is mainly used to increase the spatial domain to improve the image resolution. Exemplarily, if the input image is 64×8×8, it can be transformed into 1×64×64. At this time, r is a value greater than 1, so that the height and width of the image become larger to improve the image resolution.
[0048] Furthermore, in some embodiments, as Figure 2 shown, based on the super-resolution network structure 10 shown in Figure 1 , the feature extraction module 12 may include a first convolution module 121, at least one super-resolution module 122, a first connection layer 123, and a second convolution module 124; among them,
[0049] The first convolution module 121 is used to perform shallow feature extraction on the first processed image to obtain a first feature map;
[0050] At least one super-resolution module 122 is used to perform deep feature extraction on the first feature map to obtain at least one second feature map;
[0051] The first connection layer 123 is used to fuse at least one second feature map to obtain a third feature map;
[0052] The second convolution module 124 is used to perform a convolution operation with reduced number of channels on the third feature map to obtain a target feature map.
[0053] It should be noted that for the first convolution module 121 and the second convolution module 124, both can be composed of convolution layers. In existing network structures, most convolution layers are the second convolution layer, that is, the Conv-3 layer. However, considering the problem of too large convolution receptive field, in the embodiments of the present application, by introducing several first convolution layers, that is, the Conv-1 layer, not only can the model calculation amount be reduced, but also the problem of too large convolution receptive field can be improved.
[0054] In the embodiments of the present application, the first convolution module 121 may include a first convolution layer, and the second convolution module 124 may include a first convolution layer and a second convolution layer; wherein, the convolution kernels of the first convolution layer and the second convolution layer are different.
[0055] Further, in a specific example, the convolution kernel of the first convolution layer is 1×1, and the convolution kernel of the second convolution layer is 3×3. In this way, the first convolution layer can be represented by the Conv-1 layer, and the second convolution layer can be represented by the Conv-3 layer.
[0056] It should also be noted that in some embodiments, as Figure 2 shown, the feature extraction module 12 may further include a first adder 125; wherein,
[0057] The first adder 125 is used to perform an addition operation on the target feature map and the first processed image to obtain a second processed image.
[0058] In the embodiments of the present application, for the feature extraction module 12, the Conv-1 layer is used to perform a convolution operation on the first processed image to implement shallow feature extraction and obtain a first feature map; then at least one super-resolution module is used to perform deep feature extraction on the first feature map to obtain at least one second feature map; then a connection layer (i.e., the Concat layer) is used to fuse at least one second feature map to obtain a third feature map; then convolution operations are sequentially performed through the Conv-1 layer and the Conv-3 layer to reduce the number of channels to obtain a target feature map; finally, the first adder 125 performs an addition operation on the obtained target feature map and the first processed image input by the feature extraction module 12 to obtain the second processed image output by the feature extraction module 12.
[0059] In a specific example, the number of super-resolution modules included in the feature extraction module 12 is three.
[0060] It should be noted that in the super-resolution network structure 10, for these three super-resolution modules, for the first super-resolution module, its input is the output of the first convolutional module 121 (i.e., the first feature map), and its output will be divided into two parts. One part of the output jumps and connects to the first connection layer 123, and the other part serves as the input of the second super-resolution module; then for the second super-resolution module, its input is the output of the first super-resolution module, and its output will also be divided into two parts. One part of the output jumps and connects to the first connection layer 123, and the other part serves as the input of the third super-resolution module; then for the third super-resolution module, its input is the output of the second super-resolution module, and its output will be directly connected to the first connection layer 123; it can also be said that the outputs of these three super-resolution modules will all be connected to the first connection layer 123 so that the three obtained feature maps can be fused.
[0061] It should also be noted that the super-resolution network structure 10 described in the embodiments of the present application can also ensure that the output channel number of all convolution operations is a multiple of 32 to improve the computing efficiency of the hardware; and keeping the output channel number at a relatively large quantity is also beneficial for subsequent pruning optimization.
[0062] In a specific example, in the first convolutional module 121, the output channel number of the first convolutional layer is 96; in at least one super-resolution module 122, the output channel number of each super-resolution module is 128; in the first connection layer 123, the output channel number is calculated by accumulating the output channel numbers of all super-resolution modules. Assuming there are three super-resolution modules, then the output channel number at this time is 128 + 128 + 128 = 384; in the second convolutional module 124, the output channel number of the first convolutional layer is 128, and the output channel number of the second convolutional layer is 64.
[0063] This embodiment provides a lightweight super-resolution network structure. The super-resolution network structure includes a first spatial scaling module, a feature extraction module, and a second spatial scaling module; wherein, the first spatial scaling module is used to perform transformation processing on the input low-resolution image according to sub-pixel operation to generate a first processed image; the feature extraction module is used to extract features from the first processed image, and generate a second processed image according to the extracted target feature map and the first processed image; the second spatial scaling module is used to perform transformation processing on the second processed image according to sub-pixel operation to generate a high-resolution image. In this way, since spatial scaling modules are connected at the beginning and end of the network structure, it can be ensured that all convolution calculations in the feature extraction module are performed in a small spatial domain, thereby reducing the computational amount and memory read / write amount of the network structure, and reducing the model inference time of the network structure to achieve the purpose of real-time inference.
[0064] In another embodiment of the present application, refer to Figure 3, which shows a schematic diagram of the composition of a super-resolution module provided by an embodiment of the present application. As Figure 3 shown, the super-resolution module 30 may include a third convolution module 31, a second connection layer 32, and a fourth convolution module 33; wherein,
[0065] The third convolution module 31 is used to perform hierarchical convolution operations on the first feature map using a plurality of convolution layers to generate multiple layers of sub-feature maps;
[0066] The second connection layer 32 is used to fuse the multiple layers of sub-feature maps to generate a fused sub-feature map;
[0067] The fourth convolution module 33 is used to perform convolution operations on the fused feature map and output a target sub-feature map.
[0068] It should be noted that, in a possible implementation manner, the super-resolution module may be a residual dense block (RRDB). However, since RRDB uses more layers and connections, it involves a very large number of connection layers (Concat), making the network structure more complex. Based on this, the embodiments of the present application also provide another possible implementation manner.
[0069] In another possible implementation manner, in order to better achieve the lightweight of the network structure, the super-resolution module 30 here may be a lightweight super-resolution block, denoted as FastIMDB. In FastIMDB, only one Concat layer (i.e., the second connection layer 32) is involved, and then the Concat layer is used to fuse the multiple layers of sub-feature maps. Due to the reduction of the number of Concat layers, the computational amount and the memory read / write amount of the network structure can also be reduced, and the model inference time of the network structure can be further reduced.
[0070] It should also be noted that, for the third convolution module 31 and the fourth convolution module 33, both of them are also composed of convolution layers. In some embodiments, the third convolution module 31 may include three first convolution layers and four second convolution layers, and the fourth convolution module 33 may include one first convolution layer; wherein, the convolution kernels of the first convolution layer and the second convolution layer are different.
[0071] Further, in a specific example, the convolution kernel of the first convolution layer is 1×1, and the convolution kernel of the second convolution layer is 3×3. In this way, the first convolution layer can be represented by the Conv-1 layer, and the second convolution layer can be represented by the Conv-3 layer.
[0072] Further, hierarchical convolution operations are performed on the first feature map using a plurality of first convolution layers and second convolution layers to generate multiple layers of sub-feature maps. In a specific example, the number of layers of sub-feature maps is four, but the embodiments of the present application do not make any limitations.
[0073] Exemplarily, taking four layers as an example, for the third convolutional module 31, its input is divided into two paths. One path enters the first layer, and the first layer only includes one Conv-3 layer. The output of the Conv-3 layer in the first layer is connected to the second connection layer 32. The other path enters the second layer, and the second layer includes one Conv-1 layer and one Conv-3 layer. After passing through the Conv-1 layer in the second layer, its output is divided into two paths again. One path enters the Conv-3 layer in the second layer, and the output of the Conv-3 layer in the second layer is connected to the second connection layer 32. The other path enters the third layer, and the third layer also includes one Conv-1 layer and one Conv-3 layer. After passing through the Conv-1 layer in the third layer, its output is divided into two paths again. One path enters the Conv-3 layer in the third layer, and the output of the Conv-3 layer in the third layer is connected to the second connection layer 32. The other path enters the fourth layer, and the fourth layer also includes one Conv-1 layer and one Conv-3 layer. The output after sequentially passing through the Conv-1 layer and the Conv-3 layer in the fourth layer is also connected to the second connection layer 32. That is to say, only one Concat layer is needed here, and this Concat layer can receive four sub-feature maps, and then it fuses these four sub-feature maps.
[0074] It should also be noted that in some embodiments, as Figure 3 shown, the super-resolution module 30 may further include a second adder 34; wherein,
[0075] The second adder 34 is used to perform an addition operation on the target sub-feature map and the first feature map to generate a second feature map.
[0076] In the embodiments of the present application, after the convolution operation of the third convolutional module 31, the obtained multi-layer sub-feature maps can be fused through the second connection layer 32 to obtain a fused sub-feature map; then through the convolution operation of the fourth convolutional module 33 (i.e., the Conv-1 layer), a target sub-feature map can be obtained; finally, the second adder 34 performs an addition operation on the obtained target sub-feature map and the first feature map input by the super-resolution module 30 to obtain the second feature map output by the super-resolution module 30.
[0077] It should also be noted that the output channel number of all convolution operations can be a multiple of 32, and the output channel number can also be maintained at a relatively large quantity to facilitate improving the operation efficiency of the hardware and subsequent pruning optimization.
[0078] In a specific example, the number of output channels of each layer of sub-feature maps is 32. The number of output channels of the second connection layer is calculated by accumulating the number of output channels of all layers of sub-feature maps. Assuming there are four layers of sub-feature maps, the number of output channels at this time is 32 + 32 + 32 + 32 = 128. In addition, the number of output channels of the super-resolution module 30 is also 128.
[0079] This embodiment provides a lightweight super-resolution module, which can be applied to the super-resolution network structure 10. In this way, for the super-resolution network structure 10, not only a spatial scaling module is connected at the head and tail of the network structure, which can ensure that all convolution calculations in the feature extraction module are performed in a small spatial domain, reducing the computational amount and memory read / write amount of the network structure; but also a lightweight super-resolution module is designed, which can further reduce the computational amount and memory read / write amount of the network structure; at the same time, since several Conv-1 layers are added to reduce the model computational amount, the problem of too large convolution receptive field caused by PixShuffle can be improved, and the convergence speed during model training can be increased, achieving the purpose of real-time inference.
[0080] In another embodiment of the present application, refer to Figure 4 , which shows a detailed composition schematic diagram of a super-resolution network structure provided by an embodiment of the present application. As Figure 4 shown, for the super-resolution network structure 10 described in the foregoing embodiment, in a specific example, the super-resolution network structure can sequentially include an 8×PixShuffle module, a Conv-1 layer, three FastIMDB modules, a Concat layer, a Conv-1 layer, a Conv-3 layer, an adder, and an 8×PixShuffle module from the start end to the end end, and its specific connection relationship is shown in detail in Figure 4 shown. Through this super-resolution network structure, the input low-resolution image can be converted into a high-resolution image for output.
[0081] In the embodiment of the present application, such a super-resolution network structure can be applied to images or videos. In other words, it can also be said that the embodiment of the present application proposes a lightweight video super-resolution model structure. According to this lightweight network structure, real-time inference can be achieved in high-resolution input (such as 4000×3000 resolution) super-resolution tasks in an electronic device (such as a smart phone).
[0082] Specifically, the main features of the super-resolution network structure proposed in the embodiment of the present application are as follows:
[0083] (1) 8-fold PixShuffle is connected at the head and tail of the network structure, which can ensure that all convolution calculations are performed in a small spatial domain, greatly reducing the model computational amount and memory read / write amount.
[0084] (2) The network structure also designs a lightweight super-resolution module - FastIMDB, specifically as Figure 5 shown, which can replace the RRDB often used in the conventional super-resolution network structure. In this way, in the FastIMDB module, only one Concat layer is designed, which uses this Concat layer to fuse the frequency information of the feature maps output by multiple layers; and multiple Conv-1 layers are added, which can not only reduce the model calculation amount, but also improve the problem of too large convolution receptive field caused by PixShuffle, and improve the convergence speed during model training.
[0085] (3) Ensuring that all convolution channel numbers are multiples of 32 can improve the operation efficiency of the hardware. Exemplarily, in Figure 4 , before the FastIMDB module, the output channel number of the Conv-1 layer is 96; after the FastIMDB module, the output channel number of the Conv-1 layer is 128, and the output channel number of the Conv-3 layer is 64. In Figure 5 , the output channel number of each layer of sub-feature maps is 32, that is, the connecting line represented by a dotted line; the output channel numbers of other layers are 128, that is, the connecting line represented by a solid line.
[0086] (4) Keeping the channel numbers at a relatively large quantity can facilitate subsequent pruning optimization.
[0087] In short, through this embodiment, the specific implementation of the foregoing embodiment is elaborated in detail. It can be seen from this that according to a lightweight super-resolution network structure provided by this embodiment, the model calculation amount of this super-resolution network structure is low, and the memory read and write amount is small, so that it can be successfully deployed on the electronic device side and achieve the purpose of real-time inference.
[0088] In another embodiment of the present application, refer to Figure 6 , which shows a schematic diagram of the composition structure of an electronic device provided by an embodiment of the present application. As Figure 6 shown, the electronic device 60 may include the super-resolution network structure 10 described in any one of the foregoing embodiments.
[0089] It should be noted that in the super-resolution network structure 10, the super-resolution module may be the super-resolution module described in any one of the foregoing embodiments (such as the FastIMDB module).
[0090] It should also be noted that the electronic device 60 can be, for example, a smart phone, a tablet computer, a laptop computer, a palm computer, a personal digital assistant (PDA), a portable media player (PMP), a navigation device, a wearable device, a server, etc., and the embodiments of the present application do not make any limitations.
[0091] Further, based on the electronic device 60, refer to Figure 7 , which shows a schematic flowchart of a super-resolution reconstruction method provided by an embodiment of the present application. As Figure 7 shown, the method may include:
[0092] S701: Determine the super-resolution network structure.
[0093] S702: Input the low-resolution image into the super-resolution network structure for operation to obtain a high-resolution image.
[0094] It should be noted that the super-resolution network structure here may include a first spatial scaling module, a feature extraction module, and a second spatial scaling module, and the first spatial scaling module is located before the feature extraction module, and the second spatial scaling module is located after the feature extraction module.
[0095] It should also be noted that in some embodiments, the inputting the low-resolution image into the super-resolution network structure for operation to obtain a high-resolution image may include:
[0096] Performing transformation processing on the input low-resolution image through the first spatial scaling module to generate a first processed image;
[0097] Performing feature extraction on the first processed image through the feature extraction module, and generating a second processed image according to the extracted target feature map and the first processed image;
[0098] Performing transformation processing on the second processed image through the second spatial scaling module to generate a high-resolution image.
[0099] Here, the first spatial scaling module specifically performs spatial reduction processing on the input low-resolution image according to the sub-pixel operation, and a first processed image can be obtained; the second spatial scaling module specifically performs spatial expansion processing on the second processed image according to the sub-pixel operation, and a high-resolution image can be obtained.
[0100] In the embodiments of the present application, the first spatial scaling module and the second spatial scaling module are PixShuffle modules with a preset multiple. In a specific example, the value of the preset multiple is equal to 8, but the embodiments of the present application do not make any limitations.
[0101] Further, in some embodiments, the feature extraction module may include a first convolution module, at least one super-resolution module, a first connection layer, a second convolution module, and a first adder.
[0102] Correspondingly, the feature extraction of the first processed image by the feature extraction module and the generation of the second processed image based on the extracted target feature map and the first processed image may include:
[0103] Performing shallow feature extraction on the first processed image through the first convolution module to obtain a first feature map;
[0104] Performing deep feature extraction on the first feature map through at least one super-resolution module to obtain at least one second feature map;
[0105] Fusing at least one second feature map through the first connection layer to obtain a third feature map;
[0106] Performing a convolution operation with reduced number of channels on the third feature map through the second convolution module to obtain a target feature map;
[0107] Performing an addition operation on the target feature map and the first processed image through the first adder to obtain a second processed image.
[0108] In the embodiments of the present application, for the first convolution module and the second convolution module, both may be composed of convolution layers. In some embodiments, the first convolution module may include a first convolution layer, and the second convolution module may include a first convolution layer and a second convolution layer; wherein, the convolution kernels of the first convolution layer and the second convolution layer are different.
[0109] Further, in a specific example, the convolution kernel of the first convolution layer is 1×1, and the convolution kernel of the second convolution layer is 3×3. In this way, the first convolution layer can be represented by the Conv-1 layer, and the second convolution layer can be represented by the Conv-3 layer.
[0110] In another specific example, the number of super-resolution modules included in the feature extraction module is three.
[0111] In this way, for the feature extraction module, specifically, first, the Conv-1 layer can be used to perform a convolution operation on the first processed image to achieve shallow feature extraction and obtain the first feature map; then, three super-resolution modules can be used to perform deep feature extraction on the first feature map to obtain three second feature maps; then, a connection layer (i.e., the Concat layer) can be used to fuse these three second feature maps to obtain the third feature map; then, convolution operations of the Conv-1 layer and the Conv-3 layer can be sequentially performed to reduce the number of channels to obtain the target feature map; finally, the first adder can perform an addition operation on the obtained target feature map and the first processed image input by the feature extraction module to obtain the second processed image output by the feature extraction module.
[0112] Furthermore, in some embodiments, the super-resolution module may include a third convolution module, a second connection layer, a fourth convolution module, and a second adder;
[0113] Correspondingly, the step of performing deep feature extraction on the first feature map through at least one super-resolution module to obtain at least one second feature map may include:
[0114] Performing hierarchical convolution operations on the first feature map through the third convolution module to generate multiple layers of sub-feature maps;
[0115] Fusing the multiple layers of sub-feature maps through the second connection layer to generate a fused sub-feature map;
[0116] Performing a convolution operation on the fused feature map through the fourth convolution module to output a target sub-feature map;
[0117] Performing an addition operation on the target sub-feature map and the first feature map through the second adder to generate the second feature map.
[0118] In the embodiments of the present application, for the third convolution module and the fourth convolution module, both can also be composed of convolution layers. In some embodiments, the third convolution module may include three first convolution layers and four second convolution layers; the fourth convolution module may include one first convolution layer; wherein, the convolution kernels of the first convolution layer and the second convolution layer are different.
[0119] It should be noted that in the third convolution module and the fourth convolution module, the convolution kernel of the first convolution layer can be 1×1, represented by the Conv-1 layer; the convolution kernel of the second convolution layer can be 3×3, represented by the Conv-3 layer.
[0120] It should also be noted that in the embodiments of the present application, the number of output channels of all convolution operations can be a multiple of 32, and the number of output channels can also be maintained at a relatively large number to facilitate improving the operation efficiency of the hardware and subsequent pruning optimization. Therefore, in some embodiments, the method may further include:
[0121] Prune and optimize the super-resolution network structure, and determine the optimized super-resolution network structure as the super-resolution network structure.
[0122] Exemplarily, in the first convolution module, the number of output channels of the first convolution layer is 96; in the second convolution module, the number of output channels of the first convolution layer is 128, and the number of output channels of the second convolution layer is 64. For the super-resolution module, the number of output channels of each layer of sub-feature maps is 32, and the number of output channels of the super-resolution module is 128. In this way, since in the super-resolution network structure described in the embodiments of the application, all the number of channels are multiples of 32, and the number of output channels is also maintained at a relatively large value, pruning optimization can be performed, further reducing the model inference time of the network structure and achieving the purpose of real-time inference.
[0123] This embodiment provides a super-resolution reconstruction method, which determines a super-resolution network structure; inputs a low-resolution image into the super-resolution network structure for calculation to obtain a high-resolution image. In this way, since the super-resolution network structure is a lightweight network structure, a spatial scaling module is connected at the beginning and end of the network structure, which can ensure that all convolution calculations in the feature extraction module are performed in a small spatial domain, reducing the computational amount and memory read / write amount of the network structure; moreover, a lightweight super-resolution module is designed, which can further reduce the computational amount and memory read / write amount of the network structure; at the same time, since several Conv-1 layers are added to reduce the model computational amount, the problem of excessive convolution receptive field caused by PixShuffle can also be improved, and the convergence speed during model training can be increased, achieving the purpose of real-time inference.
[0124] In another embodiment of the present application, the embodiments of the present application further provide a computer storage medium, which stores a computer program, and when the computer program is executed by at least one processor, the method described in the foregoing embodiments is implemented.
[0125] Based on the above super-resolution network structure 10 and computer storage medium, refer to Figure 8 , which shows a schematic structural diagram of another electronic device provided by the embodiments of the present application. As Figure 8 shown, the electronic device 80 may include a processor 801, and the processor 801 may call and run a computer program from the memory to execute the method described in any one of the foregoing embodiments.
[0126] Optionally, as Figure 8 shown, the electronic device 80 may further include a memory 802. Among them, the processor 801 may call and run a computer program from the memory 802 to execute the method described in any one of the foregoing embodiments.
[0127] Among them, the memory 802 can be a separate device independent of the processor 801, or can be integrated in the processor 801.
[0128] Optionally, as Figure 8 shown, the electronic device 80 may further include a transceiver 803. The processor 801 can control the transceiver 803 to communicate with other devices. Specifically, it can send information or data to other devices, or receive information or data sent by other devices.
[0129] In the embodiments of the present application, the electronic device 80 may specifically be the electronic device described in the foregoing embodiments, or a device deployed with the super-resolution network structure 10 described in any one of the foregoing embodiments. Here, the electronic device 80 can implement the corresponding processes of each method in the embodiments of the present application. For the sake of brevity, it will not be elaborated here.
[0130] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present application can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.
[0131] It should also be noted that in the present application, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.
[0132] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.
[0133] The methods disclosed in several method embodiments provided by the present application can be arbitrarily combined without conflict to obtain new method embodiments.
[0134] The features disclosed in several product embodiments provided by the present application can be arbitrarily combined without conflict to obtain new product embodiments.
[0135] The features disclosed in several method or device embodiments provided by the present application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0136] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims described above.
Claims
1. A super-resolution network device, characterized in that, The super-resolution network device includes a first spatial scaling module, a feature extraction module, and a second spatial scaling module; wherein, The first spatial scaling module is configured to perform transformation processing on an input low-resolution image according to sub-pixel operation to generate a first processed image; The feature extraction module is configured to extract features from the first processed image, and generate a second processed image according to the extracted target feature map and the first processed image; The second spatial scaling module is configured to perform transformation processing on the second processed image according to sub-pixel operation to generate a high-resolution image; The feature extraction module includes a first convolutional module, at least one super-resolution module, a first connection layer, and a second convolutional module; The at least one super-resolution module is configured to perform deep feature extraction on the first feature map obtained by the first convolutional module to obtain at least one second feature map; Wherein, the at least one super-resolution module includes a third convolutional module, a second connection layer, and a fourth convolutional module; wherein, The third convolutional module is configured to perform hierarchical convolution operation on the first feature map by using a plurality of convolutional layers to generate multi-layer sub-feature maps; The second connection layer is configured to fuse the multi-layer sub-feature maps to generate a fused sub-feature map; The fourth convolutional module is configured to perform convolution operation on the fused sub-feature map and output a target sub-feature map; the fourth convolutional module includes a first convolutional layer; Wherein, the convolution kernel of the first convolutional layer of the third convolutional module or the fourth convolutional module is different from the convolution kernel of the second convolutional layer of the third convolutional module; Wherein, the third convolutional module includes three first convolutional layers and four second convolutional layers; the input of the third convolutional module is divided into two paths; one path enters the first layer, the first layer includes one second convolutional layer, and the output of the second convolutional layer of the first layer is connected to the second connection layer; the other path enters the second layer, the second layer includes one first convolutional layer and one second convolutional layer; after passing through the first convolutional layer of the second layer, its output is divided into two paths again; one path enters the second convolutional layer of the second layer, and the output of the second convolutional layer of the second layer is connected to the second connection layer; the other path enters the third layer, the third layer includes one first convolutional layer and one second convolutional layer; after passing through the first convolutional layer of the third layer, its output is divided into two paths again; one path enters the second convolutional layer of the third layer, and the output of the second convolutional layer of the third layer is connected to the second connection layer; the other path enters the fourth layer, the fourth layer includes one first convolutional layer and one second convolutional layer, and the output after sequentially passing through the first convolutional layer and the second convolutional layer of the fourth layer is connected to the second connection layer.
2. The super-resolution network device according to claim 1, wherein The first spatial scaling module is located before the feature extraction module, and the second spatial scaling module is located after the feature extraction module; wherein, The first spatial scaling module is specifically configured to perform spatial reduction processing on the input low-resolution image according to sub-pixel operations to obtain the first processed image; The second spatial scaling module is specifically configured to perform spatial expansion processing on the second processed image according to sub-pixel operations to obtain the high-resolution image.
3. The super-resolution network device according to claim 1 or 2, wherein The first spatial scaling module and the second spatial scaling module are PixShuffle modules with a preset multiple.
4. The super-resolution network device according to claim 3, characterized in that, The value of the preset multiple is equal to 8.
5. The super-resolution network device according to claim 1, wherein The first convolutional module is configured to perform shallow feature extraction on the first processed image to obtain the first feature map; The first connection layer is configured to fuse the at least one second feature map to obtain a third feature map; The second convolutional module is configured to perform a convolution operation with reduced number of channels on the third feature map to obtain the target feature map.
6. The super-resolution network device according to claim 5, wherein The feature extraction module includes a first adder; wherein The first adder is configured to perform an addition operation on the target feature map and the first processed image to obtain the second processed image.
7. The super-resolution network device according to claim 5, wherein The first convolutional module includes the first convolutional layer, and the second convolutional module includes the first convolutional layer and the second convolutional layer; wherein, the convolutional kernel of the first convolutional layer of the first convolutional module or the second convolutional module is different from the convolutional kernel of the second convolutional layer of the second convolutional module.
8. The super-resolution network device according to claim 7, wherein The convolutional kernel of the first convolutional layer of the first convolutional module or the second convolutional module is 1×1, and the convolutional kernel of the second convolutional layer of the second convolutional module is 3×3.
9. The super-resolution network device according to claim 5, characterized in that, The number of the at least one super-resolution module is three.
10. The super-resolution network device according to claim 9, wherein In the first convolutional module, the number of output channels of the first convolutional layer is 96; In the at least one super-resolution module, the number of output channels of each super-resolution module is 128; In the first connection layer, the number of output channels is 384; In the second convolutional module, the number of output channels of the first convolutional layer is 128, and the number of output channels of the second convolutional layer is 64.
11. The super-resolution network device according to claim 1, characterized in that, The at least one super-resolution module further includes a second adder; wherein The second adder is configured to perform an addition operation on the target sub-feature map and the first feature map to generate a second feature map.
12. The super-resolution network device according to claim 1, wherein The convolutional kernel of the first convolutional layer of the third convolutional module or the fourth convolutional module of the at least one super-resolution module is 1×1, and the convolutional kernel of the second convolutional layer of the third convolutional module of the at least one super-resolution module is 3×3.
13. The super-resolution network device according to claim 1, wherein The number of the multi-layer sub-feature maps generated by the third convolutional module of the at least one super-resolution module is four; Wherein, the number of output channels of each layer of sub-feature map is 32, and the number of output channels of the second connection layer and the output channels of the super-resolution module are both 128.
14. An electronic device, characterized in that, The electronic device includes the super-resolution network device according to any one of claims 1 to 13.
15. A super-resolution reconstruction method, characterized in that, Applied to the electronic device according to claim 14, the method includes: Determine the super-resolution network device; Input a low-resolution image into the super-resolution network device for operation to obtain a high-resolution image; wherein, the super-resolution network device includes a first spatial scaling module, a feature extraction module, and a second spatial scaling module; Correspondingly, the step of inputting a low-resolution image into the super-resolution network device for operation to obtain a high-resolution image includes: Performing transformation processing on the input low-resolution image through the first spatial scaling module to generate a first processed image; Performing feature extraction on the first processed image through the feature extraction module, and generating a second processed image according to the extracted target feature map and the first processed image; Performing transformation processing on the second processed image through the second spatial scaling module to generate the high-resolution image; Wherein, the feature extraction module includes a first convolution module, at least one super-resolution module, a first connection layer, and a second convolution module; Correspondingly, the step of performing feature extraction on the first processed image through the feature extraction module and generating a second processed image according to the extracted target feature map and the first processed image includes: Performing deep feature extraction on the first feature map obtained by the first convolution module through the at least one super-resolution module to obtain at least one second feature map; Wherein, the at least one super-resolution module includes a third convolution module, a second connection layer, and a fourth convolution module; the third convolution module includes three first convolution layers and four second convolution layers; the fourth convolution module includes one first convolution layer; the convolution kernel of the first convolution layer of the third convolution module or the fourth convolution module is different from the convolution kernel of the second convolution layer of the third convolution module; Correspondingly, the step of performing deep feature extraction on the first feature map obtained by the first convolution module through the at least one super-resolution module to obtain at least one second feature map includes: Performing hierarchical convolution operations on the first feature map through the third convolution module to generate multiple layers of sub-feature maps; Fusing the multiple layers of sub-feature maps through the second connection layer to generate a fused sub-feature map; Performing a convolution operation on the fused sub-feature map through the fourth convolution module to output a target sub-feature map.
16. The method according to claim 15, wherein The feature extraction module further includes a first adder; Correspondingly, the step of performing feature extraction on the first processed image through the feature extraction module and generating a second processed image according to the extracted target feature map and the first processed image further includes: Performing shallow feature extraction on the first processed image through the first convolution module to obtain the first feature map; Fusing the at least one second feature map obtained by the at least one super-resolution module through the first connection layer to obtain a third feature map; Performing a convolution operation with reduced number of channels on the third feature map through the second convolution module to obtain the target feature map; Performing an addition operation on the target feature map and the first processed image through the first adder to obtain the second processed image.
17. The method according to claim 16, characterized in that, The at least one super-resolution module further includes a second adder; Correspondingly, the deep feature extraction of the first feature map by the at least one super-resolution module to obtain at least one second feature map further includes: Performing an addition operation on the target sub-feature map and the first feature map by the second adder to generate the second feature map.
18. The method according to any one of claims 15 to 17, characterized in that, The method further includes: Pruning and optimizing the super-resolution network device, and determining the optimized super-resolution network device as the super-resolution network device.
Citation Information
Patent Citations
Method, device for removing image compression noise based on deep learning and processor
CN110415190A