Image super-resolution reconstruction method and device and storage medium
By introducing packet convolutional networks into the image super-resolution reconstruction network, the problem of high computational complexity caused by excessive depth of the network model in the prior art is solved, and the effect of reducing computational complexity while maintaining a good reconstruction effect.
Patent Information
- Application Number
- CN202510451953.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing deep learning-based image super-resolution reconstruction method has too many parameters and high computational complexity, making it difficult to deploy in edge devices and real-time video image processing.
A packet convolution network is introduced, and by convolution processing of the original image and multiple packet convolution feature extraction, the number of parameters calculated by each group of convolution is reduced, and the overall calculation complexity of the network is reduced.
While maintaining good reconstruction effects, the computational complexity of the model is reduced and the reconstruction efficiency is improved, so that the super-resolution reconstruction network can be better applied to resource-constrained devices and real-time image processing scenarios.
Smart Images

Figure CN119991446A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image reconstruction technology, and in particular to an image super-resolution reconstruction method, device and storage medium. Background Art
[0002] The purpose of image super-resolution reconstruction is to restore low-resolution images (LR) to original high-resolution (HR) images, because HR images contain more detailed information and can be better applied to satellite images, medical imaging, road monitoring and other fields. There are three main existing image super-resolution methods: interpolation-based methods, reconstruction-based methods, and deep learning-based methods.
[0003] The interpolation-based method reconstructs the target pixel based on the domain pixel. Common interpolation methods include nearest neighbor interpolation, quadratic interpolation, bicubic interpolation, etc. This method is relatively simple and efficient, but because it only learns pixel information in a small range, the reconstructed image is relatively blurry; the reconstruction-based method models the LR image acquisition process, constructs the prior constraints of the HR image using the regularization method, and transforms the super-resolution reconstruction problem into an optimization problem of solving the cost function with constraints. This method can better combine the prior knowledge of the image, but when the magnification is large (such as 4 times, 8 times, etc.), the smoothness constraint term will cause the reconstructed image to be too smooth; the deep learning-based method acquires the prior knowledge of the image through learning, and uses the similarity of a large number of images in high-frequency details to guide the image super-resolution reconstruction. The high-resolution image reconstructed by this method has more details and better visual effects than the first two methods. Therefore, the deep learning-based method is currently the mainstream method for image super-resolution reconstruction. However, existing deep learning-based image super-resolution reconstruction methods usually increase the depth of the network model to obtain better reconstruction effects. Since more parameters are introduced at the same time, it is difficult to deploy and apply to edge devices. In addition, more model parameters impose a greater burden on model reasoning, making it difficult for super-resolution reconstruction networks to be applied to real-time video image processing.
[0004] Therefore, how to reduce the computational complexity of the super-resolution reconstruction network while maintaining good reconstruction effects is a problem that needs to be solved urgently. Summary of the invention
[0005] The main purpose of this application is to provide an image super-resolution reconstruction method, device and storage medium, aiming to solve the technical problem of how to reduce the computational complexity of the super-resolution reconstruction network while maintaining good reconstruction effect.
[0006] To achieve the above object, the present application proposes an image super-resolution reconstruction method, which is applied to an image super-resolution reconstruction network, in which a grouped convolutional network is configured, and the image super-resolution reconstruction method includes: Performing convolution processing on the original image to be reconstructed to obtain a first feature map of the original image; By sequentially extracting features from the first feature map through a plurality of group convolution blocks connected in series in the group convolution network, a second feature map of the original image is obtained; The second feature map is up-sampled to obtain a super-resolution reconstructed image corresponding to the original image.
[0007] In one embodiment, the step of extracting features from the first feature map in sequence through a plurality of group convolution blocks connected in series in the group convolution network to obtain a second feature map of the original image comprises: For any group convolution block in the group convolution network, perform multiple group convolutions on the input feature map to obtain multiple candidate feature maps under the target number of channels, wherein the input feature map is the first feature map or the output feature map of the previous group of convolution blocks; Concatenate the candidate feature maps to obtain an output feature map of the group of convolution blocks, and output the output feature map to the next group of convolution blocks connected in series with the group of convolution blocks; After traversing each group of convolution blocks in sequence, the output feature map of the last group of convolution blocks is used as the second feature map of the original image.
[0008] In one embodiment, the step of concatenating the candidate feature maps to obtain the output feature map of the group of convolution blocks includes: Concatenate the candidate feature maps to obtain the feature map to be output of the group of convolution blocks; Perform a residual connection between the feature map to be output and the input feature map to obtain an output feature map of the group of convolution blocks.
[0009] In one embodiment, the step of using the output feature map of the last group of convolutional blocks as the second feature map of the original image includes: Perform a residual connection on the output feature map of the last group convolution block and the first feature map to obtain a second feature map of the original image.
[0010] In one embodiment, the image super-resolution reconstruction network is further configured with a detail enhancement network, and the detail enhancement network is composed of a plurality of convolutional layers, wherein a convolution kernel is configured in any convolutional layer; The step of performing upsampling processing on the second feature map to obtain a super-resolution reconstructed image corresponding to the original image comprises: Performing upsampling processing on the second feature map to obtain a candidate super-resolution reconstructed image corresponding to the original image; The candidate super-resolution reconstructed image is convolved multiple times through the detail enhancement network to obtain a super-resolution reconstructed image corresponding to the original image, wherein the size of each convolution kernel used in the multiple convolution processes is reduced successively, and any convolution kernel is obtained by convolving a preset tensor with the weight of the corresponding convolution layer.
[0011] In one embodiment, the image super-resolution reconstruction method further includes: A training image subset is selected from a preset training image set, and blur kernel learning is performed on each target training image in the preset training image subset through a preset generative adversarial network to obtain a blur kernel corresponding to each target training image; After traversing each blur kernel, performing convolution filtering on each original training image in the training image set based on any blur kernel to obtain a blurred image corresponding to each original training image; Based on each pair of corresponding original training images and blurred images, a preset multi-layer convolutional network is trained to obtain the detail enhancement network.
[0012] In one embodiment, the image super-resolution reconstruction method further includes: Generate a fuzzy image set based on a preset training image set, and perform data enhancement on the fuzzy image set to obtain a target training set; The image super-resolution reconstruction network is trained based on the target training set to optimize and adjust model parameters in the image super-resolution reconstruction network.
[0013] In one embodiment, the step of performing data enhancement on the blurred image set includes: For any image to be enhanced in the fuzzy image set, noise is added to the image to be enhanced, and after traversing each image to be enhanced, a new fuzzy image set is obtained, so as to perform the step of data enhancement on the fuzzy image set based on the new fuzzy image set.
[0014] In addition, to achieve the above-mentioned purpose, the present application also proposes an image super-resolution reconstruction device, which is applied to an image super-resolution reconstruction network, in which a grouped convolutional network is configured, and the image super-resolution reconstruction device includes: A convolution module, used for performing convolution processing on the original image to be reconstructed to obtain a first feature map of the original image; A convolution module, configured to extract features from the first feature map in sequence through a plurality of group convolution blocks connected in series in the group convolution network, so as to obtain a second feature map of the original image; The sampling module is used to perform upsampling processing on the second feature map to obtain a super-resolution reconstructed image corresponding to the original image.
[0015] In addition, to achieve the above-mentioned purpose, the present application also proposes an electronic device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the image super-resolution reconstruction method as described above.
[0016] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the image super-resolution reconstruction method described above are implemented.
[0017] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of the image super-resolution reconstruction method described above are implemented.
[0018] One or more technical solutions proposed in this application have at least the following technical effects: The present application first performs convolution processing on the original image to be reconstructed to obtain a first feature map of the original image, so as to extract the basic features of the image by applying a group of convolution kernels on the original image, capture the shallow feature information of the original image, and provide basic data representation for subsequent feature extraction and reconstruction processes; through multiple group convolution blocks connected in series in the group convolution network, feature extraction is performed on the first feature map in turn to obtain a second feature map of the original image, and the group convolution is used to capture different features to enhance the network's ability to represent image details while reducing the number of parameters of each group of convolution calculations, thereby reducing the overall computational complexity of the network; the second feature map is upsampled to obtain a super-resolution reconstructed image corresponding to the original image, so as to increase the number of pixels of the image through upsampling and restore a clearer image with richer details, thereby realizing the conversion from low resolution to high resolution and completing super-resolution reconstruction.
[0019] In summary, the present application avoids the problems of excessive number of parameters, high computational complexity, difficulty in deployment on edge devices, and difficulty in applying to real-time video image processing in traditional deep learning super-resolution reconstruction methods due to the excessive depth of the network model by introducing a group convolutional network into the traditional super-resolution reconstruction network. While maintaining good reconstruction effects, the computational complexity of the model is reduced and the reconstruction efficiency is improved, so that the super-resolution reconstruction network can be better suited to resource-constrained devices and real-time image processing scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0021] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0022] Figure 1 A schematic diagram of a process flow provided for the first embodiment of the image super-resolution reconstruction method of the present application; Figure 2 A schematic diagram of a process flow provided for the second embodiment of the image super-resolution reconstruction method of the present application; Figure 3 A schematic diagram of a simplified process of an image super-resolution reconstruction method provided in Embodiment 2 of the present application; Figure 4 A schematic diagram of the group convolution block structure of the image super-resolution reconstruction method provided in Example 2 of the present application; Figure 5 A schematic diagram of a detail enhancement network structure of an image super-resolution reconstruction method provided in Example 2 of the present application; Figure 6 A schematic diagram of the reconstruction effect of the image super-resolution reconstruction method provided in Example 2 of the present application; Figure 7 A schematic diagram of another reconstruction effect of the image super-resolution reconstruction method provided in the second embodiment of the present application; Figure 8 This is a schematic diagram of the module structure of the image super-resolution reconstruction device according to an embodiment of the present application; Fig. 9 Schematic diagram of the device structure of the hardware operating environment involved in the image super-resolution reconstruction method in the embodiment of the present application.
[0023] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0024] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0025] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0026] The main solution of the embodiment of the present application is: performing convolution processing on the original image to be reconstructed to obtain a first feature map of the original image; extracting features of the first feature map in turn through multiple group convolution blocks connected in series in a grouped convolutional network to obtain a second feature map of the original image; upsampling the second feature map to obtain a super-resolution reconstructed image corresponding to the original image.
[0027] Since existing deep learning-based image super-resolution reconstruction methods usually increase the depth of the network model to obtain better reconstruction effects, it is difficult to deploy and apply to edge devices because more parameters are introduced at the same time. In addition, more model parameters impose a greater burden on model reasoning, making it difficult for super-resolution reconstruction networks to be applied in real-time video image processing. Therefore, how to reduce the computational complexity of super-resolution reconstruction networks while maintaining good reconstruction effects is a problem that needs to be solved urgently.
[0028] The present application provides a solution by introducing a group convolutional network into a traditional super-resolution reconstruction network, thereby avoiding the problems of excessive number of parameters, high computational complexity, difficulty in deployment on edge devices, and difficulty in application in real-time video image processing caused by the excessive depth of the network model in traditional deep learning super-resolution reconstruction methods. While maintaining good reconstruction effects, the computational complexity of the model is reduced and the reconstruction efficiency is improved, so that the super-resolution reconstruction network can be better suitable for resource-constrained devices and real-time image processing scenarios.
[0029] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device that can realize the above functions. The following takes an electronic device as an example to illustrate this embodiment and the following embodiments.
[0030] Based on this, the present invention provides an image super-resolution reconstruction method. Figure 1 , Figure 1 This is a flowchart of the first embodiment of the image super-resolution reconstruction method of the present application.
[0031] In this embodiment, the image super-resolution reconstruction method is applied to an image super-resolution reconstruction network, in which a group convolutional network is configured, and the image super-resolution reconstruction method includes steps S10 to S30: Step S10, performing convolution processing on the original image to be reconstructed to obtain a first feature map of the original image; It should be noted that the original image refers to the low-resolution image to be processed, that is, the input image that needs to be super-resolution reconstructed; the first feature map refers to the shallow feature representation extracted from the original image through convolution processing, and the data contains the basic structure and information of the original image.
[0032] It can be understood that in order to extract basic features from the original image for subsequent processing, step S10 is performed to avoid the information loss problem caused by directly processing the original pixels, thereby extracting image features that are helpful for reconstruction, laying the foundation for subsequent feature extraction and reconstruction.
[0033] Exemplarily, the original image is input into a convolution layer, which contains a convolution kernel of preset size and number of channels (e.g., a convolution kernel of 3*3 size and 64 channels), which extracts local features of the image by sliding on the original image and performing a dot product operation. After the convolution operation, nonlinearity is added through an activation function (e.g., ReLU) to obtain a feature map with a preset number of channels (e.g., a 64-channel feature map), which is the first feature map of the original image, and contains shallow information such as edges and textures of the original image.
[0034] Step S20, extracting features from the first feature map in sequence through a plurality of group convolution blocks connected in series in the group convolution network to obtain a second feature map of the original image; It should be noted that a group convolutional network refers to a deep learning architecture that processes input data (such as the first feature map) through multiple group convolution blocks, each of which is responsible for extracting features at different levels; a group convolutional block refers to a basic unit in the network, which contains a group of convolutional layers that perform group convolution operations on the input feature map to extract higher-level features; the second feature map refers to a feature map processed by multiple group convolutional blocks in the group convolutional network, which contains more abstract and richer image feature information.
[0035] It is understandable that since existing deep learning-based image super-resolution reconstruction methods usually increase the depth of the network model to obtain better reconstruction effects, thereby introducing more model parameters, it is difficult to deploy and apply to edge devices, and more model parameters impose a greater burden on model reasoning, making it difficult for the super-resolution reconstruction network to be applied to real-time video image processing. Therefore, step S20 is performed to capture different features through group convolution to enhance the network's ability to represent image details, while reducing the amount of parameters for each group of convolution calculations, thereby reducing the overall computational complexity of the network and providing rich information for the final image reconstruction.
[0036] Exemplarily, the first feature map is fed as input to a group convolution network, which is composed of multiple group convolution blocks connected in series. Each group convolution block contains multiple convolution layers and possible pooling layers, which are divided into several groups. Each group independently performs convolution operations on the input feature map, and then concatenates the results of each group. The feature map extracted by each group convolution block will be passed to the next group convolution block. After the step-by-step feature extraction of multiple group convolution blocks, a second feature map containing deeper features is finally obtained.
[0037] In a feasible implementation, step S20 may include steps S21 to S23: Step S21, for any group convolution block in the group convolution network, perform multiple group convolutions on the input feature map to obtain multiple candidate feature maps under the target number of channels, wherein the input feature map is the first feature map or the output feature map of the previous group convolution block; It should be noted that the input feature map refers to the feature map of the previous stage entering the group convolution block for processing. When the group convolution block is the first group convolution block in the grouped convolution network, the feature map is the first feature map, that is, the feature map obtained directly from the convolution processing of the original image; when the group convolution block is not the first group convolution block in the grouped convolution network (that is, the second and subsequent group convolution blocks), the feature map is the output feature map of the previous group of convolution blocks, that is, the feature map after processing by the previous group of convolution blocks.
[0038] In addition, it should be noted that the target number of channels refers to the number of channels of the feature map that the group convolution block expects to obtain after group convolution. This value determines the depth of the feature map, that is, the feature dimension of each pixel; the candidate feature map refers to the intermediate feature map obtained through multiple group convolution operations within each group convolution block. The number of channels of these feature maps is the target number of channels, but they need to be integrated into a feature map in the end; the output feature map refers to the final feature map obtained after processing the group convolution block and splicing all candidate feature maps. This feature map will be passed as input to the next group convolution block or as the final second feature map.
[0039] It is understandable that since existing deep learning-based image super-resolution reconstruction methods usually increase the depth of the network model to obtain better reconstruction effects, thereby introducing more model parameters, it is difficult to deploy and apply to edge devices, and more model parameters impose a greater burden on model reasoning, making it difficult for the super-resolution reconstruction network to be applied to real-time video image processing. Therefore, step S21 is performed to capture different features through group convolution to enhance the network's ability to represent image details, while reducing the amount of parameters for each group of convolution calculations, thereby reducing the overall computational complexity of the network and providing rich information for the final image reconstruction.
[0040] Exemplarily, for any group of convolution blocks, the input feature map is received by the group of convolution blocks and divided into several small groups, each of which performs convolution operations independently. For example, if the input feature map has 64 channels, it can be divided into 8 groups, each with 8 channels, and then different convolution kernels are applied to each group for convolution. After each group of convolution operations, a candidate feature map with a target number of channels is obtained, such as a candidate feature map with 8 channels.
[0041] Step S22, concatenating the candidate feature maps to obtain an output feature map of the group of convolution blocks, and outputting the output feature map to the next group of convolution blocks connected in series with the group of convolution blocks; It is understandable that since it is necessary to integrate multiple candidate feature maps obtained by group convolution to form a complete feature representation, performing step S22 can avoid the loss of feature information, thereby improving the dimension and information richness of the feature map, and providing more comprehensive data for subsequent feature extraction and image reconstruction.
[0042] Exemplarily, the candidate feature maps generated by each group convolution operation within the group convolution block are concatenated in the channel dimension. For example, if there are 8 candidate feature maps, each with 8 channels, the output feature map obtained after concatenation will have 64 channels. The concatenation operation can be implemented through a concatenation layer, and the concatenated output feature map is then passed to the next concatenated group convolution block for further processing.
[0043] In a feasible implementation manner, the step of concatenating the candidate feature maps in step S22 to obtain the output feature map of the group of convolution blocks may include steps S221-S222: Step S221, concatenating the candidate feature maps to obtain a to-be-output feature map of the group of convolution blocks; It should be noted that the feature map to be output refers to the feature map that has not yet been residually connected with the input feature map after the group convolution and splicing operations. This feature map contains the high-level features extracted by the group convolution block, but has not yet been combined with the information of the original input feature map.
[0044] Step S222: Perform a residual connection on the feature map to be output and the input feature map to obtain an output feature map of the group of convolution blocks.
[0045] It is understandable that due to the gradient vanishing and gradient exploding problems that are often prone to occur in deep neural networks, step S222 is performed to introduce a short-circuit path in the network for residual connection, allowing the gradient to flow directly to the deeper layers of the network, which can help alleviate the gradient vanishing in deep network training, thereby avoiding the problem of training difficulties caused by increased network depth, improving the training stability and convergence speed of the network, and also helping to maintain the original information in the input feature map, thereby retaining more details and structural information during the reconstruction process.
[0046] Exemplarily, when implementing the residual connection, first ensure that the size of the feature map to be output is the same as the size of the input feature map, which is usually achieved by adjusting the number of channels of the feature map to be output or using an appropriate convolutional layer. Next, the adjusted feature map to be output is added to the original input feature map at the element level in the channel dimension. If the number of channels of the input feature map is inconsistent with the number of channels of the feature map to be output, a 1x1 convolutional layer can be used to adjust the number of channels of the input feature map to match the feature map to be output. In this way, the output of the residual connection is the final output feature map of the group convolution block, which contains both the information directly transmitted from the input feature map and the features learned through the network.
[0047] In this embodiment, by generating output feature maps based on feature map splicing and residual connection in the group convolution block, the problems caused by incomplete feature information and gradient vanishing are avoided. Feature map splicing ensures that the feature information extracted from different group convolutions is effectively integrated, avoiding the loss of feature representation; while residual connection allows information in the network to be directly propagated, reducing the gradient vanishing problem when training deep networks. These means together achieve the improvement of the richness and efficiency of feature extraction while maintaining the stability of network training, thereby achieving higher quality image feature representation and super-resolution reconstruction effects.
[0048] Step S23, after traversing each group of convolution blocks in sequence, taking the output feature map of the last group of convolution blocks as the second feature map of the original image.
[0049] It can be understood that since the output feature map of the last group convolution block contains the deepest feature information after being processed by all the group convolution blocks, performing step S23 can avoid the problem of detail loss caused by using only shallow features for image reconstruction, thereby obtaining a feature map that is highly abstract and has rich semantic information, providing a data basis for generating high-quality super-resolution images.
[0050] For example, at the end of the network, after all the group convolution blocks are processed in sequence, the feature map output by the last group convolution block will contain the richest feature information. This output feature map does not need to be concatenated or processed additionally, and is directly used as the second feature map. For example, if the network consists of four group convolution blocks, the output feature map of the fourth group convolution block will be regarded as the second feature map, which will be used in the subsequent upsampling and reconstruction process.
[0051] In this embodiment, by performing multiple group convolutions on the input feature map and concatenating the obtained deep feature map, the problems of insufficient feature representation and low computational efficiency that may occur in the feature extraction process of the traditional convolutional network are avoided, and the representation ability of the feature map is enhanced, so that the network can learn richer and more abstract image features while improving the computational efficiency of the convolutional network. Finally, by using the output feature map of the last group convolution block as the second feature map of the original image, efficient feature extraction under the deep network structure is achieved, providing higher quality feature information for subsequent super-resolution reconstruction, thereby improving the clarity and detail expression of the reconstructed image.
[0052] Step S30: up-sample the second feature map to obtain a super-resolution reconstructed image corresponding to the original image.
[0053] It should be noted that the super-resolution reconstructed image refers to an image with a higher resolution obtained through upsampling, which is a clear version of the original low-resolution image.
[0054] It can be understood that since the extracted feature map needs to be converted back to a high-resolution space to obtain a clear reconstructed image, performing step S30 can avoid the blur and distortion problems caused by directly upsampling the original image, thereby generating a super-resolution image with rich details and better visual effects, thereby improving the visual quality and practicality of the image.
[0055] Exemplarily, when implementing the upsampling process, an upsampling technique such as interpolation or deconvolution is used to increase the resolution of the second feature map. For example, nearest neighbor interpolation, bilinear interpolation, bicubic interpolation, or a more complex deconvolution layer (also called a transposed convolution layer) can be used to enlarge the size of the feature map to the target resolution. After upsampling, the features are further refined through a series of convolution layers, and finally a super-resolution reconstructed image with the same size as the original image but a higher resolution is output. This process can effectively restore image details and improve the visual quality of the image.
[0056] This embodiment provides an image super-resolution reconstruction method. By introducing a group convolutional network into a traditional super-resolution reconstruction network, the problems of excessive number of parameters, high computational complexity, difficulty in deployment on edge devices, and difficulty in application to real-time video image processing in traditional deep learning super-resolution reconstruction methods due to the excessive depth of the network model are avoided. While maintaining a good reconstruction effect, the computational complexity of the model is reduced and the reconstruction efficiency is improved, so that the super-resolution reconstruction network can be better suitable for resource-constrained devices and real-time image processing scenarios.
[0057] In a feasible implementation manner, the step of using the output feature map of the last group convolution block as the second feature map of the original image in step S23 may include step S231: Step S231, performing a residual connection on the output feature map of the last group convolution block and the first feature map to obtain a second feature map of the original image.
[0058] Exemplarily, first, ensure that the output feature map of the last group convolution block has the same size as the first feature map, which is usually achieved by adding an appropriate convolution layer or pooling layer after the last group convolution block. Next, the output feature map of the last group convolution block after size matching is added element-wise to the first feature map in the channel dimension. If the number of channels of the two is inconsistent, a 1x1 convolution layer can be used to adjust the number of channels of the first feature map to make it the same as the number of channels of the output feature map of the last group convolution block. In this way, the result of the addition operation is the second feature map, which combines the high-level features extracted by the deep network and the initial shallow features, thereby providing richer feature information for subsequent upsampling and reconstruction steps.
[0059] In this embodiment, the output feature map is residually connected with the input feature map before the group convolution, and the high-level features extracted by the deep network are fused with the initial shallow features to retain the details and structural information of the image, thereby avoiding the gradient vanishing problem that may occur in the deep network and the problem that the feature information may be lost during the transmission process. The residual connection allows the network to learn deeper features while retaining the important information of the initial feature map, thereby improving the stability and efficiency of network training while maintaining the depth and complexity of the network, achieving richer feature representation, thereby improving the image quality of super-resolution reconstruction, and achieving clearer and more delicate image details and structures.
[0060] In a feasible implementation manner, the image super-resolution reconstruction method may further include steps S01-S02: Step S01, generating a fuzzy image set based on a preset training image set, and performing data enhancement on the fuzzy image set to obtain a target training set; It should be noted that the preset training image set refers to a set of original images with high resolution and clarity, which is used to train the super-resolution reconstruction network, and can be the DIV2K dataset (the training set of this dataset has 800 images, the validation set has 100 images, and the test set has 100 images, and the image resolution is 2K); the blurred image set refers to the low-resolution images generated by applying blur processing (such as downsampling, adding noise, etc.) to the preset training image set; the target training set refers to the blurred image set after data enhancement processing, which contains more samples and changes, and is used to improve the generalization ability of the network.
[0061] It is understandable that since the low-resolution images encountered in practical applications may be diverse and complex, and the existing super-resolution reconstruction networks usually lack real-scene training data, resulting in poor performance of the existing networks in real-scene image reconstruction. Therefore, performing step S01 can avoid the problem of insufficient network generalization ability due to lack of real-scene training data and single training data, thereby improving the network's ability to reconstruct various low-resolution images.
[0062] Exemplarily, first, a set of high-resolution, high-quality preset training image sets (such as 800 training set images in the DIV2K dataset) is selected. Then, by simulating the image degradation process in reality, such as applying downsampling operations (for example, using nearest neighbor, bilinear or bicubic downsampling methods), or applying the blur kernel learned by GAN (Generative Adversarial Networks, Generative Adversarial Networks) to process the preset training image set and add Gaussian noise to generate the corresponding low-resolution blurred image set. Next, the blurred image set is data enhanced, including but not limited to operations such as rotation, flipping, scaling, cropping, and color transformation to increase the diversity and quantity of training samples. Finally, these enhanced images constitute the target training set for subsequent network training.
[0063] In a feasible implementation manner, before the step of performing data enhancement on the fuzzy image set in step S01, step S100 may also be included: Step S100, for any image to be enhanced in the fuzzy image set, add noise to the image to be enhanced, and after traversing each image to be enhanced, obtain a new fuzzy image set, so as to perform the step of data enhancement on the fuzzy image set based on the new fuzzy image set.
[0064] It is understandable that, since blurred images are often accompanied by noise in real scenes, step S100 is performed. By adding noise to the training data, the problem of performance degradation of the network when processing noisy blurred images in real scenes can be avoided, and the network's robustness to noise can be improved, and the generalization ability and reconstruction accuracy of the network in a noisy environment can be improved, so that the super-resolution reconstruction network can better handle complex situations that may be encountered in actual shooting, thereby generating clearer reconstructed images that are more in line with real scenes.
[0065] In this implementation, by adding noise to the blurred image set, it is ensured that the trained super-resolution reconstruction network can still effectively restore image details in the presence of noise, thereby improving the practicality of the network in practical applications and the quality of the reconstructed image.
[0066] Step S02: training the image super-resolution reconstruction network based on the target training set to optimize and adjust model parameters in the image super-resolution reconstruction network.
[0067] It should be noted that model parameters refer to all learnable parameters in the network, including the weights and biases of the convolutional layers, the learning rate of the optimizer, etc.
[0068] It is understandable that since the reconstruction capability of the network needs to be optimized by learning a large number of mapping relationships from low-resolution to high-resolution image pairs, performing step S02 can avoid the problem that the network cannot effectively learn image features and reconstruct details, and can achieve the effect of improving the clarity and quality of the network reconstructed image. By training and optimizing model parameters, the network can more accurately recover the details and structure of the high-resolution image from the low-resolution image.
[0069] Exemplarily, the image super-resolution reconstruction network is trained using low-resolution images in the target training set as input and high-resolution images in the preset training image set as labels. The model training is performed using the Pytorch training framework, using the Adam optimizer, with an initial learning rate of 2*10e-4, and the learning rate is halved every 200 epochs, and training is performed for 1000 epochs. The back propagation algorithm and gradient descent method can also be used to update the model parameters in the network, including the weights and biases of the convolutional layer. Some regularization techniques, such as weight decay or dropout, may also be used during the training process to prevent overfitting. Through multiple iterative training, the model parameters are continuously adjusted until the performance of the network on the validation set reaches a preset threshold or the training reaches a predetermined number of iterations, thereby completing the optimization adjustment of the network model.
[0070] In this embodiment, by generating a blurred image set based on a preset training image set and performing data enhancement on these images, a diversified target training set is obtained, which helps the network learn a wider range of image features and changes, thereby avoiding the problem of insufficient generalization ability that the network may encounter in practical applications. Then, the network is trained based on the target training set and the model parameters are optimized, ensuring that the network can more effectively recover high-resolution details from low-resolution images, thereby improving the clarity and visual quality of the reconstructed image.
[0071] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction, and will not be repeated in the following. Figure 2 , the image super-resolution reconstruction network is also configured with a detail enhancement network, and the detail enhancement network is composed of a plurality of convolutional layers, wherein a convolution kernel is configured in any convolutional layer; Step S30 may further include steps S31-S32: Step S31, performing upsampling processing on the second feature map to obtain a candidate super-resolution reconstructed image corresponding to the original image; It should be noted that the candidate super-resolution reconstructed image refers to a preliminary high-resolution image obtained through upsampling processing, which is a preliminary reconstruction of the original low-resolution image, but may not be detailed and clear enough.
[0072] It is understandable that since the resolution of the feature map needs to be increased from low resolution to high resolution to generate an image that matches the size of the original image, performing step S31 can avoid the blur and distortion problems that may occur when the high-resolution image is directly restored from the low-resolution image. By upsampling, preliminary image size restoration can be achieved, providing a basis for subsequent detail enhancement processing, thereby achieving the effect of preliminary reconstruction from low resolution to high-resolution image.
[0073] Exemplarily, the second feature map is input into an upsampling module, which can be an upsampling layer based on an interpolation method (such as bilinear interpolation, bicubic interpolation) or a deep learning model (such as a deconvolution layer or a sub-pixel convolution layer). The upsampling operation increases the resolution of the feature map to a predetermined level, for example, by enlarging the width and height of the image by 4 times. In this way, the original low-resolution feature map is upsampled to obtain a larger candidate super-resolution reconstructed image, although the image may not be clear enough in details at this time.
[0074] Step S32, performing multiple convolutions on the candidate super-resolution reconstructed image through the detail enhancement network to obtain a super-resolution reconstructed image corresponding to the original image, wherein the size of each convolution kernel used in the multiple convolution processes is reduced successively, and any convolution kernel is obtained by convolving a preset tensor with the weight of the corresponding convolution layer.
[0075] It can be understood that since it is necessary to further refine the details of the candidate super-resolution reconstructed image and improve the clarity and quality of the image, step S32 is performed to avoid the problems of detail loss and edge blur in the reconstructed image. By performing multiple convolutions using convolution kernels of gradually decreasing sizes, the details and texture information of the image can be gradually restored. At the same time, convolution based on preset tensors and convolution layer weights can better retain the structural information of the image, thereby achieving the effect of improving the visual quality and detail expression of the reconstructed image, and obtaining a super-resolution reconstructed image that is closer to the original high-resolution image.
[0076] Exemplarily, the candidate super-resolution reconstructed image obtained in the previous step is input into the detail enhancement network. The network consists of multiple convolutional layers, and the convolution kernel size of each convolutional layer gradually decreases. For example, the convolution kernel size of the first convolutional layer is 7x7, the convolution kernel size of the second convolutional layer is 5x5, and the convolution kernel size of the third convolutional layer is 3x3, and so on. In each convolutional layer, the convolution kernel is convolved with the input image, wherein any convolution kernel is generated based on a combination of a preset tensor (such as a full 1 tensor) and the weight of the convolutional layer, so that no activation function is used after the convolution operation of each convolutional layer. Through such multiple convolution processes, the network can gradually refine the details of the image and improve the clarity of the image. Finally, the output is a super-resolution reconstructed image with detail enhancement, which is visually closer to the original high-resolution image than the candidate super-resolution reconstructed image.
[0077] In this embodiment, by using multiple convolutions in the detail enhancement network, the problem of blurring and loss of details in the reconstructed image caused by direct upsampling is avoided, and a clear and delicate reconstruction from a low-resolution image to a high-resolution image is achieved. On the basis of the initial expansion of the image size by upsampling, the detail enhancement network gradually restores the details and texture information of the image by performing multiple convolutions using convolution kernels of gradually decreasing sizes, and at the same time, combines the preset tensor with the convolution layer weights to ensure the structural fidelity and visual quality of the reconstructed image, which not only improves the clarity of the super-resolution reconstructed image, but also enhances the visual effect of the image, achieving a super-resolution reconstruction effect that is closer to the original high-resolution image.
[0078] In a feasible implementation manner, the image super-resolution reconstruction method may further include steps S301 to S303: Step S301, selecting a training image subset from a preset training image set, and performing blur kernel learning on each target training image in the preset training image subset through a preset generative adversarial network to obtain a blur kernel corresponding to each target training image; It should be noted that the training image subset refers to a part of clear images selected from the preset training image set for learning the blur process; the preset generative adversarial network refers to a specially designed neural network composed of a generator and a discriminator, which is used to learn how to generate a blur kernel through an adversarial training process; the target training image refers to any image in the selected subset, which will be used as the input for learning the blur kernel; blur kernel learning refers to the image blur process parameters learned through the generative adversarial network, which describe how the image changes from clear to blurred; the blur kernel refers to a set of parameters that define the characteristics of the image blur operation, such as the type of blur (Gaussian blur, motion blur, etc.) and the degree of blur.
[0079] It can be understood that in order to simulate the diversity of image blur in a real environment and train a super-resolution network that can handle various types of blur, step S301 is performed, which can avoid the problem that the network cannot handle various blurred images encountered in practical applications, and achieve the learning of a more general and robust blur kernel, thereby improving the generalization ability of the network.
[0080] For example, a subset is randomly selected from a preset training image set, and this subset contains a number of clear training images with diversity (i.e., target training images). Then, a designed GAN, such as KernelGAN, is used, in which the generator in the GAN is responsible for generating blur kernels, and the discriminator is responsible for distinguishing the generated blur kernels from the real blur kernels. By iteratively training the GAN, the generator gradually learns to generate corresponding blur kernels for each training image, and these blur kernels can simulate the blur effects that images may encounter in real scenes.
[0081] Step S302, after traversing each blur kernel, performing convolution filtering on each original training image in the training image set based on any blur kernel to obtain a blurred image corresponding to each original training image; It should be noted that convolution filtering refers to the process of performing convolution operations on an image using a blur kernel to simulate the blur effect of the image; the original training image refers to any image in the training image set; the blurred image refers to the image obtained after convolution filtering, which simulates the image degradation caused by various reasons in actual application.
[0082] It is understandable that, since it is necessary to generate a low-resolution blurred image that matches the actual environmental conditions for subsequent network training, performing step S302 can avoid the problem of insufficient samples that may occur when using real blurred images for training, and achieve the generation of diversified training data, thereby improving the training efficiency and reconstruction quality of the network.
[0083] Exemplarily, after obtaining the blur kernel corresponding to the target training image, for each image in the training image set (ie, the original training image), any blur kernel obtained is randomly selected to perform a convolution filtering operation:
[0084] in, is a clear original training image, To perform convolution filtering on the original training image using the blur kernel, The image is convolution filtered, and Gaussian noise is added to simulate the blurring process of the image. In this way, each original training image will generate a corresponding blurred image, which will be used for subsequent network training.
[0085] Step S303: Based on each pair of corresponding original training images and blurred images, a preset multi-layer convolutional network is trained to obtain the detail enhancement network.
[0086] It should be noted that the preset multi-layer convolutional network refers to a neural network composed of multiple convolutional layers, which is designed to restore clear image details from blurred images.
[0087] It is understandable that, since it is necessary to train a network that can recover high-resolution details from low-resolution blurred images, performing step S303 can avoid the problem that the network cannot effectively recover image details during the reconstruction process, thereby achieving the training of a detail enhancement network that can enhance image details and improve the quality of reconstructed images.
[0088] For example, the original training image and the generated blurred image pair are used as input-output pairs to train a multi-layer convolutional network. The multi-layer convolutional network consists of multiple convolutional layers and does not contain an activation function. It aims to learn how to restore the details of a clear image from a blurred image. During the training process, L1 loss is used, and the loss function is:
[0089] in, is the number of images for each training, For the i The original training images, For the iAn image processed by the detail enhancement network. The network uses gradient descent or other optimization algorithms (such as Adam, SGD, etc.) to update the weights and biases of the network according to the calculated gradients until the network can effectively reconstruct details close to the original clear image from the blurred image. For example, the Adam optimizer is used for training for 100 rounds, and the initial learning rate is set to 0.0002. Finally, the trained network is the detail enhancement network, which can be used to improve the quality of super-resolution reconstruction of images. In addition, the training process can be repeated in multiple training batches, each of which may contain different pairs of original and blurred images, until the performance of the network on the validation set is no longer significantly improved or the preset number of training iterations is reached, and finally a trained detail enhancement network is obtained.
[0090] In this embodiment, by adopting a generative adversarial network to learn blur kernels, convolution filters to generate blurred images, and training a multi-layer convolutional network, the problems of low reconstruction quality and insufficient generalization ability in image super-resolution reconstruction due to the lack of real blurred image pairs and the network's inability to effectively learn image details are avoided. Thus, a blurred image matching the actual shooting blur process can be generated, and a detail enhancement network that can effectively restore image details can be trained, thereby improving the image quality of super-resolution reconstruction and the generalization ability of the network.
[0091] For example, to help understand the implementation process of the image super-resolution reconstruction method obtained by combining this embodiment with the above-mentioned embodiment 1, please refer to Figure 3 , Figure 3 A brief flowchart of an image super-resolution reconstruction method is provided, specifically: First, the LR image, i.e., the original image to be reconstructed at a low resolution, is convolved to obtain the first feature map of the LR image. Then, multiple concatenated layers (such as Figure 3 The first feature map is extracted by four groups of convolution blocks, and the output feature map of the last group of convolution blocks is residually connected with the first feature map to obtain the second feature map of the LR image. The second feature map is then upsampled to obtain a candidate super-resolution reconstructed image corresponding to the LR image, and the candidate super-resolution reconstructed image is convolved multiple times through the detail enhancement network to obtain a high-resolution HR image corresponding to the LR image, that is, a super-resolution reconstructed image.
[0092] Among them, please refer to the structure of any group of convolution blocks Figure 4 In the group convolution blocks, g=2, g=4, and g=8 respectively represent division into 2 groups, 4 groups, and 8 groups, k=1 indicates that the convolution kernel size is 1, and the residual connection between the input feature map and the output feature map can be realized in this group of convolution blocks.
[0093] Among them, please refer to the structure of the detail enhancement network Figure 5 , the candidate HR image obtained by upsampling the second feature map, that is, the candidate super-resolution reconstructed image, is processed multiple times (such as Figure 5 In the convolution process, the size of each convolution kernel used is reduced successively to obtain a high-resolution HR image corresponding to the LR image. In addition, in the process of processing RGB images, except for the last layer of filters with 3, the number of filters in other layers can be other (such as 64).
[0094] For example, please refer to Figure 6 and Figure 7 , Figure 6 To demonstrate the reconstruction effect without configuring a detail enhancement network in the super-resolution reconstruction network, the left sub-image LR image is a low-resolution original image, and the right sub-image x4 HR image is a super-resolution reconstructed image with a 4-fold high resolution. Figure 7 To demonstrate the reconstruction effect after configuring the detail enhancement network in the super-resolution reconstruction network, the left sub-image LR image is the low-resolution original image, and the right sub-image x4 HR image is the super-resolution reconstructed image with 4 times the high resolution.
[0095] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the image super-resolution reconstruction method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.
[0096] This application also provides an image super-resolution reconstruction device, please refer to Figure 8 The image super-resolution reconstruction device is applied to an image super-resolution reconstruction network, wherein a grouped convolutional network is configured in the image super-resolution reconstruction network, and the image super-resolution reconstruction device comprises: The convolution module 10 is used to perform convolution processing on the original image to be reconstructed to obtain a first feature map of the original image; A convolution module 20, configured to extract features from the first feature map in sequence through a plurality of group convolution blocks connected in series in the group convolution network, to obtain a second feature map of the original image; The sampling module 30 is used to perform up-sampling processing on the second feature map to obtain a super-resolution reconstructed image corresponding to the original image.
[0097] Optionally, the test paper assembly module 20 is further used for: For any group convolution block in the group convolution network, perform multiple group convolutions on the input feature map to obtain multiple candidate feature maps under the target number of channels, wherein the input feature map is the first feature map or the output feature map of the previous group of convolution blocks; Concatenate the candidate feature maps to obtain an output feature map of the group of convolution blocks, and output the output feature map to the next group of convolution blocks connected in series with the group of convolution blocks; After traversing each group of convolution blocks in sequence, the output feature map of the last group of convolution blocks is used as the second feature map of the original image.
[0098] Optionally, the test paper assembly module 20 is further used for: Concatenate the candidate feature maps to obtain the feature map to be output of the group of convolution blocks; Perform a residual connection between the feature map to be output and the input feature map to obtain an output feature map of the group of convolution blocks.
[0099] Optionally, the test paper assembly module 20 is further used for: Perform a residual connection on the output feature map of the last group convolution block and the first feature map to obtain a second feature map of the original image.
[0100] Optionally, the image super-resolution reconstruction network is further configured with a detail enhancement network, and the detail enhancement network is composed of a plurality of convolutional layers, wherein a convolution kernel is configured in any convolutional layer; The sampling module 30 is also used for: Performing upsampling processing on the second feature map to obtain a candidate super-resolution reconstructed image corresponding to the original image; The candidate super-resolution reconstructed image is convolved multiple times through the detail enhancement network to obtain a super-resolution reconstructed image corresponding to the original image, wherein the size of each convolution kernel used in the multiple convolution processes is reduced successively, and any convolution kernel is obtained by convolving a preset tensor with the weight of the corresponding convolution layer.
[0101] Optionally, the training module 40 in the image super-resolution reconstruction device is used to: A training image subset is selected from a preset training image set, and blur kernel learning is performed on each target training image in the preset training image subset through a preset generative adversarial network to obtain a blur kernel corresponding to each target training image; After traversing each blur kernel, performing convolution filtering on each original training image in the training image set based on any blur kernel to obtain a blurred image corresponding to each original training image; Based on each pair of corresponding original training images and blurred images, a preset multi-layer convolutional network is trained to obtain the detail enhancement network.
[0102] Optionally, the training module 40 is further used for: Generate a fuzzy image set based on a preset training image set, and perform data enhancement on the fuzzy image set to obtain a target training set; The image super-resolution reconstruction network is trained based on the target training set to optimize and adjust model parameters in the image super-resolution reconstruction network.
[0103] Optionally, the training module 40 is further used for: For any image to be enhanced in the fuzzy image set, noise is added to the image to be enhanced, and after traversing each image to be enhanced, a new fuzzy image set is obtained, so as to perform the step of data enhancement on the fuzzy image set based on the new fuzzy image set.
[0104] The image super-resolution reconstruction device provided by the present application adopts the image super-resolution reconstruction method in the above embodiment, which can solve the technical problem of how to reduce the computational complexity of the super-resolution reconstruction network while maintaining a good reconstruction effect. Compared with the prior art, the beneficial effects of the image super-resolution reconstruction device provided by the present application are the same as the beneficial effects of the image super-resolution reconstruction method provided by the above embodiment, and the other technical features in the image super-resolution reconstruction device are the same as the features disclosed in the above embodiment method, which will not be repeated here.
[0105] The present application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the image super-resolution reconstruction method in the above-mentioned embodiment 1.
[0106] Reference below Fig. 9 , which shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present application. The electronic devices in the embodiments of the present application may include but are not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions: tablet computers), PMPs (Portable Media Players: portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Fig. 9 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0107] like Fig. 9As shown, the electronic device may include a processing device 1001 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM: Random Access Memory) 1004. In RAM1004, various programs and data required for the operation of the electronic device are also stored. The processing device 1001, ROM1002, and RAM1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows an electronic device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have alternatively.
[0108] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0109] The electronic device provided by the present application adopts the image super-resolution reconstruction method in the above embodiment, which can solve the technical problem of how to reduce the computational complexity of the super-resolution reconstruction network while maintaining a good reconstruction effect. Compared with the prior art, the beneficial effects of the electronic device provided by the present application are the same as the beneficial effects of the image super-resolution reconstruction method provided by the above embodiment, and the other technical features in the electronic device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.
[0110] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0111] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
[0112] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, the computer-readable program instructions being used to execute the image super-resolution reconstruction method in the above-mentioned embodiment.
[0113] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM: Random Access Memory), a read-only memory (ROM: Read Only Memory), an erasable programmable read-only memory (EPROM: Erasable Programmable Read Only Memory or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM: CD-Read Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency: Radio Frequency), etc., or any suitable combination of the above.
[0114] The computer-readable storage medium may be included in the electronic device, or may exist independently without being installed in the electronic device.
[0115] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by an electronic device, the electronic device: performs convolution processing on the original image to be reconstructed to obtain a first feature map of the original image; extracts features from the first feature map in sequence through multiple group convolution blocks connected in series in a group convolution network to obtain a second feature map of the original image; and performs upsampling processing on the second feature map to obtain a super-resolution reconstructed image corresponding to the original image.
[0116] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0117] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0118] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.
[0119] The readable storage medium provided in the present application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned image super-resolution reconstruction method, and can solve the technical problem of how to reduce the computational complexity of the super-resolution reconstruction network while maintaining a good reconstruction effect. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in the present application are the same as the beneficial effects of the image super-resolution reconstruction method provided in the above-mentioned embodiment, and will not be described in detail here.
[0120] The present application also provides a computer program product, including a computer program, which implements the steps of the above-mentioned image super-resolution reconstruction method when executed by a processor.
[0121] The computer program product provided by the present application can solve the technical problem of how to reduce the computational complexity of the super-resolution reconstruction network while maintaining a good reconstruction effect. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as the beneficial effects of the image super-resolution reconstruction method provided by the above embodiment, and will not be described in detail here.
[0122] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A method for super-resolution image reconstruction, characterized in that: Applied to an image super-resolution reconstruction network, the image super-resolution reconstruction network is configured with a grouped convolutional network, and the image super-resolution reconstruction method includes: Performing convolution processing on the original image to be reconstructed to obtain a first feature map of the original image; By sequentially extracting features from the first feature map through a plurality of group convolution blocks connected in series in the group convolution network, a second feature map of the original image is obtained; The second feature map is up-sampled to obtain a super-resolution reconstructed image corresponding to the original image.
2. The image super-resolution reconstruction method according to claim 1, characterized in that: The step of extracting features from the first feature map in sequence through a plurality of group convolution blocks connected in series in the group convolution network to obtain a second feature map of the original image comprises: For any group convolution block in the group convolution network, perform multiple group convolutions on the input feature map to obtain multiple candidate feature maps under the target number of channels, wherein the input feature map is the first feature map or the output feature map of the previous group of convolution blocks; Concatenate the candidate feature maps to obtain an output feature map of the group of convolution blocks, and output the output feature map to the next group of convolution blocks connected in series with the group of convolution blocks; After traversing each group of convolution blocks in sequence, the output feature map of the last group of convolution blocks is used as the second feature map of the original image.
3. The image super-resolution reconstruction method according to claim 2, characterized in that: The step of concatenating the candidate feature maps to obtain the output feature map of the group of convolution blocks includes: Concatenate the candidate feature maps to obtain the feature map to be output of the group of convolution blocks; Perform a residual connection between the feature map to be output and the input feature map to obtain an output feature map of the group of convolution blocks.
4. The image super-resolution reconstruction method according to claim 2, characterized in that: The step of using the output feature map of the last group convolution block as the second feature map of the original image comprises: Perform a residual connection on the output feature map of the last group convolution block and the first feature map to obtain a second feature map of the original image.
5. The image super-resolution reconstruction method according to claim 1, characterized in that: The image super-resolution reconstruction network is also configured with a detail enhancement network, and the detail enhancement network is composed of a plurality of convolutional layers, wherein a convolutional kernel is configured in any convolutional layer; The step of performing upsampling processing on the second feature map to obtain a super-resolution reconstructed image corresponding to the original image comprises: Performing upsampling processing on the second feature map to obtain a candidate super-resolution reconstructed image corresponding to the original image; The candidate super-resolution reconstructed image is convolved multiple times through the detail enhancement network to obtain a super-resolution reconstructed image corresponding to the original image, wherein the size of each convolution kernel used in the multiple convolution processes is reduced successively, and any convolution kernel is obtained by convolving a preset tensor with the weight of the corresponding convolution layer.
6. The image super-resolution reconstruction method according to claim 5, characterized in that: The image super-resolution reconstruction method also includes: A training image subset is selected from a preset training image set, and blur kernel learning is performed on each target training image in the preset training image subset through a preset generative adversarial network to obtain a blur kernel corresponding to each target training image; After traversing each blur kernel, performing convolution filtering on each original training image in the training image set based on any blur kernel to obtain a blurred image corresponding to each original training image; Based on each pair of corresponding original training images and blurred images, a preset multi-layer convolutional network is trained to obtain the detail enhancement network.
7. The image super-resolution reconstruction method according to claim 1, characterized in that: The image super-resolution reconstruction method also includes: Generate a fuzzy image set based on a preset training image set, and perform data enhancement on the fuzzy image set to obtain a target training set; The image super-resolution reconstruction network is trained based on the target training set to optimize and adjust model parameters in the image super-resolution reconstruction network.
8. The image super-resolution reconstruction method according to claim 7, characterized in that: The step of performing data enhancement on the fuzzy image set includes: For any image to be enhanced in the fuzzy image set, noise is added to the image to be enhanced, and after traversing each image to be enhanced, a new fuzzy image set is obtained, so as to perform the step of data enhancement on the fuzzy image set based on the new fuzzy image set.
9. An electronic device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the image super-resolution reconstruction method according to any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the image super-resolution reconstruction method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Lightweight image super-resolution reconstruction method based on double attention mechanism
CN115496658A