Low-Resolution Image Reconstruction Method and System Based on Image Encoding-Decoding
Through the image encoding-decoding method, semantic features of low-resolution images are extracted and high-resolution images are generated through the decoder, which solves the blur and distortion problems during image reconstruction in traditional methods, and achieves efficient image detail recovery and texture maintenance.
Patent Information
- Application Number
- CN202311317324.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-11
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-10-11
AI Technical Summary
Traditional low-resolution image reconstruction methods are prone to introducing blur and distortion, and cannot effectively restore the details and texture information of the image.
Using an image encoding-decoding method, the low-resolution image is encoded into a high-dimensional feature vector through an image encoder, and then the feature vector is restored to a high-resolution image through a decoder, and a semantic fusion shallow image feature map is generated to realize image reconstruction.
It effectively improves the resolution of the image, maintains the details and texture information of the image, avoids blur and distortion, and the generated high-resolution images have higher clarity and details.
Smart Images

Figure CN117274059B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent image reconstruction, and in particular, to a low-resolution image reconstruction method and system based on image encoding-decoding. Background Art
[0002] Low-resolution image reconstruction is an important problem in the field of computer vision, which refers to converting a low-resolution image into a high-resolution image. Traditional interpolation methods are prone to introducing blur and distortion during the reconstruction process. Therefore, an optimized low-resolution image reconstruction scheme is expected. Summary of the Invention
[0003] An embodiment of the present invention provides a low-resolution image reconstruction method and system based on image encoding-decoding, which obtains a low-resolution image input by a user; performs image preprocessing on the low-resolution image to obtain an enhanced low-resolution image; performs image feature extraction on the enhanced low-resolution image to obtain a semantic fusion shallow image feature map; and generates a high-resolution image based on the semantic fusion shallow image feature map. In this way, the resolution of the image can be effectively improved, and the details and texture information of the image can be maintained, avoiding the phenomena of blur and distortion.
[0004] An embodiment of the present invention also provides a low-resolution image reconstruction method based on image encoding-decoding, which includes:
[0005] Obtaining a low-resolution image input by a user;
[0006] Performing image preprocessing on the low-resolution image to obtain an enhanced low-resolution image;
[0007] Performing image feature extraction on the enhanced low-resolution image to obtain a semantic fusion shallow image feature map; and
[0008] Generating a high-resolution image based on the semantic fusion shallow image feature map.
[0009] An embodiment of the present invention also provides a low-resolution image reconstruction system based on image encoding-decoding, which includes:
[0010] An image acquisition module, configured to acquire a low-resolution image input by a user;
[0011] An image preprocessing module, configured to perform image preprocessing on the low-resolution image to obtain an enhanced low-resolution image;
[0012] An image feature extraction module, configured to perform image feature extraction on the enhanced low-resolution image to obtain a semantic fusion shallow image feature map; and
[0013] A high-resolution image generation module for generating a high-resolution image based on the semantically fused shallow image feature map. Description of the Drawings
[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. In the drawings:
[0015] Figure 1 It is a flowchart of a low-resolution image reconstruction method based on image encoding and decoding provided in an embodiment of the present invention.
[0016] Figure 2 It is a schematic diagram of the system architecture of a low-resolution image reconstruction method based on image encoding and decoding provided in an embodiment of the present invention.
[0017] Figure 3 It is a block diagram of a low-resolution image reconstruction system based on image encoding and decoding provided in an embodiment of the present invention.
[0018] Figure 4 It is an application scenario diagram of a low-resolution image reconstruction method based on image encoding and decoding provided in an embodiment of the present invention. Detailed Embodiments
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer and more understandable, the following further details the embodiments of the present invention with reference to the drawings. Herein, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but not to limit the present invention.
[0020] As shown in the present invention and the claims, unless the context clearly indicates an exceptional situation, words such as "a", "an", "one", and / or "the" are not specifically singular and may also include plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the clearly identified steps and elements, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.
[0021] Flowcharts are used in the present invention to illustrate the operations performed by the system according to the embodiments of the present invention. It should be understood that the operations before or below do not necessarily need to be executed precisely in sequence. On the contrary, as needed, they can be executed in reverse order or simultaneously. At the same time, other operations can also be added to these processes, or one or several operations can be removed from these processes.
[0022] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0023] The term "exemplary" used herein means "serving as an example, embodiment, or illustration". Any embodiment described herein as "exemplary" is not necessarily to be construed as superior or better than other embodiments.
[0024] A low-resolution image refers to an image with relatively few pixels and less image detail, which is usually caused by limitations of image acquisition devices, image compression, or losses during image transmission, etc.
[0025] Low-resolution images have the following characteristics:
[0026] Fewer pixels: Low-resolution images have relatively few pixels, so the detail and clarity of the images are low. The same object or scene may appear blurred or unclear in a low-resolution image.
[0027] Information loss: During the process of converting a high-resolution image to a low-resolution image, some image information is usually lost. This lost information may include details, textures, edges, etc.
[0028] Blur and distortion: Due to the fewer pixels and information loss, low-resolution images may exhibit blur and distortion. The details in the image may become blurred, and the edges may become unclear, resulting in a decrease in image quality.
[0029] The goal of low-resolution image reconstruction is to convert a low-resolution image into a high-resolution image by using image processing and computer vision techniques to restore the lost details and improve the clarity of the image. This is of great significance for many application fields, such as medical imaging, surveillance systems, remote sensing, etc.
[0030] Low-resolution image reconstruction refers to the process of converting a low-resolution image into a high-resolution image, which is an important problem in the field of computer vision because high-resolution images usually contain more details and better visual quality. Traditional interpolation methods are widely used in low-resolution image reconstruction, which increase the image resolution based on interpolation between pixels. However, these methods tend to introduce blur and distortion because they cannot accurately restore the lost image details.
[0031] To solve this problem, researchers have proposed various optimized low-resolution image reconstruction schemes. Among them, deep learning-based methods have made significant progress. These methods use deep neural networks to learn the mapping relationship from low-resolution images to high-resolution images to achieve more accurate reconstruction results.
[0032] Deep learning-based low-resolution image reconstruction methods generally include the following steps: First, collect a set of high-resolution images and their corresponding low-resolution versions as training data, which can be generated by downsampling or other image processing techniques. Then, design a deep neural network model, usually a convolutional neural network (CNN) or a generative adversarial network (GAN), which includes an encoder and a decoder to learn the mapping from low-resolution images to high-resolution images. Next, use the prepared training data to train the network, and adjust the network parameters to improve the reconstruction quality by minimizing the difference between the reconstructed image and the real high-resolution image. Finally, use the trained network model to reconstruct new low-resolution images, input the low-resolution images into the encoder, and generate high-resolution images through the decoder.
[0033] Through deep learning methods, low-resolution image reconstruction can more accurately restore the details and clarity of images, providing a better visual experience and image analysis capabilities, which has important application values in fields such as image enhancement, image restoration, and medical image analysis.
[0034] This application proposes a technical solution for a low-resolution image reconstruction method based on image encoding-decoding. This solution utilizes the powerful feature extraction ability of a deep convolutional neural network (DCNN) to encode low-resolution images into high-dimensional feature vectors, and then restores the feature vectors to high-resolution images through an upsampling layer and a decoding layer. This solution can not only effectively improve the resolution of images, but also maintain the details and texture information of images, avoiding the phenomena of blurring and distortion. This solution has been experimented on multiple public datasets and shows better performance and effects compared with other common image reconstruction methods.
[0035] In one embodiment of the present invention, Figure 1 is a flowchart of a low-resolution image reconstruction method based on image encoding-decoding provided in an embodiment of the present invention. Figure 2 is a schematic diagram of the system architecture of a low-resolution image reconstruction method based on image encoding-decoding provided in an embodiment of the present invention. As Figure 1 and Figure 2As shown, the low-resolution image reconstruction method based on image encoding and decoding according to an embodiment of the present invention includes: 110, obtaining a low-resolution image input by a user; 120, performing image preprocessing on the low-resolution image to obtain an enhanced low-resolution image; 130, performing image feature extraction on the enhanced low-resolution image to obtain a semantic fusion shallow image feature map; and 140, generating a high-resolution image based on the semantic fusion shallow image feature map.
[0036] In step 110, ensure that the obtained low-resolution image is the image that the user wants to reconstruct. The image can be obtained through file upload, image URL input, or other appropriate means. In step 120, before performing image preprocessing, some preprocessing operations may be required on the low-resolution image, such as image size adjustment, denoising, brightness / contrast adjustment, etc. These operations help improve the effect of subsequent steps. Through the preprocessing operations, the quality of the low-resolution image can be improved, noise and distortion can be reduced, and a better input can be provided for subsequent steps. In step 130, use an appropriate image feature extraction method, such as a convolutional neural network (CNN), to extract the semantic information of the low-resolution image. These features should be able to capture the important features of the image, such as edges, textures, shapes, etc. The semantic fusion shallow image feature map can provide a more semantically informative representation, which helps to more accurately reconstruct the high-resolution image in subsequent steps. Through feature extraction, the important details and structures of the image can be captured. In step 140, use an appropriate image reconstruction method, such as a decoder or a generative adversarial network (GAN), to convert the semantic fusion shallow image feature map into a high-resolution image. These methods should be able to restore the details and clarity of the image. Through the reconstruction based on the semantic fusion shallow image feature map, high-quality high-resolution images can be generated. These images will have more details and better visual quality, providing a better visual experience and image analysis ability.
[0037] By gradually performing steps such as image acquisition, preprocessing, feature extraction, and image reconstruction, the reconstruction from a low-resolution image to a high-resolution image can be achieved. This method can improve the image quality, restore details, and provide better image visual effects and analysis ability.
[0038] To address the above technical problems, the technical concept of this application is based on the image reconstruction idea of an image encoder + image decoder. It extracts image features from a low-resolution image and realizes the image reconstruction of the low-resolution image through the decoder. Specifically, first, the low-resolution image is input into a neural network through the image encoder. The goal of the encoder is to learn to convert the input image into a low-dimensional feature representation, which usually captures the important information of the image. In the encoder, through a multi-layer convolutional neural network (CNN) or other feature extraction methods, meaningful features are extracted from the low-resolution image, and the features can include edges, textures, colors, etc. Next, the extracted features are input into the image decoder. The task of the decoder is to remap the low-dimensional features back to the high-resolution image space. The decoder usually uses deconvolution layers or upsampling techniques to restore the details and clarity of the image. The high-resolution image generated by the decoder is the reconstruction result of the low-resolution image, and these reconstructed images usually have higher clarity and details than the original low-resolution image.
[0039] The image reconstruction method based on an image encoder and an image decoder has the following benefits: Through the encoder, this method can extract meaningful features from the low-resolution image, and these features can be used for subsequent analysis, recognition, or other image processing tasks. The decoder can remap the low-dimensional features back to the high-resolution image space, thereby reconstructing the details and clarity of the image, which helps to improve the visual quality and recognizability of the image. The deep learning-based method has strong learning ability and can be trained through a large-scale dataset, thereby improving the accuracy and effect of image reconstruction.
[0040] The image reconstruction idea based on an image encoder and an image decoder is a powerful method that can extract features from low-resolution images and achieve high-quality image reconstruction, and has broad application prospects in the fields of image enhancement, super-resolution reconstruction, etc.
[0041] Based on this, in the technical solution of this application, first, a low-resolution image input by the user is obtained. The low-resolution image contains the content of the image that the user wants to reconstruct, and this information can help determine the objects, scenes, or objects that should be restored when generating the high-resolution image. The structural information such as edges, textures, and shapes in the low-resolution image can provide clues about the overall structure and layout of the image, and this information can guide the detail restoration and shape reconstruction when generating the high-resolution image. Although the color of the low-resolution image may be limited, it still provides information about the color distribution and overall tone of the image, and this information can be used to more accurately restore the color when generating the high-resolution image. The noise and distortion in the low-resolution image may affect the visual quality of the image. By analyzing and processing these noise and distortion, the result of generating the high-resolution image can be improved.
[0042] In the process of generating a high-resolution image, by utilizing this useful information in the low-resolution image input by the user, it can help the model better restore the details, structure, and color of the image, thereby generating a higher-quality high-resolution image. This information can serve as guidance and constraints to improve the accuracy and effect of the generation process.
[0043] In one embodiment of the present application, image preprocessing is performed on the low-resolution image to obtain an enhanced low-resolution image, including: performing image enhancement based on bilateral filtering on the low-resolution image to obtain the enhanced low-resolution image.
[0044] Performing image enhancement based on bilateral filtering on the low-resolution image to obtain an enhanced low-resolution image. Here, bilateral filtering is a commonly used image filtering method that preserves edge information while smoothing the image. Compared with traditional linear filtering methods (such as mean filtering and Gaussian filtering), bilateral filtering takes into account the spatial distance and gray-level difference between pixels to achieve a more accurate smoothing effect.
[0045] Among them, image enhancement based on bilateral filtering is an image processing technology used to improve the quality and visual effect of an image. It combines information in the spatial domain and the gray-level domain and can preserve the details of the image while removing noise.
[0046] Traditional mean filtering or Gaussian filtering smooths the details of the image while removing noise, resulting in the image becoming blurred. Bilateral filtering, by considering the spatial distance between pixels and the gray-level difference between pixels, can better preserve the edges and details of the image.
[0047] The core idea of bilateral filtering is to use a weighted average to calculate the new value of each pixel, where the weight is determined by two factors: the spatial distance weight, the closer the spatial distance between pixels, the greater the weight, which means that pixels closer to the current pixel have a greater impact on it, thus preserving the spatial structure of the image. The gray-level difference weight, the smaller the gray-level difference between pixels, the greater the weight, which means that pixels with gray levels similar to the current pixel have a greater impact on it, thus preserving the details of the image.
[0048] By adjusting the parameters of the spatial distance weight and the gray-level difference weight, the degree of filtering can be controlled. Larger weight parameters will produce a stronger smoothing effect, while smaller weight parameters will better preserve details. Image enhancement based on bilateral filtering can be applied to low-resolution images to improve the clarity and quality of the image by removing noise and preserving details. This method has a wide range of applications in the fields of image restoration, image enhancement, image denoising, etc.
[0049] Bilateral filtering can effectively suppress noise in images, including high-frequency and low-frequency noise. By removing the noise, the clarity and detail recovery effect of the image can be improved. Bilateral filtering takes into account the spatial distance between pixels and the gray-scale difference between pixels during the filtering process. This method can preserve the detail information in the image and avoid the loss of details caused by over-smoothing. Bilateral filtering protects the edge regions, ensuring the clarity and accuracy of the edges, which helps to improve the contour and shape recovery effect of the image. Bilateral filtering can maintain the color information of the image and avoid color offset or distortion, which is crucial for accurate color restoration when generating high-resolution images.
[0050] By image enhancement based on bilateral filtering, the quality of low-resolution images can be improved, noise and distortion can be reduced, and the details, edges, and colors of the image can be preserved. These beneficial effects will provide a better input for subsequent steps and help generate higher-quality enhanced low-resolution images.
[0051] In an embodiment of the present application, image feature extraction is performed on the enhanced low-resolution image to obtain a semantic fusion shallow image feature map, including: extracting the shallow image features, middle image features, and deep image features of the enhanced low-resolution image to obtain a shallow image feature map, a middle image feature map, and a deep image feature map; and fusing the shallow image feature map, the middle image feature map, and the deep image feature map to obtain the semantic fusion shallow image feature map.
[0052] Then, the shallow image features, middle image features, and deep image features of the enhanced low-resolution image are extracted to obtain a shallow image feature map, a middle image feature map, and a deep image feature map.
[0053] It should be understood that the shallow image features are very sensitive to the detail information of the image and can provide the texture and edge details of the image; the middle image features are feature representations between the shallow features and the deep features. Compared with the shallow image features, these features are more abstract and can capture the semantic information and structure of the image. In low-resolution image enhancement, extracting the middle image features helps to restore the overall structure and semantic content of the image; the deep features are high-level feature representations extracted by a deep convolutional neural network. These features have higher-level semantic information and can capture the abstract concepts and high-level semantic content of the image.
[0054] In a specific example of the present application, the implementation manner of extracting the shallow image features, middle image features, and deep image features of the enhanced low-resolution image to obtain a shallow image feature map, a middle image feature map, and a deep image feature map is: passing the enhanced low-resolution image through an image feature extractor based on a pyramid network to obtain a shallow image feature map, a middle image feature map, and a deep image feature map.
[0055] It should be understood that the pyramid network is a multi-scale image processing technology used to extract image features at different scales. It mimics the structure of a pyramid, and through multiple levels of filtering and downsampling operations on the input image, it obtains image features at different scales. The pyramid network usually consists of multiple resolution levels, and each level performs filtering and downsampling operations on the input image. In each level, the resolution of the image decreases, but the receptive field of the features increases. In this way, it can capture image details and structural information at different scales.
[0056] In the pyramid network, the shallow image feature maps correspond to the higher-resolution image levels, while the deep image feature maps correspond to the lower-resolution image levels. Shallow feature maps usually contain more details and texture information, while deep feature maps contain more advanced semantic information. The main advantage of the pyramid network is its ability to capture multi-level features of an image at different scales and provide rich context information, which makes the pyramid network very useful in many computer vision tasks, such as object detection, image segmentation, image enhancement, etc.
[0057] By processing a low-resolution image through an image feature extractor based on the pyramid network, shallow image feature maps, middle-layer image feature maps, and deep image feature maps can be obtained, thus providing rich feature representations and multi-scale information, and providing better performance and effects for image enhancement and other image processing tasks. The pyramid network can extract image features at different scales. By using filters and pooling operations at multiple scales, it can capture different levels of details and structural information of the image. Shallow feature maps usually contain more edge and texture information, while deep feature maps contain more advanced semantic information. The pyramid network can utilize receptive fields at different scales to capture the context information of the image. Larger receptive fields can capture a broader context, which helps to understand the global structure and semantic relationships in the image and improve the semantic understanding and feature expression ability of the image. By extracting features at different levels in the pyramid network, these features can be used for image reconstruction and enhancement. Shallow feature maps can be used to restore the details and texture of the image, middle-layer feature maps can be used to enhance the contours and shapes of the image, and deep feature maps can be used to improve the semantic understanding and content expression of the image. The pyramid network has a multi-level structure and can extract features from different levels. Therefore, it has a certain degree of robustness and stability to changes and noises in the image, which helps to improve the effect of image processing and reduce the adverse effects caused by noises and changes.
[0058] Through an image feature extractor based on a pyramid network, rich image features can be obtained at different levels, and then used for image reconstruction, enhancement, and other further processing tasks, which helps to improve the quality, clarity, and semantic understanding ability of the image.
[0059] Next, fuse the shallow image feature map, the middle-layer image feature map, and the deep image feature map to obtain the semantic fusion shallow image feature map. That is, by comprehensively using shallow features, middle-layer features, and deep features, the feature information at different levels can be fully utilized, thereby achieving comprehensive enhancement of the image. More specifically, shallow features provide details, middle-layer features provide semantics, and deep features provide high-level semantics. They complement each other and jointly promote the quality improvement of low-resolution images.
[0060] In a specific example of the present application, the encoding process of fusing the shallow image feature map, the middle-layer image feature map, and the deep image feature map to obtain the semantic fusion shallow image feature map includes: first, fusing the middle-layer image feature map and the deep image feature map to obtain a multi-scale semantic image feature map; then using a joint semantic propagation module to fuse the multi-scale semantic image feature map and the shallow image feature map to obtain the semantic fusion shallow image feature map.
[0061] The middle-layer feature map and the deep feature map usually have different spatial resolutions. Before fusion, it is necessary to ensure that they have the same size. Interpolation or convolution operations can be used to adjust the size of the feature map to make it match. Fusion methods include element-wise addition, element-wise multiplication, concatenation, etc. The choice of an appropriate fusion method depends on the specific task and the nature of the features.
[0062] The multi-scale semantic image feature map usually has a lower spatial resolution and more advanced semantic information, while the shallow feature map has a higher spatial resolution and more detailed information. Before fusion, it is necessary to ensure that they have the same size and consider their semantic differences. The joint semantic propagation module can be used to fuse the multi-scale semantic image feature map and the shallow feature map. Semantic information can be propagated from the multi-scale feature map to the shallow feature map through an attention mechanism, convolution operations, or other methods, thereby achieving semantic fusion.
[0063] More specifically, in an embodiment of the present application, the encoding process of using a joint semantic propagation module to fuse the multi-scale semantic image feature map and the shallow image feature map to obtain a semantically fused shallow image feature map includes: first, upsampling the multi-scale semantic image feature map to obtain a resolution-reconstructed feature map; subsequently, performing point convolution, batch normalization operation, and ReLU-based non-activation function operation on the global mean feature vector obtained after global mean pooling of the resolution-reconstructed feature map to obtain a global semantic vector; then, performing point convolution, batch normalization operation, and ReLU-based non-activation function operation on the resolution-reconstructed feature map to obtain a local semantic vector; next, performing point addition of the global semantic vector and the local semantic vector to obtain a semantic weight vector; further, using the semantic weight vector as a weight vector to perform weighted processing on the shallow image feature map to obtain a semantic joint feature map; finally, fusing the shallow image feature map and the semantic joint feature map to obtain the semantically fused shallow image feature map.
[0064] In an embodiment of the present application, based on the semantically fused shallow image feature map, generating a high-resolution image includes: passing the semantically fused shallow image feature map through an image reconstruction model based on a decoder to generate a high-resolution image.
[0065] Passing the semantically fused shallow image feature map through an image reconstruction model based on a decoder to generate a high-resolution image. The semantically fused shallow image feature map contains multi-scale semantic information and detailed information. Through the decoder model, these feature maps can be converted into high-resolution images, thereby improving the details and clarity of the images. The generated high-resolution images can restore more details, making the images more real and vivid.
[0066] Some details and information may be lost during the loss compression and downsampling processes of low-resolution images. Through the image reconstruction model based on a decoder, an attempt can be made to restore these lost information. The decoder model utilizes the semantic information and other context information in the semantically fused shallow image feature map for reconstruction and compensation, thereby improving the quality and integrity of the generated images.
[0067] The generated high-resolution images can provide better visual perception and image analysis capabilities. In many computer vision tasks, such as object detection, image segmentation, image recognition, etc., high-resolution images can usually provide more accurate and reliable results. Generating high-resolution images through an image reconstruction model based on a decoder can improve the performance and effects of these tasks.
[0068] Processing the semantic fusion shallow image feature map using a decoder-based image reconstruction model can improve the details and clarity of the image, restore lost information, and enhance visual perception and image analysis capabilities, which helps to improve the image quality, provide a better visual experience, and achieve better results in various computer vision tasks.
[0069] In one embodiment of the present application, the low-resolution image reconstruction method based on image encoding-decoding further includes a training step: for training the image feature extractor based on the pyramid network, the joint semantic propagation module, and the decoder-based image reconstruction model; wherein, the training step includes: obtaining training data, the training data includes training low-resolution images, and real generated high-resolution images; performing image enhancement based on bilateral filtering on the training low-resolution images to obtain training enhanced low-resolution images; passing the training enhanced low-resolution images through the image feature extractor based on the pyramid network to obtain training shallow image feature maps, training middle layer image feature maps, and training deep layer image feature maps; fusing the training middle layer image feature maps and the training deep layer image feature maps to obtain training multi-scale semantic image feature maps; using the joint semantic propagation module to fuse the training multi-scale semantic image feature maps and the training shallow image feature maps to obtain training semantic fusion shallow image feature maps; performing probability density convergence optimization with feature scale constraints on each feature matrix of the training semantic fusion shallow image feature maps to obtain optimized training semantic fusion shallow image feature maps; passing the optimized training semantic fusion shallow image feature maps through the decoder-based image reconstruction model for decoding regression to obtain a decoding loss function value; and training the image feature extractor based on the pyramid network, the joint semantic propagation module, and the decoder-based image reconstruction model based on the decoding loss function value and through the propagation in the direction of gradient descent.
[0070] In the technical solution of the present application, when passing the training enhanced low-resolution image through the image feature extractor based on the pyramid network, the training shallow image feature map, the training middle image feature map, and the training deep image feature map can express the image semantic features at different depths and different scales based on the pyramid network. Thus, when using the joint semantic propagation module to fuse the training multi-scale semantic image feature map and the training shallow image feature map to obtain the semantic fusion shallow image feature map, each feature matrix of the training shallow image feature map is weighted by the global semantic feature vector of the training multi-scale semantic image feature map to obtain each feature matrix of the training semantic fusion shallow image feature map. Therefore, when taking the shallow image semantic feature representation of each feature matrix of the training semantic fusion shallow image feature map as the foreground object feature representation, while performing the image semantic feature fusion representation at different depths and different scales, the heterogeneous distribution of the image semantic space of the high-dimensional features of each feature matrix of the training semantic fusion shallow image feature map based on different feature scales and feature depths will also be introduced, thereby causing the image semantic space probability density mapping error of the image semantic features of each feature matrix of the training semantic fusion shallow image feature map, which affects the image quality of the generated high-resolution image obtained by the training semantic fusion shallow image feature map through the image reconstruction model based on the decoder.
[0071] Here, the applicant of the present application further discovers that this imbalance is largely related to the feature expression scale, that is, the image semantic feature expression scale of the source image domain of the feature matrix and the multi-dimensional channel correlation distribution scale among each feature matrix. For example, it can be understood that relative to the scale of the multi-dimensional channel correlation distribution, the more unbalanced the image semantic feature distribution of the source image domain is, the more unbalanced the overall expression of the training semantic fusion shallow image feature map will be. Therefore, preferably, for each feature matrix of the training semantic fusion shallow image feature map, denoted as M for example k Perform probability density convergence optimization for feature scale constraint.
[0072] In one embodiment of the present invention, probability density convergence optimization of feature scale constraints is performed on each feature matrix of the training semantic fusion shallow image feature map to obtain an optimized training semantic fusion shallow image feature map, including: respectively calculating a first probability density convergence weight of a feature vector composed of global feature means of each feature matrix of the training semantic fusion shallow image feature map, and a sequence of second probability density convergence weights of each feature matrix of the training semantic fusion shallow image feature map; weighting the training semantic fusion shallow image feature map along channels with the first probability density convergence weight to obtain a weighted training semantic fusion shallow image feature map; weighting each feature matrix of the weighted training semantic fusion shallow image feature map with the sequence of second probability density convergence weights to obtain the optimized training semantic fusion shallow image feature map.
[0073] In a specific embodiment of the present invention, respectively calculating a first probability density convergence weight of a feature vector composed of global feature means of each feature matrix of the training semantic fusion shallow image feature map, and a sequence of second probability density convergence weights of each feature matrix of the training semantic fusion shallow image feature map, includes: respectively calculating a first probability density convergence weight of a feature vector composed of global feature means of each feature matrix of the training semantic fusion shallow image feature map, and a sequence of second probability density convergence weights of each feature matrix of the training semantic fusion shallow image feature map according to the following formula;
[0074] wherein, the formula is:
[0075]
[0076]
[0077] wherein, M k is the k-th feature matrix of the training semantic fusion shallow image feature map, L is the number of channels of the training semantic fusion shallow image feature map, v k is the global feature mean of the feature matrix M k , V is the feature vector composed of v k , represents the square of the two-norm of the feature vector V, S is the scale of the feature matrix M k , that is, width multiplied by height, and represents the square of the Frobenius norm of the feature matrix M k , m i,j is the eigenvalue at the (i, j) position in the feature matrix M k , w1 is the first probability density convergence weight, w 2kis the k-th probability density convergence weight of the sequence of the second probability density convergence weights.
[0078] Here, the probability density convergence optimization of the feature scale constraint can be performed by a tail distribution strengthening mechanism of a class of standard Cauchy distributions to perform correlation constraints on the multi-level distribution structure of the feature probability density distribution in the high-dimensional feature space based on the feature scale, so that the probability density distributions of high-dimensional features with different scales are uniformly expanded in the overall probability density space, thereby compensating for the heterogeneity of probability density convergence caused by feature scale deviation. In this way, during the training process, the above weight w1 is used to weight the training semantic fusion shallow image feature map along the channels, and the above weight w 2k weights each feature matrix M of the training semantic fusion shallow image feature map k to improve the convergence of the optimized training semantic fusion shallow image feature map in the predetermined probability density distribution domain, thereby improving the image quality of the generated high-resolution image obtained by the image reconstruction model based on the decoder.
[0079] In summary, the low-resolution image reconstruction method based on image encoding-decoding according to the embodiments of the present invention is clarified. Based on the image reconstruction idea of an image encoder + image decoder, it extracts image features from a low-resolution image and realizes the image reconstruction of the low-resolution image through a decoder.
[0080] In an embodiment of the present invention, Figure 3 is a block diagram of a low-resolution image reconstruction system provided in the embodiments of the present invention. As Figure 3 shown, the low-resolution image reconstruction system 200 based on image encoding-decoding according to the embodiments of the present invention includes: an image acquisition module 210 for acquiring a low-resolution image input by a user; an image preprocessing module 220 for performing image preprocessing on the low-resolution image to obtain an enhanced low-resolution image; an image feature extraction module 230 for performing image feature extraction on the enhanced low-resolution image to obtain a semantic fusion shallow image feature map; and a high-resolution image generation module 240 for generating a high-resolution image based on the semantic fusion shallow image feature map.
[0081] In the low-resolution image reconstruction system based on image encoding-decoding, the image preprocessing module is used to: perform image enhancement based on bilateral filtering on the low-resolution image to obtain the enhanced low-resolution image.
[0082] In the low-resolution image reconstruction system based on image encoding and decoding, the image feature extraction module includes: a feature map extraction unit for extracting the shallow image features, middle image features, and deep image features of the enhanced low-resolution image to obtain a shallow image feature map, a middle image feature map, and a deep image feature map; and a fusion unit for fusing the shallow image feature map, the middle image feature map, and the deep image feature map to obtain the semantically fused shallow image feature map.
[0083] Here, those skilled in the art can understand that the specific functions and operations of each unit and module in the above low-resolution image reconstruction system based on image encoding and decoding have been described in detail above with reference to Figures 1 to 2 the description of the low-resolution image reconstruction method based on image encoding and decoding, and therefore, the repeated description thereof will be omitted.
[0084] As described above, the low-resolution image reconstruction system 200 according to the embodiment of the present invention can be implemented in various terminal devices, such as a server for low-resolution image reconstruction based on image encoding and decoding. In one example, the low-resolution image reconstruction system 200 according to the embodiment of the present invention can be integrated into the terminal device as a software module and / or a hardware module. For example, the low-resolution image reconstruction system 200 can be a software module in the operating system of the terminal device, or can be an application program developed for the terminal device; of course, the low-resolution image reconstruction system 200 can also be one of the numerous hardware modules of the terminal device.
[0085] Alternatively, in another example, the low-resolution image reconstruction system 200 and the terminal device can also be separate devices, and the low-resolution image reconstruction system 200 based on image encoding and decoding can be connected to the terminal device through a wired and / or wireless network and transmit interaction information in accordance with a predefined data format.
[0086] Figure 4 This is an application scenario diagram of a low-resolution image reconstruction method provided in an embodiment of the present invention. As Figure 4 shown, in this application scenario, first, a low-resolution image input by the user is obtained (for example, as shown in C in Figure 4 ); then, the obtained low-resolution image is input to a server deployed with a low-resolution image reconstruction algorithm based on image encoding and decoding (for example, as shown in Figure 4In the S) illustrated, where the server can process the low-resolution image based on an image coding-decoding low-resolution image reconstruction algorithm to generate a high-resolution image.
[0087] In the specific embodiments described above, the objectives, technical solutions, and beneficial effects of the present invention have been further described in detail. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A low-resolution image reconstruction method based on image encoding-decoding, characterized in that, Including: Obtain a low-resolution image input by a user; Perform image preprocessing on the low-resolution image to obtain an enhanced low-resolution image; Extract image features from the enhanced low-resolution image to obtain a semantically fused shallow image feature map; Generate a high-resolution image based on the semantically fused shallow image feature map; Among them, it also includes that in the training step, for each feature matrix of the training semantically fused shallow image feature map, perform probability density convergence optimization on feature scale constraints to obtain an optimized training semantically fused shallow image feature map, including: Calculate the first probability density convergence weight of the feature vector composed of the global feature means of each feature matrix of the training semantically fused shallow image feature map and the sequence of the second probability density convergence weights of each feature matrix of the training semantically fused shallow image feature map respectively; Weight the training semantically fused shallow image feature map along the channels with the first probability density convergence weight to obtain a weighted training semantically fused shallow image feature map; Weight each feature matrix of the weighted training semantically fused shallow image feature map with the sequence of the second probability density convergence weights to obtain an optimized training semantically fused shallow image feature map; Among them, calculating the first probability density convergence weight of the feature vector composed of the global feature means of each feature matrix of the training semantically fused shallow image feature map and the sequence of the second probability density convergence weights of each feature matrix of the training semantically fused shallow image feature map respectively includes: calculating the first probability density convergence weight of the feature vector composed of the global feature means of each feature matrix of the training semantically fused shallow image feature map and the sequence of the second probability density convergence weights of each feature matrix of the training semantically fused shallow image feature map respectively with the following formula; Among them, the formula is: Among them, M k is the k-th feature matrix for training the shallow image feature map of semantic fusion, L is the number of channels for training the shallow image feature map of semantic fusion, and v k is the global feature mean of the feature matrix M k , V is the feature vector composed of v k , represents the square of the two-norm of the feature vector V, S is the scale of the feature matrix M k , and represents the square of the Frobenius norm of the feature matrix M k , m i,j is the eigenvalue at the (i, j) position in the feature matrix M k , w1 is the first probability density convergence weight, and w 2k is the k-th probability density convergence weight in the sequence of the second probability density convergence weights.
2. The low-resolution image reconstruction method based on image encoding and decoding according to claim 1, characterized in that Performing image preprocessing on the low-resolution image to obtain an enhanced low-resolution image includes: Perform image enhancement based on bilateral filtering on the low-resolution image to obtain the enhanced low-resolution image.
3. The low-resolution image reconstruction method based on image encoding and decoding according to claim 2, wherein Extracting image features from the enhanced low-resolution image to obtain a semantically fused shallow image feature map includes: Extract the shallow image features, middle image features, and deep image features of the enhanced low-resolution image to obtain a shallow image feature map, a middle image feature map, and a deep image feature map; and Fuse the shallow image feature map, the middle image feature map, and the deep image feature map to obtain the semantically fused shallow image feature map.
4. The low-resolution image reconstruction method based on image encoding and decoding according to claim 3, characterized in that Extracting the shallow image features, middle image features, and deep image features of the enhanced low-resolution image to obtain a shallow image feature map, a middle image feature map, and a deep image feature map includes: Pass the enhanced low-resolution image through an image feature extractor based on a pyramid network to obtain the shallow image feature map, the middle image feature map, and the deep image feature map.
5. The low-resolution image reconstruction method based on image encoding and decoding according to claim 4, characterized in that Fusing the shallow image feature map, the middle image feature map, and the deep image feature map to obtain the semantically fused shallow image feature map includes: Fuse the middle image feature map and the deep image feature map to obtain a multi-scale semantic image feature map; and Use a joint semantic propagation module to fuse the multi-scale semantic image feature map and the shallow image feature map to obtain the semantically fused shallow image feature map.
6. The low-resolution image reconstruction method based on image encoding and decoding according to claim 5, wherein, Based on the semantically fused shallow image feature map, generate a high-resolution image, including: Pass the semantically fused shallow image feature map through an image reconstruction model based on a decoder to generate a high-resolution image.
7. The method for reconstructing a low-resolution image based on image encoding and decoding according to claim 6, characterized in that It also includes a training step: for training the image feature extractor based on the pyramid network, the joint semantic propagation module, and the image reconstruction model based on the decoder; Among them, the training step includes: Obtain training data, which includes training low-resolution images and real generated high-resolution images; Perform image enhancement based on bilateral filtering on the training low-resolution images to obtain training enhanced low-resolution images; Pass the training enhanced low-resolution images through the image feature extractor based on the pyramid network to obtain training shallow image feature maps, training middle-layer image feature maps, and training deep-layer image feature maps; Fuse the training middle-layer image feature map and the training deep-layer image feature map to obtain a training multi-scale semantic image feature map; Use the joint semantic propagation module to fuse the training multi-scale semantic image feature map and the training shallow image feature map to obtain a training semantically fused shallow image feature map; Perform probability density convergence optimization with feature scale constraints on each feature matrix of the training semantically fused shallow image feature map to obtain an optimized training semantically fused shallow image feature map; Pass the optimized training semantically fused shallow image feature map through the image reconstruction model based on the decoder for decoding regression to obtain a decoding loss function value; Based on the decoding loss function value and through the direction propagation of gradient descent, train the image feature extractor based on the pyramid network, the joint semantic propagation module, and the image reconstruction model based on the decoder.
8. An image coding-decoding based low-resolution image reconstruction system for the image coding-decoding based low-resolution image reconstruction method according to any one of claims 1-7, characterized in that, It includes: An image acquisition module for acquiring the low-resolution image input by the user; An image preprocessing module for performing image preprocessing on the low-resolution image to obtain an enhanced low-resolution image; An image feature extraction module for performing image feature extraction on the enhanced low-resolution image to obtain a semantically fused shallow image feature map; And A high-resolution image generation module for generating a high-resolution image based on the semantically fused shallow image feature map.
Citation Information
Patent Citations
Multi-branch image semantic segmentation method and system based on AM and feature fusion
CN116681889A