Image processing method, electronic device and computer program product

By combining a multi-level deep dual encoding/decoding model with an attention residual module, the problem of feature information loss during image magnification is solved, and high-quality super-resolution image reconstruction is achieved.

CN120689210APending Publication Date: 2025-09-23CHINA MOBILE GROUP ZHEJIANG +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510867488.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

In the prior art, feature information is easily lost during image magnification, resulting in a lack of texture and edge details in the reconstructed super-resolution image and poor display effect.

Method used

A multi-level deep dual encoder-decoder model is used to extract deep features of the original image, and an attention residual module is added between adjacent encoders and decoders to obtain high-resolution images through convolution processing and upsampling.

Benefits of technology

It effectively alleviates the problem of information loss during continuous encoding of images, restores deep features, and improves the display effect and quality of images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689210A_ABST
    Figure CN120689210A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method, electronic equipment and a computer program product, and relates to the technical field of image processing. The image processing method comprises the following steps: acquiring a to-be-processed original image; sampling processing is carried out on the original image, depth feature extraction is carried out on the original image through a multi-stage depth dual coding and decoding model in the sampling process, initial features and depth features are obtained, the multi-stage depth dual coding and decoding model is a model with a plurality of symmetrical coding processes and decoding processes, and the initial features and the depth features are extracted through the multi-stage depth dual coding and decoding model. Attention residual modules are added between adjacent encoders and between adjacent decoders. And convolution processing and up-sampling processing are carried out on the initial features and the depth features to obtain a target image, and the resolution of the target image is higher than that of the original image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to an image processing method, electronic equipment, and computer program product. Background Art

[0002] As an important carrier of visual information, image resolution directly determines the richness and granularity of visual information. In practical applications, the resolution of raw images captured by cameras is sometimes insufficient to meet application requirements due to limitations such as manufacturing costs, equipment power consumption, transmission bandwidth, and imaging conditions. Low-resolution images result in the loss of significant detail, which not only reduces the visual quality of the image but also limits the performance of subsequent image processing. Image super-resolution reconstruction technology uses algorithms at the software level to achieve high-resolution images. Due to its flexibility, cost-effectiveness, and applicability, this technology has great application value in both daily life and production activities.

[0003] However, most video image applications in related technologies are to directly and violently magnify to reconstruct the image video, such as Figure 1 As shown, this method easily loses feature information during the image enlargement process, resulting in the reconstructed super-resolution image lacking texture and edge details, and the displayed image effect is blurred, affecting the user experience. Summary of the Invention

[0004] The embodiments of the present application provide an image processing method, an electronic device, and a computer program product to solve the problem in the related art that feature information is easily lost during image magnification and the display effect is poor.

[0005] In a first aspect, an embodiment of the present application provides an image processing method, comprising: Get the original image to be processed; Sampling the original image, and extracting depth features from the original image using a multi-level deep dual encoding and decoding model during the sampling process to obtain initial features and depth features, wherein the multi-level deep dual encoding and decoding model is a model with symmetrical encoding and decoding processes, and an attention residual module is added between adjacent encoders and between adjacent decoders; The initial features and the depth features are convolved and up-sampled to obtain a target image, where the resolution of the target image is higher than that of the original image.

[0006] In a second aspect, an embodiment of the present application provides an electronic device comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the method described in the first aspect.

[0007] In a third aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0008] In a fourth aspect, an embodiment of the present application provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions, which, when executed by a computer, implement the steps of the method described in the first aspect.

[0009] In an embodiment of the present application, the original image to be processed is first obtained, and then the original image is sampled. During the sampling process, a multi-level deep dual encoding and decoding model is used to extract deep features of the original image to obtain initial features and deep features, wherein the multi-level deep dual encoding and decoding model is a model with multiple symmetrical encoding and decoding processes, and an attention residual module is added between adjacent encoders and adjacent decoders. Finally, the initial features and the deep features are convolved and upsampled to obtain a target image, and the resolution of the target image is higher than that of the original image. The embodiment of the present application utilizes multiple encoding and decoding to obtain more internal feature information of the image, and adds residual blocks between adjacent encoders. Deep feature extraction and recovery are performed during continuous encoding and decoding, which can alleviate the problem of multiple continuous encoders and decoders losing information during continuous encoding and the difficulty in recovering deep features during decoding, reduce feature loss, and improve the final display effect of the image. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 is a schematic diagram of imaging effects of different resolutions provided by an embodiment of the present application; Figure 2 is a flowchart of the image processing method provided in an embodiment of the present application; Figure 3 This is the overall system architecture for image processing provided by the embodiments of the present application; Figure 4 This is a diagram of the encoding and decoding model provided in the embodiment of the present application; Figure 5 It is the continuous encoding and decoding model provided in the embodiment of the present application; Figure 6 Schematic diagram of the deep dual encoding and decoding model provided by the embodiment of the present application; Figure 7 This is a schematic diagram of the structure of the attention residual module of the deep dual encoding and decoding provided in an embodiment of the present application; Figure 8 This is a schematic diagram of the deep compression network framework provided by an embodiment of the present application; Figure 9 This is a schematic diagram of the image super-resolution reconstruction algorithm model framework of the deep dual codec provided in the embodiment of the application Figure 10 This is a schematic diagram of data preprocessing and model training provided in an embodiment of the present application; Figure 11 is a schematic diagram of an image processing device provided in an embodiment of the present application; Figure 12 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0011] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0012] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects and are not used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments of this application can be implemented in an order other than those illustrated or described herein. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0013] The following is combined with Figures 1 to 12 , an image processing method, electronic device and computer program product provided by the embodiments of the present application are described in detail through specific embodiments and their application scenarios.

[0014] like Figure 2 As shown in FIG, it is a flow chart of an image processing method provided by an embodiment of the present application. Figure 2 As shown, the image processing method may include: contents shown in S101 to S103.

[0015] In S101 , an original image to be processed is obtained.

[0016] In S102, the original image is sampled and processed, and during the sampling process, the multi-level deep dual encoding and decoding model is used to extract the depth features of the original image to obtain initial features and depth features, wherein the multi-level deep dual encoding and decoding model is a model with multiple symmetrical encoding and decoding processes, and an attention residual module is added between adjacent encoders and adjacent decoders.

[0017] In S103 , the initial features and the depth features are convolved and up-sampled to obtain a target image, where the resolution of the target image is higher than that of the original image.

[0018] In an embodiment of the present application, the original image to be processed is first obtained, and then the original image is sampled. During the sampling process, a multi-level deep dual encoding and decoding model is used to extract deep features of the original image to obtain initial features and deep features, wherein the multi-level deep dual encoding and decoding model is a model with multiple symmetrical encoding and decoding processes, and an attention residual module is added between adjacent encoders and adjacent decoders. Finally, the initial features and the deep features are convolved and upsampled to obtain a target image, and the resolution of the target image is higher than that of the original image. The embodiment of the present application utilizes multiple encoding and decoding to obtain more internal feature information of the image, and adds residual blocks between adjacent encoders. Deep feature extraction and recovery are performed during continuous encoding and decoding, which can alleviate the problem of multiple continuous encoders and decoders losing information during continuous encoding and the difficulty in recovering deep features during decoding, reduce feature loss, and improve the final display effect of the image.

[0019] This application can be used as Figure 3 The architecture shown in the figure is used for image processing. This mainly includes data modules, model training, and algorithm model application. In terms of algorithm optimization, through continuous training, optimization of neural network models, and effective learning strategies, the system can ultimately produce super-resolution reconstructed images with high efficiency and quality.

[0020] The system is divided into three main layers: application layer, model layer and data source layer.

[0021] The data source layer consists of general and specific data sources. General data sources are further divided into general image libraries and network image libraries. The general image library provides a broad range of basic high-definition images, which are more adaptable to the robustness of the model system and can improve the generalization performance of the model. The network image library includes copyright-free images from various fields that are publicly available online. Furthermore, the system decouples the data processing modules and implements preprocessing such as data cleaning, image upsampling, downsampling, and interpolation calculations to provide direct data support for model training.

[0022] The model layer proposes a variety of model algorithms to achieve different capabilities, and can also combine various models to achieve more powerful functions.

[0023] First, the image super-resolution reconstruction algorithm model of deep dual coding and decoding is presented. Specifically, this embodiment adopts a continuous encoding and decoding model to obtain feature information with continuous correlation within the image. Then, based on the continuous encoding and decoding model, a deep dual coding and decoding model is adopted. This model has a multi-level encoding and decoding structure that can respectively obtain feature information at different levels within the image. It also adopts the attention residual module of deep dual coding and decoding, which generates different weights for different levels of decoded features and recalibrates multi-level image features. Finally, a deep dual coding and decoding attention residual network is constructed to reconstruct high-resolution images. This network model is mainly divided into three parts: extracting the initial features of the low-resolution image through the initial convolutional layer, extracting deep multi-level features through multiple end-to-end connected deep dual coding and decoding attention residual modules, and reconstructing the super-resolution image through the upsampling module and reconstruction convolutional layer.

[0024] In related technologies, a Convolutional Neural Networks (CNN) encoder-decoder architecture is generally used for super-resolution reconstruction. This network architecture uses CNN as both an encoder and a decoder. The encoder encodes the image content and detail features, while the convolutional network filters out some noise; the decoder restores the image content and details. However, due to the limitations of the convolution kernel, this algorithm can only consider the information around the local receptive field of each pixel and lacks contextual information. In order to solve this problem, an Encoder-Decoder architecture based on dilated convolution can be used. This structure introduces dilated convolution, increases the receptive field of the convolution layer, and can obtain more global content information and some implicit correlation features between each receptive field. However, this will also cause the model parameters to explode and the reasoning speed to be slow. In order to solve the above problems, this application constructs a deep dual encoding and decoding model to obtain image features, and its schematic diagram is shown as follows. Figure 4 shown.

[0025] Assuming the input of the encoder and decoder model is X and the output is Y, the operation is as follows:

[0026] in, Represents the logical calculation of the encoder, Represents the logical calculation of the decoder, and Y represents the input feature X of size H x Wx C (Height×Width×Channel, height×width×number of channels) passing through the encoder The logical calculation becomes a feature of size H x W x C / r (compression factor r), which is then passed through the decoder The image features of size H x W x C are restored.

[0027] Since some images are rich in content, their features may be scattered and complex, and images of different scenes may also have continuous correlation and sparsity, continuous encoding and decoding are required to obtain these detailed features, which is very important for image reconstruction. A single-layer encoding and decoding architecture can only obtain the basic information of the shallow layer of the image. Therefore, this application can adopt a multi-layer encoding and decoding architecture to better handle the complexity and continuous correlation of the internal feature information of the image. The schematic diagram is shown as follows: Figure 5 As shown. Residual blocks are added between adjacent encoders to extract and recover deep features during the continuous encoding and decoding process. This alleviates the problem of information loss during the continuous encoding process and the difficulty in recovering deep features during decoding caused by multiple continuous encoders and decoders. Because the encoding and decoding processes are symmetrical, this application refers to this solution as a network architecture dual encoding and decoding model.

[0028] Assuming the input of the model is X and the output is Y, its operation is as follows:

[0029] Among them, the i-th calculation of the encoder is , the calculation of the i-th decoder is , Represents the calculation of the residual block. The input image feature X passes through the combined architecture of the encoder and the residual block, and multiple consecutive operations are performed to obtain a size of H x W x C / The intermediate image representation is usually a feature vector, which is also reversely calculated by the decoder multiple times to obtain the final encoding feature Y of size H x W x C.

[0030] In a possible embodiment of the present application, the use of a multi-level deep dual codec model to extract depth features from the original image during the sampling process to obtain initial features and depth features may include: using the first level deep dual codec in the multi-level deep dual codec model to extract depth features from the original image during the sampling process to obtain initial features; and using the initial features using multiple consecutive deep dual codecs and attention residual operations in the multi-level deep dual codec model to obtain depth features.

[0031] like Figure 9 As shown in the figure, the model framework mainly consists of feature extraction, convolutional layer upsampling and reconstruction operations of the deep dual encoding and decoding module.

[0032] When the required reconstructed image is input to the algorithm model, it first passes through the feature extraction of the deep dual encoding and decoding module, the input of this module is a low-resolution image, and the output is the image feature with 64 channels. Assume that the low-resolution image is , the calculation process is as follows:

[0033] in, The initial features of the low-resolution image are obtained by the deep dual encoding and decoding module. Represents the convolution operation, where 3 represents the convolution kernel size of 3x3. After K consecutive deep dual encoding and decoding attention residual modules, deep features are extracted and fused through a convolution layer with a convolution kernel of 3x3. The deep image features with 64 channels are output. The process of calculating deep feature extraction is as follows:

[0034] in, Represented as the extracted deep features, arrive This represents the computation of K consecutive attention residual modules for deep dual encoding and decoding. The image's deep features are combined with the initial features, and then all image information is reconstructed through a sub-pixel convolutional upsampling module and a reconstruction convolutional layer. The result is a high-definition, high-resolution image. The reconstructed image is a generic image with 3 channels. The upsampling and reconstruction operations are as follows:

[0035] in, Represented as the operation of the upsampling module; Represented as a reconstructed super-resolution image. Using the deep dual codec super-resolution reconstruction algorithm model, a low-resolution image can be quickly transformed into a high-resolution image. Similarly, a video is the result of a rapid reconstruction of multiple consecutive video frames.

[0036] Specifically, the convolution processing and upsampling processing of the initial features and the depth features to obtain the target image may include: combining the initial features and the rematched decoding features through local residual connections to obtain updated features; upsampling the updated features to obtain upsampled features; convolution processing of the upsampled features to obtain a reconstructed super-resolution image, wherein the reconstructed super-resolution image is the target image.

[0037] In a possible embodiment of the present application, the initial features are obtained by using multiple consecutive deep dual codecs and attention residual operations in a multi-level deep dual codec model to obtain deep features, which can include: subjecting the initial features to a first-level codec and residual block operation to obtain a first deep feature; subjecting the first deep feature to a second-level codec and residual block operation to obtain a second deep feature; subjecting the second deep feature to a third-level codec and residual block operation to obtain a third deep feature, until all deep features are obtained through multi-level codecs and residual block operations.

[0038] In this embodiment, since the image is composed of pixels, the difficulty of pixel reconstruction is related to the specific features. It is relatively easy to reconstruct image blocks and structures with slow color changes, while some details, such as textures and edge features, are often lost in the reconstruction process. Therefore, a deep dual encoding and decoding model is designed based on the continuous encoding and decoding model, as shown in the schematic diagram. Figure 6 shown.

[0039] This model adds depth on the basis of continuous encoding and decoding. It consists of three different levels of encoding and decoding and residual blocks. Different levels can obtain different image feature information. The first-level encoding and decoding blocks in this model can obtain some shallow image features such as image color distribution and color block size; the second-level continuous encoding and decoding and residual blocks extract complex image features, and the third-level continuous encoding and decoding and residual blocks extract deep abstract features of the image.

[0040] Specifically, the first-level encoding and decoding and residual block calculation process is as follows:

[0041]

[0042] in, represents the first-level encoding features, represents the encoder calculation at the first level; represents the first-level decoding features, Represents the first-level decoder calculation.

[0043] The second-level continuous encoding and decoding and residual block calculation process is as follows:

[0044]

[0045] The second level is to perform secondary continuous encoding and decoding on the image coding features based on the first level coding calculation results. 、 and Represent the second-level encoding features, encoder calculation and decoding features respectively; Represents the operation of two consecutive decoders in the second stage.

[0046] The third level is a continuous encoding and decoding and residual block calculation based on the second level encoding. The second level encoding features are further encoded three times, and then the third level encoding features are continuously decoded by the third level continuous decoder. The calculation process is as follows:

[0047]

[0048] in, represents the third-level coding feature, represents the decoding features of the third level; represents the calculation of the third-level encoder; Represents the computational process of three consecutive decoders at the third level.

[0049] In a possible embodiment of the present application, the method may further include: fusing multiple deep features to obtain fused features; performing global average pooling on the fused features to obtain compressed features; processing the compressed features through an activation function to obtain fully connected layer excitation features; and performing weight distribution processing on the fully connected layer excitation features to obtain reconfigured decoding features.

[0050] The embodiment of the present application adopts the attention residual module of deep dual encoding and decoding, and its structural diagram is as follows: Figure 7 As shown in the figure, the network dynamically allocates learning weights during training, focusing on important parts of the image, based on the difficulty of extracting image features and the amount of computation required. This is similar to the role of attention mechanisms in natural language processing. This allows the network model to extract as many rich features as possible from the image, which is crucial for image super-resolution reconstruction.

[0051] This module assigns different weights to the three levels of feature extraction based on the deep dual encoding and decoding model. Specifically, the module first assigns different weights to the three levels of decoding features ( , and ) is calculated using a 1x1 convolutional layer to extract fusion features. The calculation process is as follows:

[0052] in, represents a 1x1 convolutional layer, Represents the fused decoded features. The fused decoded features undergo attention calculations through three processes: compression, excitation, and reassignment, and then different weights are assigned based on the attention results.

[0053] Compression is to use the CNN global average pooling layer to combine the decoded features of multiple channels ( ) information is compressed into one dimension for representation, and the calculation process is as follows:

[0054] in, represents the global average pooling calculation, Represents decoded image features.

[0055] The excitation operation performs a ReLU transformation on each feature through the convolutional layer to form a fully connected layer, and then generates different weights for the three levels of decoding features through the Sigmoid activation function. The calculation process is as follows:

[0056] in, and They represent the different weights of the two 1x1 convolutional layers, δ represents the ReLU transformation (Rectified Linear Unit, ReLU), and σ represents the Sigmoid activation function.

[0057] Reassignment is the process of redistributing the weights of the three-level decoding features based on the results of the excitation operation. The calculation process is as follows:

[0058] in, Represents the decoding characteristics after reconfiguration.

[0059] Finally, this module combines the input image features with the re-matched decoding features through local residual connections to update the image features. The operation process is as follows: ,in Represents the updated image features.

[0060] In a possible embodiment of the present application, the use of a multi-level deep dual codec model to extract deep features from the original image during the sampling process to obtain initial features and depth features may include: using a deep compression network to compress and restore the original image during the sampling process to obtain deep compressed image features; and using a multi-level deep dual codec model to extract deep features from the deep compressed image features to obtain initial features and depth features.

[0061] Among them, the use of a deep compression network to compress and restore the original image during the sampling process to obtain deep compressed image features can include: performing residual processing and upsampling processing on the original image to obtain first upsampling data; performing residual processing and downsampling processing on the first upsampling data to obtain compressed deep abstract image features; performing residual processing and upsampling processing on the deep abstract image features to obtain restored abstract image features; and obtaining the deep compressed image features after multiple feature compression and restoration.

[0062] In the deep dual encoding and decoding module, this embodiment uses a lightweight deep compression network to keep the network model lightweight. The network framework diagram is as follows: Figure 8 shown.

[0063] The network mainly consists of two upsampling and downsampled twice to Composition, top layer and the lower layer Both upsampling and downsampling are model calculation units composed of residual calculation and upsampling (downsampling) calculation connected in sequence. It consists of N residual calculation models and N+1 CNN downsampling modules. The downsampling module compresses the image, and the subsequent residual calculation module extracts features from the compressed image. The compression and feature extraction of the step-by-step model are repeated N times. The final result is obtained by the CNN downsampling block operation and is used as the first downsampling. The input of the deep compression network is X, which is first up-sampled to , the first downsampling to , the outputs obtained are and :

[0064]

[0065] Image feature X undergoes the first upsampling To extract features from the residual network module, the model is first downsampled to Calculate the compressed deep abstract image features; secondly, upsample the model to The restored abstract image features are calculated, and the process of feature extraction and restoration is repeated twice, and the output is finally calculated. .

[0066]

[0067] This application uses a deep compression network in the deep dual encoding and decoding module, which not only improves the computing speed of the model, enabling the model to perform real-time super-resolution image and video reconstruction in the metaverse, but also ensures the quality of image reconstruction.

[0068] In one embodiment, the model training scenario diagram is as follows Figure 10 shown.

[0069] Because the original data type is incorrect and unsuitable for model training, the data is first preprocessed. The original dataset is cleaned, Gaussian blurred, upsampled, downsampled, and vectorized to obtain a standard dataset that the model can recognize. The standard dataset is then split into training and test sets in an 8:2 ratio. Image patches at corresponding locations in the corresponding high-resolution images are used as ground truth images for loss function calculation.

[0070] During the training process, the Adam optimizer is used to optimize the network. First, the deep dual encoder-decoder network model is iteratively trained with an initial learning rate of 0.00006, which is reduced to 0.1 after every 20 iterations. After that, the super-resolution reconstruction network is trained as a whole, and the loss function is:

[0071] in, The difference between the super-resolution result and the high-resolution image truth value Loss. The smaller the result of the loss function, the more the model is converging. When it tends to be stable and is small enough, it means that the model training is basically completed, and we can also add data for incremental training.

[0072] By continuously training the model to see if the results have converged and whether they meet the current requirements, the parameters are continuously adjusted to output the final prediction model. This application uses the test set to predict the prediction model and determines whether it meets the standards based on the final results.

[0073] like Figure 11 As shown in FIG, a schematic diagram of an image processing device provided in an embodiment of the present application is shown. Figure 11 As shown, the image processing device may include: an acquisition module 1101 , a feature extraction module 1102 and a processing module 1103 .

[0074] Among them, the acquisition module 1101 is used to obtain the original image to be processed; the feature extraction module 1102 is used to sample the original image, and use the multi-level deep dual encoding and decoding model to extract deep features of the original image during the sampling process to obtain initial features and depth features, wherein the multi-level deep dual encoding and decoding model is a model with multiple symmetrical encoding processes and decoding processes, and an attention residual module is added between adjacent encoders and adjacent decoders; the processing module 1103 is used to convolve and upsample the initial features and depth features to obtain a target image, and the resolution of the target image is higher than the resolution of the original image.

[0075] In an embodiment of the present application, first, the acquisition module 1101 acquires the original image to be processed, and then the feature extraction module 1102 samples the original image. During the sampling process, the multi-level deep dual encoding and decoding model is used to extract deep features of the original image to obtain initial features and deep features, wherein the multi-level deep dual encoding and decoding model is a model with multiple symmetrical encoding and decoding processes, and an attention residual module is added between adjacent encoders and adjacent decoders. Finally, the processing module 1103 performs convolution processing and upsampling processing on the initial features and deep features to obtain a target image, and the resolution of the target image is higher than the resolution of the original image. The embodiment of the present application uses multiple encoding and decoding to obtain more internal feature information of the image, and adds residual blocks between adjacent encoders. Deep feature extraction and recovery are performed during continuous encoding and decoding. This can alleviate the problem of information loss during continuous encoding of multiple continuous encoders and decoders and the difficulty in recovering deep features during decoding, reduce feature loss, and improve the final display effect of the image.

[0076] In one possible embodiment of the present application, the feature extraction module 1102 is used to: extract depth features of the original image using the first-level deep dual codec in the multi-level deep dual codec model during the sampling process to obtain initial features; and use the initial features to obtain depth features using multiple consecutive deep dual codecs and attention residual operations in the multi-level deep dual codec model.

[0077] In one possible implementation of the present application, the feature extraction module 1102 is used to: subject the initial features to a first level of encoding and decoding and residual block operations to obtain a first depth feature; subject the first depth features to a second level of encoding and decoding and residual block operations to obtain a second depth feature; subject the second depth features to a third level of encoding and decoding and residual block operations to obtain a third depth feature, until all depth features are obtained through multiple levels of encoding and decoding and residual block operations.

[0078] In one possible implementation of the present application, the feature extraction module 1102 is used to: fuse multiple deep features to obtain fused features; perform global average pooling on the fused features to obtain compressed features; process the compressed features through an activation function to obtain fully connected layer excitation features; and perform weight distribution on the fully connected layer excitation features to obtain reconfigured decoding features.

[0079] In one possible implementation of the present application, the processing module 1103 is used to: combine the initial features and the reconfigured decoded features through local residual connections to obtain updated features; upsample the updated features to obtain upsampled features; and perform convolution on the upsampled features to obtain a reconstructed super-resolution image, where the reconstructed super-resolution image is the target image.

[0080] In a possible embodiment of the present application, the feature extraction module 1102 is used to: compress and restore the original image using a deep compression network during the sampling process to obtain deep compressed image features; and perform deep feature extraction on the deep compressed image features using a multi-level deep dual encoding and decoding model to obtain initial features and deep features.

[0081] In one possible embodiment of the present application, the feature extraction module 1102 is used to: perform residual processing and upsampling processing on the original image to obtain first upsampling data; perform residual processing and downsampling processing on the first upsampling data to obtain compressed deep abstract image features; perform residual processing and upsampling processing on the deep abstract image features to obtain restored abstract image features; and obtain deep compressed image features after multiple feature compression and restoration.

[0082] The image processing device of this application has the following functions: Figures 1 to 10 The method embodiment shown is described in detail, so for any details not fully described in this embodiment, please refer to the relevant descriptions in the aforementioned embodiments and will not be repeated here.

[0083] like Figure 12 As shown, an embodiment of the present application also provides an electronic device 1200, including a processor 1201, a memory 1202, and a program or instruction stored in the memory 1202 and executable on the processor 1201. When the program or instruction is executed by the processor 1201, each process of the above-mentioned image processing method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.

[0084] Optionally, an embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program implements the various processes of the above-described image processing method embodiment and can achieve the same technical effects. To avoid repetition, the details are not described here. The computer-readable storage medium may be, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0085] Optionally, an embodiment of the present application also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the various processes of the above-mentioned image processing method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, they will not be repeated here.

[0086] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or system. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or system comprising the element.

[0087] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a more preferred embodiment. Based on this understanding, the technical solution of this application, or the part that contributes to the existing technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of this application.

[0088] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

Claims

1. An image processing method, characterized in that: include: Get the original image to be processed; Sampling the original image, and extracting depth features from the original image using a multi-level deep dual encoding and decoding model during the sampling process to obtain initial features and depth features, wherein the multi-level deep dual encoding and decoding model is a model with symmetrical encoding and decoding processes, and an attention residual module is added between adjacent encoders and between adjacent decoders; The initial features and the depth features are convolved and up-sampled to obtain a target image, where the resolution of the target image is higher than that of the original image.

2. The method according to claim 1, characterized in that The method of extracting depth features from the original image using a multi-level depth dual encoding / decoding model during the sampling process to obtain initial features and depth features includes: During the sampling process, the first level deep dual codec in the multi-level deep dual codec model is used to extract deep features from the original image to obtain initial features; The initial features are used to obtain deep features by using multiple consecutive deep dual encoding and decoding and attention residual operations in a multi-level deep dual encoding and decoding model.

3. The method according to claim 2, characterized in that The initial features are obtained by using a plurality of continuous deep dual encoding and decoding and attention residual operations in a multi-level deep dual encoding and decoding model to obtain deep features, including: The initial features are subjected to first-level encoding and decoding and residual block operations to obtain first depth features; The first depth feature is subjected to a second-level encoding and decoding and residual block operation to obtain a second depth feature; The second depth feature is subjected to a third level of encoding and decoding and residual block operation to obtain a third depth feature, and all depth features are obtained after multiple levels of encoding and decoding and residual block operation.

4. The method according to claim 3, characterized in that The method further comprises: Fuse multiple deep features to obtain fused features; Performing global average pooling processing on the fused features to obtain compressed features; The compression feature is processed by an activation function to obtain a fully connected layer excitation feature; The fully connected layer excitation features are weighted to obtain re-weighted decoding features.

5. The method according to claim 4, characterized in that The convolution processing and up-sampling processing of the initial features and the depth features to obtain a target image includes: Combining the initial features and the reconfigured decoding features through a local residual connection to obtain an updated feature; Performing upsampling processing on the updated features to obtain upsampled features; Convolution processing is performed on the up-sampled features to obtain a reconstructed super-resolution image, where the reconstructed super-resolution image is the target image.

6. The method according to claim 1, wherein The method of extracting depth features from the original image using a multi-level depth dual encoding / decoding model during the sampling process to obtain initial features and depth features includes: During the sampling process, the original image is compressed and restored using a deep compression network to obtain deep compressed image features; A multi-level deep dual encoding and decoding model is used to extract deep features from the deep compressed image features to obtain initial features and deep features.

7. The method according to claim 6, characterized in that The method of compressing and restoring the original image using a deep compression network during the sampling process to obtain deep compressed image features includes: Performing residual processing and upsampling processing on the original image to obtain first upsampling data; performing residual processing and downsampling processing on the first upsampled data to obtain compressed deep abstract image features; Performing residual processing and upsampling processing on the deep abstract image features to obtain restored abstract image features; After multiple feature compression and restoration, the deep compressed image features are obtained.

8. An electronic device, characterized in that: The method comprises a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method according to claims 1 to 7.

10. A computer program product, characterized in that The computer program product comprises a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, which implement the steps of the method according to claims 1 to 7 when executed by a computer.