Image processing method and device, electronic equipment, and storage medium
By processing image feature information through multiple coding modules and densely connected networks, the problems of high computational load and poor reconstruction effect in existing technologies are solved, and efficient and accurate image reconstruction is achieved.
Patent Information
- Application Number
- CN202210789728.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-06
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2042-07-06
AI Technical Summary
Existing technologies struggle to fully and adequately extract feature information during image reconstruction, resulting in high computational costs and poor performance, leading to inaccurate image reconstruction.
Multiple coding modules are used to process image features, high-frequency information is determined through residual connections, and dense connection networks are used for image reconstruction, which reduces the amount of computation and improves computational efficiency and accuracy.
It improves the accuracy and precision of image reconstruction, reduces computational load, expands the application scope of image reconstruction, and is suitable for a variety of smart devices.
Smart Images

Figure CN115063319B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of image processing, and in particular, to an image processing method, an image processing apparatus, an electronic device, and a computer readable storage medium. BACKGROUND
[0002] In the image processing process, image reconstruction is a common way to improve image quality. For example, a low-resolution image can be converted into a high-resolution image through image super-resolution reconstruction to enhance the detail information.
[0003] In the related art, when performing image reconstruction, it is difficult to fully and sufficiently extract and process the features in the image, which results in a large amount of calculation and affects the effect of image reconstruction.
[0004] It should be noted that the information disclosed in the above background section is only used to strengthen the understanding of the background of the present disclosure, and therefore can include information that does not constitute prior art known to those of ordinary skill in the art. SUMMARY
[0005] The purpose of the present disclosure is to provide an image processing method and apparatus, an electronic device, and a storage medium, thereby at least partially overcoming the problem of large amount of calculation and poor image reconstruction effect caused by the limitations and defects of the related art.
[0006] Other characteristics and advantages of the present disclosure will become apparent from the following detailed description, or will be learned by practice of the present disclosure.
[0007] According to a first aspect of the present disclosure, an image processing method is provided, comprising: obtaining a to-be-processed image, and performing feature extraction on the to-be-processed image to obtain feature information of the to-be-processed image; performing feature processing on the feature information through a plurality of encoding modules respectively to determine intermediate information, and determining high-frequency information output by each encoding module based on a residual connection of the feature information and the intermediate information; performing image reconstruction on the to-be-processed image according to the high-frequency information to obtain a reconstructed image of the to-be-processed image.
[0008] According to a second aspect of the present disclosure, an image processing apparatus is provided, comprising: a feature extraction module configured to obtain a to-be-processed image, and perform feature extraction on the to-be-processed image to obtain feature information of the to-be-processed image; a high-frequency information determination module configured to perform feature processing on the feature information through a plurality of encoding modules respectively to determine intermediate information, and determine high-frequency information output by each encoding module based on a residual connection of the feature information and the intermediate information; and an image reconstruction module configured to perform image reconstruction on the to-be-processed image according to the high-frequency information to obtain a reconstructed image of the to-be-processed image.
[0009] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform the image processing method of the first aspect and possible implementation manners thereof by executing the executable instructions.
[0010] According to a fourth aspect of the present disclosure, a computer readable storage medium is provided, having stored thereon a computer program, which, when executed by a processor, implements the image processing method of the first aspect and possible implementation manners thereof.
[0011] In the image processing method, the image processing device, the electronic device and the computer readable storage medium provided in the embodiments of the present disclosure, on one hand, the intermediate information is obtained by respectively processing the feature information of the to-be-processed image through the plurality of encoding modules, and then the high-frequency information of each encoding module is determined according to the residual connection of the feature information and the intermediate information, which can transmit the low-frequency information while focusing on calculating the high-frequency information output by each encoding module, thereby improving the accuracy of calculating the high-frequency information of each encoding module, and also improving the accuracy and comprehensiveness of the extracted high-frequency information of the to-be-processed image, enhancing the accuracy of the high-frequency detail information of the image, and further improving the precision and accuracy of image reconstruction. On the other hand, the high-frequency information of each encoding module is determined according to the residual connection of the feature information and the intermediate information, which can utilize the pre-learned feature information, reduce the amount of calculation, improve the calculation efficiency, avoid the limitation caused by large amount of calculation, and realize efficient landing, thereby increasing the application range of image reconstruction.
[0012] It should be understood that the general description above and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0013] The accompanying drawings, which are incorporated into and form part of the specification, illustrate one embodiment consistent with the present disclosure and, together with the specification, serve to explain the principles of the present disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without creative labor based on these drawings.
[0014] Figure 1 A schematic diagram of a system architecture to which the image processing method of the embodiments of the present disclosure can be applied is shown.
[0015] Figure 2 A schematic diagram of an image processing method in the embodiments of the present disclosure is shown.
[0016] Figure 3 A schematic diagram of image blocking in the embodiments of the present disclosure is shown.
[0017] Figure 4 A structural schematic diagram of a visual reconstruction model in an embodiment of the present disclosure is shown.
[0018] Figure 5 A flowchart of determining high frequency information of each coding module in an embodiment of the present disclosure is shown.
[0019] Figure 6 A structural schematic diagram of each coding module in an embodiment of the present disclosure is shown.
[0020] Figure 7 A schematic diagram of multiple coding modules in an embodiment of the present disclosure is shown.
[0021] Figure 8 A flowchart of obtaining a target image in an embodiment of the present disclosure is shown.
[0022] Figure 9 A flowchart of reconstructing a to-be-processed image in an embodiment of the present disclosure is shown.
[0023] Figure 10 A specific flowchart of image reconstruction in an embodiment of the present disclosure is shown.
[0024] Figure 11 A block diagram of an image processing apparatus in an embodiment of the present disclosure is shown.
[0025] Figure 12 A block diagram of an electronic device in an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0026] Example implementations will now be described more fully with reference to the accompanying drawings. Example implementations may, however, be implemented in many different forms and should not be construed as limited to the implementations set forth herein; rather, these implementations are provided so that this disclosure will be thorough and complete, and will fully convey the scope of example implementations to those skilled in the art. The described features, structures, or characteristics can be combined in one or more implementations. In the following description, numerous specific details are provided to give a thorough understanding of example implementations. One skilled in relevant art will recognize, however, that the implementations can be practiced without one or more of the specific details, or with other methods, components, materials, and so forth. In other instances, well-known structures have not been described in detail to avoid obscuring the understanding of this description. It will be appreciated that the scope of example implementations of the present disclosure includes both the particular examples discussed and ones that are similar in spirit but not necessarily in details of the disclosure.
[0027] In addition, the accompanying drawings are only schematic and are non-limiting illustrative of the present disclosure. Identical reference signs denote identical or similar parts throughout the figures. Some of the blocks in the drawings are functional entities that may not necessarily have a corresponding physical or logical entity in an implantation. These functional entities may be implemented by software or in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0028] In the image super-resolution reconstruction process of the related technology, an IPT (Image Processing Transformer) first transforms an image into a feature map through a head structure, then performs cutting and flattening operations on the feature map, inputs the deformed feature vector into a deep self-attention transformation network Transformer for processing to obtain an output feature with the same dimension; then, after reshaping and splicing operations, the feature map is restored to a feature map with the same dimension as the input; and then the feature map is input into a tail structure to restore the required image. The head structure and the tail structure are used for feature extraction and restoration of the image, and the Transformer model is used for feature processing. The model is first trained on a large data set, and then adjusted for individual tasks. In the above scheme, the IPT is based on the Transformer model proposed for natural language, without model structure adjustment for image tasks. Instead, a two-dimensional image is converted into a one-dimensional vector for operation, which cannot utilize the spatial information of the image. The entire image is directly input into the Transformer model, which has a large amount of calculation and cannot obtain local features and local correlations, cannot take into account all the information, has certain limitations, and is difficult to land and apply. The Transformer model is directly migrated from the natural language task to the image task without targeted optimization, so the effect is poor. Moreover, the amount of calculation is large, resulting in low calculation efficiency.
[0029] To solve the technical problems in the related technology, an image processing method provided in the embodiments of the present disclosure can be applied to the application scenario of obtaining a high-resolution image by performing image super-resolution reconstruction on a low-resolution image.
[0030] Figure 1 A schematic diagram of an application scenario to which the image processing method and device of the present disclosure can be applied is shown.
[0031] As Figure 1As shown, the client 101 can be a smart device with image processing function, for example, can be a smart phone, a computer, a tablet computer, a smart speaker, a smart watch, a vehicle-mounted device, a wearable device, etc. The image to be processed can be any type of low-resolution image in various scenes, can be an image taken, or an image obtained from a network or other terminal, and the type of image is not limited.
[0032] In the embodiment of the present disclosure, the client 101 obtains the image to be processed 102, and extracts features of the image to be processed to obtain feature information of the image to be processed. Further, the feature information is processed by a plurality of encoding modules to determine intermediate information, and the intermediate information and the feature information are processed by a residual network to obtain high-frequency information of each encoding module, and the high-frequency information output by each encoding module is densely connected to determine the high-frequency information of each image to be processed. Further, based on the high-frequency information of the image to be processed obtained by dense connection, each image block represented by the image to be processed is reconstructed to obtain a reconstructed image 103 corresponding to the image to be processed, and then a target image 104 of the original image to which the image to be processed belongs is generated.
[0033] It should be noted that the image processing method provided by the embodiment of the present disclosure can be executed by the client 101. Accordingly, the image processing method can be set in the client 101 by a program or the like. The image processing method provided by the embodiment of the present disclosure can also be executed by a server, which can be a background system providing image processing related services in the embodiment of the present disclosure, and can include a portable computer, a desktop computer, a smart phone, etc. One electronic device with computing function or a cluster formed by multiple electronic devices. In the embodiment of the present disclosure, the image processing method is executed by the client as an example.
[0034] Next, with reference to Figure 2 The image processing method in the embodiment of the present disclosure is described in detail.
[0035] In step S210, an image to be processed is obtained, and features of the image to be processed are extracted to obtain feature information of the image to be processed.
[0036] In the embodiments of the present disclosure, the to-be-processed image can be a complete image or a partial image in the complete image, i.e., an image block. Here, taking the to-be-processed image as an image block as an example for description. When the to-be-processed image is an image block, the to-be-processed image can be obtained by performing block processing on an original image representing a complete image. The original image can be an image of a first resolution, and the original image can be a complete image or a defective image with partial loss. The original image can be various types of images, which can be a low-resolution image captured by a camera or a low-resolution image obtained from a network or other storage devices.
[0037] In some embodiments, after obtaining the original image, the original image can be subjected to block processing to obtain image blocks corresponding to the original image, and the image blocks are taken as to-be-processed images. Each image block has the same size, and each image block has the same number of channels as the original image. Based on this, the number of image blocks represented by the to-be-processed image can be determined according to the size of the original image and the size of each to-be-processed image. It should be noted that for original images of different sizes, the sizes of the divided image blocks can be the same or different, which is limited according to actual needs.
[0038] Reference Figure 3 As shown in FIG. 3, the original image 301 can be subjected to block processing to obtain a plurality of to-be-processed images, i.e., image blocks 302. The vector size of the original image 301 can be represented as H×W×C, where H represents the image height of the original image, W represents the image width of the original image, and C represents the image channel number of the original image. Based on this, assuming that the length and width of the to-be-processed image 302 are (P, P), the number of to-be-processed images obtained by block processing of the original image can be represented as: N=H×W / P 2 . When each to-be-processed image is expanded into a one-dimensional vector, the vector size of each to-be-processed image can be represented as P×P×C, and the vector size of the original image can be transformed as H×W×C=N×(P×P×C). In the embodiments of the present disclosure, a plurality of to-be-processed images are obtained by block processing of the original image, and then image processing is performed based on the plurality of to-be-processed images, so that local information and local correlation can be obtained, the completeness and comprehensiveness of the information are improved, and the problem of large amount of calculation caused by inputting the original image as a whole into the model is avoided, the amount of calculation is reduced, and the calculation efficiency is improved.
[0039] After obtaining the image patches representing the image to be processed, super-resolution reconstruction can be performed based on each partitioned image. For example, a visual reconstruction model can be used to perform image super-resolution reconstruction on each image to obtain the target image corresponding to the original image to which the image to be processed belongs. The visual reconstruction model can be, for example, DVSR (Dense ViT Super-Resolution, a super-resolution model based on ViT and dense connections). The DVSR model can be a model structure for image super-resolution tasks, and it can be implemented by structurally adjusting the ViT (Vision Transformer) model. For example, the position encoding and classification modules in the ViT model can be removed, and multiple encoding modules and dense connection networks can be added to the ViT model to fuse the ViT model, dense connections, and cascaded residuals, ultimately obtaining a visual reconstruction model that includes a feature extraction network, multiple encoding modules, and a dense connection network, forming a feature processing network and an image reconstruction network, thereby better recovering the high-frequency information of the image.
[0040] Figure 4 The diagram illustrates the structure of the visual reconstruction model. (See reference) Figure 4 As shown, the visual reconstruction model 400 mainly includes the following modules: a feature extraction network 401, a feature processing network 402, and an image reconstruction network 403. The feature extraction network extracts features from each image (image patch) to obtain feature information. The feature processing network includes multiple encoding modules and a densely connected network. Each encoding module includes multiple residual networks. The feature processing network processes the feature information of each image to determine the high-frequency information output by each encoding module. The image reconstruction network obtains the high-frequency information of the image to be processed based on the high-frequency information of each encoding module, and performs image super-resolution reconstruction based on the high-frequency information of the image to be processed, thereby obtaining a high-resolution image of the image to be processed as the reconstructed image.
[0041] based on Figure 4 The visual reconstruction model in the text first uses a feature extraction network to extract features from the image pairs to be processed, thereby obtaining the feature information of the images to be processed. In some embodiments, the feature extraction network can consist of a single convolutional layer. Therefore, feature extraction can be achieved by performing a convolution operation on each image to be processed through a single convolutional layer, thus obtaining the feature information corresponding to the image to be processed. This feature information can be represented as a two-dimensional vector.
[0042] Continue to refer to Figure 2As shown, in step S220, the feature information is processed by multiple encoding modules to determine intermediate information, and the high-frequency information output by each encoding module is determined based on the feature information and the residual connection of the intermediate information.
[0043] In this embodiment, the encoding module can be a module included in the feature processing network for acquiring high-frequency information. The feature processing network can be used to process the received feature information to obtain the information required by the target task. The features to be extracted differ depending on the target task. In this embodiment, since the target task is image super-resolution, which requires high-frequency information from the image, the features to be extracted are the high-frequency information of the image. Based on this, the feature processing network can be considered as being used to acquire the high-frequency information of the image to be processed. High-frequency information refers to the difference between a clear image and a blurred image obtained by directly enlarging the image.
[0044] Continue to refer to Figure 4 The model structure described above uses a feature processing network to process the feature information of the image to be processed, thereby obtaining the high-frequency information of each encoding module. Furthermore, it generates the high-frequency information of each image to be processed based on the high-frequency information of each encoding module. The feature processing network can be a Vision Transformer module, which can be composed of multiple stacked encoding modules, and these multiple encoding modules can be Transformer blocks.
[0045] Figure 5 The flowchart illustrating the determination of high-frequency information for each encoding module is shown in the figure. (Refer to...) Figure 5 As shown, the main steps include:
[0046] In step S510, the feature information is processed by the residual network in each encoding module to determine the intermediate information.
[0047] In this embodiment of the disclosure, each encoding module may include multiple residual networks. It should be noted that the number of residual networks in each encoding module can be the same, and the type of network layers in the residual networks within each encoding module can also be the same, but the network parameters of the network layers in each residual network can be different. The principle of a residual network is as follows: the input data of the next layer is the output data of the previous layer, plus the input data of the previous layer, to achieve residual connections.
[0048] Figure 6 The diagram illustrates the structure of each encoding module. (See reference...) Figure 6As shown in the middle, each encoding block is composed of a normalization layer Norm, a multi-head attention mechanism layer Multi-Head Attention and a multi-layer perception MLP layer, and there is a short cut connection between layers, thereby forming a residual network. A plurality of residual networks can be included in each encoding block, for example, the plurality of residual networks can include a first residual network and a second residual network. Exemplarily, the first residual network includes a normalization layer and a multi-head attention mechanism layer, and the second residual network includes a normalization layer and a multi-layer perception. Residual processing of the feature information by the first residual network and the second residual network can enable each encoding block to generate intermediate information by processing the feature information, and the intermediate information can be used to generate high-frequency information of the image to be processed.
[0049] In some embodiments, the feature information can be processed by the first residual network of the encoding block to generate a residual result, and the residual result and the feature information can be processed by the second residual network to determine the intermediate information. The intermediate information obtained by each encoding block can be the same or different, which is determined according to the network parameters of the residual networks included in the encoding block. In each encoding block, the input of each residual network can be determined by the feature information, for example, the first residual network takes the feature information as input, and the second residual network takes the output of the last residual network (the first residual network) adjacent to it and the input of the first residual network as input. The input of the second residual network is mapped to a residual by all network layers included in the residual network, and the sum of the input and the residual is taken as the output of the second residual network, i.e., the intermediate information. Referring to Figure 6 As shown in the middle, in the encoding block, the feature information is normalized by the first residual network, and the normalized result is convoluted using the multi-head attention mechanism to map the feature information to obtain a residual result. Further, the feature information and the residual result are normalized by the second residual network, and the normalized result is convoluted using the multi-layer perception to map the feature information and the residual result to obtain the intermediate information.
[0050] In step S520, the feature information is connected to the intermediate information to perform residual processing on the feature information and the intermediate information, and determine the high-frequency information output by each encoding block.
[0051] In this embodiment of the disclosure, after obtaining intermediate information, residual processing can be performed between feature information and intermediate information to determine the high-frequency information of each encoding module. In some embodiments, residual processing can be performed between the input of the encoding module and the output of the last residual network in the residual network contained in each encoding module to obtain the high-frequency information output by each encoding module. Based on this, since the inputs of the residual networks are directly connected to the subsequent network layers, the subsequent network layers can directly learn the residuals. Furthermore, since the feature information is directly connected to the output of the last residual network of each encoding module, low-frequency information can flow through short connections within the same encoding module, allowing the encoding module to focus more on learning the high-frequency information of the image to be processed, avoiding the influence of other information, thereby improving the accuracy and comprehensiveness of the high-frequency information.
[0052] It should be noted that, since the residual networks in the encoding modules may be different, the high-frequency information of the image to be processed output by each encoding module may be the same or different, depending on the specific model.
[0053] Building upon this foundation, after obtaining the high-frequency information of each encoding module, the feature information input to the first encoding module can be connected to the high-frequency information output by any of the multiple encoding modules. This creates residual connections between the feature information and the high-frequency information of multiple encoding modules, achieving long connections between the input and output of different encoding modules. The number of residual connections between different encoding modules can be multiple, i.e., the number of long connections can be multiple, determined according to actual needs. For example, residual connections can be made between the input feature information and the high-frequency information output by the third encoding module, between the input feature information and the high-frequency information output by the (L-3)th encoding module, and between the input feature information and the high-frequency information output by the Lth encoding module. By using residual networks to directly connect the input feature information to the output of any encoding module to achieve long connections across modules, shallow features can be passed to deep networks, taking into account all feature information, improving information completeness, allowing the network to learn better, and avoiding overfitting. Furthermore, the entire network can obtain high-frequency information represented by the difference between the input and output without needing other information, reducing the complexity.
[0054] Continue to refer to Figure 2 As shown, in step S230, the image to be processed is reconstructed based on the high-frequency information to obtain the reconstructed image of the image to be processed.
[0055] In the embodiments of the present disclosure, after obtaining the high-frequency information of each encoding module, image reconstruction can be performed according to the high-frequency information of each encoding module. Since each to-be-processed image can be calculated by multiple encoding modules, the high-frequency information of each encoding module can be processed to obtain the high-frequency information of the to-be-processed image. For example, the high-frequency information of each encoding module is spliced to obtain the high-frequency information of the to-be-processed image, and the to-be-processed image is reconstructed according to the high-frequency information of the to-be-processed image to obtain the reconstructed image of the to-be-processed image.
[0056] In some embodiments, the high-frequency information of each encoding module can be spliced in the channel dimension. For example, the high-frequency information of each encoding module can be spliced in the channel dimension by dense connection. The dense connection is used to connect the high-frequency information output by different encoding modules in the channel dimension, so that the output of the encoding module in front can be directly transmitted into the network layer of the encoding module behind. For example, the encoding module A is connected with the encoding module B, and through the dense connection, the encoding module A can be directly connected with all the encoding modules and network layers behind the encoding module B.
[0057] Since the visual reconstruction model includes multiple encoding modules, the high-frequency information output by each encoding module in the visual reconstruction model can be spliced in the channel dimension by dense connection. For example, for each encoding module in the multiple encoding modules, the high-frequency information output by each encoding module can be connected with the high-frequency information output by the remaining encoding modules except for each encoding module itself, so that the high-frequency information output by the first encoding module can be transmitted to all the encoding modules, and the high-frequency information output by all the encoding modules can be spliced.
[0058] For example, if there are L encoding modules in the visual reconstruction model, the high-frequency information of the first encoding module is y1, the high-frequency information of the second encoding module is y2, the high-frequency information of the third encoding module is y3, and the high-frequency information of the Lth encoding module is yL. When performing dense connection, y1 can be connected with y2, y1 can be connected with y3, y1 can be connected with the high-frequency information of all the encoding modules behind the first encoding module, and so on until y1 is connected with yL; at the same time, y2 can be connected with y3, y2 can be connected with the high-frequency information of all the encoding modules behind the second encoding module, and so on until y2 is connected with yL; and y3 can be connected with the high-frequency information of all the encoding modules behind the third encoding module, and so on until y3 is connected with yL. The above steps are repeated until the high-frequency information of all the encoding modules is connected with the high-frequency information of all the encoding modules behind the encoding module, so as to realize the dense connection between the high-frequency information output by each encoding module.
[0059] In some embodiments, when the high-frequency information of each encoding module is spliced in the channel dimension, the high-frequency information of all encoding modules can be connected in parallel in the same dimension to obtain the high-frequency information of the image to be processed. By splicing the features in the channel dimension through dense connection, the calculated feature information will be retained to enter all subsequent network layers to reuse the feature information of the network, retain the shallow layer features to the deep network for further learning, which can ensure the propagation of the gradient, and at the same time, the low-dimensional features and high-dimensional features can be used for joint calculation, and accurate high-frequency information can be obtained through a network with fewer layers. This connection method makes the effectiveness of feature and gradient transmission, and simplifies the operation difficulty.
[0060] It should be noted that the multiple encoding modules of the visual reconstruction model can be connected in series, that is, the input of the next encoding module is determined according to the output of the previous encoding module. Figure 7 The schematic diagram of the multiple encoding modules is shown in FIG. 1, which is described with reference to Figure 7 In FIG. 1, the input (feature information) is transmitted to the first encoding module, and the feature information is processed by the multiple residual networks in the first encoding module to generate the high-frequency information y1 of the first encoding module; the input of the second encoding module is determined according to the output of the first encoding module, and the input is processed by the multiple residual networks in the first encoding module to generate the high-frequency information y2 of the second encoding module; the above steps are repeated until the high-frequency information yL of the Lth encoding module is obtained. Further, each of the high-frequency information y1 to yL of each encoding module is connected through dense connection. And the input of the first encoding module and the high-frequency information output by any module of the L encoding modules are long-connected through the cascaded residual network, for example, the input of the first encoding module and the high-frequency information output by the third encoding module are long-connected, and the input of the first encoding module and the high-frequency information output by the Lth encoding module are long-connected. Through multiple long connections and dense connections, the cross-module connection between the multiple encoding modules is realized, and the entire feature processing process is completed.
[0061] Through the dense connection and residual connection in FIG. 1, the shallow features are transmitted to the deep features to determine the high-frequency information, so that more accurate high-frequency information can be obtained for image reconstruction, and overfitting can be avoided, and the accuracy and reliability are improved. Figure 7
[0062] Further, the image reconstruction can be performed on the to-be-processed image based on the high-frequency information of the to-be-processed image obtained by the splicing, to obtain a reconstructed image of the to-be-processed image. The reconstructed image can include the high-frequency information corresponding to each image block represented by the to-be-processed image. When performing super-resolution reconstruction, the super-resolution reconstruction can be first performed on each image block represented by the to-be-processed image, and then the reconstructed images of each image block can be fused to obtain the reconstructed result of the original image to which the to-be-processed image belongs, i.e., the target image.
[0063] The super-resolution reconstruction refers to reconstructing a low-resolution image into a high-resolution image, so as to recover the detailed part in the image. Based on this, the reconstructed image of the to-be-processed image can be a second-resolution image, and the second resolution is greater than the first resolution.
[0064] Referring to the model structure shown in Figure 3 Based on the model structure shown in, the image reconstruction can be performed on the to-be-processed image by an image reconstruction network based on the high-frequency information of the to-be-processed image, to obtain a reconstructed image corresponding to the to-be-processed image. The reconstructed image can be a part of a high-resolution image corresponding to the original image. In some embodiments, the image reconstruction network can include multiple convolution layers, activation layers, and pixelshuffle layers. The multiple convolution layers are used to restore the high-frequency information into an image; the activation layers are used to perform an activation operation, and the activation function can be RELU; and the pixelshuffle layers are used to perform spatial conversion, such as conversion from depth to space. Based on this, the convolution operation can be performed on the high-frequency information of the to-be-processed image to restore the high-frequency information of the to-be-processed image into a reference image, and then the spatial conversion is performed on the reference image by the pixelshuffle layer, to convert the reference image from depth to space, so as to obtain the reconstructed result of each to-be-processed image as the reconstructed image of the to-be-processed image.
[0065] The steps S210 to S230 are repeatedly performed, to extract the feature information of each to-be-processed image, further perform the feature processing on the feature information to determine the intermediate information, connect the residual of the feature information and the intermediate information to obtain the high-frequency information of each encoding module for processing each to-be-processed image, and reconstruct each to-be-processed image according to the high-frequency information output by each encoding module to obtain the reconstructed image of the to-be-processed image, so as to obtain the reconstructed images of all image blocks represented by all to-be-processed images.
[0066] After obtaining the reconstructed image of each to-be-processed image, the to-be-processed image and the reconstructed image of the to-be-processed image can be fused to obtain the target image of the original image to which the to-be-processed image belongs. Figure 8 The flowchart for obtaining the target image is shown in, and referring to Figure 8 The flowchart for obtaining the target image is shown in, and referring to
[0067] In step S810, the to-be-processed image is fused with the reconstructed image of the to-be-processed image to obtain a fused image of the to-be-processed image.
[0068] In step S820, the fused image is fused according to the arrangement order of the to-be-processed image to generate a target image of an original image to which the to-be-processed image belongs.
[0069] In the embodiments of the present disclosure, the to-be-processed image can be up-sampled, and the up-sampling result can be fused with the reconstructed image of the to-be-processed image to obtain a fused image of the to-be-processed image. The fused image herein refers to an image block of a second resolution corresponding to the to-be-processed image, which is generated by super-resolution reconstruction of the to-be-processed image of a first resolution. The to-be-processed image can be up-sampled by a preset multiple, for example, 2 times or 4 times, and the like, which is not specially limited herein.
[0070] After obtaining the fused images of all the to-be-processed images, a high-resolution image of each image block is obtained, and the fused images of all the to-be-processed images can be spliced and fused according to the arrangement order to obtain a target image of an original image to which the to-be-processed image belongs. The original image refers to a complete image corresponding to the image block represented by the to-be-processed image, and the target image can be an image of a second resolution corresponding to an original image of a first resolution. In some embodiments, the fused images of a plurality of to-be-processed images can be fused according to the arrangement order of the to-be-processed image to obtain a second-resolution image as a target image of the overall image. The arrangement order can be the arrangement order of the image block, which is determined according to the image cutting manner. It should be noted that the target image and the to-be-processed image contain basically the same content, only the resolution is different.
[0071] Figure 9 FIG. 1 schematically shows a whole flowchart of image processing, which is described in detail with reference to Figure 9As shown in the figure, the visual reconstruction model mainly includes a feature extraction network 901, a feature processing network 902, and an image reconstruction network 903. Among them, the feature processing network contains a residual network, and the image reconstruction network contains a dense connection network. Based on this, the to-be-processed image 904 can be input into the visual reconstruction model, and first, the feature extraction network 901 is used to extract the features of the to-be-processed image to obtain its feature information; then the feature information is input into the feature processing network 902, and the residual connection of the feature information and the intermediate information of each encoding module in the visual reconstruction model is determined through the residual network in the feature processing network. The high-frequency information output by each encoding module is determined, and the high-frequency information output by all encoding modules is spliced to obtain the high-frequency information of the to-be-processed image through the dense connection network. Further, the high-frequency information of the to-be-processed image can be input into the image reconstruction network 903, and the image reconstruction network is used to reconstruct the to-be-processed image to obtain the reconstructed image 9041.
[0072] At the same time, each to-be-processed image can be reconstructed through the feature extraction network 901, the feature processing network 902, and the image reconstruction network 903 respectively to obtain the reconstructed images of all to-be-processed images. Further, the reconstructed images of each to-be-processed image can be fused to obtain the target image of the original image to which the to-be-processed image belongs.
[0073] Figure 10 The specific flowchart of image reconstruction is shown in the figure, and the specific flowchart of image reconstruction is shown in the figure. Figure 10 As shown in the figure, the following steps are mainly included:
[0074] In step S1001, the original image is divided into blocks to obtain a plurality of to-be-processed images. The to-be-processed image is an image block.
[0075] In step S1002, the to-be-processed image is input as an input into the feature extraction network to obtain the feature information corresponding to the to-be-processed image.
[0076] In step S1003, each to-be-processed image is input into the feature processing network respectively to obtain the high-frequency information output by each encoding module, and the high-frequency information of each encoding module is spliced to obtain the high-frequency information of each to-be-processed image.
[0077] For example, the feature processing network can have a plurality of encoding modules, and each encoding module can contain a plurality of residual networks, such as a first residual network and a second residual network. For each encoding module, the feature information is processed by the first residual network to generate a residual result, and the residual result and the feature information are processed by the second residual network to determine the intermediate information. And the feature information and the intermediate information are connected in residual to obtain the high-frequency information.
[0078] The above steps are repeatedly performed to obtain high-frequency information output by each encoding module through the residual network in each encoding module. Further, the input of the first encoding module is connected in residual with the high-frequency information of any encoding module, and the high-frequency information of each encoding module is spliced through dense connection, so as to obtain spliced high-frequency information as high-frequency information of the to-be-processed image.
[0079] In step S1004, the high-frequency information of the to-be-processed image is input into the image reconstruction network to obtain a reconstructed image corresponding to the to-be-processed image. Illustratively, the image reconstruction is performed through convolution operation and conversion operation.
[0080] In step S1005, a target image of the original image is generated according to the reconstructed image of the to-be-processed image. Illustratively, the reconstructed images of each to-be-processed image are fused in the arrangement order to obtain a high-definition image of the original image.
[0081] In the embodiments of the present disclosure, the cascade residual network and the dense connection network are added on the basis of the ViT model, the accuracy of the ViT model in processing image features is utilized, the cascade residual effectively ensures the flow of low-frequency information between the residual blocks, so that the model uses the low-frequency information of the pre-learning to enable the visual reconstruction model to focus on learning high-frequency information, improve the accuracy of the determined high-frequency information, and better restore the high-resolution image and improve the image super-resolution quality. At the same time, the visual reconstruction model with the dense connection network and the cascade residual network reduces the amount of parameters required to determine the high-frequency information, reduces the calculation amount, reduces the calculation load, and improves the calculation speed compared with the related art, and is suitable for landing scenarios. At the same time, the gradient propagation is ensured, so that the model is easier to train. The ViT is adopted to obtain local information and reduce the calculation amount, which also conforms to the practice of the traditional image super-resolution model, avoids the appearance of artificial lines in the high-resolution image, and improves the image quality.
[0082] The technical solutions in the embodiments of the present disclosure can also be used in other high-definition image editing algorithms, such as fog removal, reflection removal, image editing on a low-resolution image, and then image reconstruction. It can also be applied to some high-resolution cameras to realize obstacle removal and impurity removal in high-resolution photos of mobile phones.
[0083] The embodiments of the present disclosure provide an image processing apparatus, which, as shown in Figure 11 The image processing apparatus 1100 can include:
[0084] The feature extraction module 1101 is configured to acquire a to-be-processed image and perform feature extraction on the to-be-processed image to obtain feature information of the to-be-processed image.
[0085] The high-frequency information determination module 1102 is configured to determine intermediate information by performing feature processing on the feature information by using a plurality of encoding modules respectively, and determine high-frequency information output by each encoding module based on the feature information and a residual connection of the intermediate information.
[0086] The image reconstruction module 1103 is configured to perform image reconstruction on the to-be-processed image according to the high-frequency information, and obtain a reconstructed image of the to-be-processed image.
[0087] In an example embodiment of the present disclosure, the high-frequency information determination module comprises: an intermediate information calculation module configured to determine the intermediate information by performing feature processing on the feature information by using a residual network in each encoding module; and a high-frequency calculation module configured to connect the feature information to the intermediate information to perform residual processing on the feature information and the intermediate information, and determine high-frequency information output by each encoding module.
[0088] In an example embodiment of the present disclosure, the residual network comprises a first residual network and a second residual network, and the intermediate information calculation module comprises: a residual processing module configured to perform residual processing on the feature information by using the first residual network to generate a residual result, and perform residual processing on the residual result and the feature information by using the second residual network to determine the intermediate information.
[0089] In an example embodiment of the present disclosure, the device further comprises: an information combination module configured to connect the feature information to high-frequency information output by any one of the plurality of encoding modules to combine the feature information and the high-frequency information output by the any one of the plurality of encoding modules.
[0090] In an example embodiment of the present disclosure, the image reconstruction module comprises: a splicing module configured to splice high-frequency information output by each of the plurality of encoding modules to determine high-frequency information of the to-be-processed image, and perform image reconstruction on the to-be-processed image according to the high-frequency information of the to-be-processed image to obtain the reconstructed image of the to-be-processed image.
[0091] In an example embodiment of the present disclosure, the splicing module comprises: a splicing control module configured to connect, in parallel, high-frequency information output by each of the plurality of encoding modules and high-frequency information output by remaining encoding modules except the each of the plurality of encoding modules in the same channel.
[0092] In an example embodiment of the present disclosure, the image reconstruction module comprises: an image recovery module configured to perform convolution operation on the high-frequency information of the to-be-processed image to recover the high-frequency information of the to-be-processed image into a reference image, and perform spatial conversion on the reference image to obtain the reconstructed image of the to-be-processed image.
[0093] In an example embodiment of the present disclosure, the apparatus further comprises an image fusion module configured to fuse the to-be-processed image and a reconstructed image of the to-be-processed image to generate a target image corresponding to an original image to which the to-be-processed image belongs.
[0094] In an example embodiment of the present disclosure, the image fusion module comprises a fusion module configured to fuse the to-be-processed image and a reconstructed image of the to-be-processed image to obtain a fused image of the to-be-processed image; and an arrangement module configured to fuse the fused image according to an arrangement order of the to-be-processed image to generate a target image of an original image to which the to-be-processed image belongs.
[0095] In an example embodiment of the present disclosure, the high-frequency information determination module is configured to determine the intermediate information by residual processing of the feature information through a plurality of encoding modules in a visual reconstruction model; the visual reconstruction model comprises a feature extraction network, a plurality of encoding modules, and a dense connection network.
[0096] It should be noted that the specific details of the modules in the above image processing apparatus have been described in detail in the corresponding image processing method, and thus will not be described here.
[0097] An example embodiment of the present disclosure further provides an electronic device. The electronic device can be the client 101 described above, or can be a server. Generally, the electronic device can comprise a processor and a memory, the memory being configured to store executable instructions of the processor, and the processor being configured to execute the above image processing method by executing the executable instructions.
[0098] The following will take the mobile terminal 1200 in Figure 12 as an example to exemplarily describe the structure of the electronic device. Those skilled in the art should understand that, in addition to the components specially used for mobile purposes, Figure 12 the structure in can also be applied to devices of a fixed type.
[0099] As shown in Figure 12 , the mobile terminal 1200 can specifically comprise a processor 1201, a memory 1202, a bus 1203, a mobile communication module 1204, an antenna 1, a wireless communication module 1205, an antenna 2, a display screen 1206, a camera module 1207, an audio module 1208, a power module 1209, and a sensor module 1210.
[0100] The processor 1201 can include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, an encoder, a decoder, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. The image processing method in the example embodiment can be executed by the AP, the GPU, or the DSP, and when the method involves neural-network related processing, the NPU can be used to execute the method, for example, the NPU can load the neural-network parameters and execute the neural-network related algorithm instructions. For example, the network parameters of the visual reconstruction model can be loaded, the feature information of the image to be processed is determined by the residual connection of each coding module in the residual network, the high-frequency information output by each coding module is determined according to the residual connection of the feature information and the intermediate information, and thus the high-frequency information of the image to be processed is determined.
[0101] An encoder can encode (i.e., compress) an image or a video to reduce the data size for storage or transmission. A decoder can decode (i.e., decompress) the encoded data of an image or a video to restore the image or video data. The mobile terminal 1200 can support one or more encoders and decoders, such as JPEG (Joint Photographic Experts Group), PNG (Portable Network Graphics), BMP (Bitmap), etc. image formats, MPEG (Moving Picture Experts Group) 1, MPEG 10, H.1063, H.1064, HEVC (High Efficiency Video Coding), etc. video formats.
[0102] The processor 1201 can be connected to the memory 1202 or other components through the bus 1203.
[0103] The memory 1202 can be used to store computer-executable program codes including instructions. The processor 1201 executes various functional applications of the mobile terminal 1200 and image processing by running the instructions stored in the memory 1202. The memory 1202 can also store application data, such as storing images, videos, etc. files.
[0104] The communication function of mobile terminal 1200 can be implemented through mobile communication module 1204, antenna 1, wireless communication module 1205, antenna 2, modem processor, and baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Mobile communication module 1204 can provide 3G, 4G, 5G and other mobile communication solutions for mobile terminal 1200. Wireless communication module 1205 can provide wireless communication solutions such as wireless LAN, Bluetooth, and near-field communication for mobile terminal 1200.
[0105] The display screen 1206 is used to implement display functions, such as displaying the user interface, images, and videos. The camera module 1207 is used to implement shooting functions, such as capturing images and videos. The audio module 1208 is used to implement audio functions, such as playing audio and capturing voice. The power module 1209 is used to implement power management functions, such as charging the battery, supplying power to the device, and monitoring battery status. The sensor module 1210 may include one or more sensors to implement corresponding sensing and detection functions. For example, the sensor module 1210 may include an inertial sensor, which is used to detect the motion posture of the mobile terminal 1200 and output inertial sensing data.
[0106] It should be noted that the present disclosure also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or it may exist alone and not assembled into the electronic device.
[0107] Computer-readable storage media can be, for example—but not limited to—electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0108] A computer-readable storage medium can be sent, propagated, or transmitted for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted using any suitable medium, including but not limited to: wireless, wireline, optical fiber, RF, etc., or any suitable combination thereof.
[0109] The computer-readable storage medium carries one or more programs which, when executed by one of the electronic devices, cause the device to perform methods described in the following embodiments.
[0110] Those skilled in the art will easily understand from the above description of the embodiments that the example embodiments described herein can be implemented by software, or by software in combination with necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash disk, a mobile hard disk, or the like) or on a network, and includes a number of instructions to make a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) execute the methods according to the embodiments of the present disclosure.
[0111] In addition, the above-described figures are only schematic illustrations of the processes included in the methods according to the example embodiments of the present disclosure, and are not intended to be limiting. It is easy to understand that the processes shown in the above-described figures do not indicate or limit the time sequence of the processes. In addition, it is also easy to understand that the processes can be executed synchronously or asynchronously, for example, in multiple modules.
[0112] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such a division is not mandatory. Indeed, according to the embodiments of the present disclosure, the features and functionalities of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functionalities of one module or unit described above can be further divided into embodied by multiple modules or units.
[0113] Other embodiments of the present disclosure will be apparent to those skilled in the art with the consideration of the specification and practice of the disclosed content. The present application intends to cover any variations, uses, or adaptive changes of the present disclosure that follow the general principles of the present disclosure and include common knowledge or conventional technical means in the technical field that are not disclosed by the present disclosure. The specification and embodiments are only considered as examples, and the true scope and spirit of the present disclosure are indicated by the claims. It should be understood that the present disclosure is not limited to the precise structures already described and shown in the drawings, and various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is only limited by the appended claims.
Claims
1. An image processing method, characterized in that, include: The image to be processed is acquired, and feature extraction is performed on the image to be processed to obtain the feature information of the image to be processed; the image to be processed is an image patch; The feature information is processed by multiple encoding modules to determine intermediate information, and the high-frequency information output by each encoding module is determined based on the feature information and the residual connections of the intermediate information. The high-frequency information output by each encoding module is connected with the high-frequency information output by all encoding modules following it in the channel dimension to form dense connections. The input of the first encoding module is connected with the high-frequency information output by any module to form long connections. Cross-module connections between multiple encoding modules are realized through multiple long connections and dense connections to determine the high-frequency information of the image to be processed. Each encoding module includes multiple residual networks. The high-frequency information is convolved to restore the high-frequency information of the image to be processed to a reference image, and the reference image is spatially transformed to obtain the reconstructed image of the image to be processed; the image to be processed is upsampled, and a fusion image is determined based on the upsampling result and the reconstructed image of the image to be processed; the fusion image is stitched and fused according to the arrangement order of the image blocks to obtain the target image of the original image to which the image to be processed belongs.
2. The image processing method according to claim 1, characterized in that, The residual network includes a first residual network and a second residual network. The step of performing feature processing on the feature information through the residual network in each encoding module to determine the intermediate information includes: The feature information is processed by a first residual network to generate a residual result, and the residual result and the feature information are processed by a second residual network to determine the intermediate information.
3. The image processing method according to claim 1, characterized in that, The method further includes: The feature information is connected to the high-frequency information output by any one of the multiple encoding modules, so as to combine the feature information and the high-frequency information of any one of the encoding modules.
4. The image processing method according to claim 1, characterized in that, The step of reconstructing the image to be processed based on the high-frequency information to obtain the reconstructed image of the image to be processed includes: The high-frequency information output by each encoding module is spliced to determine the high-frequency information of the image to be processed, and the image to be processed is reconstructed based on the high-frequency information of the image to be processed to obtain the reconstructed image of the image to be processed.
5. The image processing method according to claim 4, characterized in that, The concatenation of high-frequency information from each encoding module includes: Using the same channel, the high-frequency information output by each of the multiple encoding modules is connected in parallel with the high-frequency information output by the remaining encoding modules other than each of the individual encoding modules.
6. The image processing method according to claim 1, characterized in that, The process of determining the fused image based on the upsampling result and the reconstructed image of the image to be processed, and then stitching and fusing the fused image according to the arrangement order of image blocks to obtain the target image of the original image to which the image to be processed belongs, includes: The upsampling result of the image to be processed is fused with the reconstructed image of the image to be processed to obtain the fused image of the image to be processed; The fused images are fused according to the order of the images to be processed to generate a target image of the original image to which the images to be processed belong.
7. The image processing method according to claim 1, characterized in that, The step of performing feature processing on the feature information through multiple encoding modules to determine intermediate information includes: The intermediate information is determined by residual processing of the feature information through multiple encoding modules in the visual reconstruction model; the visual reconstruction model includes a feature extraction network, multiple encoding modules, and a dense connection network.
8. An image processing apparatus, characterized in that, include: The feature extraction module is used to acquire the image to be processed and to extract features from the image to be processed to obtain the feature information of the image to be processed; the image to be processed is an image block; A high-frequency information determination module is used to perform feature processing on the feature information through multiple encoding modules to determine intermediate information, and to determine the high-frequency information output by each encoding module based on the feature information and the residual connections of the intermediate information; the high-frequency information output by each encoding module is connected with the high-frequency information output by all encoding modules following the first encoding module in the channel dimension to form dense connections, and the input of the first encoding module is connected with the high-frequency information output by any module in a long connection. Cross-module connections between multiple encoding modules are realized through multiple long connections and dense connections to determine the high-frequency information of the image to be processed; each encoding module includes multiple residual networks; The image reconstruction module is used to perform convolution operations on the high-frequency information to restore the high-frequency information of the image to be processed into a reference image, and to perform spatial transformation on the reference image to obtain the reconstructed image of the image to be processed; to upsample the image to be processed, and to determine a fusion image based on the upsampling result and the reconstructed image of the image to be processed; and to stitch and fuse the fusion image according to the arrangement order of the image blocks to obtain the target image of the original image to which the image to be processed belongs.
9. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the image processing method according to any one of claims 1-7 by executing the executable instructions.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the image processing method according to any one of claims 1-7.
Citation Information
Patent Citations
Image processing method and device, electronic equipment and computer readable storage medium
CN110782395A
License plate image super-resolution reconstruction model and method based on multi-scale features
CN111915490A
Image enhancement method and device and electronic equipment
CN112150400A
Image processing method and device, computer equipment and storage medium
CN113763243A