Image processing method and device based on multilayer convolutional neural network, and storage medium

Through multi-layer convolutional neural network, image feature extraction, amplification and detail enhancement are gradually carried out, and global optimization is carried out, solving the problems of image blurring and detail loss in the existing image amplification method, and achieving high definition and rich details image amplification effect.

CN120107071APending Publication Date: 2025-06-06BEIJING THUNDERSTONE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510186652.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing image amplification methods such as interpolation methods can easily lead to blurred images, serious loss of details, and unclear edges when the magnification ratio is large.

Method used

The image processing method based on multi-layer convolutional neural network is adopted to gradually extract, amplify and enhance the details through low-level, middle-level and high-level networks, and finally global optimization is carried out to generate clear and detailed images.

Benefits of technology

The problems of blurred details or excessive amplification in traditional methods are effectively avoided. The generated images have reached a high level in clarity, richness of details and visual consistency, and are flexible to adapt to different types of image processing tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107071A_ABST
    Figure CN120107071A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and provides an image processing method based on a multi-layer convolutional neural network, and the method comprises the steps: carrying out the feature extraction of a target image through a lower-layer network of the multi-layer convolutional neural network, and obtaining a feature map of the target image; the low-layer network carries out amplification and detail enhancement on the feature map of the target image to obtain a preliminary detail-enhanced image; a middle-layer network of the multi-layer convolutional neural network carries out re-amplification and detail re-enhancement on the preliminary detail enhanced image to obtain an intermediate image; a high-level network of the multilayer convolutional neural network amplifies the intermediate image for three times to obtain a final amplified image; and the high-level network globally optimizes the final magnified image. According to the technical scheme, details of the image can be kept when the image is magnified, and the method is suitable for different types of image processing tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an image processing method, device and storage medium based on a multi-layer convolutional neural network. Background Art

[0002] Image magnification (also known as super-resolution) is an important research topic in computer vision and is widely used in medical imaging, satellite image analysis, security monitoring, video processing and other fields. In these applications, image magnification not only needs to improve the resolution, but also needs to retain or enhance the details as much as possible for better analysis and processing. Existing image magnification methods are mainly interpolation methods, including nearest neighbor interpolation, bilinear interpolation and bicubic interpolation, etc. These methods insert new pixels at the position of the magnified pixels and estimate the value of the new pixels by the values ​​of the adjacent pixels. However, the above-mentioned traditional methods such as interpolation for image magnification have the disadvantage that they usually lead to blurred images, especially when the magnification ratio is large, the details are severely lost and the edges are not clear. Summary of the invention

[0003] The present application provides an image processing method, device and storage medium based on a multi-layer convolutional neural network, which can maintain the details of the image when the image is enlarged and can be used for different types of image processing tasks.

[0004] On the one hand, the present application provides an image processing method based on a multi-layer convolutional neural network, the method comprising:

[0005] The lower layer network of the multi-layer convolutional neural network extracts features of the target image to obtain a feature map of the target image;

[0006] The low-level network amplifies and enhances the details of the feature map of the target image to obtain a preliminary detail-enhanced image;

[0007] The middle layer network of the multi-layer convolutional neural network re-enlarges and re-enhances the details of the preliminary detail-enhanced image to obtain an intermediate image;

[0008] The high-level network of the multi-layer convolutional neural network amplifies the intermediate image three times to obtain a final amplified image;

[0009] The high-level network globally optimizes the final enlarged image.

[0010] On the other hand, the present application provides an image processing device based on a multi-layer convolutional neural network, the device comprising:

[0011] A feature extraction module, used to extract features of a target image using a low-layer network of a multi-layer convolutional neural network to obtain a feature map of the target image;

[0012] A low-level processing module, used to use the low-level network to amplify and enhance the details of the feature map of the target image to obtain a preliminary detail-enhanced image;

[0013] A middle-layer processing module, used for re-enlarging and re-enhancing the details of the preliminary detail-enhanced image using the middle-layer network of the multi-layer convolutional neural network to obtain an intermediate image;

[0014] An enlargement module, used to enlarge the intermediate image three times using a high-level network of the multi-layer convolutional neural network to obtain a final enlarged image;

[0015] An optimization module is used to globally optimize the final enlarged image using the high-level network.

[0016] In a third aspect, the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the technical solution of the image processing method based on a multi-layer convolutional neural network as described above when executing the computer program.

[0017] In a fourth aspect, the present application provides a storage medium storing a computer program, which, when executed by a processor, implements the steps of the technical solution of the above-mentioned image processing method based on a multi-layer convolutional neural network.

[0018] From the technical solution provided by the present application, it can be seen that the low-level network of the multi-layer convolutional neural network performs feature extraction, magnification and detail enhancement on the target image to obtain a preliminary detail-enhanced image, the middle-level network of the multi-layer convolutional neural network re-enlarges and re-enhances the preliminary detail-enhanced image to obtain an intermediate image, and the high-level network of the multi-layer convolutional neural network amplifies and optimizes the intermediate image three times to output the final enlarged image. Compared with the prior art, the technical solution of the present application gradually refines the image information and optimizes step by step from local features to global consistency, thereby avoiding the problem of blurred details or over-enlargement that is prone to occur in traditional single convolutional networks, so that the generated image reaches a high level in clarity, detail richness and visual consistency. Furthermore, this layered processing solution can be adjusted according to different application requirements to adapt to different types of image processing tasks, and therefore, it also has considerable flexibility. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0020] Figure 1 is a flowchart of an image processing method based on a multi-layer convolutional neural network provided in an embodiment of the present application;

[0021] Figure 2 is a structural schematic diagram of an image processing device based on a multi-layer convolutional neural network provided in an embodiment of the present application;

[0022] Figure 3 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0023] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0024] In this specification, adjectives such as first and second may be used only to distinguish one element or action from another element or action, without necessarily requiring or implying any actual such relationship or order. Where circumstances permit, reference to an element or component or step (etc.) should not be interpreted as being limited to only one of the elements, components, or steps, but may be one or more of the elements, components, or steps, etc.

[0025] In this specification, for the convenience of description, the sizes of various parts shown in the drawings are not drawn according to the actual proportional relationship.

[0026] Image magnification (also known as super-resolution) is an important research topic in computer vision and is widely used in medical imaging, satellite image analysis, security monitoring, video processing and other fields. In these applications, image magnification not only needs to improve the resolution, but also needs to retain or enhance the details as much as possible for better analysis and processing. Existing image magnification methods are mainly interpolation methods, including nearest neighbor interpolation, bilinear interpolation and bicubic interpolation, etc. These methods insert new pixels at the position of the magnified pixels and estimate the value of the new pixels by the values ​​of the adjacent pixels. However, the above-mentioned traditional methods such as interpolation for image magnification have the disadvantage that they usually lead to blurred images, especially when the magnification ratio is large, the details are severely lost and the edges are not clear.

[0027] In view of the above problems in the prior art, this application proposes an image processing method based on a multi-layer convolutional neural network, the flow chart of which is shown in the attached figure. Figure 1 As shown, it mainly includes steps S101 to S105, which are described in detail as follows:

[0028] Step S101: The lower layer network of the multi-layer convolutional neural network extracts features of the target image to obtain a feature map of the target image.

[0029] Multi-level Convolutional Neural Network (MCNN) is a special convolutional neural network (CNN), which improves the effect of image processing by introducing a multi-level hierarchical structure, and has a wide range of applications in the fields of image classification, target detection, face recognition and image enhancement. The low-level network of a multi-layer convolutional neural network refers to the convolutional neural network closest to the input layer (or including the input layer), which contains a small number of convolutional layers and pooling layers, and is mainly used to extract low-level features of the image. Although the number of layers of the low-level network is relatively small or the depth is low, the problem of gradient disappearance or gradient explosion in deep network training is still unavoidable. In order to solve the above problems, as an embodiment of the present application, the low-level convolutional network of the multi-layer convolutional neural network extracts features from the target image, and the feature map of the target image obtained can be: the convolution layer of the low-level network extracts features from the target image to obtain a preliminary feature map of the target image; the first residual block of the low-level network obtains the residual corresponding to the preliminary feature map; the preliminary feature map is added to the residual corresponding to the preliminary feature map obtained by the first residual block to obtain the feature map of the target image.

[0030] In the embodiment of the present application, the residual block of the low-level network (referred to as the first residual block here to distinguish it from the residual block of the subsequent high-level network) can be composed of two convolutional layers, and the output is the residual of the input plus the original input. The model is as follows:

[0031] R(x)=F(x,{W i})+x

[0032] Where x is the input image (or feature map), F(x,{W i}) is the feature map after convolution and nonlinear activation function (such as ReLU), W i is the convolution kernel parameter in the network, and R(x) is the output residual map. From the residual block model, we can see that the residual block learns the difference (i.e., residual) between the input image and its reconstructed image, rather than directly learning the mapping from low resolution to high resolution.

[0033] The training of the residual block model can be to input the low-resolution image into the deep neural network. After the network is processed, the network outputs a predicted detail part The predicted details should be as close as possible to the true details R i ,Right now:

[0034]

[0035] In order to optimize the parameters of the network, the output details are as close as possible to the real details R i , we need to define a loss function, such as mean square error (MSE) loss, to measure the difference between the predicted details and the true details:

[0036]

[0037] Among them, L(θ) is the loss function, θ is the parameter of the deep neural network, and N is the number of samples in the training set. By minimizing the loss function, the parameter θ of the deep neural network will gradually be optimized, making the detailed part of the prediction Closer to the true residual R i , thereby better enhancing the details of the image.

[0038] From the above training of the residual block model, it can be seen that the reason why the first residual block can obtain the residual corresponding to the preliminary feature map is that when it was trained in the early stage, it has learned the relationship between the reconstructed image and the preliminary feature map. Therefore, after obtaining the preliminary feature map of the target image, the feature map of the target image can be obtained by adding the preliminary feature map to the residual corresponding to the target image obtained by the first residual block. It is precisely because the difference between the input and the target is directly learned instead of the target, so the gradient disappearance or gradient explosion problems in deep network training can be avoided, and the network can be helped to focus on the recovery of details. It should be noted that the preliminary feature map here can be interpreted in a broad sense, that is, it can refer to both the target image input by the input layer and the feature map that has been obtained by the network layer before the current layer.

[0039] Step S102: The low-level network enlarges and enhances the details of the feature map of the target image to obtain a preliminary detail-enhanced image.

[0040] Specifically, the low-level network enlarges and enhances the details of the feature map of the target image, and the preliminary detail-enhanced image can be obtained by: the low-level network upsamples the feature map of the target image to obtain a preliminary enlarged image of the target image; the preliminary enlarged image is transformed in the frequency domain and the spatial domain respectively to obtain the corresponding first detail-enhanced image and the second detail-enhanced image; the preliminary enlarged image, the first detail-enhanced image and the second detail-enhanced image are fused to obtain the preliminary detail-enhanced image. Considering that the Fourier transform has the advantages of fine edge enhancement and global processing capability, and the wavelet transform has the advantages of local detail enhancement and can adaptively enhance details of different sizes and directions at different scales, in order to complement each other, the frequency domain transform of the above embodiment can be a Fourier transform, and the spatial domain transform can be a wavelet transform. As for the weight coefficients of the first detail-enhanced image and the second detail-enhanced image when the preliminary enlarged image, the first detail-enhanced image and the second detail-enhanced image are fused, they can be optimized by minimizing the loss function after using the loss function to measure the difference between the image after detail enhancement and the real image, or these weight coefficients can be dynamically adjusted by the gradient descent method to ensure the effect of image detail enhancement. Specifically, the weight coefficients of the first detail enhancement map and the second detail enhancement map may be implemented through steps S1021 to S1024, which are described in detail as follows:

[0041] Step S1021: using the weight coefficients of the first detail enhancement map and the second detail enhancement map to represent a preliminary detail enhanced image.

[0042] First, according to the previous steps, the preliminary detail enhancement image I Final It can be expressed as:

[0043] I Final =I HR +λ 1 *H HighFreq +λ 2 *H Enhance

[0044] Among them, I HR is the upsampled image, i.e., the initial enlarged image of the target image, H HighFreq is the high frequency part in the frequency domain obtained by Fourier transform, i.e., the first detail enhancement image. Enhance is the detail part enhanced by wavelet transform, i.e., the second detail enhancement map, λ 1 and λ 2 They are H HighFreq and H Enhance The weight coefficient is also a hyperparameter that controls the strength of detail enhancement and determines the impact of the detail enhancement part on the final image.

[0045] Step S1022: Determine the comprehensive loss function.

[0046] To optimize λ 1 and λ 2 , a loss function is needed to measure the effect of image detail enhancement, that is, the difference between the preliminary detail enhanced image and the real high-resolution image. In the embodiment of the present application, perceptual loss (Perceptual Loss) and structural similarity loss (SSIM Loss) are used as loss functions, among which perceptual loss (Perceptual Loss) is used to measure the similarity of images in the feature space, usually using a pre-trained convolutional neural network (such as VGG) to extract features and then calculate the loss, and structural similarity loss (SSIM Loss) is used to measure the structural consistency of the image. The comprehensive loss function L is defined here total for:

[0047] L total =α*L perceptual +β*L SSIM +γ*L MSE

[0048] Among them, L perceptual is the perceptual loss, which is used to measure the similarity of high-level features, L SSIM is the structural similarity loss, which is used to measure the structural consistency of the image. MSE is the mean square error loss, which is used to measure the difference at the pixel level. α, β, and γ are the weight hyperparameters of the loss function, which are used to control the influence of each part of the loss.

[0049] Specifically, the perceptual loss L perceptual The image quality is measured by calculating the difference between the original image and the generated image in some intermediate layer features. Generally speaking, the pre-trained VGG network is used to extract features, and the perceptual loss is calculated as:

[0050]

[0051] Among them, φ i (·) represents the feature map of the VGG network at layer i, and N is the number of selected feature layers.

[0052] Structural similarity loss (SSIM) evaluates image quality by measuring the similarity of images in terms of brightness, contrast, and structure. SSIM loss can be expressed as:

[0053] L SSIM =1-SSIM(I Final ,I HR )

[0054] Among them, SSIM is the structural similarity index, and the value range is from -1 to 1. The closer the value is to 1, the more similar the structure is.

[0055] The mean square error (MSE) loss measures the pixel-level error of the image and can be used as an auxiliary loss to help optimize detail recovery:

[0056]

[0057] Where N is the number of pixels in the image, I Final (i) and I HR (i) represent the preliminary detail enhanced image I Final and the value of the initial magnified image of the target image at the i-th pixel.

[0058] Step S1023: Optimizing the weight coefficients of the first detail enhancement map and the second detail enhancement map.

[0059] To determine λ 1 and λ 2 , we need to transform the comprehensive loss function L total Therefore, the optimization goal is to adjust λ by gradient descent method. 1 and λ 2 , so that the comprehensive loss function L total Minimize. The optimization formula is:

[0060]

[0061] The gradient descent update rule is:

[0062]

[0063] Where η is the learning rate, and They are the comprehensive loss function L total Relative to λ 1 and λ 2 gradient.

[0064] Step S1024: Dynamically adjust λ 1 and λ 2 .

[0065] During the training process, the gradient descent algorithm will be based on the comprehensive loss function L total Feedback dynamically adjusts λ 1 and λ 2 The value of the comprehensive loss function L total Converges to the optimal value. At this time, λ 1 and λ 2 It represents the optimal detail enhancement strength parameter, which can maximize the restoration of image details while maintaining the structural consistency of the image.

[0066] Step S103: The middle layer network of the multi-layer convolutional neural network re-enlarges and re-enhances the details of the preliminary detail-enhanced image to obtain an intermediate image.

[0067] In the embodiment of the present application, the middle-layer network of the multi-layer convolutional neural network re-enlarges and re-enhances the details of the preliminary detail-enhanced image. It can refer to the residual idea used by the low-layer network of the multi-layer convolutional neural network in the aforementioned embodiment to extract features of the target image, and can also refer to the idea of ​​fusing the detail enhancement map obtained by transforming the frequency domain and spatial domain when the low-layer network in the aforementioned embodiment enlarges and enhances the details of the feature map of the target image. It will not be elaborated here.

[0068] Step S104: The high-level network of the multi-layer convolutional neural network amplifies the intermediate image three times to obtain a final amplified image.

[0069] As an embodiment of the present application, the high-level network of the multi-layer convolutional neural network magnifies the intermediate image three times, and obtaining the final magnified image can be achieved by the following steps S1041 and S1042, which are described in detail as follows:

[0070] Step S1041: The high-level network extracts features of the intermediate image according to the convolution kernel of the high-level network and the number of channels of the intermediate image to obtain an intermediate feature map.

[0071] Specifically, each convolution kernel of the first convolution kernel with the same number of channels as the intermediate image can be used to perform convolution operations on each channel of the intermediate image, that is, use each convolution kernel to perform convolution operations on each channel of the intermediate image to obtain a feature map after convolution; then, use each second convolution kernel of C' second convolution kernels to perform convolution on all channels of the convolution feature map to obtain an intermediate feature map with C' channels. In the above embodiment, using each convolution kernel of the first convolution kernel with the same number of channels as the intermediate image to perform convolution operations on each channel of the intermediate image is actually a deep convolution operation, that is, if the size of the intermediate image is H×W×C, the size of the first convolution kernel is K×K×1, the stride of the convolution operation is s, and the padding is p, then each convolution kernel of the first convolution kernel performs convolution operations on each channel of the intermediate image, which can be expressed as follows:

[0072]

[0073] Where I(i·s+m,j·s+m,c) is the value of the intermediate image at position (i·s+m,j·s+m) and channel c, k c (m,n) is the value of the cth convolution kernel at position (m,n), F d(i, j, c) is the value of the convolutional feature map at (i, j) and channel c. The size of the convolutional feature map is

[0074] Using each of the C' second convolution kernels to convolve all channels of the convolved feature map to obtain C' intermediate feature maps is a point-by-point convolution operation. That is, if the size of the convolved feature map is H'×W'×C, the size of the second convolution kernel is 1×1×C, and there are C' second convolution kernels in total, then using each of the C' second convolution kernels to convolve all channels of the convolved feature map can be expressed as follows:

[0075]

[0076] Among them, I(i,j,c) is the value of the convolution feature map of size H'×W'×C at position (i,j) and channel c, k c' (0,0,c) is the value of the c'th second convolution kernel at position (0,0) and channel c, F p (i, j, c') is the value of the intermediate feature map at position (i, j) and channel c' obtained after point-by-point convolution. p From the expression of (i,j,c'), we can see that each 1×1×C second convolution kernel is convolved with all channels of I(i,j,c), and a total of C' intermediate feature maps of size H'×W'×C' can be obtained.

[0077] In the above embodiment, on the one hand, each convolution kernel of the first convolution kernel with the same number of channels as the intermediate image is used to perform a convolution operation on each channel of the intermediate image respectively. Since the convolution operation is performed on each channel of the intermediate image separately, the local features can be better retained. Using each of the C' second convolution kernels to convolve all channels of the convolved feature map is essentially a linear combination between channels, so the global features can be captured. The combination of the two can better represent the features of the intermediate image. On the other hand, since the number of parameters is reduced, the combination of the two convolution methods can reduce the complexity of the model, thereby reducing the risk of overfitting, which means that good generalization ability can be demonstrated even on small data sets.

[0078] Step S1042: Rearrange pixels of the intermediate feature map according to the upsampling factor and the target number of channels to obtain a final enlarged image.

[0079] After step S1042, the intermediate feature map obtained is still a low-resolution feature map. Rearranging the pixels of the intermediate feature map according to the upsampling factor and the target number of channels is essentially converting the channel dimension of the low-resolution feature map into a spatial dimension, thereby improving the resolution of the image. Specifically, if the size of the intermediate feature map is H'×W'×C', where H' is the longitudinal size of the intermediate feature map, i.e., the height, W is the lateral size of the intermediate feature map, i.e., the width, and C' is the number of channels of the intermediate feature map, C'=r 2 ×C new , r is the upsampling factor, C new is the target number of channels. First, the channel dimensions of the intermediate feature map are rearranged so that H'×W'×r×r×C new , which means that the feature vector at each position is split into r×r small blocks, each of which corresponds to a new channel. Then, the rearranged feature map is reorganized into (H'×r)×(W'×r)×C new , that is, each r×r small block is expanded and spliced ​​into the new high-resolution feature map to obtain the final enlarged image.

[0080] As another embodiment of the present application, the high-level network of the multi-layer convolutional neural network magnifies the intermediate image three times to obtain the final magnified image through the following steps S'1041 to S'1045, which are described in detail as follows:

[0081] Step S'1041: The high-level network extracts features from the intermediate image to obtain a high-level feature map.

[0082] Here, the high-level network can perform feature extraction on the intermediate image using any existing feature extraction scheme, which will not be elaborated herein.

[0083] Step S'1042: Fuse the high-level feature map with the intermediate image through skip connections to obtain a fused feature map.

[0084] In the embodiment of the present application, the skip connection can be a direct connection between the outputs of the high-level network and the middle-level network of the multi-layer convolutional neural network without passing through the layer outputs of the middle-level network, that is, directly connecting the low-level features output by the middle-level network with the high-level features output by the high-level network, ensuring that the low-level detail information can affect the processing results of the higher level, thereby avoiding the loss of details. In other words, the skip connection can be a low-level feature output by the middle-level network (assuming that F i Represented) and high-level features output by high-level networks (assuming F jA connection is added between the [representation] such that the output of the previous layer can be directly passed to the next layer. Assume that there is a layer i in the middle-level network and a layer j in the high-level network (where i < j). In the conventional way, the feature map output by the i-th layer needs to go through several layers of operations in the middle-level network before being passed to the j-th layer of the high-level network. However, in the embodiments of the present application, a skip connection can be established between the i-th layer of the middle-level network and the j-th layer of the high-level network in the following way:

[0085]

[0086] where F i is the low-level feature output by the i-th layer of the middle-level network, and F j is the high-level feature output by the j-th layer of the high-level network. represents the fused feature map obtained after the skip connection. Through the skip connection, the high-level network can receive more information from the middle-level network, especially the detailed information, which helps to restore the details of the image.

[0087] Step S’1043: Upsample the fused feature map to obtain a high-resolution image.

[0088] Specifically, upsampling the fused feature map can be performed by using a deconvolution (or transposed convolution) operation on the fused feature map to obtain a high-resolution image.

[0089] Step S’1044: The second residual block obtains the residual corresponding to the high-resolution image.

[0090] Step S’1045: Add the preliminary enlarged image to the residual corresponding to the high-resolution image obtained by the second residual block to obtain the final enlarged image.

[0091] Step S’1045 is similar to adding the preliminary feature map to the residual corresponding to the preliminary feature map obtained by the first residual block in the foregoing embodiments. Both use the residual idea, that is, the model learns the detailed differences between the low-resolution image and the high-resolution image, and then adds the preliminary enlarged image to the residual corresponding to the high-resolution image obtained by the second residual block to obtain the final enlarged image.

[0092] Step S105: The high-level network globally optimizes the final enlarged image.

[0093] On the one hand, considering that traditional convolutional networks are usually limited to local receptive fields, which means that they can only perceive information in local areas, the self-attention mechanism calculates global dependencies so that the representation of each pixel can depend on other parts of the image. In this way, areas far away in the image can also affect the feature representation of the target area, thereby helping the network capture global contextual information; on the other hand, in image super-resolution or generation tasks, the restoration of details needs to consider the global structure of the image. For example, some details in the image may need to be coordinated through contextual information during the magnification process to avoid blurred or inconsistent areas. The self-attention mechanism calculates the similarity between pixels to help the model understand and adjust these details, so that the output image maintains consistency and detail coordination during the magnification process. Therefore, in an embodiment of the present application, the high-level network can globally optimize the final enlarged image through the self-attention mechanism.

[0094] From the above attached Figure 1 From the example of the image processing method based on the multi-layer convolutional neural network, it can be seen that the low-level network of the multi-layer convolutional neural network performs feature extraction, magnification and detail enhancement on the target image to obtain a preliminary detail-enhanced image, the middle-level network of the multi-layer convolutional neural network re-enlarges and re-enhances the details of the preliminary detail-enhanced image to obtain an intermediate image, and the high-level network of the multi-layer convolutional neural network amplifies and optimizes the intermediate image three times to output the final enlarged image. Compared with the prior art, the technical solution of the present application gradually refines the image information and optimizes step by step from local features to global consistency, thereby avoiding the problem of blurred details or over-enlargement that is prone to occur in traditional single convolutional networks, so that the generated image reaches a high level in clarity, detail richness and visual consistency. Furthermore, this layered processing solution can be adjusted according to different application requirements to adapt to different types of image processing tasks, and therefore, it also has considerable flexibility.

[0095] Please see attached Figure 2 , is an image processing device based on a multi-layer convolutional neural network provided in an embodiment of the present application, the device may include a feature extraction module 201, a low-level processing module 202, a middle-level processing module 203, an amplification module 204 and an optimization module 205, which are described in detail as follows:

[0096] The feature extraction module 201 is used to extract features of the target image using the lower layer network of the multi-layer convolutional neural network to obtain a feature map of the target image;

[0097] A low-level processing module 202 is used to amplify and enhance the details of the feature map of the target image using the low-level network to obtain a preliminary detail-enhanced image;

[0098] A middle layer processing module 203 is used to re-enlarge and re-enhance the details of the preliminary detail-enhanced image using a middle layer network of a multi-layer convolutional neural network to obtain an intermediate image;

[0099] A magnification module 204 is used to magnify the intermediate image three times using a high-level network of a multi-layer convolutional neural network to obtain a final magnified image;

[0100] The optimization module 205 is used for the high-level network to perform global optimization on the final enlarged image.

[0101] From the above attached Figure 2 It can be seen from the example of an image processing device based on a multi-layer convolutional neural network that the low-level network of the multi-layer convolutional neural network performs feature extraction, magnification, and detail enhancement on the target image to obtain a preliminary detail-enhanced image, the middle-level network of the multi-layer convolutional neural network re-enlarges and re-enhances the details of the preliminary detail-enhanced image to obtain an intermediate image, and the high-level network of the multi-layer convolutional neural network amplifies and optimizes the intermediate image three times to output the final enlarged image. Compared with the prior art, the technical solution of the present application gradually refines the image information and optimizes step by step from local features to global consistency, thereby avoiding the problem of blurred details or over-enlargement that is prone to occur in traditional single convolutional networks, so that the generated image reaches a high level in clarity, detail richness, and visual consistency. Furthermore, this layered processing solution can be adjusted according to different application requirements to adapt to different types of image processing tasks, and therefore, it also has considerable flexibility.

[0102] Figure 3 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Figure 3 As shown, the electronic device 3 of this embodiment mainly includes: a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30, such as a program of an image processing method based on a multi-layer convolutional neural network. When the processor 30 executes the computer program 32, the steps in the above-mentioned image processing method based on a multi-layer convolutional neural network are implemented, such as Figure 1 Alternatively, when the processor 30 executes the computer program 32, the functions of each module / unit in the above-mentioned device embodiments are implemented, for example Figure 2 The functions of the feature extraction module 201, the low-level processing module 202, the middle-level processing module 203, the amplification module 204 and the optimization module 205 are shown.

[0103] Exemplarily, the computer program 32 of the image processing method based on the multi-layer convolutional neural network mainly includes: the low-layer network of the multi-layer convolutional neural network extracts features of the target image to obtain a feature map of the target image; the low-layer network amplifies and enhances the details of the feature map of the target image to obtain a preliminary detail-enhanced image; the middle-layer network of the multi-layer convolutional neural network re-enlarges and re-enhances the details of the preliminary detail-enhanced image to obtain an intermediate image; the high-layer network of the multi-layer convolutional neural network amplifies the intermediate image three times to obtain a final enlarged image; the high-layer network globally optimizes the final enlarged image. The computer program 32 can be divided into one or more modules / units, one or more modules / units are stored in the memory 31, and are executed by the processor 30 to complete the present application. One or more modules / units can be a series of computer program instruction segments that can complete specific functions, and the instruction segments are used to describe the execution process of the computer program 32 in the electronic device 3. For example, the computer program 32 can be divided into the functions of a feature extraction module 201, a low-level processing module 202, a middle-level processing module 203, an amplification module 204 and an optimization module 205 (modules in the virtual device), and the specific functions of each module are as follows: the feature extraction module 201 is used to use the low-level network of the multi-layer convolutional neural network to extract features of the target image and obtain a feature map of the target image; the low-level processing module 202 is used to use the low-level network to amplify and enhance the details of the feature map of the target image to obtain a preliminary detail-enhanced image; the middle-level processing module 203 is used to use the middle-level network of the multi-layer convolutional neural network to re-enlarge and re-enhance the details of the preliminary detail-enhanced image to obtain an intermediate image; the amplification module 204 is used to use the high-level network of the multi-layer convolutional neural network to amplify the intermediate image three times to obtain a final amplified image; the optimization module 205 is used to use the high-level network to globally optimize the final amplified image.

[0104] The electronic device 3 may include but is not limited to a processor 30 and a memory 31. Those skilled in the art will appreciate that Figure 3 It is only an example of the electronic device 3 and does not constitute a limitation of the electronic device 3. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.

[0105] The processor 30 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0106] The memory 31 may be an internal storage unit of the electronic device 3, such as a hard disk or memory of the electronic device 3. The memory 31 may also be an external storage device of the electronic device 3, such as a plug-in hard disk, a smart memory card (SmartMedia Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device 3. Further, the memory 31 may also include both an internal storage unit of the electronic device 3 and an external storage device. The memory 31 is used to store computer programs and other programs and data required by the electronic device. The memory 31 may also be used to temporarily store data that has been output or is to be output.

[0107] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0108] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0109] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0110] In the embodiments provided in the present application, it should be understood that the disclosed devices / equipment and methods can be implemented in other ways. For example, the device / equipment embodiments described above are only schematic, for example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0111] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0112] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0113] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program of the image processing method based on the multi-layer convolutional neural network can be stored in a storage medium. When the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments, that is, the low-level network of the multi-layer convolutional neural network extracts features of the target image to obtain a feature map of the target image; the low-level network amplifies and enhances the details of the feature map of the target image to obtain a preliminary detail-enhanced image; the middle-level network of the multi-layer convolutional neural network re-enlarges and re-enhances the details of the preliminary detail-enhanced image to obtain an intermediate image; the high-level network of the multi-layer convolutional neural network amplifies the intermediate image three times to obtain a final enlarged image; the high-level network performs global optimization on the final enlarged image. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. Storage media may include: any entity or device capable of carrying computer program code, recording media, USB flash drives, mobile hard disks, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content contained in the storage medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, storage media do not include electric carrier signals and telecommunication signals.

[0114] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application is described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some of the technical features therein by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application. The specific implementation methods described above further describe the purpose, technical solutions and beneficial effects of the present application in detail. It should be understood that the above description is only the specific implementation method of the present application, and is not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the protection scope of the present invention.

Claims

1. An image processing method based on a multi-layer convolutional neural network, characterized in that: The method comprises: The lower layer network of the multi-layer convolutional neural network extracts features of the target image to obtain a feature map of the target image; The low-level network amplifies and enhances the details of the feature map of the target image to obtain a preliminary detail-enhanced image; The middle layer network of the multi-layer convolutional neural network re-enlarges and re-enhances the details of the preliminary detail-enhanced image to obtain an intermediate image; The high-level network of the multi-layer convolutional neural network amplifies the intermediate image three times to obtain a final amplified image; The high-level network globally optimizes the final enlarged image.

2. The image processing method based on a multi-layer convolutional neural network as claimed in claim 1, characterized in that: The lower-layer network includes a first residual block, and the lower-layer network of the multi-layer convolutional neural network extracts features of the target image to obtain a feature map of the target image, including: The convolution layer of the low-level network performs feature extraction on the target image to obtain a preliminary feature map of the target image; The first residual block obtains the residual corresponding to the preliminary feature map; The preliminary feature map is added to the residual corresponding to the preliminary feature map obtained by the first residual block to obtain a feature map of the target image.

3. The image processing method based on a multi-layer convolutional neural network as claimed in claim 2, characterized in that: The low-level network amplifies and enhances the details of the feature map of the target image to obtain a preliminary detail-enhanced image, including: The low-level network upsamples the feature map of the target image to obtain a preliminary enlarged image of the target image; Performing frequency domain transformation and spatial domain transformation on the preliminary enlarged image respectively to obtain a corresponding first detail enhanced image and a second detail enhanced image; The preliminary enlarged image, the first detail-enhanced image and the second detail-enhanced image are fused to obtain the preliminary detail-enhanced image.

4. The image processing method based on a multi-layer convolutional neural network as claimed in claim 3, characterized in that: The high-level network of the multi-layer convolutional neural network amplifies the intermediate image three times to obtain a final amplified image, including: The high-level network extracts features from the intermediate image according to the convolution kernel of the high-level network and the number of channels of the intermediate image to obtain an intermediate feature map; The intermediate feature map is rearranged in pixels according to the upsampling factor and the target number of channels to obtain the final enlarged image.

5. The image processing method based on a multi-layer convolutional neural network as claimed in claim 4, characterized in that: The step of extracting features from the intermediate image according to the convolution kernel of the high-level network and the number of channels of the intermediate image to obtain an intermediate feature map includes: Using each convolution kernel of the first convolution kernel, the number of which is the same as the number of channels of the intermediate image, to perform a convolution operation on each channel of the intermediate image, respectively, to obtain a convolved feature map; Each of the C' second convolution kernels is used to convolve all channels of the convolved feature map to obtain the intermediate feature map with C' channels.

6. The image processing method based on a multi-layer convolutional neural network according to any one of claims 1 to 5, characterized in that: The high-level network includes a second residual block, and the high-level network of the multi-layer convolutional neural network amplifies the intermediate image three times to obtain a final amplified image, including: The high-level network extracts features from the intermediate image to obtain a high-level feature map; The high-level feature map is fused with the intermediate image through a jump connection to obtain a fused feature map; Upsampling the fused feature map to obtain a high-resolution image; The second residual block obtains the residual corresponding to the high-resolution image; The preliminary enlarged image is added to the residual corresponding to the high-resolution image obtained by the second residual block to obtain the final enlarged image.

7. The image processing method based on a multi-layer convolutional neural network as claimed in claim 6, characterized in that: The high-level network globally optimizes the final enlarged image, including: The high-level network globally optimizes the final enlarged image through a self-attention mechanism.

8. An image processing device based on a multi-layer convolutional neural network, characterized in that: The device comprises: A feature extraction module, used to extract features of a target image using a low-layer network of a multi-layer convolutional neural network to obtain a feature map of the target image; A low-level processing module, used to use the low-level network to amplify and enhance the details of the feature map of the target image to obtain a preliminary detail-enhanced image; A middle-layer processing module, used for re-enlarging and re-enhancing the details of the preliminary detail-enhanced image using the middle-layer network of the multi-layer convolutional neural network to obtain an intermediate image; An enlargement module, used to enlarge the intermediate image three times using a high-level network of the multi-layer convolutional neural network to obtain a final enlarged image; An optimization module is used to globally optimize the final enlarged image using the high-level network.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.