Image processing method and device based on deep learning, equipment and storage medium

By extracting image features through convolutional neural networks and combining multi-scale feature fusion and attention mechanisms, the problem of detail loss during image magnification in existing technologies is solved, achieving simultaneous improvement in image resolution and detail.

CN121053010APending Publication Date: 2025-12-02BEIJING THUNDERSTONE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511122247.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

Existing image magnification techniques cannot effectively restore high-frequency details, resulting in blurred images or jagged edges. Furthermore, traditional interpolation algorithms and CNN-based methods have limitations in detail enhancement, making it difficult to maintain detail fidelity while improving resolution.

Method used

The image features edge, color distribution and texture features are extracted by convolutional neural network. The image features are enhanced in detail by combining multi-scale feature fusion and attention mechanism. The features are then magnified by interpolation and convolution operations, and finally post-processing optimization is performed.

Benefits of technology

It achieves the goal of maintaining detail and improving resolution while reducing noise and edge blurring during image magnification, thus enhancing visual quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121053010A_ABST
    Figure CN121053010A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and provides an image processing method based on deep learning, and the method comprises the steps: extracting the image features of a to-be-processed image through a convolutional neural network, the image features including an edge feature, a color distribution feature and a texture feature; inputting the image features of the to-be-processed image into the deep learning model, and performing detail enhancement on the image features of the to-be-processed image through multi-scale feature fusion and an attention mechanism to generate a detail enhanced image; amplifying the detail enhanced image through interpolation and convolution operation to obtain an amplified image and retaining the details of the image; and carrying out post-processing optimization on the amplified image. According to the technical scheme of the invention, the overall improvement of the image visual quality is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to an image processing method, apparatus, device, and storage medium based on deep learning. Background Technology

[0002] Image upscaling methods are widely used in medical imaging, satellite image analysis, security monitoring, video processing, and other fields. In these applications, image upscaling not only needs to increase resolution but also needs to preserve or enhance details as much as possible for better analysis and processing. Existing image upscaling techniques mainly rely on traditional interpolation algorithms (such as bicubic interpolation) or super-resolution methods based on convolutional neural networks (CNNs). However, traditional interpolation algorithms cannot recover high-frequency details during upscaling, leading to image blurring or jagged edges. While CNN-based methods can improve resolution, they have limitations in detail enhancement, such as insufficient sensitivity to texture and color distribution and a tendency to introduce artifacts. Summary of the Invention

[0003] This application provides an image processing method, apparatus, device, and storage medium based on deep learning, which can achieve an overall improvement in image visual quality.

[0004] On the one hand, this application provides an image processing method based on deep learning, the method comprising:

[0005] Image features of the image to be processed are extracted using a convolutional neural network, including edge features, color distribution features, and texture features.

[0006] The image features of the image to be processed are input into a deep learning model, and the image features of the image to be processed are enhanced in detail through multi-scale feature fusion and attention mechanism to generate a detail-enhanced image;

[0007] The detail-enhancing image is magnified by interpolation and convolution operations to obtain a magnified image while preserving the image details.

[0008] The magnified image is then post-processed and optimized.

[0009] On the other hand, this application provides a deep learning-based image processing apparatus, the apparatus comprising:

[0010] The extraction module is used to extract image features of the image to be processed through a convolutional neural network. The image features include edge features, color distribution features, and texture features.

[0011] The generation module is used to input the image features of the image to be processed into a deep learning model, and to enhance the details of the image features of the image to be processed through multi-scale feature fusion and attention mechanism to generate a detail-enhanced image;

[0012] The magnification module is used to magnify the detail-enhancing image through interpolation and convolution operations to obtain a magnified image while preserving the image details;

[0013] An optimization module is used to perform post-processing optimization on the magnified image.

[0014] Thirdly, this application provides an electronic device, the device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the technical solution of the deep learning-based image processing method described above.

[0015] Fourthly, this application provides a storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described deep learning-based image processing method.

[0016] As can be seen from the technical solution provided in this application, on the one hand, by extracting edge features, color distribution features, and texture features of the image to be processed through a convolutional neural network, the low-level details and high-level semantic information of the image can be fully captured, providing multi-dimensional feature support for subsequent processing; on the other hand, by combining multi-scale feature fusion and attention mechanism to enhance the weight of key details (such as high-frequency edges and complex textures) of the image features of the image to be processed, the problem of detail loss caused by a single feature layer is avoided. Then, the detail-enhanced image is magnified through interpolation and convolution operations. That is, by designing the magnification and detail enhancement steps in series, it is ensured that the magnified image can still maintain the authenticity of details when the resolution is improved; on the third hand, the noise, edge blurring, and color distortion problems that may exist in the magnified image are optimized in stages, and finally the overall visual quality is improved. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart of the deep learning-based image processing method provided in the embodiments of this application;

[0019] Figure 2 This is a schematic diagram of the structure of the deep learning-based image processing device provided in the embodiments of this application;

[0020] Figure 3 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0021] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0022] In this specification, adjectives such as "first" and "second" are used only to distinguish one element or action from another, without necessarily requiring or implying any actual such relationship or order. Where circumstances permit, reference to an element, component, or step (etc.) should not be construed as limited to only one element, component, or step, but may include one or more of the elements, components, or steps, etc.

[0023] For ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn to actual scale.

[0024] Image upscaling methods are widely used in medical imaging, satellite image analysis, security monitoring, and video processing. In these applications, image upscaling not only needs to increase resolution but also needs to preserve or enhance details as much as possible for better analysis and processing. Existing image upscaling techniques mainly rely on traditional interpolation algorithms (such as bicubic interpolation) or super-resolution methods based on convolutional neural networks (CNNs). However, traditional interpolation algorithms cannot recover high-frequency details during upscaling, leading to image blurring or jagged edges. While CNN-based methods can improve resolution, they have limitations in detail enhancement, such as insufficient sensitivity to texture and color distribution and a tendency to introduce artifacts. Furthermore, existing techniques typically separate image upscaling from post-processing (such as denoising and sharpening), resulting in redundant processing and difficulty in maintaining detail fidelity while improving resolution. In recent years, some research has attempted to optimize detail enhancement through multi-scale feature fusion or attention mechanisms, but the following problems still exist:

[0025] 1) Single feature extraction level: Most methods rely on feature maps of only a single level, which cannot effectively integrate low-level details (such as edges) with high-level semantic information (such as texture), resulting in insufficient detail enhancement;

[0026] 2) The contradiction between detail preservation and magnification: Super-resolution reconstruction networks are prone to losing local details of the original image during the magnification process, especially in complex texture areas;

[0027] 3) Post-processing is disconnected from the model: Post-processing operations such as denoising and sharpening are not optimized in coordination with the deep learning model, which may destroy the enhanced details.

[0028] To address the aforementioned problems in existing technologies, this application proposes a deep learning-based image processing method, the flowchart of which is attached. Figure 1 As shown, the main steps include S101 to S104, which are detailed below:

[0029] Step S101: Extract image features of the image to be processed using a convolutional neural network, wherein the image features include edge features, color distribution features, and texture features.

[0030] By extracting image features such as edges, color distribution, and texture from the image to be processed using convolutional neural networks, it is possible to comprehensively capture low-level details and high-level semantic information of the image, providing multi-dimensional feature support for subsequent processing. Simultaneously, considering the potential inconsistencies in pixel value distribution between different images due to differences in lighting and contrast, a unified numerical range can avoid gradient oscillations caused by differences in feature dimensions during the early stages of model training, significantly accelerating convergence. Furthermore, data diversity can prevent the model from overfitting to the limited samples in the training set, improving its generalization ability to unknown images. In other words, data augmentation forces the model to learn the diversity of data, making subsequent steps (such as detail enhancement and super-resolution reconstruction) less sensitive to spatial transformations of the image, thus focusing more on extracting orientation-independent detail features (such as texture and edge continuity). Therefore, before extracting image features from the image to be processed using convolutional neural networks, some preprocessing can be performed on the image, specifically including: normalizing the image to be processed, linearly mapping pixel values ​​to the [0, 1] interval; and randomly flipping or rotating the normalized image to generate enhanced training data. Normalizing the image to be processed can be achieved by: calculating the pixel mean and standard deviation of the image to be processed; and standardizing the image to be processed based on the pixel mean and standard deviation.

[0031] In the above embodiments, the convolutional neural network can be a neural network with an alternating stacked structure of "multiple convolutional layers + multiple pooling layers". That is, a convolutional layer module composed of multiple convolutional layers and a pooling layer module composed of multiple pooling layers are alternately connected in series to gradually generate feature maps at different levels of abstraction. The convolutional layer module includes multiple convolutional layers connected in series. The size of the convolutional kernel of each convolutional layer can be variable or constant. Convolutional operations are performed. Each convolutional layer is followed by a ReLU activation function. Multiple max pooling layers are connected after the convolutional layer module as pooling layer modules. For example, convolutional layer module 1 + pooling layer module 1 + convolutional layer module 1 + pooling layer module 2 + ... + convolutional layer module n + pooling layer module n. Here, convolutional layer module 1 includes m convolutional layers and ReLU activation functions, pooling layer module 1 includes k max pooling layers, convolutional layer module 2 includes p convolutional layers and ReLU activation functions, pooling layer module 2 includes q max pooling layers, ..., and so on. As an embodiment of this application, the extraction of image features of the image to be processed by the convolutional neural network can be achieved through steps S1011 and S1012, as detailed below:

[0032] Step S1011: For the image to be processed, perform convolution operations with a kernel of k×k in multiple convolutional layers in sequence and pass them through the ReLU activation function to obtain the feature map after ReLU activation.

[0033] In convolution operations, the choice of convolution kernel size is closely related to factors such as computational efficiency, receptive field, and task adaptability. Specifically, larger convolution kernels (e.g., 5×5) increase the number of parameters and computational cost, easily leading to overfitting. Conversely, smaller convolution kernels (e.g., 1×1) are mainly used for channel dimension adjustment and cannot effectively extract spatial features. In lossless image upscaling tasks, edge continuity and local texture details are crucial. A convolution kernel of suitable size can effectively capture pixel correlations within the local neighborhood (e.g., edge direction, color gradation, etc.) while avoiding the blurring problem introduced by large convolution kernels. Based on the balance between computational efficiency and receptive field, and considering the lossless image upscaling requirements of this application, a k×k convolution kernel can specifically be 3×3.

[0034] Step S1012: The feature map after ReLU activation is downsampled through a max pooling layer with multiple pooling layers to gradually extract edge features, color distribution features and texture features at different levels.

[0035] It should be noted that if average pooling is used on the feature map after ReLU activation, high-frequency details may be weakened, resulting in blurred image edges. Other pooling methods, such as random pooling, introduce randomness, which may disrupt the consistency of image details. Therefore, in this embodiment, the feature map after ReLU activation is downsampled using a max pooling layer with multiple pooling layers. This preserves local maxima, ensuring that high-frequency information such as edges and textures are still highlighted after downsampling, while suppressing background noise.

[0036] Considering that when the derivative of the activation function has a small range, the gradient is multiplied layer by layer in backpropagation using the chain rule. If the gradient of each layer is less than 1, the gradient approaches zero after multiple multiplications. The more layers a convolutional neural network has, the more times the gradient is multiplied, and the more significant the decay becomes. This is known as vanishing gradient, where the gradient value decays exponentially with the number of layers during backpropagation in a deep neural network, leading to slow or even stagnant weight updates in deep networks. Specifically, in this application, vanishing gradient will prevent convolutional neural networks from effectively learning high-frequency details (such as edges, textures, etc.), resulting in insufficient extracted features. Subsequent multi-scale fusion and attention mechanisms will also fail to effectively enhance details due to low feature quality. Therefore, in this embodiment, between convolutional layers and pooling layers, the output feature map of the current convolutional layer and the output feature map of the previous convolutional layer can be added element-wise through a residual connection. The resulting residual feature map is then activated by the ReLU activation function and input into subsequent pooling or convolutional layers.

[0037] Step S102: Input the image features of the image to be processed into the deep learning model, and enhance the details of the image features of the image to be processed through multi-scale feature fusion and attention mechanism to generate a detail-enhanced image.

[0038] In this embodiment, the deep learning model is a specialized structure integrating multi-scale feature fusion and dual attention mechanisms. It mainly includes a feature pyramid network (or its variants), residual modules (i.e., skip connections), pooling layers, channel attention modules, and spatial attention modules, etc. The channel attention module includes the ReLU activation function and the Sigmoid function, while the spatial attention module includes convolutional layers and the Sigmoid function. To enhance the weights of key details (e.g., high-frequency edges, complex textures, etc.) and avoid detail loss caused by a single feature layer, this embodiment utilizes multi-scale feature fusion and attention mechanisms to enhance the image features of the image to be processed, generating a detail-enhanced image. Specifically, as an embodiment of this application, the image features of the image to be processed are input into the deep learning model. Multi-scale feature fusion and attention mechanisms are used to enhance the image features of the image to be processed, generating a detail-enhanced image. This can be achieved through steps S1021 to S1024, detailed as follows:

[0039] Step S1021: Input the edge features, color distribution features, and texture feature maps of the image to be processed into the pyramid structure to generate a multi-scale feature map containing low-level details and high-level semantics.

[0040] In this embodiment, the pyramid structure refers to a multi-scale feature extraction architecture, namely a Feature Pyramid Network (FPN) or a variant thereof. Its function is to upsample deep feature maps and add them element-wise to shallow feature maps, generating a pyramid feature sequence with decreasing spatial resolution but increasing semantic information. In the multi-scale feature maps generated by the pyramid structure, low-level details refer to high-resolution, low-abstraction features extracted by the shallow layers of the network (i.e., layers close to the network input), including edge sharpness, pixel-level texture, and local color gradients, such as a cat's whiskers, fabric fibers, etc. High-level semantics refer to low-resolution, high-abstraction features extracted by the deep layers of the network (i.e., layers far from the network input). These features often have a large receptive field, helping to understand the global context of the image, including object outlines and semantic regions, such as a cat's head, car tires, etc.

[0041] Step S1022: By using skip connections, the multi-scale feature map containing low-level details and high-level semantics is concatenated with the basic feature map extracted from the initial convolutional layer of the convolutional neural network to generate a fused feature map.

[0042] Specifically, the resolution of the multi-scale feature map is first adjusted by performing bilinear upsampling to align it with the resolution of the base feature map; that is, UP(f i ) = resize(f isize = (H B W B Then, a 1×1 convolution operation is performed on the base feature map to correct the number of channels in the base feature map (because the number of channels in the initial convolutional layer may not be mismatched), that is, B' = Conv 1×1 (B); Concatenate the tensor along the channel axis to obtain the fused feature map, i.e., Among them, H B and W B Here, C represents the height and width of the base feature map (in pixels), C is the number of channels, and f is the width and height of the base feature map. i Let UP(f) be the i-th layer pyramid feature map of the multi-scale feature map. i F is the resolution-aligned multi-scale feature map, B is the base feature map, B' is the channel-corrected base feature map, and F is the multi-scale feature map. mix This is the fused feature map.

[0043] Step S1023: Calculate the global average pooling value of each channel of the fused feature map in the channel dimension, thereby generating channel attention weights through a fully connected layer, and adjusting the weights of each channel based on the channel attention weights.

[0044] Specifically, the fused feature map F is calculated. mix The global average pooling value for each channel can be obtained by compressing the fused feature map in the spatial dimension to generate a channel statistical vector, i.e., the global average pooling value of the c-th channel. And output the channel descriptor Therefore, the channel attention weight w generated by the fully connected layer can be expressed as w = σ(W2δ(W1z)), where, These are the weights of the first fully connected layer, used for dimensionality reduction. δ represents the weights of the second fully connected layer, used to recover the dimension; δ is the ReLU activation function; σ is the Sigmoid function, aimed at compressing the weights to the interval [0, 1]; r is the compression ratio of the fully connected layer, used to reduce computational cost; finally, the fused feature map F... mix The c-th channel is multiplied by the corresponding attention weight, i.e. Therefore, the weights of each channel were adjusted based on the channel attention weights.

[0045] Step S1024: Calculate spatial attention weights for each pixel position of the fused feature map with adjusted channel weights, and highlight the detailed features of important regions based on the spatial attention weights.

[0046] Specifically, targeting Channel compression is performed, that is, feature information is aggregated along the channel dimension to generate a two-dimensional spatial response map. The goal is to eliminate channel dimension differences and preserve spatial distribution characteristics; then, a convolution operation is performed on M(i,j) using a convolution kernel of size k×k (e.g., k=7). conv (i,j) = Conv(M), capturing local region correlations (e.g., edge continuity, texture patches); generating a spatial weight map S using the Sigmoid function. weight (i,j)=σ(M conv (i,j)), S weight (i,j) represents the spatial attention weight at pixel position (i,j); finally, pixel-level weight adjustment F is performed on the fused feature map whose channel weights have been adjusted. out (i,j,c)=S weight (i,j)×F attn (i,j,c), where if S weight (i,j) takes a large value, for example, S weight If (i,j) = 0.9, then the detailed features of important regions can be highlighted based on spatial attention weights. It should be noted here that the so-called important regions can be pixel regions in the fused feature map after the channel weights have been adjusted where the local texture complexity is higher than a preset threshold, or pixel positions where the gradient changes drastically.

[0047] Step S103: Magnify the detail-enhancing image by interpolation and convolution operations to obtain the magnified image while preserving the image details.

[0048] As one embodiment of this application, the detailed enhancement image is magnified through interpolation and convolution operations to obtain a magnified image while preserving the image details. This can be achieved through steps S1031 to S1034, as detailed below:

[0049] Step S1031: Interpolate the image to be processed to obtain a low-resolution version of the interpolated image.

[0050] First, the image to be processed is reduced in size according to a preset target reduction ratio to obtain a low-resolution image I(i,j); then, the image is processed according to the formula... Perform bilinear interpolation and output a low-resolution interpolated image, i.e., the low-resolution version of the interpolation result I. LR (x',y'), where i and j are the pixel coordinates of image I(i,j), x=x'·s and y=y'·s are the image coordinates of image I(i,j), and s is the preset target scaling ratio.

[0051] Step S1032: Perform a transposed convolution operation on the feature map of the detail enhancement image through a transposed convolutional layer to obtain a magnified feature map with a resolution of the target size.

[0052] Specifically, for the feature map of the detail-enhanced image The transposed convolution operation can be performed on it according to the following formula to obtain a magnified feature map with a resolution of the target size:

[0053]

[0054] in, For learnable transposed convolution kernels, χ map For F detail The position mapping function determines how the input values ​​are filled into the output space.

[0055] Step S1033: Perform local detail reconstruction on the magnified feature map using a convolutional layer with an l×l kernel.

[0056] Specifically, local detail reconstruction of the magnified feature map using a convolutional layer with an l×l kernel can be achieved by: inserting a ReLU activation function and a residual block containing two convolutional layers and a skip connection between the transposed convolutional layer and the convolutional layer with an l×l kernel; passing the magnified feature map output from the transposed convolutional layer sequentially through the convolutional layer with an l×l kernel and the ReLU activation function to generate an intermediate feature map; and adding the intermediate feature map and the magnified feature map output from the transposed convolutional layer element-wise in the residual block, with the output being the result of local detail reconstruction. It should be noted that the element-wise addition of the intermediate feature map and the magnified feature map output from the transposed convolutional layer refers to adding corresponding elements of these two feature maps. Specifically, if F... mid (x,y,c) represents the intermediate feature map, then the element-wise summation can be expressed as F recon (x,y,c)=F upscaled (x,y,c)+F mid (x,y,c), only the values ​​at the same position (i,j,c) are added, and there is no calculation across positions, F recon (x,y,c) represents the result of local detail reconstruction. Thus, during the image upscaling stage, by combining interpolation and convolution operations, the resolution is improved while the deep learning model is used to locally reconstruct the features of the detail-enhanced image, reducing edge blurring or artifacts caused by simple interpolation.

[0057] Step S1034: The result of local detail reconstruction is superimposed with the interpolation result of the low-resolution version of the image to be processed.

[0058] As can be seen from steps S1031 to S1034 of the above embodiment, by designing the magnification and detail enhancement steps in series, it is ensured that the magnified image can still maintain the authenticity of details when the resolution is increased.

[0059] To suppress redundant features and preserve effective details, feature calibration can be performed after the image features of the image to be processed are enhanced through multi-scale feature fusion and attention mechanisms, and before the enhanced image is magnified through interpolation and convolution operations. That is, the enhanced image is multiplied element-wise with the edge features extracted by the convolutional neural network to obtain a new image, which is then used as the object for magnification through interpolation and convolution operations.

[0060] Step S104: Perform post-processing optimization on the magnified image.

[0061] To further improve image quality, in this embodiment, the magnified image can be post-processed and optimized, mainly including denoising, sharpening, and color correction, etc. Denoising the magnified image can be performed by: performing wavelet transform on the magnified image to separate high-frequency noise components from low-frequency image components; performing threshold filtering on the high-frequency components to preserve edge information while suppressing noise; and performing inverse wavelet transform on the processed high-frequency and low-frequency components to generate a denoised image. Sharpening the magnified image can be performed by: performing Laplacian filtering on the denoised image to extract edge gradient information; and weighted superposition of the edge gradient information and the denoised image to enhance detail contrast. Color correction of the magnified image can be performed by: converting the sharpened image generated by weighted superposition from the RGB color space to the LAB color space; performing histogram equalization on the luminance channel (L) and Gaussian smoothing filtering on the chrominance channels (A, B); and converting the processed LAB image back to the RGB color space to generate a color-corrected image. As can be seen from the above embodiments, by performing a series of post-processing operations such as denoising, sharpening, and color correction on the magnified image, the noise introduced during the image magnification process is gradually eliminated, the edge contrast is enhanced, and the color distribution is adjusted, ultimately achieving a balance between detail preservation and visual quality.

[0062] From the above appendix Figure 1 As can be seen from the example of the deep learning-based image processing method, on the one hand, by extracting edge features, color distribution features, and texture features of the image to be processed through convolutional neural networks, it can comprehensively capture the low-level details and high-level semantic information of the image, providing multi-dimensional feature support for subsequent processing; on the other hand, by combining multi-scale feature fusion and attention mechanisms to enhance the weight of key details (such as high-frequency edges and complex textures) of the image features to be processed, the problem of detail loss caused by a single feature layer is avoided. Then, the enhanced image is magnified through interpolation and convolution operations. That is, by designing the magnification and detail enhancement steps in a series, it is ensured that the magnified image can still maintain the authenticity of details when the resolution is increased; thirdly, the noise, edge blurring, and color distortion problems that may exist in the magnified image are optimized in stages, ultimately achieving an overall improvement in visual quality.

[0063] Please see the appendix Figure 2 This application provides an image processing device based on deep learning, which may include an extraction module 201, a generation module 202, a magnification module 203, and an optimization module 204, as detailed below:

[0064] The extraction module 201 is used to extract image features of the image to be processed through a convolutional neural network, wherein the image features include edge features, color distribution features and texture features;

[0065] The generation module 202 is used to input the image features of the image to be processed into the deep learning model, and to enhance the details of the image features of the image to be processed through multi-scale feature fusion and attention mechanism to generate a detail-enhanced image;

[0066] The magnification module 203 is used to magnify the detail-enhancing image through interpolation and convolution operations to obtain the magnified image while preserving the image details;

[0067] The optimization module 204 is used to perform post-processing optimization on the magnified image.

[0068] From the above appendix Figure 2 As can be seen from the example of the deep learning-based image processing device, on the one hand, by extracting edge features, color distribution features, and texture features of the image to be processed through a convolutional neural network, it can comprehensively capture the low-level details and high-level semantic information of the image, providing multi-dimensional feature support for subsequent processing; on the other hand, by combining multi-scale feature fusion and attention mechanisms to enhance the weight of key details (such as high-frequency edges and complex textures) of the image features to be processed, the problem of detail loss caused by a single feature layer is avoided. Then, the enhanced image is magnified through interpolation and convolution operations. That is, by designing the magnification and detail enhancement steps in a series, it is ensured that the magnified image can still maintain the authenticity of details when the resolution is increased; thirdly, the noise, edge blurring, and color distortion problems that may exist in the magnified image are optimized in stages, ultimately achieving an overall improvement in visual quality.

[0069] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. For example... Figure 3 As shown, the electronic device 3 in this embodiment mainly includes: a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30, such as a program for a deep learning-based image processing method. When the processor 30 executes the computer program 32, it implements the steps described in the deep learning-based image processing method embodiment, for example... Figure 1 The steps S101 to S104 are shown. Alternatively, when the processor 30 executes the computer program 32, it implements the functions of each module / unit in the above-described device embodiments, for example... Figure 2 The functions of the extraction module 201, generation module 202, amplification module 203, and optimization module 204 are shown.

[0070] For example, the computer program 32 of the deep learning-based image processing method mainly includes: extracting image features of the image to be processed through a convolutional neural network, wherein the image features include edge features, color distribution features, and texture features; inputting the image features of the image to be processed into a deep learning model, performing detail enhancement on the image features of the image to be processed through multi-scale feature fusion and attention mechanisms to generate a detail-enhanced image; enlarging the detail-enhanced image through interpolation and convolution operations to obtain an enlarged image while retaining the details of the image; and performing post-processing optimization on the enlarged image. The computer program 32 can be divided into one or more modules / units, one or more modules / units are stored in memory 31 and executed by processor 30 to complete this application. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions, which are used to describe the execution process of the computer program 32 in the electronic device 3. For example, computer program 32 can be divided into the functions of extraction module 201, generation module 202, magnification module 203, and optimization module 204 (modules in the virtual device). The specific functions of each module are as follows: extraction module 201 is used to extract image features of the image to be processed through a convolutional neural network, wherein the image features include edge features, color distribution features, and texture features; generation module 202 is used to input the image features of the image to be processed into a deep learning model, and to enhance the details of the image features of the image to be processed through multi-scale feature fusion and attention mechanism to generate a detail-enhanced image; magnification module 203 is used to magnify the detail-enhanced image through interpolation and convolution operations to obtain a magnified image while retaining the details of the image; optimization module 204 is used to perform post-processing optimization on the magnified image.

[0071] Electronic device 3 may include, but is not limited to, processor 30 and memory 31. Those skilled in the art will understand that... Figure 3 This is merely an example of electronic device 3 and does not constitute a limitation on electronic device 3. It may include more or fewer components than shown, or combine certain components, or different components. For example, electronic devices may also include input / output devices, network access devices, buses, etc.

[0072] The processor 30 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0073] The memory 31 can be an internal storage unit of the electronic device 3, such as a hard disk or RAM. The memory 31 can also be an external storage device of the electronic device 3, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, the memory 31 can include both internal and external storage units of the electronic device 3. The memory 31 is used to store computer programs and other programs and data required by the electronic device. The memory 31 can also be used to temporarily store data that has been output or will be output.

[0074] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed. That is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above-described device can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0075] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0076] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0077] In the embodiments provided in this application, it should be understood that the disclosed apparatus / device and method can be implemented in other ways. For example, the apparatus / device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0078] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0079] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0080] If integrated modules / units are implemented as software functional units and sold or used as independent products, they can be stored in a storage medium. Based on this understanding, all or part of the processes in the above-described embodiments can also be implemented by a computer program instructing related hardware. The computer program for the deep learning-based image processing method can be stored in a storage medium. When executed by a processor, this computer program can implement the steps of the various method embodiments described above, namely: extracting image features of the image to be processed through a convolutional neural network, wherein the image features include edge features, color distribution features, and texture features; inputting the image features of the image to be processed into a deep learning model, performing detail enhancement on the image features of the image to be processed through multi-scale feature fusion and attention mechanisms to generate a detail-enhanced image; enlarging the detail-enhanced image through interpolation and convolution operations to obtain an enlarged image while retaining image details; and performing post-processing optimization on the enlarged image. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. Storage media can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the contents of storage media can be appropriately added to or removed according to the requirements of legislation and patent practice in a jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, storage media may not include electrical carrier signals and telecommunication signals.

[0081] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. These modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application. The specific embodiments described above further illustrate the purpose, technical solutions, and beneficial effects of this application. It should be understood that the above descriptions are merely specific embodiments of this application and are not intended to limit the protection scope of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this invention.

Claims

1. A deep learning-based image processing method, characterized in that, The method includes: Image features of the image to be processed are extracted using a convolutional neural network, including edge features, color distribution features, and texture features. The image features of the image to be processed are input into a deep learning model, and the image features of the image to be processed are enhanced in detail through multi-scale feature fusion and attention mechanism to generate a detail-enhanced image; The detail-enhancing image is magnified by interpolation and convolution operations to obtain a magnified image while preserving the image details. The magnified image is then post-processed and optimized.

2. The image processing method based on deep learning according to claim 1, characterized in that, The step of extracting image features from the image to be processed using a convolutional neural network includes: For the image to be processed, a convolution operation with a kernel of k×k is performed sequentially in multiple convolutional layers and then passed through the ReLU activation function to obtain a feature map after ReLU activation; The feature map after ReLU activation is downsampled through a max pooling layer with multiple pooling layers to gradually extract edge features, color distribution features and texture features at different levels.

3. The image processing method based on deep learning according to claim 2, characterized in that, Between the convolutional layer and the pooling layer, the output feature map of the current convolutional layer is added element-wise with the output feature map of the previous convolutional layer through residual connection, and the residual feature map obtained by element-wise addition is activated by the ReLU activation function and then input into the subsequent pooling layer or convolutional layer.

4. The image processing method based on deep learning according to claim 1, characterized in that, The step of inputting the image features of the image to be processed into a deep learning model, and performing detail enhancement on the image features of the image to be processed through multi-scale feature fusion and attention mechanisms to generate a detail-enhanced image includes: The edge features, color distribution features, and texture feature maps are input into the pyramid structure to generate a multi-scale feature map containing low-level details and high-level semantics. By using skip connections, the multi-scale feature map is concatenated with the basic feature map extracted by the initial convolutional layer of the convolutional neural network to generate a fused feature map. In the channel dimension, the global average pooling value of each channel of the fused feature map is calculated, thereby generating channel attention weights through a fully connected layer, and adjusting the weights of each channel based on the channel attention weights; For each pixel position of the fused feature map with adjusted channel weights, spatial attention weights are calculated, and detailed features of important regions are highlighted based on these spatial attention weights.

5. The image processing method based on deep learning according to claim 1, characterized in that, The process of magnifying the detail-enhanced image through interpolation and convolution operations to obtain a magnified image while preserving its details includes: Interpolate the image to be processed to obtain a low-resolution version of the interpolated image; The feature map of the detail enhancement image is transposed and convolved by a transposed convolutional layer to obtain a magnified feature map with a resolution of the target size. The magnified feature map is reconstructed using a convolutional layer with an l×l kernel; The result of the local detail reconstruction is superimposed with the interpolation result of the low-resolution version of the image to be processed using residual superposition.

6. The image processing method based on deep learning according to claim 5, characterized in that, The process of reconstructing local details in the magnified feature map using a convolutional layer with an l×l kernel includes: Between the transposed convolutional layer and the convolutional layer with a kernel size of l×l, a ReLU activation function and a residual block containing two convolutional layers and a skip connection are inserted. The magnified feature map is passed sequentially through a convolutional layer with an l×l kernel and a ReLU activation function to generate an intermediate feature map; The residual block adds the intermediate feature map to the magnified feature map element by element, and the output result is used as the result of the local detail reconstruction.

7. The method according to claim 1, characterized in that, After enhancing the image features of the image to be processed through multi-scale feature fusion and attention mechanisms, and before magnifying the enhanced image through interpolation and convolution operations, the method further includes: The detail-enhanced image is multiplied element-wise with the edge features to obtain a new image, which is then used as the object for magnification through interpolation and convolution operations.

8. An image processing device based on deep learning, characterized in that, The device includes: The extraction module is used to extract image features of the image to be processed through a convolutional neural network. The image features include edge features, color distribution features, and texture features. The generation module is used to input the image features of the image to be processed into a deep learning model, and to enhance the details of the image features of the image to be processed through multi-scale feature fusion and attention mechanism to generate a detail-enhanced image; The magnification module is used to magnify the detail-enhancing image through interpolation and convolution operations to obtain a magnified image while preserving the image details; An optimization module is used to perform post-processing optimization on the magnified image.

9. An electronic device, the device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.

10. A storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.