Image sharpening method and device based on neural network, equipment and storage medium

By simplifying and quantifying the neural network model, the problem of difficult application of existing defuzzy methods in edge devices is solved, and the model is simplified and real-time reasoning capabilities are improved.

CN120070250APending Publication Date: 2025-05-30HAIWEI ZHIZAO TECH (WUHAN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510050605.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing defuzzing method based on convolutional neural networks is too complex, resulting in too high computational performance and storage space requirements, making it difficult to effectively apply in edge devices.

Method used

A simplified neural network model is adopted and quantified through a deep learning inference engine to obtain a preset model for clear image processing of fuzzy images.

Benefits of technology

Simplifies model complexity, makes it suitable for edge deployment, and improves the model's real-time inference capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070250A_ABST
    Figure CN120070250A_ABST
Patent Text Reader

Abstract

The invention discloses an image sharpening method and device based on a neural network, equipment and a storage medium, and relates to the technical field of image processing, and the method comprises the steps: obtaining an image data set; constructing an initial model based on a preset network; based on the image data set, training the initial model to obtain a reference model; quantifying the reference model through a deep learning inference engine to obtain a preset model; and performing image sharpening processing on the fuzzy image through the preset model to obtain a clear image. According to the method, the neural network model is simplified, the model is quantified based on the deep learning inference engine, the complexity of the model is simplified, the method is suitable for edge deployment, and the real-time inference capability of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and particularly to an image sharpening method, device, equipment and storage medium based on neural network. Background Art

[0002] At present, the deblurring methods based on convolutional neural networks are mainly applied to defocus deblurring and motion deblurring. Through the training of the data set, the neural network can learn a large number of parameters for end-to-end processing of images. This type of method can effectively improve the clarity of images and videos, and can also be used for detail enhancement and other aspects. Although in the academic community, the deblurring network algorithm has achieved very excellent results based on various new network structures, the stacking of a large number of structures has led to an extremely large model, resulting in the need for powerful computing performance and storage space for model inference. Limited by the limited computing power and storage performance of edge devices, it is very difficult to apply various network models with excellent performance in the papers in practice. At the same time, the model parameters are fixed and cannot meet the deblurring requirements of different scales. Summary of the Invention

[0003] The main purpose of this application is to provide an image sharpening method, device, equipment and storage medium based on neural network, aiming to solve the technical problems of how to simplify the model complexity, be suitable for edge deployment and improve the real-time inference ability of the model.

[0004] To achieve the above object, this application proposes an image sharpening method based on neural network, and the method includes:

[0005] Obtain an image data set;

[0006] Build an initial model based on a preset network;

[0007] Train the initial model based on the image data set to obtain a reference model;

[0008] Quantize the reference model through a deep learning inference engine to obtain a preset model;

[0009] Perform image sharpening processing on the blurred image through the preset model to obtain a sharp image.

[0010] In an embodiment, the step of building an initial model based on a preset network includes:

[0011] Build a first basic model based on a preset network;

[0012] Set a U-shaped network structure with a preset number of layers and a preset number of downsamplings in the first basic model to obtain a second basic model;

[0013] The retention layer, encoding layer, and decoding layer of the second basic model are constructed by a convolutional layer, an activation function layer, a class residual structure, a batch normalization layer, and a pixel shuffle structure to obtain an initial model.

[0014] In one embodiment, the step of training the initial model based on the image dataset to obtain a reference model includes:

[0015] Cut the data of the image dataset through the initial model to obtain cut data;

[0016] Perform forward propagation training on the cut data to obtain forward propagation data;

[0017] Based on the forward propagation data, train the initial model through a preset optimizer to obtain a reference model.

[0018] In one embodiment, the step of training the initial model through a preset optimizer based on the forward propagation data to obtain a reference model includes:

[0019] Calculate the loss function of the model based on the forward propagation data;

[0020] Calculate the parameter gradient of the initial model through the loss function;

[0021] Based on the parameter gradient, train the initial model through a preset optimizer to obtain a reference model.

[0022] In one embodiment, the step of quantifying the reference model through a deep learning inference engine to obtain a preset model includes:

[0023] Convert the reference model into an open neural network exchange format to obtain a format conversion model;

[0024] Convert the format conversion model into a tensor real-time processing format through a deep learning inference engine to obtain a preset model.

[0025] In one embodiment, the step of performing image sharpening processing on a blurred image through the preset model to obtain a sharp image includes:

[0026] Input the blurred image into the preset model;

[0027] Perform downsampling and upsampling on the blurred image through the preset model to obtain sampling data;

[0028] Perform post-processing on the sampling data to obtain a sharp image.

[0029] In one embodiment, the step of obtaining the image dataset includes:

[0030] Obtain a real blurred data set;

[0031] Sharpen the clear images in the real blurred data set to obtain sharpened images;

[0032] Use the blurred images and the sharpened images in the real blurred data set as an image data set.

[0033] In addition, to achieve the above object, the present application also proposes an image sharpening device based on a neural network, and the device includes:

[0034] A data acquisition module, configured to acquire an image data set;

[0035] A model construction module, configured to construct an initial model based on a preset network;

[0036] A model training module, configured to train the initial model based on the image data set to obtain a reference model;

[0037] A model quantization module, configured to quantize the reference model through a deep learning inference engine to obtain a preset model;

[0038] An image processing module, configured to perform image sharpening processing on a blurred image through the preset model to obtain a clear image.

[0039] In addition, to achieve the above object, the present application also proposes an image sharpening device based on a neural network, and the device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the image sharpening method based on a neural network as described above.

[0040] In addition, to achieve the above object, the present application also proposes a storage medium, and the storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the image sharpening method based on a neural network as described above are implemented.

[0041] In addition, to achieve the above object, the present application also provides a computer program product, and the computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the image sharpening method based on a neural network as described above are implemented.

[0042] One or more technical solutions proposed by the present application have at least the following technical effects:

[0043] By adopting a simplified neural network model and quantizing the model based on a deep learning inference engine, the technical problem that the network model structure of the deblurring network algorithm is stacked, resulting in an extremely large model and being unable to be actually applied due to the limited computing power and storage performance of edge devices, is solved. Compared with the prior art, the complexity of the model is simplified, it is suitable for edge deployment, and the real-time inference ability of the model is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The drawings herein are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0046] Figure 1 It is a schematic flowchart provided for Embodiment 1 of the image clarity improvement method based on a neural network of the present application;

[0047] Figure 2 It is a schematic flowchart provided for Embodiment 2 of the image clarity improvement method based on a neural network of the present application;

[0048] Figure 3 It is a schematic diagram of the model details of the image clarity improvement method based on a neural network provided for Embodiment 2 of the present application;

[0049] Figure 4 It is a corresponding input-output diagram of the RELU6 function of the image clarity improvement method based on a neural network provided for Embodiment 2 of the present application;

[0050] Figure 5 It is a schematic diagram of the overall model structure without hyperparameters of the image clarity improvement method based on a neural network provided for Embodiment 2 of the present application;

[0051] Figure 6 It is a schematic diagram of the overall model structure with hyperparameters of the image clarity improvement method based on a neural network provided for Embodiment 2 of the present application;

[0052] Figure 7 It is a schematic flowchart provided for Embodiment 3 of the image clarity improvement method based on a neural network of the present application;

[0053] Figure 8 It is a schematic flowchart provided for Embodiment 4 of the image clarity improvement method based on a neural network of the present application;

[0054] Figure 9Schematic flowchart provided for Embodiment 5 of the image clarity improvement method based on neural network in this application;

[0055] Figure 10 Schematic diagram showing the effect of image clarity improvement processing of the image clarity improvement method based on neural network provided for Embodiment 5 of this application;

[0056] Figure 11 Schematic diagram showing the effect of image clarity improvement processing of the image clarity improvement method based on neural network provided for Embodiment 5 of this application;

[0057] Figure 12 Schematic flowchart provided for Embodiment 6 of the image clarity improvement method based on neural network in this application;

[0058] Figure 13 Schematic diagram of the module structure of the image clarity improvement device based on neural network according to an embodiment of this application;

[0059] Figure 14 Schematic diagram of the device structure of the hardware operating environment involved in the image clarity improvement method based on neural network in an embodiment of this application.

[0060] The implementation, functional features, and advantages of this application will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners

[0061] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application.

[0062] To better understand the technical solutions of this application, the following will be described in detail with reference to the accompanying drawings of the specification and specific implementation manners.

[0063] The main solution of the embodiment of this application is: obtaining an image data set; constructing an initial model based on a preset network; training the initial model based on the image data set to obtain a reference model; quantifying the reference model through a deep learning inference engine to obtain a preset model; performing image clarity improvement processing on a blurred image through the preset model to obtain a clear image.

[0064] In this embodiment, for the convenience of description, the following will be described with the internal actuator of the image clarity improvement system based on neural network as the execution subject.

[0065] Due to the stacking of the network model structures of the existing deblurring network algorithms, the model is extremely large, and it cannot be actually applied due to the limited computing power and storage performance of edge devices.

[0066] The present application provides a solution, which adopts a simplified neural network model and quantizes the model based on a deep learning inference engine, achieving the simplification of the model complexity, suitability for edge deployment, and improvement of the model's real-time inference ability.

[0067] As can be seen from the above embodiments, since the present application adopts a simplified neural network model and quantizes the model based on a deep learning inference engine, it solves the technical problem that the network model structure of the deblurring network algorithm is stacked, resulting in an extremely large model, and it cannot be actually applied due to the limited computing power and storage performance of edge devices, achieving the simplification of the model complexity, suitability for edge deployment, and improvement of the model's real-time inference ability.

[0068] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of implementing the above functions. Hereinafter, taking the internal actuator of an image clarity system based on a neural network as an example, this embodiment and the following embodiments will be described.

[0069] Based on this, the embodiments of the present application provide an image clarity method based on a neural network, referring to Figure 1 , Figure 1 is a schematic flowchart of the first embodiment of the image clarity method based on a neural network of the present application.

[0070] In this embodiment, the image clarity method based on a neural network includes steps S10 to S40:

[0071] Step S10, obtaining an image data set.

[0072] It should be noted that the image data set is an image data set obtained after processing the real blur data set REDS. The real blur data set REDS contains clear images and blurred images. In order to make the images generated by the model clearer, the clear images are sharpened and combined with the blurred images to form the image data set.

[0073] Additionally, it should be noted that the Real Blur Dataset (REDS) is a dataset for computer vision research. It focuses on blurring situations in real-world scenarios and contains a large amount of video and image data, which are captured by professional equipment in various real-world scenarios such as natural landscapes, urban streets, and indoor environments. The dataset includes high-resolution and low-resolution video sequences, as well as corresponding blurred and sharp image pairs. The image pairs consist of sharp images and blurred images. The blurring of the blurred images is mainly caused by defocus and jitter during shooting, as well as JPEG compression losses during video transmission. It covers various situations that cause blurring, such as camera motion blur and object motion blur. This dataset plays an important role in the research, model training, and evaluation of computer vision tasks such as video super-resolution, depth estimation, image and video deblurring, and also provides other relevant data such as the camera response function.

[0074] Step S20: Construct an initial model based on a preset network.

[0075] It should be noted that the preset network is the Unet network. The name of the Unet network comes from its U-shaped structure, which mainly consists of a contracting path and an expanding path. The contracting path is used for downsampling and can extract the features of the image. It usually includes multiple convolutional layers and pooling layers. As the downsampling progresses, the size of the feature map gradually decreases, but the number of channels gradually increases, which helps to capture high-level semantic information in the image. The expanding path is used for upsampling to restore the feature map to the size of the original image. It is symmetrical to the contracting path and includes transposed convolutional layers or upsampling layers and convolutional layers. The size of the feature map is gradually increased through upsampling, while the number of channels is reduced.

[0076] The structure of the initial model is modified based on the Unet network. The model includes a branch for retaining the features of the original image and a Unet-like structure. The model first performs downsampling to reduce the size of the feature map and increase the number of channels, and then gradually scales back to the original size through the decoding layer.

[0077] Step S30: Train the initial model based on the image dataset to obtain a reference model.

[0078] Input the image dataset into the initial model. The initial model will use the data in the image dataset to train the model to optimize the model's parameters, etc. Model training is a process in which the model learns the internal laws and features of the data from the initial state. During the training process, based on the loss function, calculate the difference between the model's prediction result and the true result, and then through the backpropagation algorithm, backpropagate this error signal from the output layer to each hidden layer, and then adjust the parameters such as the weights and biases of the model. This process is like a sculptor constantly carving the raw material according to the expected shape of the work, so that the parameters of the model gradually converge to the optimal value or a local optimal value, thereby improving the accuracy of the model. A good training process will enable the model not only to remember the training data, but also to make reasonable predictions for unseen data. By adopting appropriate data augmentation techniques during the training process, such as rotation, flipping, adding noise, etc., and using a validation set to monitor and adjust the model, overfitting of the model can be prevented. The model trained in this way can still accurately complete tasks in various actual scenarios, such as under different lighting conditions, shooting angles, object deformations, etc.

[0079] Step S40, quantize the reference model through a deep learning inference engine to obtain a preset model.

[0080] It should be noted that the deep learning inference engine is TensorRT of NVIDIA. TensorRT is a high-performance deep learning inference engine launched by NVIDIA. It can optimize deep learning models, such as performing layer fusion, fusing convolutional layers, bias layers, activation function layers, etc., reducing the amount of computation and the number of memory accesses, thereby improving the inference speed. In TensorRT, quantization is an important optimization method, which can reduce the model size and accelerate the inference process, especially on hardware platforms such as GPUs or DPUs. For example, converting from FP32 (single-precision floating-point number) to FP16 (half-precision floating-point number) or even INT8 (8-bit integer) can further accelerate the inference process and reduce memory occupancy. It supports the import of models trained by various deep learning frameworks, such as TensorFlow, PyTorch, etc., facilitating developers to quickly deploy existing models to different hardware platforms for efficient inference. At the same time, TensorRT is deeply optimized for NVIDIA's GPUs, giving full play to the parallel computing power of GPUs. Utilizing its large number of computing cores and high-bandwidth memory, it can achieve extremely fast inference speeds, especially when dealing with large-scale data and complex models. In practical applications, TensorRT is widely used in fields such as intelligent security, autonomous driving, and industrial inspection. For example, in an intelligent security system, it quickly processes the image data transmitted by surveillance cameras to determine whether there are abnormal situations; in an autonomous driving vehicle, it efficiently processes the data of vehicle-mounted sensors and cameras to achieve real-time environmental perception and decision-making.

[0081] Step S50: Perform image sharpening processing on the blurred image through the preset model to obtain a sharp image.

[0082] When performing image sharpening processing on a blurred image through a model to obtain a sharp image, first perform the same preprocessing operations on the blurred image to be processed to make it meet the requirements of the model input, and then input it into the trained model. The model will analyze and process the blurred features in the blurred image based on the previously learned features and mapping relationships, gradually reconstruct the image through the decoder part, and restore the details, textures, etc. that may have been lost originally, so as to output a relatively sharp image. Finally, further post-processing can be performed on the output image, such as appropriate sharpening, contrast adjustment, etc., to further improve the sharpness and visual effect of the image, and finally obtain a sharp image that meets the expectations, which can be used in many actual application scenarios such as medical image diagnosis, remote sensing image analysis, security monitoring image viewing, etc.

[0083] This embodiment provides an image sharpening method based on a neural network. By adopting a simplified neural network model and quantizing the model based on a deep learning inference engine, it solves the technical problem that the network model structure of the deblurring network algorithm is stacked, resulting in an extremely large model, and it cannot be actually applied due to the limited computing power and storage performance of edge devices, realizing the simplification of the model complexity, suitability for edge deployment, and improvement of the real-time inference ability of the model.

[0084] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 2 , step S20 includes steps S21 to S23:

[0085] Step S21: Construct a first basic model based on a preset network.

[0086] Build the first basic model based on the Unet network. First, construct the overall architecture, which is in a U shape, including a contracting path (encoder), an expanding path (decoder), and skip connections between them. The contracting path receives image data from the input layer, extracts local features through multiple convolutional layers with 3×3 convolutional kernels, combined with batch normalization and ReLU activation functions, and then downsamples through a 2×2 max pooling layer. This process is repeated continuously to make the feature map smaller while increasing high-level semantic information. The expanding path uses upsampling methods such as transposed convolution or bilinear interpolation to restore the size of the feature map. After each upsampling, it is refined and adjusted using a convolutional layer with batch normalization and ReLU activation functions, and the feature maps of the corresponding layers in the contracting path are fused with the help of skip connections. Finally, an output layer is set, and the number of its neurons is determined according to the specific task. For example, in an image segmentation task, it is related to the number of target categories and is used to output the probability of the pixel belonging to each category.

[0087] Step S22: Set a U-shaped network structure with a preset number of layers and a preset number of downsamplings in the first basic model to obtain a second basic model.

[0088] It should be noted that the U-shaped network structure with a preset number of layers is a 5-layer Unet structure. The preset number of downsamplings is 4 times (downsamplings in encoder layers 0, 1, 2, and 3).

[0089] For the first basic model, when the downsampling depth of the Unet network is too large, a series of problems will occur. On the one hand, subject to the requirement of model lightweight, as the downsampling goes deeper, the number of model parameters and the amount of computation will gradually increase, which may make it difficult to deploy and run in some resource-constrained environments. On the other hand, since the number of channels in the deep layers of the model is relatively low, it means that after deep downsampling, the information carried by the feature maps may become relatively less, thus affecting the effective extraction and expression of image features by the model. In this case, after model training, there is an easy phenomenon of local non-processing, that is, the image features in some areas cannot be fully processed and analyzed, thus affecting the final output result. To solve this problem, the number of downsampling times of the model can be controlled at 4 times, corresponding to the encoder0, 1, 2, and 3 layers respectively. By restricting the number of downsampling times, the number of model parameters and the amount of computation can be reduced to a certain extent, and at the same time, the problem of information loss caused by the relatively low number of channels in the deep layers can be avoided. Such a design not only meets the requirements of model lightweight but also can ensure the effective extraction of image features by the model to a certain extent. At the same time, in order to enable the model to have a deeper expressive ability, a 5-layer Unet structure is still adopted. However, for the last encoder4 layer, no downsampling operation is performed. The advantage of doing this is that while maintaining the overall network depth, the negative impact brought by excessive downsampling is avoided. In this way, the model can use more layers to extract and fuse image features without increasing too much computation, thereby improving the performance and accuracy of the model. The complexity of the model is reduced, and thus it can be applied to the limited computing power and storage performance of edge devices for edge deployment.

[0090] Step S23, construct the retention layer, encoding layer, and decoding layer of the second basic model through a convolutional layer, activation function layer, class residual structure, batch normalization layer, and pixel shuffle structure to obtain an initial model.

[0091] Set a convolutional layer with a stride of 2 for the retention layer of the second basic model for downsampling to preserve the original image information, followed by an activation function layer (RELU layer) for data constraint. The encoding layer (encoder layer) includes a class residual structure little_res_block, which is a streamlined residual network, and at the same time, a batch normalization layer (bn layer) and an activation function layer (RELU layer) are added. The convolutional layer conv1 in it is used to adjust the number of channels, and the stride of the convolutional layer conv2 becomes 2 for downsampling. The structure of the decoding layer (decoder layer) is similar to that of the encoding layer and also includes a convolutional layer, an activation function layer (RELU layer), and a residual network little_res_block with a class residual structure. The difference is that the decoding layer does not use convolution to implement the upsampling of the feature layer but uses a pixel shuffle structure (pixelSuffle structure) for upsampling.

[0092] A schematic diagram of the model details of the initial model is as follows Figure 3 shown, where keep (keep layer) includes conv_s2_p1 (convolution layer) and RELU (activation function layer); encoder (encoding layer) includes conv1_s1 (convolution layer), little_res_block (residual-like structure), conv2_s2_p1 (convolution layer), bn (batch normalization layer), and RELU (activation function layer), where little_res_block (residual-like structure) includes conv1_s1 (convolution layer), RELU (activation function layer), conv2_s1 (convolution layer), bn (batch normalization layer), conv3_s1 (convolution layer), add (addition), and conv4_s1 (convolution layer); decoder (decoding layer) includes conv1_s1 (convolution layer), RELU (activation function layer), pixelSuffle (pixel shuffle structure), and little_res_block (residual-like structure).

[0093] When constructing the model details, the traditional residual network structure was streamlined to a certain extent to obtain the residual-like structure little_res_block to adapt to this image processing task. At the same time, most of the bn layers were removed because the normalization effect of the bn layer would seriously affect the expression of the image processing model, but completely removing the bn layer would make the parameter distribution of the model unstable and affect model quantization.

[0094] The RELU function selects the RELU6 function. The RELU6 function has a truncation characteristic, which is mainly reflected in setting an upper limit on the output value, helping to limit the output range of neurons and making the output of the network more stable and predictable. The RELU6 function will control the parameter between 0 and 6 to avoid errors due to an overly large range during model quantization. The expression of the RELU6 function can be represented by the following formula

[0095] ReLU6 = min(6, max(0, x))

[0096] where min() represents returning the minimum value in a set of numbers, and max() represents returning the maximum value in a set of numbers.

[0097] The input-output correspondence diagram of the RELU6 function is as follows Figure 4 shown, where the horizontal axis represents the input value of the function, and the vertical axis represents the output value of the function. When the input value is less than or equal to 0, the output value is 0. When the input value is between 0 and 6, the output value is equal to the input value. When the input value is greater than or equal to 6, the output value is 6.

[0098] After adjusting the retention layer, encoding layer, and decoding layer of the second basic model, the obtained model is the initial model. The adjusted initial model can use hyperparameters to adjust the model effect. For the initial model, the overall structural schematic diagram of the model without hyperparameters is as shown in Figure 5 Figure [0000220], including: input (input layer), keep (retention layer), intro (input preprocessing layer), encoder (encoding layer), torch.cat (tensor concatenation function), middle (middle layer), decoder (decoding layer), ending (ending layer), output (output layer). The overall structural schematic diagram of the model with hyperparameters is as shown in Figure 6 Figure [0000221], including: input (input layer), keep (retention layer), intro (input preprocessing layer), encoder (encoding layer), torch.cat (tensor concatenation function), middle (middle layer), decoder (decoding layer), ending (ending layer), output (output layer), where A is the input hyperparameter, with a range of 1 - 2. 1 represents a slight effect, and 2 represents the strongest effect. For example, 101 represents a hyperparameter with a poor image sharpening effect, and 200 represents a hyperparameter with a good image sharpening effect.

[0099] This embodiment provides an image sharpening method based on a neural network. A first basic model is constructed based on a preset network; a U-shaped network structure with a preset number of layers and a preset number of downsamplings are set in the first basic model to obtain a second basic model; the retention layer, encoding layer, and decoding layer of the second basic model are constructed through a convolutional layer, activation function layer, class residual structure, batch normalization layer, and pixel shuffle structure to obtain an initial model, which simplifies the model complexity and is suitable for edge deployment.

[0100] Based on the first embodiment of the present application, in the third embodiment of the present application, the same or similar content as in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 7 Figure [0000226], step S30 includes steps S31 - S33:

[0101] Step S31, the data in the image dataset is cut through the initial model to obtain cut data.

[0102] During model training, it is necessary to preprocess the image dataset by cutting the data into 256x256 blocks to accelerate training. After the data is cut, the amount of computation per single calculation will decrease. When the model processes these small data blocks, the computational resources and time required for operations such as feature extraction and gradient calculation are less compared to directly processing the entire large-scale data. More training rounds can be completed per unit time, promoting the acceleration of the training process. At the same time, it is convenient to utilize multiple computing units for parallel computing. Different computing units can independently process different data blocks, and operations such as forward propagation and backward propagation can be carried out simultaneously, giving full play to the parallel computing ability of the hardware, improving the training efficiency, and accelerating the model convergence speed. Additionally, it can optimize the memory occupancy situation. Loading and training large-sized data easily causes memory tension or even overflow. After cutting into small blocks, it can be flexibly loaded and processed in batches according to the memory capacity, avoiding training interruption caused by insufficient memory and ensuring the continuous and efficient progress of training.

[0103] Step S32: Perform forward propagation training on the cut data to obtain forward propagation data.

[0104] The cut data blocks are sequentially input into the model. Based on the network structure set inside it, such as the neurons and convolutional kernels in each layer, the model will perform layer-by-layer computational processing on the input data blocks according to the corresponding operation rules. Starting from the input layer, the data undergoes continuous linear transformation and non-linear transformation of activation functions and other operations through each hidden layer, and finally, the corresponding output results are obtained at the output layer. These output results are the forward propagation data. Through such a forward propagation process, the model can initially extract and map the input data features. Subsequently, the loss value can be calculated based on the obtained forward propagation data, and then the model parameters can be adjusted through backpropagation, so that the model is continuously optimized and develops in the direction of more accurately processing data and achieving the training goal.

[0105] Step S33: Based on the forward propagation data, train the initial model through a preset optimizer to obtain a reference model.

[0106] It should be noted that the preset optimizer is the AdaW (Adaptive Weight Decay) optimizer. The AdaW optimizer is an adaptive optimization algorithm that can dynamically adjust the learning rate of each parameter according to factors such as the importance of the parameters during the training process. Compared with some traditional optimizers, it can more effectively handle the different requirements of different parameters during optimization, enabling the model parameters to be updated more reasonably towards the optimal solution, which helps to improve the training efficiency and the final performance of the model.

[0107] In model training, the AdaW optimizer is selected and the initial learning rate is set to 0.001. The cosine annealing function is used to dynamically adjust the learning rate. As the training progresses, the learning rate will gradually change according to the variation law of the cosine function, enabling the model to explore the parameter space with an appropriate step size in the initial stage and then finely adjust the parameters slowly to avoid situations such as falling into local optimal solutions. At the same time, the maximum number of training times is set to 400,000 times, allowing the model to continuously perform operations such as forward propagation, loss calculation, and parameter update through backpropagation until this upper limit is reached. During this period, the PSNR loss calculation function is used to calculate a value based on the pixel value difference between the model output and the real target, construct a loss with this value, and promote the model to continuously reduce the loss through backpropagation, improving the processing ability of the input data, making the model converge more efficiently and accurately during the entire training process, and finally obtaining a reference model.

[0108] Additionally, it should be noted that the cosine annealing function is a strategy for adjusting the learning rate. Its basic principle is based on the characteristics of the cosine function. During training, the learning rate will change according to the curve shape of the cosine function. The PSNR (Peak Signal-to-Noise Ratio) loss calculation function is often used to measure the difference between the model output and the real target, especially widely used in related fields such as images. It calculates a value based on the pixel value difference between the output image (the image constructed from the data output by the model after forward propagation) and the real image. The larger this value, the closer the output image is to the real image, that is, the better the prediction effect of the model.

[0109] In a feasible implementation manner, step S33 includes steps S331 to S333:

[0110] Step S331, calculate the loss function of the model based on the forward propagation data.

[0111] When calculating the model loss function using the PSNR loss calculation function based on the forward propagation data, first calculate the PSNR further according to the mean square error (MSE). The MSE is obtained by summing the squares of the pixel value differences of the forward propagation data output by the model and the corresponding real data in the corresponding dimensions and taking the average. Its formula is to sum the squares of the differences between the model output data and the real data at all pixel points and then divide by the product of the total number of rows and the total number of columns of the data. Then, substitute the MSE and the maximum possible value of the image pixel value (for example, this value is 255 for an 8-bit grayscale image) into the calculation formula of PSNR to calculate the PSNR value. The higher the PSNR value, the closer the model output is to the real situation. This value can be used to measure the quality of the model's forward propagation output this time, and then a loss function can be constructed based on this to reflect the deviation degree of the model from the ideal state, providing a basis for adjusting the model parameters through backpropagation later.

[0112] Step S332, calculate the parameter gradients of the initial model through a loss function.

[0113] When calculating the parameter gradients of a model based on a loss function, it is mainly achieved by using the method of taking derivatives. First of all, the loss function essentially describes the degree of difference between the model output and the true target, and it is a function of each parameter of the model. For example, in a common neural network, the model contains many parameters such as weights and biases, and the value of the loss function depends on the values of these parameters. Then, using the derivative rules in calculus, the partial derivatives of the loss function with respect to each parameter are calculated separately, and the values of these partial derivatives are the gradients of the corresponding parameters. Specifically, in the backpropagation algorithm, starting from the output layer where the loss function is located, according to the chain rule, the error (i.e., the gradient information of the loss function for the parameters of each layer) is gradually passed back to the previous hidden layers and the input layer, and the gradients corresponding to each parameter in each layer are calculated in turn. These calculated parameter gradients reflect the rate of change of the loss function as each parameter changes. The positive or negative sign indicates whether the parameter should be increased or decreased to make the loss function value decrease, and its magnitude reflects the sensitivity of the parameter change to the loss function. Subsequently, the parameter gradients can be used to update the parameters of the model to continuously optimize the model and make it develop in the direction of continuously decreasing the loss function value and improving performance.

[0114] Step S333, based on the parameter gradients, train the initial model through a preset optimizer to obtain a reference model.

[0115] After obtaining the parameter gradients of the model, the AdaW optimizer first determines the initial update step size for each parameter based on the set initial learning rate of 0.001 and the currently obtained parameter gradient information. Since a cosine annealing function is used to dynamically adjust the learning rate, as the number of training rounds increases, the learning rate changes according to the variation law of the cosine function, and then the corresponding update step sizes of the parameters will also be dynamically adjusted. The AdaW optimizer itself also has the characteristic of adaptively adjusting the update strength of each parameter according to factors such as the importance of the parameter. For those parameters that have a greater impact on the model loss function and are more critical, it will more reasonably allocate the update weights and adjust them in a direction that is more conducive to reducing the loss and optimizing the model performance. In each round of training, the current parameter gradients are combined with the dynamically changing learning rate and the adaptive mechanism of the optimizer to update each parameter in the model. This process is continuously repeated, and the model continuously cycles among operations such as forward propagation, calculating the loss, backpropagation to obtain gradients, and using the AdaW optimizer to update parameters until the set maximum number of training times of 400,000 times is reached. After repeated iterative training through such a complete set of processes, the model continuously optimizes its own parameters and gradually converges to a better state, and finally obtains a reference model that can better process the input data and meet the expected goals.

[0116] The model performance is improved by continuously optimizing the model through calculating the loss function.

[0117] This embodiment provides an image sharpening method based on a neural network. The data of the image dataset is cut through the initial model to obtain cut data; forward propagation training is performed on the cut data to obtain forward propagation data; based on the forward propagation data, the initial model is trained through a preset optimizer to obtain a reference model, thereby improving the performance of the model.

[0118] Based on the first embodiment of the present application, in the fourth embodiment of the present application, the same or similar content as that in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 8 , step S40 includes steps S41 to S42:

[0119] Step S41, converting the reference model into an open neural network exchange format to obtain a format conversion model.

[0120] It should be noted that the Open Neural Network Exchange (ONNX) is an open format for representing deep learning models, aiming to achieve interoperability between different deep learning frameworks (such as PyTorch, TensorFlow, MXNet, etc.). By defining unified computational graphs and node operation specifications, etc., models can be shared, deployed, and optimized across different frameworks. After converting the model to the ONNX format, it can be conveniently used on various platforms and tools that support ONNX. For example, it can be deployed to mobile devices, embedded devices, or used for model inference acceleration, etc.

[0121] For example, when converting the reference model to the ONNX format, ensure that the trained, structurally complete, and parameter-loaded model is in a usable state. In PyTorch, use the torch.onnx.export function to specify the model, example data that meets the input requirements, the output file path, and relevant configuration parameters to convert it to the ONNX format. In TensorFlow, first save it in a suitable format and then use the tf2onnx tool to convert it as required.

[0122] In step S42, convert the format conversion model into a tensor real-time processing format through a deep learning inference engine to obtain a preset model.

[0123] It should be noted that the TensorRT format is a model format introduced by NVIDIA for optimizing deep learning model inference. It is mainly used to accelerate the inference process of models in the NVIDIA GPU hardware environment. As deep learning models become more and more complex, the inference speed in practical applications becomes a key factor. TensorRT effectively improves the inference performance of models by performing a series of optimizations on the model and converting the model to the TRT format, thereby leveraging the parallel computing power of the GPU and specific hardware acceleration technologies.

[0124] Select the quantization precision according to the model characteristics and application scenarios, that is, convert 32-bit floating-point numbers (float32) to 16-bit floating-point numbers (float16) and low-precision 8-bit integer format (int8). Then use the API provided by NVIDIA's TensorRT to load the model from the ONNX file to build the network definition, configure the build parameters to enable the quantization mode (such as setting INT8 quantization-related parameters, etc.) and associate a custom calibrator class to implement the calibration dataset loading and calibration logic to build a quantized TensorRT engine and save it as a.trt file. Finally, use the TensorRT runtime environment to load the quantized.trt model, compare the input data with the output results for functional verification, and at the same time test performance metrics such as inference speed. If the performance is not ideal, further adjust the quantization strategy, recalibrate, etc. to complete the model quantization process and achieve model optimization deployment and performance improvement.

[0125] Quantizing the model can reduce the model size, lower memory occupancy, and accelerate calculations because low-precision operations are usually faster than high-precision operations. Int8 quantization of the model requires data calibration. Use the officially provided code to prepare a set of quantization datasets to generate the quantization weight file, and then use this file to perform int8 quantization on the model. Performing int8 quantization of the model on the Nvidia platform improves the inference speed of the model while maintaining the model effect. The inference speed on the RTX3060 laptop platform reaches 150 frames / s for 720p, and the inference speed on the nvidia-Orin development board reaches more than 30 frames, which can be used for actual deployment.

[0126] This embodiment provides an image clarity method based on a neural network, which converts the reference model into an open neural network exchange format to obtain a format conversion model; and converts the format conversion model into a tensor real-time processing format through a deep learning inference engine to obtain a preset model, realizing the reduction of the model complexity and the improvement of the real-time inference ability of the model.

[0127] Based on the first embodiment of this application, in the fifth embodiment of this application, the same or similar content as in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 9 , step S50 includes steps S51 to S53:

[0128] Step S51, input the blurred image into the preset model.

[0129] After the model is trained and quantized, the blurred image can be input into the model for image sharpening. First, ensure that the format of the blurred image meets the input requirements of the preset model. If not, perform preprocessing operations such as scaling, cropping, and format conversion to make it adaptable. At the same time, use the corresponding image loading library to read the image into an array form that can be processed by the program and load it into the program environment. Then, ensure that the preset model is in a usable state. For example, for a trained model, load the parameters, or for an untrained model, define the model structure and initialize the parameters. Switch the model to the evaluation mode if necessary. After that, pass the processed image to the input layer of the model in the form of a tensor to achieve image input.

[0130] Step S52: Perform downsampling and upsampling on the blurred image through the preset model to obtain sampling data.

[0131] The downsampling process reduces the amount of data in the input blurred image. The downsampling operation in the model is usually achieved by means of specific layers or algorithms. Convolution operations with a convolution kernel stride greater than 1 are used to perform downsampling. By increasing the convolution stride, the convolution kernel moves to cover a larger range each time, thereby reducing the size of the output feature map and achieving the purpose of downsampling. After downsampling, the upsampling operation is performed. Upsampling aims to restore or expand the data with a smaller size after downsampling to a size close to the original image size or the desired size. The PixelShuffle structure is used for upsampling. The PixelShuffle structure realizes upsampling based on the principles of channel rearrangement and pixel position adjustment. It receives the input of a feature map with a specific number of channels and rearranges the pixel information in the channels according to the set rules (such as according to the upsampling factor r, reorganizing the number of channels C of the input feature map into r groups, corresponding to the pixel information in the r×r small areas in the output image), expanding the image size and improving the resolution. In models commonly used for tasks such as image super-resolution in deep learning convolutional neural networks, in cooperation with other layers, after the previous convolutional layer extracts features, it performs upsampling to restore the features to a high-resolution form, and then followed by a convolutional layer for refinement and optimization. Moreover, compared with methods such as transposed convolution, it can avoid artifacts such as checkerboard effects and more naturally and evenly distribute pixels to generate high-quality upsampled images. After this series of operations of downsampling and upsampling, the finally obtained sampling data combines the features refined during downsampling and the information such as the expansion and restoration of features during upsampling.

[0132] Step S53: Post-process the sampling data to obtain a sharp image.

[0133] After operations such as downsampling and upsampling on the sampled data, its numerical range may not meet the conventional requirements for display or subsequent use. At this time, data normalization adjustment is required. A common method is to normalize the pixel values to the range of 0 to 1 or 0 to 255. For example, if the pixel values in the sampled data originally fall within a relatively wide and irregular interval, by finding the minimum and maximum values in the data and then mapping each pixel value to the desired normalized interval according to the linear transformation method, it ensures that visual attributes such as the brightness and contrast of the image are within a reasonable range, laying a foundation for generating a clear image subsequently. If it is a color image, there may be a problem of color deviation after the above operations on the sampled data, and color correction is required. Based on known color standards or reference images, the true color can be restored by adjusting parameters such as the gain and bias of the red, green, and blue channels. At the same time, the sampled data can also be processed for detail restoration and noise removal to make the obtained clear image have a better effect.

[0134] A schematic diagram of the image clarity processing effect is as Figure 10 and Figure 11 shown, where the left picture is the picture before image clarity processing, and the right picture is the picture after image clarity processing.

[0135] This embodiment provides a method for image clarity based on a neural network. The reference model is converted into an open neural network exchange format to obtain a format conversion model; the format conversion model is converted into a tensor real-time processing format through a deep learning inference engine to obtain a preset model, achieving a reduction in the complexity of the model.

[0136] Based on the first embodiment of the present application, in the sixth embodiment of the present application, the same or similar content as in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 12 , step S10 includes steps S11 to S13:

[0137] Step S11, obtaining a real blurred data set.

[0138] It should be noted that the real blurred data set is the REDS (Real-time and Edge-aware Super-resolution) data set. The REDS data set is a well-known data set for image processing tasks such as super-resolution. The real blurred data set REDS contains clear images and blurred images.

[0139] Step S12, performing sharpening processing on the clear images in the real blurred data set to obtain sharpened images.

[0140] To sharpen the clear images in the REDS dataset, a sharpening algorithm based on a convolution kernel can be used, such as Laplacian sharpening. Its core lies in designing a suitable Laplacian convolution kernel. Commonly, the central pixel value is positive, and the weight values corresponding to the surrounding adjacent pixels are negative. For example, a 3×3 Laplacian convolution kernel can be [[0, -1, 0], [-1, 4, -1], [0, -1, 0]]. Convolve this convolution kernel with the pixel matrix corresponding to the clear images in the REDS dataset. By calculating the difference relationship between the central pixel and the surrounding pixels, the contrast of the image edges and details is enhanced, making the originally relatively soft edges sharper, thereby highlighting the main body and key features in the image and achieving the sharpening effect to obtain sharpened images. Secondly, a sharpening method based on deep learning can also be adopted. For example, a small convolutional neural network (CNN) can be constructed. Taking the clear images in the REDS dataset as input, features are gradually extracted and edge information is strengthened through structures such as convolutional layers and activation layers in the network. After the network is trained on a training set composed of pairs of labeled sharpened images and clear images, it can perform sharpening processing on other clear images in the REDS dataset and output sharpened images that meet the requirements. This method can better adapt to the sharpening needs of different types of images and can learn more effective sharpening strategies in complex scenarios. Regardless of which sharpening method is used, it is necessary to reasonably adjust the sharpening parameters according to the characteristics of the images in the REDS dataset, such as the resolution, content scene, color mode, etc. of the images, such as the coefficients of the convolution kernel, the cut-off frequency of the high-pass filter, the hyperparameters of the deep learning network, etc., to ensure that sharpened images with good visual effects, which not only highlight the key details but also do not cause problems such as artifacts due to oversharpening, are obtained.

[0141] Step S13, using the blurred images and the sharpened images in the real blurred dataset as an image dataset.

[0142] Based on the original annotation attributes in the REDS dataset or by using image analysis algorithms to judge, such as looking at features like edge sharpness and high-frequency information content, select blurred images. Unify the format of the blurred images and the sharpened images (such as uniformly converting them to a suitable format like PNG format), standardize the size (adjust to a specific size using an image scaling algorithm), and ensure that the color mode is consistent (such as uniformly being the RGB mode). Then establish a label system to label and classify the blurred images and the sharpened images, making them an image dataset for subsequent use and processing.

[0143] This embodiment provides an image sharpening method based on a neural network, which obtains a real blurred data set; performs sharpening processing on the clear images in the real blurred data set to obtain sharpened images; and uses the blurred images and the sharpened images in the real blurred data set as an image data set, thereby realizing the acquisition of blurred images that can be used for model processing.

[0144] It should be noted that the above examples are only for understanding this application and do not constitute a limitation to the image sharpening method based on a neural network of this application. Any simple transformation in more forms based on this technical concept is within the protection scope of this application.

[0145] This application also provides an image sharpening device based on a neural network. Please refer to Figure 13 , and the device includes:

[0146] A data acquisition module 10, configured to acquire an image data set;

[0147] A model construction module 20, configured to construct an initial model based on a preset network;

[0148] A model training module 30, configured to train the initial model based on the image data set to obtain a reference model;

[0149] A model quantization module 40, configured to quantize the reference model through a deep learning inference engine to obtain a preset model;

[0150] An image processing module 50, configured to perform image sharpening processing on a blurred image through the preset model to obtain a clear image.

[0151] In one embodiment, the data acquisition module 10 is further configured to construct a first basic model based on a preset network; set a U-shaped network structure with a preset number of layers and a preset number of downsamplings in the first basic model to obtain a second basic model; and construct a retention layer, an encoding layer, and a decoding layer of the second basic model through a convolutional layer, an activation function layer, a class residual structure, a batch normalization layer, and a pixel shuffle structure to obtain an initial model.

[0152] In one embodiment, the model construction module 20 is further configured to cut the data of the image data set through the initial model to obtain cut data; perform forward propagation training on the cut data to obtain forward propagation data; and train the initial model based on the forward propagation data through a preset optimizer to obtain a reference model.

[0153] In one embodiment, the model training module 30 is further configured to calculate a loss function of the model based on the forward propagation data; calculate a parameter gradient of the initial model through the loss function; and train the initial model through a preset optimizer based on the parameter gradient to obtain a reference model.

[0154] In one embodiment, the model training module 30 is further configured to convert the reference model into an Open Neural Network Exchange (ONNX) format to obtain a format conversion model; and convert the format conversion model into a TensorRT format through a deep learning inference engine to obtain a preset model.

[0155] In one embodiment, the model quantization module 40 is further configured to input a blurred image into the preset model; perform downsampling and upsampling on the blurred image through the preset model to obtain sampled data; and perform post-processing on the sampled data to obtain a clear image.

[0156] In one embodiment, the image processing module 50 is further configured to obtain a real blurred data set; perform sharpening processing on the clear images in the real blurred data set to obtain sharpened images; and use the blurred images and the sharpened images in the real blurred data set as an image data set.

[0157] The image sharpening device based on a neural network provided in this application adopts the image sharpening method based on a neural network in the above embodiment, and can solve the technical problems of how to simplify the model complexity, be applicable to edge deployment, and improve the real-time inference ability of the model. Compared with the prior art, the beneficial effects of the image sharpening device based on a neural network provided in this application are the same as those of the image sharpening method based on a neural network provided in the above embodiment, and other technical features in the image sharpening device based on a neural network are the same as the features disclosed in the above embodiment method, which will not be elaborated herein.

[0158] This application provides an image sharpening device based on a neural network. The image sharpening device based on a neural network includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the image sharpening method based on a neural network in the first embodiment above.

[0159] Next, refer to Figure 14, which shows a schematic structural diagram of a neural network-based image sharpening device suitable for implementing the embodiments of the present application. The neural network-based image sharpening device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 14 The shown neural network-based image sharpening device is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0160] As Figure 14 shown, the neural network-based image sharpening device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the neural network-based image sharpening device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the neural network-based image sharpening device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a neural network-based image sharpening device having various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be implemented or had alternatively.

[0161] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by a processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.

[0162] The image sharpening device based on a neural network provided by the present application adopts the image sharpening method based on a neural network in the above embodiments, and can solve the technical problems of how to simplify the model complexity, be suitable for edge deployment, and improve the real-time inference ability of the model. Compared with the prior art, the beneficial effects of the image sharpening device based on a neural network provided by the present application are the same as those of the image sharpening method based on a neural network provided in the above embodiments, and other technical features in the image sharpening device based on a neural network are the same as the features disclosed in the method of the previous embodiment, and will not be elaborated here.

[0163] It should be understood that the various parts disclosed in the present application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0164] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0165] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the image sharpening method based on a neural network in the above embodiments.

[0166] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0167] The above computer-readable storage medium can be included in the neural network-based image sharpening device; it can also exist independently without being assembled into the neural network-based image sharpening device.

[0168] The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by the neural network-based image sharpening device, the neural network-based image sharpening device is caused to: obtain an image data set; construct an initial model based on a preset network; train the initial model based on the image data set to obtain a reference model; quantize the reference model through a deep learning inference engine to obtain a preset model; and perform image sharpening processing on the blurred image through the preset model to obtain a sharp image.

[0169] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN: Local Area Network) or a wide area network (WAN: Wide Area Network), or it can be connected to an external computer (for example, by connecting through the Internet using an Internet service provider).

[0170] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0171] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.

[0172] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned image clarity improvement method based on neural networks, and can solve the technical problems of how to simplify the model complexity, be applicable to edge deployment, and improve the real-time inference ability of the model. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the image clarity improvement method based on neural networks provided by the above embodiments, and will not be elaborated here.

[0173] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the above-mentioned image clarity improvement method based on a neural network.

[0174] The computer program product provided by the present application can solve the technical problems of how to simplify the model complexity, be applicable to edge deployment, and improve the real-time inference ability of the model. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the image clarity improvement method based on a neural network provided in the above embodiments, and will not be elaborated herein.

[0175] The above are only partial embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. A method for image sharpening based on a neural network, characterized in that: The method comprises: Get image dataset; Build an initial model based on the preset network; Based on the image data set, training the initial model to obtain a reference model; Quantifying the reference model through a deep learning inference engine to obtain a preset model; The preset model is used to perform image clarity processing on the blurred image to obtain a clear image.

2. The method according to claim 1, characterized in that The step of building an initial model based on the preset network includes: Building a first basic model based on a preset network; A U-shaped network structure with a preset number of layers and a preset number of downsamplings are set in the first basic model to obtain a second basic model; The holding layer, encoding layer and decoding layer of the second basic model are constructed through convolutional layers, activation function layers, residual-like structures, batch normalization layers and pixel shuffling structures to obtain an initial model.

3. The method according to claim 1, characterized in that The step of training the initial model based on the image data set to obtain a reference model comprises: Cutting the data of the image data set by using the initial model to obtain cut data; Performing forward propagation training on the cutting data to obtain forward propagation data; Based on the forward propagation data, the initial model is trained by a preset optimizer to obtain a reference model.

4. The method according to claim 3, characterized in that The step of training the initial model based on the forward propagation data by a preset optimizer to obtain a reference model comprises: Based on the forward propagation data, calculating the loss function of the model; Calculating the parameter gradient of the initial model through a loss function; Based on the parameter gradient, the initial model is trained by a preset optimizer to obtain a reference model.

5. The method according to claim 1, characterized in that The step of quantizing the reference model by a deep learning inference engine to obtain a preset model comprises: Converting the reference model into an open neural network exchange format to obtain a format conversion model; The format conversion model is converted into a tensor real-time processing format through a deep learning inference engine to obtain a preset model.

6. The method according to claim 1, characterized in that The step of performing image clarity processing on the blurred image by using the preset model to obtain a clear image comprises: Inputting the blurred image into the preset model; Down-sampling and up-sampling the blurred image by using the preset model to obtain sampling data; The sampling data is post-processed to obtain a clear image.

7. The method according to claim 1, characterized in that The step of obtaining the image data set comprises: Get the real fuzzy data set; Performing sharpening processing on the clear image in the real fuzzy data set to obtain a sharpened image; The blurred image in the real blurred data set and the sharpened image are used as image data sets.

8. An image sharpening device based on a neural network, characterized in that: The device comprises: A data acquisition module, used for acquiring an image data set; Model building module, used to build an initial model based on a preset network; A model training module, used to train the initial model based on the image data set to obtain a reference model; A model quantization module, used to quantize the reference model through a deep learning inference engine to obtain a preset model; The image processing module is used to perform image clarity processing on the blurred image through the preset model to obtain a clear image.

9. An image sharpening device based on a neural network, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the neural network-based image sharpening method according to any one of claims 1 to 7.

10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the neural network-based image sharpening method as described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Model quantification method and device, electronic equipment and storage medium

    CN121145947A