An image enhancement method and related device
By adopting a multi-layer network structure on devices with weak computing capabilities to process images of different resolutions, the problem of excessive computing overhead when processing ultra-high-definition images in the prior art is solved, and high-quality image enhancement effect is achieved.
Patent Information
- Application Number
- CN202110221711.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-27
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-02-27
AI Technical Summary
The prior art is difficult to effectively enhance image on devices with weak computing power, especially when processing ultra-high-definition resolution images, complex neural networks can lead to excessive computing overhead and power consumption.
Using a multi-layer network structure, the upper and lower networks respectively process images with incremental resolution layers, the upper network processes low-resolution images to obtain a larger receptive field, and the lower network processes high-resolution images to further process the output of the upper network, thus taking into account both global information and local information.
By this method, the network size can be significantly reduced while maintaining image enhancement quality, so that image enhancement networks can be deployed and run on devices with weak computing power.
Smart Images

Figure CN113066018B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to an image enhancement method and related devices. Background Art
[0002] Artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0003] Deep learning methods have been a key driving force in the development of the field of artificial intelligence in recent years and have achieved remarkable results in various tasks of computer vision. In the field of image enhancement (also known as image quality enhancement), deep learning-based methods have surpassed traditional methods.
[0004] Deep learning methods usually process images through complex neural networks. Complex neural networks usually have a large receptive field, thus bringing better image enhancement quality. However, with the popularization of ultra-high-definition resolution images / videos, complex neural networks usually bring an exponentially increasing computational overhead and power consumption overhead. For devices with relatively weak computing power (such as consumer terminal products like mobile phones or tablets), it is usually difficult to use complex deep learning methods for image enhancement due to factors such as running speed, running video memory, and power consumption.
[0005] Therefore, there is an urgent need for a method that can perform image enhancement on devices with relatively weak computing power. Summary of the Invention
[0006] This application provides an image enhancement method that adopts a multi-layer network structure, and the upper-layer network and the lower-layer network respectively process images with gradually increasing resolutions. Moreover, the output of the upper-layer network is passed to the lower-layer network for further processing by the lower-layer network. Since the upper-layer network processes low-resolution images, it can obtain a larger receptive field than the lower-layer network, and the lower-layer network that processes high-resolution images further processes the output of the upper-layer network, thereby achieving both global information and local information, and ensuring the quality of image enhancement. By processing images with different resolutions through multiple layers of networks, it is possible to effectively reduce the depth of the network on the basis of obtaining a large receptive field, thereby significantly reducing the scale of the network, enabling the network to be deployed and run on devices with relatively weak computing power.
[0007] The first aspect of the present application provides an image enhancement method, including: the terminal obtains a first image and a second image, where the second image is obtained by performing downsampling processing on the first image. The first image is an image to be enhanced, such as an image that needs to be deblurred. The terminal can obtain the second image by performing downsampling processing on the first image, and the resolution of the second image is lower than that of the first image. The terminal processes the second image through a first network to obtain a first feature and a third image. The first feature is an intermediate feature extracted by the first network, and the resolution of the third image is the same as that of the first image. The third image is an image obtained after performing image enhancement on the second image. The terminal generates a fourth image according to the third image and the first image. For example, the terminal adds the third image and the first image to obtain the fourth image. The terminal processes the fourth image and the first feature through a second network to obtain a target image. The first network and the second network can be networks for performing image enhancement, such as convolutional neural networks.
[0008] In this solution, since downsampling processing is performed on the first image and the second image obtained by the downsampling processing is processed by the first network, the feature map extracted by the first network can have a larger receptive field. After the third image output by the first network is fused with the first image and then input into the second network for processing, the final output image of the second network can have a better image enhancement effect, ensuring the quality of the enhanced image.
[0009] Optionally, in a possible implementation manner, the method further includes: the terminal processes a fifth image through the third network to obtain a sixth image, where the fifth image is an image obtained by performing downsampling processing on the second image. The terminal processing the second image through the first network specifically includes: the terminal fuses the sixth image with the second image, for example, adds the sixth image and the second image to obtain a seventh image; the terminal processes the seventh image through the first network to obtain the third image. That is to say, there are other networks above the first network, and the terminal can process images with different resolutions layer by layer through multiple different networks.
[0010] Optionally, in a possible implementation, the first network includes a first sub-network and a second sub-network, and the second network includes a third sub-network and a fourth sub-network. The terminal processes the second image through the first network to obtain a first feature and a third image, including: the terminal processes the second image through the first sub-network to obtain a first feature; the terminal processes the first feature through the second sub-network to obtain the third image. The terminal processes the fourth image through the second network to obtain a target image, including: the terminal processes the fourth image through the third sub-network to obtain a second feature; the terminal generates a third feature according to the first feature and the second feature; the terminal processes the third feature through the fourth sub-network to obtain the target image.
[0011] Among them, the terminal may fuse the first feature and the second feature to obtain a third feature. For example, the terminal performs an upsampling process on the first feature to obtain the upsampled first feature; then adds the upsampled first feature and the second feature to obtain the third feature. Since the first feature is extracted based on the second image with a lower resolution, and the second feature is extracted based on the fourth image with a higher resolution, the terminal may perform an upsampling process on the first feature before fusing the first feature and the second feature. For example, the terminal performs an upsampling process on the first feature to obtain an upsampled feature, and then the terminal adds the upsampled feature and the second feature to obtain a third feature.
[0012] In this solution, by fusing the feature-level information extracted by the upper network and the lower network, as well as the image-level information between the upper and lower networks, the lower network can take into account the feature information obtained by the upper network as much as possible while processing local information, and promote the depth design of each layer of the network to be as simple as possible.
[0013] Optionally, in a possible implementation, the first sub-network and the second sub-network include a convolutional neural network.
[0014] Optionally, in a possible implementation, the first network may be implemented using an encoder-decoder structure. Specifically, the first sub-network includes an encoder, and the first sub-network is used to obtain encoded features. The second sub-network includes a decoder, and the second sub-network is used to obtain a decoded image. Among them, the encoder can be used to perform feature extraction and compression, by performing feature extraction on the input image and compressing the extracted features to obtain encoded features. The decoder can be used to perform a feature restoration operation, by performing upsampling and feature restoration on the encoded features to obtain a decoded image.
[0015] Optionally, in a possible implementation, the encoder includes a first convolutional layer and a channel-spatial attention module, and the decoder includes a second convolutional layer and a residual module.
[0016] Optionally, in a possible implementation, the terminal processes the second image through the first sub-network to obtain a first feature, including: the terminal divides the second image in the spatial dimension to obtain a plurality of non-overlapping sub-images. The terminal splices the plurality of sub-images in the channel dimension to obtain a spliced image, and the number of channels of the spliced image is N times the number of channels of the second image. The terminal processes the spliced image through a first convolutional neural network to obtain a processed feature. The terminal divides the processed feature in the channel dimension to obtain a plurality of sub-processed features. Wherein, the number of sub-processed features is the same as the number of sub-images. The terminal splices the plurality of sub-processed features in the spatial dimension to obtain the first feature.
[0017] Optionally, in a possible implementation, the terminal processes the first feature through the second sub-network to obtain the third image, including: the terminal divides the first feature in the spatial dimension to obtain a plurality of sub-features; the terminal splices the plurality of sub-features in the channel dimension to obtain a spliced feature; the terminal processes the spliced feature through a second convolutional neural network to obtain a processed image; the terminal divides the processed image in the channel dimension to obtain a plurality of sub-processed images; the terminal splices the plurality of sub-processed images in the spatial dimension to obtain the third image.
[0018] In this embodiment, by introducing a partition acceleration method to accelerate the processing speed of the encoder and the decoder, the operation efficiency of the neural network can be improved on the premise of ensuring that the enhanced image quality is not affected, and the effect equivalent to the original method in terms of enhancement quality and meeting real-time processing can be achieved in a specific scenario. In addition, after introducing the partition acceleration method, the partitions for processing different sub-images have self-learning differential convolutional kernels, which can realize differential learning, so that the local enhancement of the image is better.
[0019] Optionally, in a possible implementation, the image enhancement method is used to implement at least one of the following image enhancement tasks: image super-resolution reconstruction, image denoising, image dehazing, image deblurring, image contrast enhancement, image demosaicing, image deraining, image color enhancement, image brightness enhancement, image detail enhancement, and image dynamic range enhancement.
[0020] The second aspect of the present application provides a method for training a model, including: obtaining an image sample pair, where the image sample pair includes a first image and an enhanced image corresponding to the first image; performing downsampling processing on the first image to obtain a second image; processing the second image through a first network to obtain a first feature and a third image, where the first feature is an intermediate feature extracted by the first network, and the resolution of the third image is the same as that of the first image; generating a fourth image according to the third image and the first image; processing the fourth image through a second network to obtain a target image; obtaining a loss function according to the enhanced image corresponding to the first image and the target image, where the loss function is used to indicate the difference between the enhanced image and the target image in the image sample pair; training the first network and the second network according to the loss function to obtain a trained first network and a trained second network; where the first network and the second network are used for image enhancement.
[0021] Optionally, in a possible implementation manner, the first network includes a first sub-network and a second sub-network, and the second network includes a third sub-network and a fourth sub-network; processing the second image through the first network to obtain a first feature and a third image includes: processing the second image through the first sub-network to obtain the first feature; processing the first feature through the second sub-network to obtain the third image; processing the fourth image through the second network to obtain a target image includes: processing the fourth image through the third sub-network to obtain a second feature; generating a third feature according to the first feature and the second feature; processing the third feature through the fourth sub-network to obtain the target image.
[0022] Optionally, in a possible implementation manner, the first sub-network and the second sub-network include convolutional neural networks.
[0023] Optionally, in a possible implementation manner, the first sub-network includes an encoder, and the first sub-network is used to obtain encoded features; the second sub-network includes a decoder, and the second sub-network is used to obtain decoded images.
[0024] Optionally, in a possible implementation manner, the encoder includes a first convolutional layer and a channel-spatial attention module, and the decoder includes a second convolutional layer and a residual module.
[0025] Optionally, in a possible implementation, the processing of the second image by the first sub-network to obtain a first feature includes: segmenting the second image in the spatial dimension to obtain a plurality of sub-images; splicing the plurality of sub-images in the channel dimension to obtain a spliced image; processing the spliced image through a first convolutional neural network to obtain a processed feature; segmenting the processed feature in the channel dimension to obtain a plurality of sub-processed features; and splicing the plurality of sub-processed features in the spatial dimension to obtain the first feature.
[0026] Optionally, in a possible implementation, the terminal processes the first feature through the second sub-network to obtain the third image, including: the terminal segmenting the first feature in the spatial dimension to obtain a plurality of sub-features; the terminal splicing the plurality of sub-features in the channel dimension to obtain a spliced feature; the terminal processing the spliced feature through a second convolutional neural network to obtain a processed image; the terminal segmenting the processed image in the channel dimension to obtain a plurality of sub-processed images; and the terminal splicing the plurality of sub-processed images in the spatial dimension to obtain the third image.
[0027] Optionally, in a possible implementation, the generating of the fourth image according to the third image and the first image includes: performing an addition process on the third image and the first image to obtain the fourth image.
[0028] Optionally, in a possible implementation, the training method of the model is used to implement at least one of the following image enhancement tasks: image super-resolution reconstruction, image denoising, image dehazing, image deblurring, image contrast enhancement, image demosaicing, image deraining, image color enhancement, image brightness enhancement, image detail enhancement, and image dynamic range enhancement.
[0029] A third aspect of the present application provides an image processing apparatus, including an acquisition unit and a processing unit. The acquisition unit is configured to acquire a first image and a second image, where the second image is obtained by performing downsampling processing on the first image; the processing unit is configured to process the second image through a first network to obtain a first feature and a third image, where the first feature is an intermediate feature extracted by the first network, and the resolution of the third image is the same as the resolution of the first image; the processing unit is further configured to generate a fourth image according to the third image and the first image; the processing unit is further configured to process the fourth image through a second network to obtain a target image; where the first network and the second network are used for image enhancement.
[0030] Optionally, in a possible implementation, the first network includes a first sub-network and a second sub-network, and the second network includes a third sub-network and a fourth sub-network; the processing unit is further configured to process the second image through the first sub-network to obtain a first feature; the processing unit is further configured to process the first feature through the second sub-network to obtain the third image; the processing unit is further configured to process the fourth image through the third sub-network to obtain a second feature; the processing unit is further configured to generate a third feature according to the first feature and the second feature; the processing unit is further configured to process the third feature through the fourth sub-network to obtain the target image.
[0031] Optionally, in a possible implementation, the generating a third feature according to the first feature and the second feature includes: performing upsampling processing on the first feature to obtain the upsampled first feature; adding the upsampled first feature and the second feature to obtain the third feature.
[0032] Optionally, in a possible implementation, the first sub-network includes an encoder, and the first sub-network is configured to obtain encoded features; the second sub-network includes a decoder, and the second sub-network is configured to obtain a decoded image.
[0033] Optionally, in a possible implementation, the encoder includes a first convolutional layer and a channel-spatial attention module, and the decoder includes a second convolutional layer and a residual module.
[0034] Optionally, in a possible implementation, the processing unit is further configured to segment the second image in the spatial dimension to obtain a plurality of sub-images; the processing unit is further configured to splice the plurality of sub-images in the channel dimension to obtain a spliced image; the processing unit is further configured to process the spliced image through an encoder to obtain encoded features; the processing unit is further configured to segment the encoded features in the channel dimension to obtain a plurality of sub-encoded features; the processing unit is further configured to splice the plurality of sub-encoded features in the spatial dimension to obtain the first feature.
[0035] Optionally, in a possible implementation, the processing unit is further configured to add the third image and the first image to obtain a fourth image.
[0036] Optionally, in a possible implementation, the image processing device is configured to implement at least one of the following image enhancement tasks: image super-resolution reconstruction, image denoising, image dehazing, image deblurring, image contrast enhancement, image demosaicing, image deraining, image color enhancement, image brightness enhancement, image detail enhancement, and image dynamic range enhancement.
[0037] The fourth aspect of the present application provides a model training device, including an acquisition unit and a processing unit. The acquisition unit is configured to acquire an image sample pair, where the image sample pair includes a first image and an enhanced image corresponding to the first image; the processing unit is configured to perform downsampling processing on the first image to obtain a second image; the processing unit is further configured to process the second image through a first network to obtain a first feature and a third image, where the first feature is an intermediate feature extracted by the first network, and the resolution of the third image is the same as that of the first image; the processing unit is further configured to generate a fourth image according to the third image and the first image; the processing unit is further configured to process the fourth image through a second network to obtain a target image; the processing unit is further configured to obtain a loss function according to the enhanced image corresponding to the first image and the target image, where the loss function is used to indicate the difference between the enhanced image and the target image in the image sample pair; the processing unit is further configured to train the first network and the second network according to the loss function to obtain a trained first network and a trained second network; where the first network and the second network are used for image enhancement.
[0038] Optionally, in a possible implementation, the processing unit is further configured to process the second image through the first network to obtain the third image and the first feature, where the first feature is an intermediate feature extracted by the first network; the processing unit is further configured to process the fourth image and the first feature through a second network to obtain a target image.
[0039] Optionally, in a possible implementation, the first network includes a first sub-network and a second sub-network, and the second network includes a third sub-network and a fourth sub-network; the processing unit is further configured to process the second image through the first sub-network to obtain the first feature; the processing unit is further configured to process the first feature through the second sub-network to obtain the third image; the processing unit is further configured to process the fourth image through the third sub-network to obtain a second feature; the processing unit is further configured to generate a third feature according to the first feature and the second feature; and process the third feature through the fourth sub-network to obtain the target image.
[0040] Optionally, in a possible implementation, generating a third feature according to the first feature and the second feature includes: the processing unit is further configured to: perform upsampling processing on the first feature to obtain the upsampled first feature; and add the upsampled first feature and the second feature to obtain the third feature.
[0041] Optionally, in a possible implementation, the first sub-network includes an encoder, and the first sub-network is used to obtain encoded features; the second sub-network includes a decoder, and the second sub-network is used to obtain a decoded image.
[0042] Optionally, in a possible implementation, the encoder includes a first convolutional layer and a channel-spatial attention module, and the decoder includes a second convolutional layer and a residual module.
[0043] Optionally, in a possible implementation, the processing unit is further configured to: segment the second image in the spatial dimension to obtain a plurality of sub-images; splice the plurality of sub-images in the channel dimension to obtain a spliced image; process the spliced image through an encoder to obtain encoded features; segment the encoded features in the channel dimension to obtain a plurality of sub-encoded features; and splice the plurality of sub-encoded features in the spatial dimension to obtain the first feature.
[0044] Optionally, in a possible implementation, the processing unit is further configured to add the third image and the first image to obtain a fourth image.
[0045] Optionally, in a possible implementation, the model training device is used to implement at least one of the following image enhancement tasks: image super-resolution reconstruction, image denoising, image dehazing, image deblurring, image contrast enhancement, image demosaicing, image deraining, image color enhancement, image brightness enhancement, image detail enhancement, and image dynamic range enhancement.
[0046] A fifth aspect of the present application provides an image processing device, which may include a processor. The processor is coupled to a memory, and the memory stores program instructions. When the program instructions stored in the memory are executed by the processor, the methods described in the first aspect or the second aspect are implemented. For the steps in each possible implementation manner of the processor executing the first aspect or the second aspect, reference may be specifically made to the first aspect, and details are not described herein again.
[0047] A sixth aspect of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program runs on a computer, the computer is caused to execute the methods described in the first aspect or the second aspect.
[0048] A seventh aspect of the present application provides a circuit system, and the circuit system includes a processing circuit configured to execute the methods described in the first aspect or the second aspect.
[0049] The eighth aspect of this application provides a computer program product which, when running on a computer, enables the computer to execute the method described in the first aspect or the second aspect above.
[0050] The ninth aspect of this application provides a chip system. The chip system includes a processor for supporting a server or a threshold value acquisition device to implement the functions involved in the first aspect or the second aspect above. For example, it is used to send or process the data and / or information involved in the above method. In a possible design, the chip system further includes a memory for storing the necessary program instructions and data of the server or the communication device. The chip system can be composed of chips or can include chips and other discrete devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 A structural schematic diagram of an artificial intelligence main framework;
[0052] Figure 2 A schematic diagram of a convolutional neural network provided by an embodiment of this application;
[0053] Figure 3 A schematic diagram of a convolutional neural network provided by an embodiment of this application;
[0054] Figure 4 A schematic diagram of a system architecture provided by an embodiment of this application;
[0055] Figure 5 A flowchart of an image enhancement method provided by an embodiment of this application;
[0056] Figure 6 An example diagram of an image enhancement method provided by an embodiment of this application;
[0057] Figure 7 Another example diagram of an image enhancement method provided by an embodiment of this application;
[0058] Figure 8 A structural schematic diagram of an RCSA provided by an embodiment of this application;
[0059] Figure 9 A structural schematic diagram of an encoder provided by an embodiment of this application;
[0060] Figure 10 A structural schematic diagram of a residual module provided by an embodiment of this application;
[0061] Figure 11 A structural schematic diagram of a decoder provided by an embodiment of this application;
[0062] Figure 12Schematic flowchart of a partition acceleration method provided by an embodiment of the present application;
[0063] Figure 13 Schematic diagram of a network architecture provided by an embodiment of the present application;
[0064] Figure 14 Schematic diagram of a partition acceleration enhancement module 500 provided by an embodiment of the present application;
[0065] Figure 15 Schematic diagram of a network architecture for performing image deblurring provided by an embodiment of the present application;
[0066] Figure 16 Schematic diagram for comparing the effects of an image enhancement method provided by an embodiment of the present application;
[0067] Figure 17 Schematic diagram for comparing the effects of an image deblurring provided by an embodiment of the present application;
[0068] Figure 18 Another schematic diagram for comparing the effects of an image deblurring provided by an embodiment of the present application;
[0069] Figure 19 Another schematic diagram for comparing the effects of an image deblurring provided by an embodiment of the present application;
[0070] Figure 20 Schematic flowchart of a model training method provided by an embodiment of the present application;
[0071] Figure 21 Schematic diagram of the structure of an image processing device provided by an embodiment of the present application;
[0072] Figure 22 Schematic diagram of the structure of a model training device provided by an embodiment of the present application;
[0073] Figure 23 Schematic diagram of the structure of an execution device provided by an embodiment of the present application;
[0074] Figure 24 Schematic diagram of the structure of a chip provided by an embodiment of the present application. Detailed implementation manners
[0075] The embodiments of the present invention will be described below with reference to the accompanying drawings in the embodiments of the present invention. The terms used in the embodiments of the present invention are only for explaining the specific embodiments of the present invention, and are not intended to limit the present invention.
[0076] The embodiments of the present application will be described below in conjunction with the accompanying drawings. As can be known to those of ordinary skill in the art, with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0077] The terms "first", "second", etc. in the description and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing when describing objects with the same attributes in the embodiments of the present application. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device comprising a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.
[0078] For ease of understanding, the technical terms involved in the embodiments of the present application will be explained below.
[0079] Peak Signal-to-Noise Ratio (PSNR): An engineering term representing the ratio of the maximum possible power of a signal to the power of the destructive noise that affects its representation accuracy. In the field of image processing, PSNR is a method for evaluating the quality of signal reconstruction. PSNR can usually be simply defined by the mean square error. Generally speaking, the higher the PSNR of the reconstructed image, the smaller the gap between the reconstructed image and the true image.
[0080] Structural Similarity (SSIM): An index used to measure the similarity between two images. The higher the SSIM of the reconstructed image, the more similar the structure of the reconstructed image is to the true image.
[0081] Image enhancement: Refers to the technology of processing the brightness, color, contrast, saturation, dynamic range, etc. of an image to meet a certain specific index.
[0082] Image resolution: The resolution of an image is represented by the number of horizontal pixels and vertical pixels of the image. For example: 480P represents 640×480 pixels, and 960P represents 1280×960 pixels.
[0083] Receptive Field: In a convolutional neural network, the receptive field refers to the size of the area on the input image that the pixels on the feature map output by each layer of the convolutional neural network are mapped to. That is, the points on the feature map are calculated from all the pixels in the receptive field area of the input image. The larger the value of the receptive field, the larger the range of the original image that the points on the feature map can touch, which also means that more global and higher-level semantic features can be obtained. On the contrary, the smaller the value of the receptive field, the more local and detailed the features contained in the points on the feature map tend to be.
[0084] First, the overall workflow of the artificial intelligence system is described. Please refer to Figure 1 , Figure 1 which shows a schematic structural diagram of the artificial intelligence main framework. The above artificial intelligence theme framework is elaborated from two dimensions: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). Among them, the "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it can be the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, data undergoes the refinement process of "data - information - knowledge - wisdom". The "IT value chain" reflects the value brought by artificial intelligence to the information technology industry from the underlying infrastructure of artificial intelligence, information (provision and processing technology implementation) to the industrial ecosystem of the system.
[0085] (1) Infrastructure.
[0086] The infrastructure provides computing power support for the artificial intelligence system, enables communication with the external world, and is supported through the basic platform. It communicates with the external world through sensors; the computing power is provided by intelligent chips (such as CPU, NPU, GPU, ASIC, FPGA and other hardware acceleration chips); the basic platform includes relevant platform guarantees and supports such as distributed computing frameworks and networks, and can include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the external world to obtain data, and these data are provided to the intelligent chips in the distributed computing system provided by the basic platform for calculation.
[0087] (2) Data.
[0088] The data at the upper layer of the infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voices, texts, and also involves Internet of Things data of traditional devices, including business data of existing systems and perception data such as force, displacement, liquid level, temperature, humidity, etc.
[0089] (3) Data processing.
[0090] Data processing generally includes data training, machine learning, deep learning, search, inference, decision-making, etc.
[0091] Among them, machine learning and deep learning can perform symbolic and formal intelligent information modeling, extraction, preprocessing, training, etc. on data.
[0092] Inference refers to the process of simulating the intelligent reasoning mode of humans in a computer or intelligent system, and using formal information to perform machine thinking and solve problems according to the inference control strategy. The typical function is search and matching.
[0093] Decision-making refers to the process of making decisions after intelligent information is inferred, and usually provides functions such as classification, sorting, prediction, etc.
[0094] (4) General capabilities.
[0095] After the data undergoes the above-mentioned data processing, some general capabilities can be further formed based on the results of the data processing. For example, it can be an algorithm or a general system. For example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.
[0096] (5) Intelligent products and industry applications.
[0097] Intelligent products and industry applications refer to the products and applications of artificial intelligence systems in various fields. It is the encapsulation of the overall artificial intelligence solution, which productizes intelligent information decision-making and realizes practical applications. Its application fields mainly include: intelligent terminals, intelligent transportation, intelligent healthcare, autonomous driving, smart cities, etc.
[0098] The method provided by this application will be described below from the model training side and the model application side:
[0099] The model training method provided by the embodiments of this application can be specifically applied to data processing methods such as data training, machine learning, and deep learning, perform symbolic and formal intelligent information modeling, extraction, preprocessing, training, etc. on the training data, and finally obtain a trained neural network model (such as the target neural network model in the embodiments of this application); and the target neural network model can be used for model inference. Specifically, the input data can be input into the target neural network model to obtain output data.
[0100] Since the embodiments of this application involve a large number of applications of neural networks, for the sake of easy understanding, the relevant terms and related concepts such as neural networks involved in the embodiments of this application will be introduced below.
[0101] (1) Neural network.
[0102] A neural network can be composed of neural units. A neural unit can refer to an operation unit that takes xs (i.e., input data) and an intercept of 1 as inputs. The output of this operation unit can be:
[0103] where s = 1, 2, ……, n, n is a natural number greater than 1, Ws is the weight of xs, b is the bias of the neural unit, and f is the activation function of the neural unit, which is used to introduce non - linear characteristics into the neural network to convert the input signal in the neural unit into an output signal. The output signal of this activation function can be used as the input of the next convolutional layer, and the activation function can be a sigmoid function. A neural network is a network formed by connecting multiple such single neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, and the local receptive field can be a region composed of several neural units.
[0104] (2) A Convolutional Neural Network (CNN) is a deep neural network with a convolutional structure. A convolutional neural network contains a feature extractor composed of convolutional layers and subsampling layers. This feature extractor can be regarded as a filter, and the convolution process can be regarded as performing convolution on an input image or a convolutional feature plane using a trainable filter. A convolutional layer refers to the layer of neurons in a convolutional neural network that performs convolution processing on the input signal (such as the first convolutional layer and the second convolutional layer in this embodiment). In the convolutional layer of a convolutional neural network, a neuron can only be connected to some adjacent - layer neurons. In a convolutional layer, there are usually several feature planes, and each feature plane can be composed of some rectangular - arranged neural units. The neural units in the same feature plane share weights, and the shared weight here is the convolutional kernel. Sharing weights can be understood as a way of extracting image information that is independent of position. The underlying principle here is that the statistical information of a certain part of an image is the same as that of other parts. That is to say, the image information learned in one part can also be used in another part. Therefore, for all positions on the image, we can use the same learned image information. In the same convolutional layer, multiple convolutional kernels can be used to extract different image information. Generally, the more the number of convolutional kernels, the richer the image information reflected by the convolution operation.
[0105] The convolutional kernel can be initialized in the form of a matrix with a random size, and during the training process of the convolutional neural network, the convolutional kernel can learn to obtain reasonable weights. Additionally, the direct benefit brought by sharing weights is to reduce the connections between the layers of the convolutional neural network and at the same time reduce the risk of overfitting.
[0106] Specifically, as Figure 2 shown, the convolutional neural network (CNN) 100 may include an input layer 110, a convolutional layer / pooling layer 120, where the pooling layer is optional, and a neural network layer 130.
[0107] Among them, the structure composed of the convolutional layer / pooling layer 120 and the neural network layer 130 may be the first convolutional layer and the second convolutional layer described in this application. The input layer 110 is connected to the convolutional layer / pooling layer 120, the convolutional layer / pooling layer 120 is connected to the neural network layer 130, and the output of the neural network layer 130 can be input to the activation layer, and the activation layer can perform non-linear processing on the output of the neural network layer 130.
[0108] The convolutional layer / pooling layer 120. Convolutional layer: As Figure 2 shown, the convolutional layer / pooling layer 120 may include layers such as examples 121 - 126. In one implementation, layer 121 is a convolutional layer, layer 122 is a pooling layer, layer 123 is a convolutional layer, layer 124 is a pooling layer, 125 is a convolutional layer, and 126 is a pooling layer; in another implementation, 121 and 122 are convolutional layers, 123 is a pooling layer, 124 and 125 are convolutional layers, and 126 is a pooling layer. That is, the output of the convolutional layer can be used as the input of the subsequent pooling layer or as the input of another convolutional layer to continue the convolutional operation.
[0109] Taking the convolutional layer 121 as an example, the convolutional layer 121 can include many convolutional operators, which are also called kernels. Their role in image processing is equivalent to a filter that extracts specific information from the input image matrix. Essentially, a convolutional operator can be a weight matrix, which is usually predefined. During the process of performing convolution operations on an image, the weight matrix usually processes the input image pixel by pixel (or two pixels by two pixels... depending on the value of the stride) along the horizontal direction, thus completing the work of extracting specific features from the image. The size of this weight matrix should be related to the size of the image. It should be noted that the depth dimension of the weight matrix and the depth dimension of the input image are the same. During the convolution operation, the weight matrix extends to the entire depth of the input image. Therefore, convolving with a single weight matrix will produce a convolved output with a single depth dimension. However, in most cases, instead of using a single weight matrix, multiple weight matrices with the same dimension are applied. The outputs of each weight matrix are stacked to form the depth dimension of the convolved image. Different weight matrices can be used to extract different features from the image. For example, one weight matrix is used to extract edge information of the image, another weight matrix is used to extract specific colors of the image, and yet another weight matrix is used to blur the unwanted noise in the image... These multiple weight matrices have the same dimension, and the feature maps extracted by these multiple weight matrices with the same dimension also have the same dimension. Then, the multiple feature maps with the same dimension that are extracted are combined to form the output of the convolution operation.
[0110] The weight values in these weight matrices need to be obtained through a large amount of training in actual applications. Each weight matrix formed by the weight values obtained through training can extract information from the input image, thus helping the convolutional neural network 100 to make correct predictions.
[0111] When the convolutional neural network 100 has multiple convolutional layers, the initial convolutional layer (such as 121) often extracts more general features, which can also be called low-level features; as the depth of the convolutional neural network 100 increases, the later convolutional layers (such as 126) extract more and more complex features, such as high-level semantic features. The higher the semantic features, the more suitable they are for the problem to be solved.
[0112] Pooling layer: Since it is often necessary to reduce the number of training parameters, a pooling layer is often introduced periodically after the convolutional layer, that is, as shown in Figure 2 in Figure 120, each layer of 121 - 126 can be a convolutional layer followed by a pooling layer, or multiple convolutional layers followed by one or more pooling layers.
[0113] Neural network layer 130: After being processed by the convolutional layer / pooling layer 120, the convolutional neural network 100 is still not sufficient to output the required output information. As mentioned before, the convolutional layer / pooling layer 120 only extracts features and reduces the parameters brought by the input image. However, in order to generate the final output information (the required class information or other relevant information), the convolutional neural network 100 needs to use the neural network layer 130 to generate one or a group of outputs with the number of classes required. Therefore, the neural network layer 130 may include multiple hidden layers (such as Figure 2 131, 132 to 13n shown) and an output layer 140. The parameters contained in the multiple hidden layers can be pre-trained according to the relevant training data of the specific task type. For example, the task type may include image recognition, image classification, image super-resolution reconstruction, etc.
[0114] After the multiple hidden layers in the neural network layer 130, that is, the last layer of the entire convolutional neural network 100 is the output layer 140. The output layer 140 has a loss function similar to categorical cross-entropy, which is specifically used to calculate the prediction error. Once the forward propagation of the entire convolutional neural network 100 (such as Figure 2 the propagation from 110 to 140 is the forward propagation) is completed, the backward propagation (such as Figure 2 the propagation from 140 to 110 is the backward propagation) will start to update the weight values and biases of the previously mentioned layers to reduce the loss of the convolutional neural network 100 and the error between the result output by the convolutional neural network 100 through the output layer and the ideal result.
[0115] It should be noted that, as Figure 2 shown, the convolutional neural network 100 is only an example of a convolutional neural network. In specific applications, the convolutional neural network may also exist in the form of other network models. For example, as Figure 3 shown, multiple convolutional layers / pooling layers are parallel, and the features extracted separately are all input to the fully neural network layer 130 for processing.
[0116] (3) Deep neural network.
[0117] A deep neural network (DNN), also known as a multi-layer neural network, can be understood as a neural network with many hidden layers. Here, "many" does not have a specific measurement standard. Dividing the DNN according to the positions of different layers, the neural network inside the DNN can be divided into three categories: the input layer, the hidden layer, and the output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the middle layers are all hidden layers. The layers are fully connected, that is to say, any neuron in the i-th layer must be connected to any neuron in the (i + 1)-th layer. Although the DNN looks very complex, in terms of the work of each layer, it is actually not complex. Simply put, it is the following linear relationship expression: Among them, is the input vector, is the output vector, is the bias vector, W is the weight matrix (also known as the coefficient), and α() is the activation function. Each layer simply performs the following simple operation on the input vector to obtain the output vector Since the DNN has many layers, the number of coefficients W and bias vectors is also very large. The definitions of these parameters in the DNN are as follows: Taking the coefficient W as an example: Suppose in a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer where the coefficient W is located, and the subscripts correspond to the index 2 of the output third layer and the index 4 of the input second layer. In summary: The coefficient from the k-th neuron in the (L - 1)-th layer to the j-th neuron in the L-th layer is defined as It should be noted that the input layer does not have the W parameter. In a deep neural network, more hidden layers enable the network to better depict complex situations in the real world. Theoretically speaking, the more parameters a model has, the higher its complexity and the larger its "capacity", which means it can complete more complex learning tasks. Training a deep neural network is also the process of learning the weight matrix, and its ultimate goal is to obtain the weight matrices of all layers of the trained deep neural network (the weight matrix formed by vectors W of many layers).
[0118] (4) Loss function.
[0119] During the process of training a deep neural network, since it is desired that the output of the deep neural network be as close as possible to the value that is truly desired to be predicted, the weight vectors of each layer of the neural network can be updated by comparing the predicted value of the current network with the truly desired target value and then according to the difference between the two. (Of course, there is usually an initialization process before the first update, that is, parameters are pre-configured for each layer in the deep neural network). For example, if the predicted value of the network is too high, the weight vector is adjusted to make it predict lower, and continuous adjustment is made until the deep neural network can predict the truly desired target value or a value very close to the truly desired target value. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or the objective function. They are important equations used to measure the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. Then the training of the deep neural network becomes a process of minimizing this loss as much as possible.
[0120] (5) Backpropagation algorithm.
[0121] The convolutional neural network can use the error backpropagation (BP) algorithm to correct the magnitudes of the parameters in the initial super-resolution model during the training process, making the reconstruction error loss of the super-resolution model smaller and smaller. Specifically, the forward propagation of the input signal until the output will generate an error loss, and the initial super-resolution model parameters are updated by backpropagating the error loss information, so that the error loss converges. The backpropagation algorithm is a backpropagation movement dominated by the error loss, aiming to obtain the optimal parameters of the super-resolution model, such as the weight matrix.
[0122] (6) Linear operation.
[0123] Linearity refers to the relationship between quantities in proportion and in a straight line. Mathematically, it can be understood as a function with a constant first derivative. Linear operations can include but are not limited to addition operations, no-operation, identity operation, convolution operation, batch normalization BN operation, and pooling operation. Linear operations can also be called linear mappings. Linear mappings need to satisfy two conditions: homogeneity and additivity. If either condition is not met, it is non-linear.
[0124] Among them, homogeneity means f(ax) = af(x); additivity means f(x + y) = f(x) + f(y); for example, f(x) = ax is linear. It should be noted that x, a, and f(x) here are not necessarily scalars, but can be vectors or matrices, forming a linear space of any dimension. If x and f(x) are n-dimensional vectors, when a is a constant, it equivalently satisfies homogeneity, and when a is a matrix, it equivalently satisfies additivity. Relatively speaking, a function graph that is a straight line does not necessarily conform to a linear mapping. For example, f(x) = ax + b neither satisfies homogeneity nor additivity, so it belongs to a non-linear mapping.
[0125] In the embodiments of the present application, the composition of multiple linear operations can be called a linear operation, and each linear operation included in the linear operation can also be called a sub-linear operation.
[0126] Figure 4 is a schematic diagram of a system architecture provided by an embodiment of the present application. In Figure 4 it, the execution device 110 is configured with an input / output (I / O) interface 112 for data interaction with external devices, and the user can input data to the I / O interface 112 through the client device 140.
[0127] During the preprocessing of the input data by the execution device 120, or during the relevant processing such as calculation by the calculation module 111 of the execution device 120 (such as implementing the functions of the neural network in the present application), the execution device 120 can call data, code, etc. in the data storage system 150 for corresponding processing, or store the data, instructions, etc. obtained from the corresponding processing into the data storage system 150.
[0128] Finally, the I / O interface 112 returns the processing result to the client device 140 and thus provides it to the user.
[0129] Optionally, the client device 140 can be, for example, a control unit in an autonomous driving system, a functional algorithm module in a mobile phone terminal. For example, the functional algorithm module can be used to implement related tasks.
[0130] It is worth noting that the training device 120 can generate corresponding target models / rules (such as the target neural network model in this embodiment) based on different targets or different tasks and different training data. The corresponding target models / rules can be used to achieve the above targets or complete the above tasks, so as to provide the required results for the user.
[0131] In Figure 4In the situation shown, the user can manually input data, and this manual input can be operated through the interface provided by the I / O interface 112. In another situation, the client device 140 can automatically send input data to the I / O interface 112. If the client device 140 is required to automatically send input data and user authorization is needed, the user can set the corresponding permissions in the client device 140. The user can view the results output by the execution device 110 on the client device 140, and the specific presentation forms can be display, sound, actions, etc. The client device 140 can also be used as a data acquisition terminal to collect the input data input to the I / O interface 112 and the output results of the output I / O interface 112 as new sample data, and store them in the database 130. Of course, it is also possible not to collect data through the client device 140, but for the I / O interface 112 to directly use the input data input to the I / O interface 112 and the output results of the output I / O interface 112 as new sample data and store them in the database 130.
[0132] It should be noted that Figure 4 This is only a schematic diagram of a system architecture provided by an embodiment of the present application. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, in Figure 4 the data storage system 150 is an external memory relative to the execution device 110. In other cases, the data storage system 150 can also be placed in the execution device 110.
[0133] Deep learning methods, especially the convolutional neural network method based on CNN, are the key driving forces for the development in the field of artificial intelligence in recent years and have achieved remarkable results in various tasks of computer vision. In the field of image enhancement, deep learning-based methods have surpassed traditional methods.
[0134] Currently, deep learning methods usually process images through complex neural networks. The more layers a neural network contains and the more channels in each layer, the more complex this neural network is. Complex neural networks usually have a larger receptive field, thus bringing better image enhancement quality. However, with the popularization of ultra-high-definition resolution images / videos, complex neural networks usually bring exponentially increasing computational overhead and power consumption overhead. For devices with relatively weak computing capabilities (such as consumer terminal products like mobile phones or tablets), it is usually difficult to use complex deep learning methods for image enhancement due to factors such as running speed, running video memory, and power consumption.
[0135] In view of this, the embodiments of the present application provide a network architecture that can simultaneously take into account a large receptive field and locally optimal convolution processing, and the network scale is much smaller than that of traditional networks. Specifically, the network architecture adopts a multi-layer network structure, and the upper-layer network and the lower-layer network are respectively used to process low-resolution images and high-resolution images. Moreover, the output of the upper-layer network is passed to the lower-layer network for further processing. Since the upper-layer network processes low-resolution images, a large receptive field can be obtained, and the output of the upper-layer network is further processed by the lower-layer network that processes high-resolution images, so as to simultaneously take into account global information and local information and ensure the quality of image enhancement. Through the design of the multi-layer network, on the basis of obtaining a large receptive field, the depth of the network can be effectively reduced, thus significantly reducing the network scale, enabling the network to be deployed and run on devices with weak computing power.
[0136] The image enhancement method provided by the embodiments of the present application can be applied to a terminal, especially a terminal with weak computing power. Exemplarily, the terminal can be, for example, a digital camera, a surveillance camera device, a mobile phone, a personal computer (PC), a laptop computer, a server, a tablet computer, a smart TV, a mobile internet device (MID), a wearable device, a virtual reality (VR) device, an augmented reality (AR) device, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical surgery, a wireless terminal in a smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, a wireless terminal in a smart home, etc. For the convenience of description, hereinafter, the image enhancement method provided by the embodiments of the present application will be introduced by taking the application of the image enhancement method to a terminal as an example.
[0137] It should be understood that the images in the embodiments of the present application can be static images (or called static pictures) or dynamic images (or called dynamic pictures), such as RGB images, black-and-white images or grayscale images, etc. For the convenience of description, in the following embodiments of the present application, static images or dynamic images will be uniformly referred to as images.
[0138] It can be referred to Figure 5 and Figure 6 , Figure 5Schematic flowchart of an image enhancement method provided by an embodiment of the present application. Figure 6 Example diagram of an image enhancement method provided by an embodiment of the present application. As Figure 5 shown, the image enhancement method includes the following steps 501-504.
[0139] Step 501, the terminal obtains a first image and a second image, where the second image is obtained by performing downsampling processing on the first image.
[0140] In this embodiment, the image enhancement method can be applied to implement at least one of the following image enhancement tasks: image super-resolution reconstruction, image denoising, image dehazing, image deblurring, image contrast enhancement, image demosaicing, image deraining, image color enhancement, image brightness enhancement, image detail enhancement, and image dynamic range enhancement.
[0141] Among them, the first image is the image to be enhanced, such as an image that needs to be deblurred. The terminal can obtain the second image by performing downsampling processing on the first image, and the resolution of the second image is lower than that of the first image. There are various ways for the terminal to perform downsampling processing, including but not limited to methods such as bilinear interpolation and trilinear interpolation. In practical applications, the terminal selects the downsampling multiple according to actual needs, such as multiples of 2 times, 3 times, or 4 times, etc. Exemplarily, when the resolution of the first image is 1000×1000, and the downsampling multiple selected by the terminal is 2 times, after the terminal performs downsampling processing on the first image, the resolution of the obtained second image is 500×500.
[0142] Step 502, the terminal processes the second image through a first network to obtain a first feature and a third image, where the first feature is an intermediate feature extracted by the first network, and the resolution of the third image is the same as that of the first image.
[0143] Among them, the first network can be a network for performing image enhancement, such as a convolutional neural network.
[0144] Step 503, the terminal generates a fourth image according to the third image and the first image.
[0145] In this embodiment, after obtaining the third image based on the first network, the terminal can fuse the information in the third image and the first image to generate a fourth image, which includes the feature information of the third image and the feature information of the first image.
[0146] Optionally, the resolution of the third image obtained by the terminal through the first network may be the same as that of the first image. The terminal may perform an addition process on the third image and the first image to obtain a fourth image. Exemplarily, when the resolutions of the first image and the third image are the same, the terminal adds the pixel values at the same positions on the first image and the third image to obtain a new pixel value at that position. By adding the pixel values at each position of the first image and the third image, the fourth image is obtained.
[0147] Step 504, the terminal processes the fourth image and the first feature through the second network to obtain a target image.
[0148] Wherein, the second network may be a network for performing image enhancement, such as a convolutional neural network. Optionally, the structure of the second network may be similar to or the same as the structure of the first network.
[0149] In this embodiment, since the first image is subjected to downsampling processing and the second image obtained by the downsampling processing is processed by the first network, the feature map extracted by the first network can have a larger receptive field, that is, the first network can capture more global feature information. After the image information output by the first network is fused with the first image and then input into the second network for processing, the second network can capture local feature information based on the global feature information of the first network, so as to achieve both global feature information and local feature information, and ensure the quality of the enhanced image.
[0150] Generally speaking, in traditional networks, a larger receptive field is obtained by increasing the network depth horizontally. This method usually requires a very deep network depth (such as hundreds of convolutional layers) to obtain an ideal receptive field. In this embodiment, a larger receptive field is obtained by increasing the network levels vertically. Only a few layers of networks are required to obtain a larger receptive field, which can effectively reduce the network depth and significantly reduce the network scale, enabling the deployment and operation of a network for image enhancement on devices with weak computing capabilities.
[0151] Exemplarily, reference may be made to Figure 7 , Figure 7 which is another example diagram of an image enhancement method provided by an embodiment of the present application. As Figure 7 shown, the first network includes a first sub-network and a second sub-network. The process of the terminal processing the second image through the first network includes: the terminal processes the second image through the first sub-network to obtain a first feature, and the first feature is the intermediate feature extracted by the first sub-network; the terminal processes the first feature through the second sub-network to obtain the third image.
[0152] The second network includes a third sub-network and a fourth sub-network. The process of the terminal performing image processing on the fourth image through the second network includes: the terminal processes the fourth image through the third sub-network to obtain a second feature, and the second feature is the sub-feature extracted by the third sub-network. The terminal generates a third feature according to the first feature and the second feature; the terminal processes the third feature through the fourth sub-network to obtain the target image. Among them, the terminal may fuse the first feature and the second feature to obtain the third feature. Since the first feature is extracted based on the second image with a lower resolution, and the second feature is extracted based on the fourth image with a higher resolution, the terminal may perform upsampling processing on the first feature before fusing the first feature and the second feature. For example, the terminal performs upsampling processing on the first feature to obtain an upsampled feature, and then the terminal performs an addition operation on the upsampled feature and the second feature to obtain the third feature.
[0153] In this solution, by fusing the feature-level information extracted by the upper layer network and the lower layer network, as well as the image-level information between the upper and lower layer networks, the lower layer network can take into account the feature information obtained by the upper layer network as much as possible while processing local information, prompting the depth design of each layer of the network to be as simple as possible.
[0154] Optionally, the first network may be implemented using an encoder-decoder structure. Specifically, the first sub-network includes an encoder, and the first sub-network is used to obtain encoded features through the encoder; the second sub-network includes a decoder, and the second sub-network is used to obtain a decoded image through the decoder. Generally speaking, an encoder can be used to perform feature extraction and compression. By performing feature extraction on the input image and compressing the extracted features, encoded features are obtained. A decoder can be used to perform feature restoration operations. By performing upsampling and feature restoration on the encoded features, a decoded image is obtained.
[0155] Exemplarily, the encoder may include a first convolutional layer and a Residual Channel Spatial Attention (RCSA) module. For example, in the encoder, the first convolutional layer and the RCSA connected in sequence form an encoding unit, and multiple encoding units are connected in sequence to form the encoder. Reference may be made to Figure 8 and Figure 9 , Figure 8 which is a schematic structural diagram of an RCSA provided by an embodiment of the present application; Figure 9 which is a schematic structural diagram of an encoder provided by an embodiment of the present application. As Figure 9 shown, the encoder includes three encoding units connected in sequence, and each encoding unit includes a first convolutional layer and an RCSA.
[0156] Exemplarily, the decoder includes a second convolutional layer and a residual module. For example, in the decoder, the second convolutional layer and the residual module connected in sequence form a decoding unit, and multiple decoding units are connected in sequence to form the decoder. Refer to Figure 10 and Figure 11 , Figure 10 which is a schematic structural diagram of a residual module provided by an embodiment of the present application; Figure 11 which is a schematic structural diagram of a decoder provided by an embodiment of the present application. As Figure 11 shown, the decoder includes three decoding units connected in sequence, and each decoding unit includes a residual module and a second convolutional layer.
[0157] In a possible embodiment, in order to improve the execution efficiency of the first sub-network and the second sub-network, a partition acceleration method is also provided in this embodiment, which can accelerate the processing speed of the first sub-network and the second sub-network. In this partition method, the input is divided into multiple non-overlapping regions in the spatial dimension, the divided regions are concatenated in the channel dimension, and then sent to the encoder or decoder for parallel computing. After obtaining the encoded features output by the encoder or the decoded features output by the decoder, they are then decomposed in the channel dimension and then dimensionally restored in the spatial dimension. By performing parallel computing on multiple regions in multiple partitions, the running speed of the network is reduced.
[0158] Exemplarily, the terminal processes the second image through the first sub-network to obtain a first feature, including: the terminal divides the second image in the spatial dimension to obtain a plurality of non-overlapping sub-images. The terminal concatenates the plurality of sub-images in the channel dimension to obtain a concatenated image, and the number of channels of the concatenated image is N times the number of channels of the second image. The terminal processes the concatenated image through the encoder to obtain encoded features. The terminal divides the encoded features in the channel dimension to obtain a plurality of sub-encoded features. Among them, the number of sub-encoded features is the same as the number of sub-images. The terminal concatenates the plurality of sub-encoded features in the spatial dimension to obtain the first feature.
[0159] For example, assume that the second image input to the first sub-network is represented as H×W×C. Among them, H and W belong to the spatial dimension, H represents the width of the second image (Width), and W represents the height of the second image (Height); C belongs to the channel dimension, and C represents the number of channels of the second image. For example, when the second image is an RGB image, the number of channels of the second image is 3.
[0160] Specifically, refer to Figure 12 , Figure 12 which is a schematic flow diagram of a partition acceleration method provided by an embodiment of the present application. As Figure 12As shown, the process of the terminal processing the second image H×W×C through the partition acceleration method includes the following steps S1 - S5.
[0161] S1. The terminal performs non - overlapping segmentation on the second image H×W×C in the spatial dimension (H, W dimensions) to obtain a plurality of non - overlapping sub - images. For example, the second image H×W×C is divided into N = 2 n regions, the H dimension of the second image is segmented times, and the W dimension of the second image is segmented times. For example, as Figure 12 shown, the terminal divides the image into four sub - images to obtain sub - image 1, sub - image 2, sub - image 3, and sub - image 4.
[0162] where floor represents rounding down, N = N h ×N w . In this way, the dimension of each sub - image obtained by segmentation becomes H / N h ×W / N w ×C.
[0163] S2. The terminal splices the obtained multiple sub - images in the channel dimension (i.e., C dimension) to obtain a spliced image. Among them, since the spliced image is obtained by splicing N sub - images in the channel dimension, the dimension of the spliced image is H / N h ×W / N w ×NC.
[0164] S3. The terminal sends the spliced image into an encoder for encoding processing to obtain encoded features, and the dimension of the encoded features is where the dimension of the encoded features output by the encoder is related to the structure of the encoder. In practical applications, the structure of the encoder can be set according to actual needs, so as to determine the dimension of the encoded features.
[0165] S4. The terminal evenly divides the encoded features output by the encoder into N sub - encoded features in the channel dimension, and the dimension of each sub - encoded feature is
[0166] S5. The terminal splices the N sub - encoded features obtained by segmentation in the spatial dimension to obtain a first feature. The dimension of the first feature obtained after splicing is
[0167] In this embodiment, by introducing a partition acceleration method to accelerate the processing speed of the encoder and decoder, it is possible to improve the operating efficiency of the neural network on the premise of ensuring that the enhanced image quality is not affected, and achieve an effect equivalent to the enhanced quality of the original method and meeting real-time processing requirements in specific scenarios. In addition, after introducing the partition acceleration method, the partitions for processing different sub-images have self-learning differential convolution kernels, which can achieve differential learning, thereby making the local enhancement of the image better.
[0168] In the above-described embodiments, the image enhancement method provided by the embodiments of the present application is introduced by taking the network architecture including two layers of networks (i.e., the first network and the second network) as an example. In actual applications, the number of layers of the networks included in the network architecture can also be greater than 2 layers, such as 3 layers or 4 layers, etc. For ease of understanding, the following will introduce the specific process of the image enhancement method provided by the embodiments of the present application when the network architecture includes n layers (n > 2) of networks.
[0169] Reference can be made to Figure 13 , Figure 13 which is a schematic diagram of a network architecture provided by an embodiment of the present application.
[0170] As Figure 13 shown, the network architecture for executing the image enhancement method includes an input image processing unit 100, an encoder unit group 200, an encoded feature processing unit group 300, and a decoder unit group 400. Among them, the encoder unit group 200 and the decoder unit group 400 also include adapted partition acceleration enhancement modules 500. Reference can be made to Figure 14 , Figure 14 which is a schematic diagram of a partition acceleration enhancement module 500 provided by an embodiment of the present application. As Figure 14 shown, the partition acceleration enhancement module 500 includes a spatial dimension splitting unit, a channel dimension splicing unit, an encoder / decoder, a channel dimension splitting unit, and a spatial dimension splicing unit.
[0171] In Figure 14 , the encoder unit group 200 includes multiple encoder units, the encoded feature processing unit group 300 includes multiple encoded feature processing units, and the decoder unit group 400 includes multiple decoder units. Each layer of the network includes an encoder unit, an encoded feature processing unit, and a decoder unit; therefore, the encoder unit group 200, the encoded feature processing unit group 300, and the decoder unit group 400 constitute a multi-layer network.
[0172] During the process of performing image enhancement, the input image processing unit 100 receives the original image I and performs downsampling on the original image I at different multiples to obtain a multi-resolution input image group {I2, I3, …, In}. Among them, the resolutions of images I_2 to I_n gradually decrease. For example, assuming the resolution of the original image I is H×W, the resolution of image I_2 can be H / 2×W / 2, and the resolution of image I_n can be H / n×W / n. For each layer of the network architecture, the resolutions of the input images of each layer are different. Moreover, from the top layer network to the bottom layer network, the resolutions of the input images gradually increase.
[0173] As Figure 14 shown, the input of the top layer network is image In. Image In sequentially passes through the encoder unit, the encoded feature processing unit, and the decoder unit in the top layer network, and outputs image On. Among them, the resolution of the output image On is H / (n - 1)×W / (n - 1), that is, the resolution of the output image On is the same as the resolution of the corresponding image In-1 of the next layer network.
[0174] After the top layer network outputs image On, the input image processing unit 100 adds the output image On and the input image In-1, and inputs the resulting residual image into the second layer network. In the second layer network, after the encoder unit processes the input image, it obtains encoded features and passes them to the encoded feature processing unit of the second layer network. The encoded feature processing unit in the second layer network obtains the encoded features obtained by the encoder processing unit of the second layer network and the encoded features obtained by the encoder processing unit of the previous layer network (i.e., the top layer network), and after performing upsampling on the encoded features obtained by the previous layer, adds the features obtained by the upsampling process and the encoded features obtained in the current layer to obtain the added residual features. The decoder unit in the second layer network receives and processes the added encoded features to obtain the output image On-1.
[0175] Similarly, in the networks other than the top layer network, the input of the encoder unit in this layer network is the residual image obtained by adding the input image of this layer network and the output image of the previous layer network. The input of the encoded feature processing unit in this layer network is the residual feature obtained by adding the encoded features output by the encoder unit in this layer network and the encoded features output by the encoder unit in the previous layer network and subjected to upsampling processing.
[0176] In this way, by sequentially performing image processing on each layer of the network until the final output image O is obtained at the bottom layer network.
[0177] To test the enhancement effect of the image enhancement method provided in the embodiments of the present application, in the embodiments of the present application, image deblurring is used as the image enhancement task, and the image deblurring effect is tested by using the image enhancement method provided in the embodiments of the present application.
[0178] Reference can be made to Figure 15 , Figure 15 which is a schematic diagram of a network architecture for performing image deblurring provided in the embodiments of the present application. Among them, Figure 15 the positions of the networks in each layer are opposite to those in Figure 13 but the processing flow is the same. Figure 13 In Figure 15 image processing is performed in the direction from the upper network to the lower network,
[0179] while in Figure 15 image processing is performed in the direction from the lower network to the upper network, but actually both are performed in the direction from the low-resolution image to the high-resolution image.
[0180] In addition, except for the top network that processes the original resolution image, partition acceleration enhancement modules are included in the encoder units and decoder units of each of the other networks to accelerate the speed of feature processing. Among them, since the image processed in the top network is of the original resolution, in order to ensure the coherence of image processing and avoid fragmentation in image processing caused by image block division, the partition acceleration enhancement module may not be used in the top network.
[0181] Reference can be made to Figure 16 , Figure 16 which is a schematic diagram for comparing the effects of an image enhancement method provided in the embodiments of the present application. Figure 16 The comparison results between the image enhancement method provided in this embodiment and multiple mainstream algorithms in the related art are given in Figure 16 . In Figure 16It can be seen that the PSNR corresponding to the image enhancement method provided in this embodiment is higher than that of other comparison algorithms, and the running time is also much less than that of other comparison algorithms.
[0182] In addition, this test was conducted on multiple data sets. Specifically, reference can be made to Figures 17 - 19 , Figure 17 which is a schematic diagram for comparing the effects of image deblurring provided by an embodiment of this application; Figure 18 which is another schematic diagram for comparing the effects of image deblurring provided by an embodiment of this application; Figure 19 which is yet another schematic diagram for comparing the effects of image deblurring provided by an embodiment of this application.
[0183] As Figure 17 shown, Figure 17 in (a) represents the input image, Figure 17 in (b) represents the ground truth image, Figure 17 in (c)-(j) represents the images obtained by the algorithms of related technologies; Figure 17 in (k) represents the image obtained by the algorithm of this embodiment. By comparing the images in Figure 17 , it can be seen that the algorithm processing provided by this embodiment has a better deblurring effect, the text edges and walls in the obtained image are clearer, and the image has higher PSNR and SSIM metrics.
[0184] Similarly, Figure 18 in (a) represents the input image, Figure 18 in (b) represents the ground truth image, Figure 18 in (c)-(j) represents the images obtained by the algorithms of related technologies; Figure 18 in (k) represents the image obtained by the algorithm of this embodiment. By comparing the images in Figure 18 , it can be seen that the algorithm processing provided by this embodiment has a better deblurring effect, the wall texture details in the obtained image are more prominent and clear, and the image has higher PSNR and SSIM metrics.
[0185] As Figure 19 shown, Figure 19 in (a) represents the input image, Figure 19 in (b) represents the ground truth image, Figure 19 in (c)-(j) represents the images obtained by the algorithms of related technologies; Figure 19 in (k) represents the image obtained by the algorithm of this embodiment. By comparing the images in Figure 19 , it can be seen that the algorithm processing provided by this embodiment has a better deblurring effect, the roof tile texture in the obtained image is clearer, the lines are more real, and the wall details are more prominent and clear.
[0186] Reference can be made to Figure 20 , Figure 20 which is a schematic flowchart of a model training method provided by an embodiment of the present application. As Figure 20 shown, a model training method provided by an embodiment of the present application includes steps 2001 to 2007.
[0187] Step 2001: Obtain an image sample pair, where the image sample pair includes a first image and an enhanced image corresponding to the first image.
[0188] In this embodiment, before the image training device performs model training, an image sample pair for training can be obtained. Among them, the first image and the enhanced image corresponding to the first image are two images in the same scene, and the image quality of the enhanced image corresponding to the first image is higher than that of the first image. Image quality refers to one or more of color, brightness, saturation, contrast, dynamic range, resolution, texture details, sharpness, etc. For example, the first image is a blurred image, and the enhanced image corresponding to the first image is a de-blurred image.
[0189] Step 2002: Perform downsampling processing on the first image to obtain a second image.
[0190] Step 2003: Process the second image through a first network to obtain a first feature and a third image, where the first feature is an intermediate feature extracted by the first network, and the resolution of the third image is the same as that of the first image.
[0191] Step 2004: Generate a fourth image according to the third image and the first image.
[0192] Step 2005: Process the fourth image through a second network to obtain a target image.
[0193] Among them, steps 2002 to 2005 are similar to steps 501 to 504 above. For details, reference can be made to steps 501 to 504, which will not be elaborated here.
[0194] Step 2006: Obtain a loss function according to the enhanced image corresponding to the first image and the target image, where the loss function is used to indicate the difference between the enhanced image and the target image in the image sample pair.
[0195] In this embodiment, after obtaining the target image, the loss function corresponding to the enhanced image and the target image corresponding to the first image can be obtained to determine the difference between the enhanced image and the target image in the image sample pair.
[0196] In a possible implementation, the loss function corresponding to the enhanced image and the target image can be obtained based on the reconstruction loss function and the gradient loss function to ensure that the enhanced image can meet the requirements of objective and subjective metrics. Exemplarily, the reconstruction loss function can use the L1 norm. The gradient loss function can be the loss representing the average gradient of the enhanced image and the target image in the x / y directions.
[0197] Step 2007: Train the first network and the second network according to the loss function to obtain the trained first network and the trained second network.
[0198] In this embodiment, after obtaining the loss function, the model parameters of the image processing model to be trained can be updated based on the loss function until the model training conditions are met (for example, the value of the loss function is less than a preset value), to obtain the trained first network and the trained second network. Among them, the trained first network and the trained second network can refer to Figure 5 the descriptions in the corresponding embodiments and will not be elaborated here.
[0199] Optionally, in a possible implementation, the first network includes a first sub-network and a second sub-network, and the second network includes a third sub-network and a fourth sub-network; processing the second image through the first network to obtain a first feature and a third image includes: processing the second image through the first sub-network to obtain a first feature; processing the first feature through the second sub-network to obtain the third image; processing the fourth image through the second network to obtain a target image includes: processing the fourth image through the third sub-network to obtain a second feature; generating a third feature according to the first feature and the second feature; processing the third feature through the fourth sub-network to obtain the target image.
[0200] Optionally, in a possible implementation, the first sub-network includes an encoder, and the first sub-network is used to obtain encoded features; the second sub-network includes a decoder, and the second sub-network is used to obtain a decoded image.
[0201] Optionally, in a possible implementation, the encoder includes a first convolutional layer and a channel-spatial attention module, and the decoder includes a second convolutional layer and a residual module.
[0202] Optionally, in a possible implementation, the processing of the second image by the first sub-network to obtain a first feature includes: segmenting the second image in the spatial dimension to obtain a plurality of sub-images; splicing the plurality of sub-images in the channel dimension to obtain a spliced image; processing the spliced image through an encoder to obtain an encoded feature; segmenting the encoded feature in the channel dimension to obtain a plurality of sub-encoded features; and splicing the plurality of sub-encoded features in the spatial dimension to obtain the first feature.
[0203] Optionally, in a possible implementation, the generating of a fourth image according to the third image and the first image includes: performing an addition process on the third image and the first image to obtain the fourth image.
[0204] Optionally, in a possible implementation, the training method of the model is used to implement at least one of the following image enhancement tasks: image super-resolution reconstruction, image denoising, image dehazing, image deblurring, image contrast enhancement, image demosaicing, image deraining, image color enhancement, image brightness enhancement, image detail enhancement, and image dynamic range enhancement.
[0205] Reference can be made to Figure 21 , Figure 21 which is a schematic structural diagram of an image processing device provided by an embodiment of the present application. As Figure 21 shown, an image processing device provided by an embodiment of the present application includes: an acquisition unit 2101 and a processing unit 2102. The acquisition unit 2101 is configured to acquire a first image and a second image, where the second image is obtained by performing downsampling processing on the first image; the processing unit 2102 is configured to process the second image through a first network to obtain a first feature and a third image; the processing unit 2102 is further configured to generate a fourth image according to the third image and the first image; the processing unit 2102 is further configured to process the fourth image through a second network to obtain a target image; wherein, the first network and the second network are used for image enhancement.
[0206] Optionally, in a possible implementation, the processing unit 2102 is further configured to process the second image through the first network to obtain the third image and a first feature, where the first feature is an intermediate feature extracted by the first network; the processing unit 2102 is further configured to process the fourth image and the first feature through a second network to obtain a target image.
[0207] Optionally, in a possible implementation, the first network includes a first sub-network and a second sub-network, and the second network includes a third sub-network and a fourth sub-network; the processing unit 2102 is further configured to process the second image through the first sub-network to obtain a first feature; the processing unit 2102 is further configured to process the first feature through the second sub-network to obtain the third image; the processing unit 2102 is further configured to process the fourth image through the third sub-network to obtain a second feature; the processing unit 2102 is further configured to generate a third feature according to the first feature and the second feature; the processing unit 2102 is further configured to process the third feature through the fourth sub-network to obtain the target image.
[0208] Optionally, in a possible implementation, the first sub-network includes an encoder, and the first sub-network is configured to obtain an encoded feature; the second sub-network includes a decoder, and the second sub-network is configured to obtain a decoded image.
[0209] Optionally, in a possible implementation, the encoder includes a first convolutional layer and a channel-spatial attention module, and the decoder includes a second convolutional layer and a residual module.
[0210] Optionally, in a possible implementation, the processing unit 2102 is further configured to segment the second image in the spatial dimension to obtain a plurality of sub-images; the processing unit 2102 is further configured to splice the plurality of sub-images in the channel dimension to obtain a spliced image; the processing unit 2102 is further configured to process the spliced image through an encoder to obtain an encoded feature; the processing unit 2102 is further configured to segment the encoded feature in the channel dimension to obtain a plurality of sub-encoded features; the processing unit 2102 is further configured to splice the plurality of sub-encoded features in the spatial dimension to obtain the first feature.
[0211] Optionally, in a possible implementation, the processing unit 2102 is further configured to add the third image and the first image to obtain a fourth image.
[0212] Optionally, in a possible implementation, the image processing device is configured to implement at least one of the following image enhancement tasks: image super-resolution reconstruction, image denoising, image dehazing, image deblurring, image contrast enhancement, image demosaicing, image deraining, image color enhancement, image brightness enhancement, image detail enhancement, and image dynamic range enhancement.
[0213] Reference can be made to Figure 22 , Figure 22 which is a schematic structural diagram of a model training device provided by an embodiment of this application. AsFigure 22 As shown in Figure 22 , a model training device provided by an embodiment of the present application includes: an acquisition unit 2201 and a processing unit 2202. The acquisition unit 2201 is configured to acquire an image sample pair, where the image sample pair includes a first image and an enhanced image corresponding to the first image; the processing unit 2202 is configured to perform downsampling processing on the first image to obtain a second image; the processing unit 2202 is further configured to process the second image through a first network to obtain a first feature and a third image, where the first feature is an intermediate feature extracted by the first network, and the resolution of the third image is the same as that of the first image; the processing unit 2202 is further configured to generate a fourth image according to the third image and the first image; the processing unit 2202 is further configured to process the fourth image through a second network to obtain a target image; the processing unit 2202 is further configured to obtain a loss function according to the enhanced image corresponding to the first image and the target image, where the loss function is used to indicate the difference between the enhanced image and the target image in the image sample pair; the processing unit 2202 is further configured to train the first network and the second network according to the loss function to obtain a trained first network and a trained second network; where the first network and the second network are used for image enhancement.
[0214] Optionally, in a possible implementation manner, the processing unit 2202 is further configured to process the second image through the first network to obtain the third image and the first feature, where the first feature is an intermediate feature extracted by the first network; the processing unit 2202 is further configured to process the fourth image and the first feature through a second network to obtain a target image.
[0215] Optionally, in a possible implementation manner, the first network includes a first sub-network and a second sub-network, and the second network includes a third sub-network and a fourth sub-network; the processing unit 2202 is further configured to process the second image through the first sub-network to obtain a first feature; the processing unit 2202 is further configured to process the first feature through the second sub-network to obtain the third image; the processing unit 2202 is further configured to process the fourth image through the third sub-network to obtain a second feature; the processing unit 2202 is further configured to generate a third feature according to the first feature and the second feature; and process the third feature through the fourth sub-network to obtain the target image.
[0216] Optionally, in a possible implementation manner, the first sub-network includes an encoder, and the first sub-network is configured to obtain an encoded feature; the second sub-network includes a decoder, and the second sub-network is configured to obtain a decoded image.
[0217] Optionally, in a possible implementation, the encoder includes a first convolutional layer and a channel-spatial attention module, and the decoder includes a second convolutional layer and a residual module.
[0218] Optionally, in a possible implementation, the processing unit 2202 is further used to: segment the second image in the spatial dimension to obtain multiple sub-images; splice the multiple sub-images in the channel dimension to obtain a spliced image; process the spliced image through an encoder to obtain a coding feature; segment the coding feature in the channel dimension to obtain multiple sub-coding features; and splice the multiple sub-coding features in the spatial dimension to obtain the first feature.
[0219] Optionally, in a possible implementation, the processing unit 2202 is further configured to perform addition processing on the third image and the first image to obtain a fourth image.
[0220] Optionally, in one possible implementation, the model training device is used to implement at least one of the following image enhancement tasks: image super-resolution reconstruction, image denoising, image dehazing, image deblurring, image contrast enhancement, image demosaicing, image deraining, image color enhancement, image brightness enhancement, image detail enhancement, and image dynamic range enhancement.
[0221] Next, an execution device provided by an embodiment of the present application is introduced. Figure 23 , Figure 23 This is a schematic diagram of a structure of an execution device provided in an embodiment of the present application. The execution device 2300 can be specifically a mobile phone, a tablet, a laptop, a smart wearable device, a server, etc., which is not limited here. Among them, the execution device 2300 can be deployed with Figure 23 The data processing device described in the corresponding embodiment is used to implement Figure 23 The data processing function in the corresponding embodiment. Specifically, the execution device 2300 includes: a receiver 2301, a transmitter 2302, a processor 2303 and a memory 2304 (wherein the number of the processor 2303 in the execution device 2300 can be one or more, Figure 23 In the example of FIG. 2301 , a processor 2303 may include an application processor 23031 and a communication processor 23032. In some embodiments of the present application, the receiver 2301, the transmitter 2302, the processor 2303 and the memory 2304 may be connected via a bus or other means.
[0222] The memory 2304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 2303. A part of the memory 2304 may also include a non-volatile random access memory (NVRAM). The memory 2304 stores processor and operation instructions, executable modules or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations.
[0223] The processor 2303 controls the operation of the execution device. In a specific application, the components of the execution device are coupled together through a bus system, which may include a power bus, a control bus, a status signal bus, etc. in addition to the data bus. However, for the sake of clarity, all kinds of buses are referred to as the bus system in the figure.
[0224] The methods disclosed in the embodiments of the present application described above can be applied to the processor 2303 or implemented by the processor 2303. The processor 2303 may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above methods can be completed by the integrated logic circuit in the hardware of the processor 2303 or instructions in software form. The above-mentioned processor 2303 may be a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, and may further include an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The processor 2303 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the methods disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 2304, and the processor 2303 reads the information in the memory 2304 and combines its hardware to complete the steps of the above methods.
[0225] The receiver 2301 can be used to receive input digital or character information, and generate signal inputs related to the relevant settings and function controls of the execution device. The transmitter 2302 can be used to output digital or character information through the first interface; the transmitter 2302 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; the transmitter 2302 can also include display devices such as a display screen.
[0226] In an embodiment of the present application, in one case, the processor 2303 is used to execute Figure 6 the image enhancement method executed by the execution device in the corresponding embodiment.
[0227] An embodiment of the present application also provides a computer program product, which when running on a computer, causes the computer to execute the steps executed by the aforementioned execution device, or causes the computer to execute the steps executed by the aforementioned training device.
[0228] An embodiment of the present application also provides a computer-readable storage medium, in which a program for signal processing is stored. When running on a computer, it causes the computer to execute the steps executed by the aforementioned execution device, or causes the computer to execute the steps executed by the aforementioned training device.
[0229] The execution device, training device or terminal device provided in the embodiment of the present application may specifically be a chip, and the chip includes: a processing unit and a communication unit. The processing unit may be a processor, for example, and the communication unit may be an input / output interface, a pin or a circuit, etc. The processing unit can execute the computer execution instructions stored in the storage unit, so that the chip in the execution device executes the image enhancement method described in the above embodiment, or so that the chip in the training device executes the image enhancement method described in the above embodiment. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit may also be a storage unit outside the chip in the radio access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0230] Specifically, please refer to Figure 24 , Figure 24A schematic structural diagram of a chip provided by an embodiment of the present application. The chip can be embodied as a neural network processor NPU 2400. The NPU 2400 is mounted on the main CPU (Host CPU) as a coprocessor, and tasks are assigned by the Host CPU. The core part of the NPU is the arithmetic circuit 2403, and the arithmetic circuit 2403 is controlled by the controller 2404 to extract matrix data from the memory and perform multiplication operations.
[0231] In some implementations, the arithmetic circuit 2403 includes multiple processing units (Process Engine, PE) inside. In some implementations, the arithmetic circuit 2403 is a two-dimensional systolic array. The arithmetic circuit 2403 can also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 2403 is a general matrix processor.
[0232] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit fetches the corresponding data of matrix B from the weight memory 2402 and caches it on each PE in the arithmetic circuit. The arithmetic circuit fetches the data of matrix A from the input memory 2401 and performs matrix operations with matrix B, and the partial results or final results of the obtained matrix are stored in the accumulator 2408.
[0233] The unified memory 2406 is used to store input data and output data. The weight data is directly transported through the Direct Memory Access Controller (DMAC) 2405, and the DMAC transports it to the weight memory 2402. The input data is also transported to the unified memory 2406 through the DMAC.
[0234] BIU is the Bus Interface Unit, that is, the bus interface unit 2424, which is used for the interaction between the AXI bus, the DMAC, and the Instruction Fetch Buffer (IFB) 2409.
[0235] The bus interface unit 2424 (Bus Interface Unit, abbreviated as BIU) is used for the instruction fetch memory 2409 to obtain instructions from the external memory, and is also used for the storage unit access controller 2405 to obtain the original data of the input matrix A or the weight matrix B from the external memory.
[0236] The DMAC is mainly used to transport the input data in the external memory DDR to the unified memory 2406, or transport the weight data to the weight memory 2402, or transport the input data to the input memory 2401.
[0237] The vector calculation unit 2407 includes multiple operation processing units, which, if necessary, further process the output of the operation circuit 2403, such as vector multiplication, vector addition, exponential operation, logarithmic operation, magnitude comparison, etc. It is mainly used for non-convolution / full connection layer network calculations in neural networks, such as Batch Normalization, pixel-level summation, upsampling of the feature plane, etc.
[0238] In some implementations, the vector calculation unit 2407 can store the processed output vector into the unified memory 2406. For example, the vector calculation unit 2407 can apply a linear function; or, a non-linear function to the output of the operation circuit 2403, such as linearly interpolating the feature plane extracted by the convolutional layer, or, for another example, the vector of the accumulated value, to generate the activation value. In some implementations, the vector calculation unit 2407 generates normalized values, pixel-level summation values, or both. In some implementations, the processed output vector can be used as the activation input to the operation circuit 2403, such as for use in subsequent layers in the neural network.
[0239] The instruction fetch buffer 2409 connected to the controller 2404 is used to store the instructions used by the controller 2404;
[0240] The unified memory 2406, the input memory 2401, the weight memory 2402, and the instruction fetch memory 2409 are all On-Chip memories. The external memory is private to this NPU hardware architecture.
[0241] Wherein, the processor mentioned anywhere above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above program.
[0242] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.
[0243] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions accomplished by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits, etc. However, for the present application, in more cases, software program implementation is a better embodiment. Based on such an understanding, the technical solution of the present application, in essence, or the part that makes contributions to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc of a computer, etc., and includes several instructions for causing a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in various embodiments of the present application.
[0244] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0245] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training device, or data center to another website, computer, training device, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a training device or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
Claims
1. An image enhancement method, characterized in that, Including: Obtain a first image and a second image, where the second image is obtained by downsampling the first image; Process the second image through a first network to obtain a first feature and a third image, where the first feature is an intermediate feature extracted by the first network, and the resolution of the third image is the same as that of the first image; Generate a fourth image based on the third image and the first image; Process the fourth image and the first feature through a second network to obtain a target image; Wherein, the first network and the second network are used for image enhancement; the process of processing the second image through the first network to obtain a first feature includes: Segment the second image in the spatial dimension to obtain a plurality of sub-images; Stitch the plurality of sub-images in the channel dimension to obtain a stitched image; Process the stitched image through a first convolutional neural network to obtain a processed feature; Segment the processed feature in the channel dimension to obtain a plurality of sub-processed features; Stitch the plurality of sub-processed features in the spatial dimension to obtain the first feature.
2. The method according to claim 1, characterized in that, The first network includes a first sub-network and a second sub-network, and the second network includes a third sub-network and a fourth sub-network; The process of processing the second image through the first network to obtain a first feature and a third image includes: Process the second image through the first sub-network to obtain the first feature; Process the first feature through the second sub-network to obtain the third image; The process of processing the fourth image through the second network to obtain a target image includes: Process the fourth image through the third sub-network to obtain a second feature; Generate a third feature based on the first feature and the second feature; Process the third feature through the fourth sub-network to obtain the target image.
3. The method according to claim 2, characterized in that, The process of generating a third feature based on the first feature and the second feature includes: Perform upsampling processing on the first feature to obtain the upsampled first feature; Add the upsampled first feature and the second feature to obtain the third feature.
4. The method according to claim 2 or 3, characterized in that, The first sub-network and the second sub-network include convolutional neural networks.
5. The method according to claim 2 or 3, characterized in that, The process of processing the first feature through the second sub-network to obtain the third image includes: Segment the first feature in the spatial dimension to obtain a plurality of sub-features; Stitch the plurality of sub-features in the channel dimension to obtain a stitched feature; Process the stitched feature through a second convolutional neural network to obtain a processed image; Segment the processed image in the channel dimension to obtain a plurality of sub-processed images; Stitch the plurality of sub-processed images in the spatial dimension to obtain the third image.
6. The method according to any one of claims 1 to 3, characterized in that, The process of generating a fourth image based on the third image and the first image includes: Perform an addition process on the third image and the first image to obtain a fourth image.
7. The method according to any one of claims 1 to 3, characterized in that, The method is used to implement at least one of the following image enhancement tasks: image super-resolution reconstruction, image denoising, image dehazing, image deblurring, image contrast enhancement, image demosaicking, image deraining, image color enhancement, image brightness enhancement, image detail enhancement, and image dynamic range enhancement.
8. A training method for a model, characterized in that, It includes: Obtaining an image sample pair, where the image sample pair includes a first image and an enhanced image corresponding to the first image; Performing downsampling processing on the first image to obtain a second image; Processing the second image through a first network to obtain a first feature and a third image, where the first feature is an intermediate feature extracted by the first network, and the resolution of the third image is the same as that of the first image; Generating a fourth image according to the third image and the first image; Processing the fourth image and the first feature through a second network to obtain a target image; Obtaining a loss function according to the enhanced image corresponding to the first image and the target image, where the loss function is used to indicate the difference between the enhanced image and the target image in the image sample pair; Training the first network and the second network according to the loss function to obtain a trained first network and a trained second network; Wherein, the first network and the second network are used for image enhancement; the processing of the second image through the first network to obtain a first feature includes: Segmenting the second image in the spatial dimension to obtain a plurality of sub-images; Stitching the plurality of sub-images in the channel dimension to obtain a stitched image; Processing the stitched image through a first convolutional neural network to obtain a processed feature; Segmenting the processed feature in the channel dimension to obtain a plurality of sub-processed features; Stitching the plurality of sub-processed features in the spatial dimension to obtain the first feature.
9. The method according to claim 8, characterized in that, The first network includes a first sub-network and a second sub-network, and the second network includes a third sub-network and a fourth sub-network; Processing the second image through the first network to obtain a first feature and a third image, including: Processing the second image through the first sub-network to obtain the first feature; Processing the first feature through the second sub-network to obtain the third image; Processing the fourth image through the second network to obtain a target image, including: Processing the fourth image through the third sub-network to obtain a second feature; Generating a third feature according to the first feature and the second feature; Processing the third feature through the fourth sub-network to obtain the target image.
10. The method according to claim 9, characterized in that, The generating a third feature according to the first feature and the second feature includes: Performing upsampling processing on the first feature to obtain an upsampled first feature; Adding the upsampled first feature and the second feature to obtain the third feature.
11. The method according to claim 9 or 10, characterized in that, The first sub-network and the second sub-network include convolutional neural networks.
12. The method according to claim 9 or 10, characterized in that, The processing of the first feature through the second sub-network to obtain the third image includes: Segmenting the first feature in the spatial dimension to obtain a plurality of sub-features; Concatenate the multiple sub - features in the channel dimension to obtain a concatenated feature; Process the concatenated feature through a second convolutional neural network to obtain a processed image; Segment the processed image in the channel dimension to obtain multiple sub - processed images; Concatenate the multiple sub - processed images in the spatial dimension to obtain the third image.
13. The method according to any one of claims 8 to 10, characterized in that, Said generating a fourth image according to the third image and the first image includes: Perform an addition process on the third image and the first image to obtain a fourth image.
14. According to the method according to any one of claims 8 to 10, characterized in that, The method is used to implement at least one of the following image enhancement tasks: image super - resolution reconstruction, image denoising, image de - fogging, image de - blurring, image contrast enhancement, image demosaicing, image de - raining, image color enhancement, image brightness enhancement, image detail enhancement, and image dynamic range enhancement.
15. An image processing apparatus, characterized in that, Comprising a memory and a processor; the memory stores code, and the processor is configured to execute the code. When the code is executed, the image processing device executes the method according to any one of claims 1 to 7.
16. A model training apparatus, characterized in that, Comprising a memory and a processor; the memory stores code, and the processor is configured to execute the code. When the code is executed, the model training device executes the method according to any one of claims 8 to 14.
17. A computer storage medium, characterized in that, The computer storage medium stores instructions, and when the instructions are executed by a computer, the computer implements the method according to any one of claims 1 to 14.
18. A computer program product, characterized in that, The computer program product stores instructions, and when the instructions are executed by a computer, the computer implements the method according to any one of claims 1 to 14.
Citation Information
Patent Citations
Image enhancement method and device and electronic equipment
CN112150400A