Image Processing Method, Apparatus, Electronic Device, and Computer-Readable Medium
By determining and removing the pending noise in the image using a pre-trained denoising model, the problem of inaccurate noise training in the prior art is solved, and a more ideal image denoising effect is achieved.
Patent Information
- Application Number
- CN202011349525.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-26
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2040-11-26
AI Technical Summary
When training noise, the existing neural network-based image noise reduction model cannot reflect the real environment noise, resulting in inaccurate model training and unsatisfactory noise reduction effect.
By acquiring a pre-trained denoising model, the distribution area of the noise to be processed in the target image is determined, and the noise extraction operation of the model is used to remove the noise to be processed in the image. This model is based on training of noisy samples images, and the noise is the real noise collected by the image acquisition device.
The training accuracy and denoising effect of the denoising model are improved, making the image processing results more ideal.
Smart Images

Figure CN112308804B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image technology, and more particularly, to an image processing method, apparatus, electronic device, and computer-readable medium. Background Art
[0002] Image denoising is a fundamental problem in the fields of computer vision and image processing. In the process of people obtaining images through image acquisition devices, due to the physical constraints of the image acquisition devices themselves and the limitations of the external light environment, the images captured inevitably contain noise, which in turn affects the imaging quality.
[0003] Currently, the process of image denoising based on neural networks is to obtain a noise-free image by passing a noisy image through a denoising model. The image samples required for training the noise of the denoising model are all artificially adding noise (for example, Gaussian white noise) to the original image to generate image samples for training the denoising model. However, the artificially added noise cannot simulate the noise in the real environment, resulting in inaccurate training of the denoising model and unsatisfactory denoising effect of the model. Summary of the Invention
[0004] This application proposes an image processing method, apparatus, electronic device, and computer-readable medium to improve the above defects.
[0005] In a first aspect, an embodiment of this application provides an image processing method, including: obtaining a target image to be processed, where the target image contains noise to be processed; determining, based on a pre-trained denoising model, a distribution area of the noise to be processed in the target image as a first area, where the denoising model is trained based on a noisy sample image, and the noise contained in the noisy sample image is the noise collected when an image acquisition device acquires an image; removing the noise to be processed in the first area according to the image in a second area in the target image to obtain the denoised target image, where the second area is the area outside the first area in the target image.
[0006] In a second aspect, an embodiment of the present application further provides an image processing apparatus, including: an acquisition unit, a determination unit, and a processing unit. The acquisition unit is configured to acquire a target image to be processed, where the target image contains noise to be processed. The determination unit is configured to determine, based on a pre-trained denoising model, a distribution area of the noise to be processed in the target image as a first area, where the denoising model is trained based on a noisy sample image, and the noise contained in the noisy sample image is the noise collected by an image acquisition device when acquiring an image. The processing unit is configured to remove the noise to be processed in the first area according to the image in a second area in the target image, to obtain the denoised target image, where the second area is an area outside the first area in the target image.
[0007] In a third aspect, an embodiment of the present application further provides an electronic device, including: one or more processors; a memory; one or more application programs, where the one or more application programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the above method.
[0008] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, where the computer-readable storage medium stores program code executable by a processor, and when the program code is executed by the processor, the processor executes the above method.
[0009] The image processing method, apparatus, electronic device, and computer-readable medium provided by the present application determine, based on a pre-trained denoising model, a distribution area of the noise to be processed in the target image as a first area. Through the noise extraction operation of the denoising model, the area of the noise to be processed in the image to be processed can be determined to implement the noise extraction operation. Then, according to the image in the second area, that is, the area outside the first area in the target image, the noise to be processed in the first area is removed. Since the second area is an area outside the noise area, the image in this second area is theoretically a noise-free image. Using this noise-free image to remove the noise in the first area can thus remove the noise in the image to be processed. In addition, the denoising model is trained based on a noisy sample image, and the noise contained in the noisy sample image is the noise collected by an image acquisition device when acquiring an image. Therefore, the noise in the sample on which the denoising model is based is not artificially added but the noise collected by the image acquisition device when acquiring an image of the real world, so that the training result of the denoising model is more accurate and the denoising of the denoising model is more ideal.
[0010] Other features and advantages of the embodiments of the present application will be described in the subsequent description, and some of them will become obvious from the description, or can be understood by implementing the embodiments of the present application. The objectives and other advantages of the embodiments of the present application can be achieved and obtained through the structures specifically pointed out in the written description, claims, and drawings. Description of the Drawings
[0011] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0012] Figure 1 A schematic diagram showing the application scenario of the embodiments of the present application;
[0013] Figure 2 A flowchart showing the method of the image processing method according to an embodiment of the present application;
[0014] Figure 3 A schematic diagram showing the distribution area of the noise to be processed in the target image according to the embodiments of the present application;
[0015] Figure 4 A schematic diagram showing the target image after denoising according to the embodiments of the present application;
[0016] Figure 5 A flowchart showing the method of the image processing method according to another embodiment of the present application;
[0017] Figure 6 A schematic diagram showing the process of removing noise according to the embodiments of the present application;
[0018] Figure 7 Shows the Figure 5 Flowchart of S570 in the image processing method shown;
[0019] Figure 8 A schematic diagram showing the first sub-region and the second sub-region according to an embodiment of the present application;
[0020] Figure 9 A schematic diagram showing the first sub-region and the second sub-region according to another embodiment of the present application;
[0021] Figure 10 A flowchart showing the method of the image processing method according to still another embodiment of the present application;
[0022] Figure 11 A block diagram showing the modules of the image processing device according to an embodiment of the present application;
[0023] Figure 12 shows a structural block diagram of an electronic device provided by an embodiment of the present application;
[0024] Figure 13 shows a storage unit for storing or carrying program codes for implementing an image processing method according to an embodiment of the present application. Detailed implementation manners
[0025] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Components in the embodiments of the present application usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application claimed, but merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.
[0026] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, terms such as "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0027] With the development of information technology and Internet technology, image processing technology has also been successfully applied to solutions including disaster relief, weather prediction, photo entertainment, face recognition, shopping quick payment, etc. However, during the processes of image capture, storage, transmission, and image processing by a camera, the image is easily affected by factors such as rainy days, foggy days and other weather, poor lighting conditions, and camera shake, resulting in unclear captured images. In order to ensure the imaging effect of the image, it is necessary to restore the unclear image to a clear image, that is, it is necessary to perform image denoising processing on the unclear image.
[0028] Traditional image denoising algorithms use image prior models for denoising, such as non-local self-similarity models, gradient models, sparse dictionary models, and Markov random field models. The classic three-dimensional block matching (BM3D) algorithm and its variant CBM3D for color images are mainly based on three-dimensional non-local similar block matching. Different from the traditional non-local mean idea (NLM), three-dimensional non-local similar block matching uses a hard threshold linear transformation to find similar blocks, then processes these three-dimensional arrays by joint filtering, and finally returns the processed result to the original image through inverse transformation to obtain the denoised image. Traditional image denoising algorithms are usually complex in dealing with optimization problems, often involve cumbersome parameter tuning during the denoising process, and the effects in terms of noise removal and detail preservation are not satisfactory.
[0029] In recent years, the image denoising process based on neural networks is to obtain a noise-free image by passing a noisy image through a denoising model. Currently, the training process of the denoising model is often as follows: a noise distribution is designed in advance. Specifically, the noise distribution is a noise distribution simulated by a computer, such as Gaussian noise. Then, the original image is obtained, which is a real-world image collected. The preset noise distribution is added to the original image to form a noisy image, and the denoising model is trained. The noisy image is input into the denoising model to obtain a denoised picture, and the trained denoising model is obtained through supervised learning.
[0030] However, the inventors found in their research that the image samples required for training the noise of the denoising model are all artificially adding noise (such as Gaussian white noise) to the original image to generate image samples for training the denoising model. However, the artificially added noise cannot reflect the noise in the real environment, making the training of the denoising model inaccurate and the denoising effect of the model unsatisfactory.
[0031] Therefore, to solve the above defects, the embodiments of the present application provide an image processing method, device, electronic device, and computer-readable medium. Through the noise extraction operation of the denoising model, the area of the to-be-processed noise in the to-be-processed image can be determined to implement the noise extraction operation. Then, based on the second area, that is, the image of the area outside the first area in the target image, the to-be-processed noise in the first area is removed. The denoising model is trained based on a noisy sample image, and the noise contained in the noisy sample image is the noise collected by the image acquisition device when acquiring an image. Therefore, the noise in the sample on which the denoising model is based is not artificially added but the noise collected by the image acquisition device when acquiring a real-world image, thus making the training result of the denoising model more accurate and the denoising of the denoising model more ideal.
[0032] As an implementation manner, the embodiments of the present application can be applied to a user terminal. That is, the user terminal can be the execution subject of the image processing method of the present application. Specifically, the execution subject can be an application installed in the user terminal. Then, the training process of the denoising model and the process of denoising the image according to the denoising model are both executed by the user terminal.
[0033] As another implementation manner, the embodiments of the present application can be applied to a server. The server can be the execution subject of the image processing method of the present application. Then, the training process of the denoising model and the process of denoising the image according to the denoising model are both executed by the server. The server can obtain the image to be processed uploaded by the user terminal and perform noise reduction processing on the image to be processed based on the denoising model.
[0034] The embodiments of the present application can be applied to an image processing system, such as Figure 1 shown. The image processing system includes a server 10 and a user terminal 20. The server 10 and the user terminal 20 are located in a wireless network or a wired network, and data interaction can be performed between the server 10 and the user terminal 20. Among them, the server 10 can be a single server or a server cluster, and can be a local server or a cloud server.
[0035] As an implementation manner, the user terminal 20 can be the terminal used by the user. The user browses images through the user terminal, and the user terminal can be the device for the user to collect images. Then, in some embodiments, an image acquisition device is provided in the user terminal. The server 10 can store the pictures in the user terminal 20. And in some embodiments, the server 10 can be used to train the models or algorithms involved in the embodiments of the present application. In addition, the server 10 can also migrate the trained models or algorithms to the user terminal. Of course, it can also be that the user terminal 20 directly trains the models or algorithms involved in the embodiments of the present application. Specifically, in the embodiments of the present application, the execution subject of each method step in the embodiments of the present application is not limited.
[0036] Please refer to Figure 2 , Figure 2 which shows an image processing method provided by the embodiments of the present application. The execution subject of this method can be the above-mentioned server or the above-mentioned user terminal. Specifically, this method includes: S201 to S203.
[0037] S201: Obtain a target image to be processed, where the target image contains noise to be processed.
[0038] As an implementation manner, the image to be processed may be an image captured by a user terminal in a camera application. Specifically, the user uses the camera application in the user terminal to capture an image. After the image is captured, the captured image serves as the image to be processed, so that when the user finishes capturing, the image captured by the user can be automatically denoised according to the image processing method of the present application, and the denoised image is obtained and saved. As an implementation manner, it is also possible to save both the denoised image and the target image to be processed, that is, the image before denoising.
[0039] As another implementation manner, the image to be processed may also be an image requested to be displayed by the user terminal. Specifically, the image to be processed may be an image downloaded by the user terminal. When the image is rendered, it is denoised based on the image processing method of the embodiment of the present application, and the denoised image is displayed. The downloaded image may be an image captured by other devices outside the user terminal.
[0040] Digital images are often affected by imaging device and external environment noise interference during the digitization and transmission processes. Specifically, due to the physical constraints of the image acquisition device itself and the limitations of the external light environment, noise inevitably exists in the captured image. As an implementation manner, the image to be processed may be the noise brought by the lighting conditions. Specifically, it may be the noise caused by insufficient light when the camera capturing the image to be processed captures the corresponding real scene of the image.
[0041] S202: Determine the distribution area of the noise to be processed in the target image based on a pre-trained denoising model as a first area.
[0042] As an implementation manner, the denoising model may be a neural network model based on deep learning. For example, it may be a Convolutional Neural Networks (CNN). Specifically, the feature layers of the neural network model can be set according to the characteristics of the noise in the image, so that the model can determine the noise area in the image based on the feature layers, that is, perform the operation of extracting the noise in the image. The parameters (such as the weights of each feature layer, etc.) in the trained denoising model are relatively reasonable, so that the denoising model can accurately determine the noise area in the image.
[0043] As an implementation, the denoising model is trained based on the noisy sample images, and the noise contained in the noisy sample images is the noise captured by the image acquisition device when acquiring images. Specifically, when setting the initial denoising model, that is, the untrained denoising model, the network parameters in the denoising model may not be reasonable enough, and the noise area determined by the initial denoising model is not accurate enough. Through continuous learning of the initial denoising model with the noisy sample images, the network parameters of the denoising model are continuously optimized, so that the noise area determined by the trained denoising model is more accurate and the extracted noise is more reasonable.
[0044] In some embodiments, the noise contained in the noisy sample images is the noise captured by the image acquisition device when acquiring images. Among them, the image acquisition device may be the camera of the above user terminal. As an implementation, the image to be processed may be an image captured by the camera of the user terminal, then the noise contained in the noisy sample images is also the noise captured by the camera of the user terminal. As an implementation, the image acquisition device may acquire multiple images as a sample image set under different shooting environments. Each image contains noise, and the noise of each image is related to the shooting environment of the image. Among them, the shooting environment may include lighting conditions, weather, shooting equipment, and shooting parameters, etc. Then, the denoising model trained based on the sample image set enables the denoising model to have more accurate noise extraction and removal capabilities for different shooting environments.
[0045] In the embodiments of the present application, compared with the artificially added noise simulated by a computer, the denoising model can be trained using the noise captured by the image acquisition device when acquiring images. Since the artificially added noise is simulated by a computer, it generally satisfies some mathematical relationships (for example, Gaussian function), belongs to a known mathematical distribution, and is ordered noise, which is different from the real noise when the image acquisition device captures images of the real world. The real noise is often disordered and is often difficult to describe with mathematical relationships. Therefore, the noise contained in the noisy sample images is the noise captured by the image acquisition device when acquiring images. Compared with training the denoising model with artificially added noise, the noise extraction of the trained denoising model is more accurate. Specifically, the training process of the denoising model will be introduced in subsequent embodiments.
[0046] S203: Remove the noise to be processed in the first region according to the image in the second region of the target image, and obtain the denoised target image.
[0047] Among them, the second region is the region outside the first region in the target image. Specifically, the first region corresponds to the noise to be processed in the image to be processed, that is, the distribution region of the noise to be processed is the first region, such as Figure 3As shown, the first region 301 is the distribution region of the noise to be processed. It can be seen that the noise in this first region is caused by defects such as insufficient light during image acquisition. It should be noted that the other distribution regions of the noise are not marked in Figure 3 the text.
[0048] Therefore, after determining the distribution region of the noise to be processed in the image to be processed, the region without distributed noise, that is, the second region, can be determined. Then, the noise to be processed in the first region can be removed according to the image in the second region. As an implementation manner, it can be to generate the image in the first region according to the image in the second region, so as to replace the image in the first region with the generated image. Specifically, there are similar or identical data in the image data in the second region and the image data in the first region in the image to be processed, where the image data can be the pixel data corresponding to the pixel units. As Figure 3 shown, the noise in the first region 301 is the noise caused by light. Specifically, the image in the first region 301 is too dark, that is, the color of the shadow of the object is too dark, and there is also an image of the shadow of the object in the second region, that is, there is also a region with a darker image. An object should have a shadow under light, so there should be an image of the shadow in the image to be processed. Thus, the pixel data in the first region can be updated according to the pixel data of the image in the second region that is closer to the first region. Of course, the pixel data in the first region can also be updated according to the pixel data in the sub-region in the second region that matches the image category of the first region.
[0049] As an implementation manner, the first region is divided into multiple sub-regions. The size of the sub-region can be composed of N thousand pixel points, where N is a positive integer. The sub-regions in the first region are named first sub-regions. Multiple sub-regions can also be divided in the second region and named second sub-regions, and the size of the first sub-region is the same as the size of the second sub-region, that is, it is only a sub-region composed of the same number of pixel points. Based on the distance between each of the first sub-regions and each of the second sub-regions, the second sub-region corresponding to each of the first sub-regions is determined. For the specific implementation manner, please refer to the subsequent embodiments.
[0050] As another implementation manner, the pixel data in the first region can be updated according to the pixel data in the sub-region in the second region that matches the image category of the first region. Considering that in the same shooting scene, the images in some regions contain noise, and the images in some regions do not contain noise, and the categories of the two regions are not very different. For example, in an image, there are multiple target objects. The images in the regions of some target objects do not contain noise, and the regions of some target objects contain noise. For example, Figure 3In the case where, affected by the illumination conditions, noise exists in the shadow of an object in the first region, while no noise exists in the shadow of an object in the second region, the image of the shadow of the object in the first region can be modified based on the image of the shadow of the object in the second region.
[0051] In some embodiments, the noise in the first region is the noise corresponding to the illumination shadow of the object. For example, if the shadow is too dark or there are noise points, etc., the noise corresponding to the illumination shadow in the first region can be named the shadow noise to be processed. It can be determined that the noise type in the first region is the illumination shadow type. Then, the image data corresponding to the specified illumination shadow region in the second region can be determined according to this illumination shadow type as the reference pixel data, and the pixel data of the illumination shadow in the first region can be modified based on this reference pixel data, so that the image of the illumination shadow in the first region is consistent with the image of the illumination shadow in the second region, thereby removing the noise in the image of the illumination shadow in the first region. As an implementation manner, the specified illumination shadow region can be the illumination shadow region in the second region that matches the size of the shadow noise to be processed.
[0052] As another implementation manner, the specified illumination shadow region can also be determined by the category of the projection object corresponding to the shadow noise to be processed. Herein, the projection object refers to the object that blocks the propagation of light and generates the shadow noise to be processed. Specifically, when light irradiates on the projection object, a shadow corresponding to the projection object is shown behind the light irradiation direction, and noise exists in this shadow, so it becomes the shadow noise to be processed. Determine the category of this projection object, where the category includes human body, animal, building, furniture, electrical appliance, etc. For example, when the projection object is an animal, the contour and characteristic information of the projection object can be collected, such as ears, horns, ears and limbs. When the projection object is a human body, face feature extraction can be performed on the projection object. The methods of face feature extraction can include knowledge-based representation algorithms or representation methods based on algebraic features or statistical learning. After determining the specified category of this projection object, search for the specified object of the same specified category in the second region, and use the illumination shadow region corresponding to this specified object as the specified illumination shadow region.
[0053] In some embodiments, when the number of identified specified objects is relatively large, it is necessary to determine a reference object among the multiple specified objects, and use the illumination shadow area corresponding to the reference object as the specified illumination shadow area. Specifically, the implementation manner of determining the reference object among the multiple specified objects may be to determine the shadow parameters of the illumination shadow area corresponding to each specified object, determine the shadow parameters that meet the preset shadow conditions as the target shadow parameters, and use the illumination shadow area corresponding to the target shadow parameters as the specified illumination shadow area. Among them, the shadow parameters may be the brightness or darkness of the shadow or the size of the shadow, etc. Then, the illumination shadow area corresponding to the to-be-processed shadow noise also corresponds to shadow parameters, denoted as specified shadow parameters. The implementation manner of determining the shadow parameters that meet the preset shadow conditions as the target shadow parameters may be to search for the shadow parameters that match the target shadow parameters among the shadow parameters of the illumination shadow area corresponding to each specified object as the shadow parameters of the preset shadow conditions. As an implementation manner, the shadow parameter is the size of the shadow. Determine the shadow size of the illumination shadow area corresponding to each specified object, and search for the illumination shadow area whose difference in size from the shadow size of the to-be-processed shadow noise is less than the specified value as the specified illumination shadow area.
[0054] Therefore, the embodiment of the present application determines the distribution area of the to-be-processed noise in the target image based on a pre-trained denoising model as the first area. Through the noise extraction operation of the denoising model, the area of the to-be-processed noise in the to-be-processed image can be determined to implement the noise extraction operation. Then, based on the second area, that is, the image of the area outside the first area in the target image, the to-be-processed noise in the first area is removed. Since the second area is the area outside the noise area, the image in this second area is theoretically a noise-free image. Using this noise-free image to remove the noise in the first area can thus remove the noise in the to-be-processed image. In addition, the denoising model is trained based on a noisy sample image, and the noise contained in the noisy sample image is the noise collected by the image acquisition device when acquiring an image. Therefore, the noise in the sample on which the denoising model is based is not artificially added but the noise collected by the image acquisition device when acquiring an image of the real world, which makes the training result of the denoising model more accurate and the denoising of the denoising model more ideal. As Figure 4 shown, compared with Figure 3 , the image in the first area 301 is no longer as Figure 3 dark as before, and the color is more realistic, that is, the noise caused by the illumination conditions is eliminated.
[0055] Please refer to Figure 5 , Figure 5The figure shows an image processing method provided by an embodiment of the present application. The execution subject of this method can be the above-mentioned server or the above-mentioned user terminal. Specifically, this method includes: S510 to S570.
[0056] S510: Denoise the noisy sample image through a denoising model to be trained, and obtain first noise data and a first image.
[0057] Among them, the first noise data includes a plurality of noise values and the pixel coordinates corresponding to each noise value. And this first noise value is the noise collected when the image acquisition device acquires the sample image. For example, the noise caused by the illumination conditions in the real environment.
[0058] The first noise data refers to the noise data extracted by the denoising model to be trained from the noisy sample image. The first image refers to the data in the original image data corresponding to the noisy sample image except for the first noise data. As an implementation, the original image data can be a pixel matrix, denoted as the original image matrix. This matrix corresponds to a plurality of pixel data, and each pixel data corresponds to a pixel coordinate. Among them, the pixel data can be the rgb value of the pixel point corresponding to each pixel coordinate. Specifically, assume that the size of the original image is M×N, where both M and N are positive integers. The width of the original image is M pixels, and the height of the original image is N pixels. Then the corresponding original image matrix is a matrix of M×N. For example, M is 6 and N is 5, then this original image matrix is a matrix with 5 rows and 6 columns, specifically:
[0059] (1)
[0060] Each element in the formula (1) corresponds to a pixel coordinate point in the original image. The value corresponding to each element in this matrix is the pixel value of the pixel coordinate point corresponding to this element. For example, a 11 is the pixel value corresponding to the pixel coordinate point (1,1).
[0061] The process of the denoising model to be trained denoising the noisy sample image can be that the denoising model extracts the first noise data from the noisy sample image, so as to be able to determine the noise area in the noisy sample image, and then be able to determine the pixel coordinates corresponding to this noise area. Denote the pixel value corresponding to the pixel coordinates of this noise area as the noise value, so as to be able to determine a plurality of noise values and the pixel coordinates corresponding to each noise value.
[0062] As an implementation, elements within the matrix corresponding to the first noise data can be determined from the original image matrix, denoted as noise elements. Then, the noise elements within this matrix can characterize the pixel coordinates of the first noise data, and the value corresponding to the noise element is the pixel value, i.e., the noise value, corresponding to the pixel coordinates of the first noise data. As an implementation, the first noise data can also be a matrix, named the noise matrix. Taking the above-mentioned original image matrix as a 5-row and 6-column matrix as an example, the noise matrix can be:
[0063] (2)
[0064] In formula (2), for each pixel coordinate point of the first noise data corresponding to the non-zeroed elements, the positions of each pixel coordinate point of the first noise data can thus be determined, and further, the region of the first noise data within the original image can be determined. As an implementation, the above zeroing operation can also be replaced by replacing it with other data, which is not limited here.
[0065] Then the first image is the image within the region outside the region corresponding to the first noise data in the original image. For example, the matrix corresponding to the first image is denoted as the noise-free image matrix. Taking the above-mentioned original image matrix as a 5-row and 6-column matrix as an example, the noise-free image matrix is:
[0066] (3)
[0067] Comparing formula (2) and (3) above, it can be seen that the elements corresponding to the first noise data are all zeroed, and the matrix obtained by superimposing formula (2) and (3) is the matrix of formula (1), i.e., the original image matrix.
[0068] The more accurate the denoising model is trained, the more accurate the extracted first noise data is, then the less noise the first image contains, and even it can be a pure noise-free image.
[0069] S520: Obtain a plurality of second noise data based on at least some of the noise values of the first noise data and the pixel coordinates corresponding to at least some of the noise values.
[0070] As an implementation manner, at least some of the noise values of the first noise data can be determined, which are alternative noise values, and the pixel coordinates corresponding to each alternative noise value are determined, denoted as noise coordinates. Based on the alternative noise values and the noise coordinates corresponding to each alternative noise value, the second noise data is determined. Since the first noise data is not artificially added or an ordered noise generated according to a preset mathematical function, but noise caused by the influence of the lighting conditions in the real environment on shooting, the second noise data determined based on the alternative noise values in the first noise data and the noise coordinates corresponding to each alternative noise value can also reflect the noise in the real environment. Moreover, by generating the second noise data from the first noise data, the sample images can be increased, the sample quantity can be improved, and thus the accuracy of extracting noise after training the denoising model can be increased. At the same time, the labor cost of collecting the noisy sample image set can be avoided.
[0071] In some embodiments, it may be to keep the noise coordinates of each alternative noise value unchanged, but change the value of the alternative noise value to obtain the second noise data. For example, set a part of the alternative noises to zero or change them to other values.
[0072] In other embodiments, the noise coordinates corresponding to the alternative noise values can also be scrambled. For example, if the pixel coordinates of the noise value A are (x1, y1) and the pixel coordinates of the noise value B are (x2, y2), then their positions can be swapped, that is, the pixel coordinates of the noise value A are (x2, y2), and the pixel coordinates of the noise value B are (x1, y1).
[0073] As an implementation manner, the area of the first noise data corresponding to the noisy sample image is denoted as the noise area, and multiple second noise data are all located within the noise area, so as to avoid the second noise data being added to the area outside the noise area, and further prevent the first image from being contaminated by noise, resulting in a worse result of training the denoising model.
[0074] Specifically, the implementation manner of obtaining multiple second noise data based on at least some of the noise values of the first noise data and the pixel coordinates corresponding to at least some of the noise values can be to perform multiple pixel position swapping operations on the first noise data to obtain multiple second noise data, where the multiple pixel position swapping operations are not all the same, and the pixel position swapping operation is to swap the pixel coordinates corresponding to at least some of the multiple noise values in the first noise data.
[0075] Taking the above noise matrix V as an example, after performing the pixel position swapping operation on the noise matrix V, the noise matrix V' of the obtained second noise data is:
[0076] (4)
[0077] It can be seen that in the noise matrix V' of the second noise data, a 23 and a 34 are swapped in position, and a 32 and a 24 are swapped in position. As an implementation, this position swapping operation can also be to perform a position swap and then swap again based on the swapped position, that is, continuously perform at least two position swaps for a certain pixel coordinate point. For example, after swapping the positions of a 23 and a 34 , then swap the new positions of a 33 and a 23 .
[0078] Based on the above position swapping operation, multiple second noise data can be obtained, and each second noise data is obtained according to the first noise data, that is, the noise values in each second noise data belong to the noise values in the first noise data, but the distribution positions of the respective noise values are not all the same.
[0079] S530: Obtain multiple second images from the multiple second noise data and the first image.
[0080] Since the first image is regarded as a noise-free image, and the second noise data is noise data generated according to the first noise data, and both the first noise data and the second noise data are within the noise region, the second noise data can be superimposed on the first image to obtain multiple second images. Specifically, the implementation of obtaining multiple second images from the multiple second noise data and the first image can be to superimpose each second noise data on the first image to obtain the multiple second images.
[0081] As an implementation, it can be to replace the pixel values of the specified pixel coordinates in the first image with the noise values in the second noise data corresponding to the specified pixel coordinates according to the pixel coordinates corresponding to each noise value in the second noise data for the region in the first image corresponding to the second noise data, so as to fuse into multiple second images.
[0082] As another implementation, the image data of the first image is denoted as the first image data, which includes multiple first pixel data and the pixel coordinates corresponding to each first pixel data, and the second noise data includes multiple second noise values and the pixel coordinates corresponding to each second noise value. Superimpose the first image data and the second noise data. Specifically, for each pixel coordinate, add the first pixel data corresponding to the pixel coordinate and the second noise value, and use the added value as the second pixel data for the pixel coordinate. Each pixel coordinate and the second pixel data constitute the second image data, so as to obtain the second image.
[0083] The matrix T' of the second image data obtained by adding the noise-free image matrix T corresponding to the above-mentioned first image and the noise matrix V' of the second noise data is as follows:
[0084] (5)
[0085] By comparing the matrix R and the matrix T', it can be seen that the original image data is different from the second image data. The second image data is also an image containing noise data, and the area where the noise is distributed has not changed. However, all the noise is different from the noise in the original image data. Therefore, although the generated second image data is also computer-synthesized, the noise it contains can characterize the noise in the real world.
[0086] S540: Training the denoising model to be trained based on the distribution gap between the noisy sample image and each of the second images to obtain a trained denoising model.
[0087] Among them, the distribution gap is used to measure the overall deviation of each pixel value between two image data. As an implementation manner, due to the characteristic of noise that noise is noise no matter where it appears in the image, and the distribution of non-noise is often logical. Randomly shuffling the distribution of pixel values in the non-noise area will disrupt the distribution of non-noise, making the pixel values with chaotic logic become noise, which may cause more noise to appear in the non-noise area.
[0088] As an implementation manner, as Figure 6 shown, considering an extreme case, if after denoising the noisy sample image through the denoising model to be trained, the first noise data obtained is all noise-free areas, that is, the pixel data of the noise-free area in the original image is used as the first noise data, and the noise area in the original image is used as the first image data. Then when performing the operation of obtaining multiple second images from multiple second noise data and the first image, the distribution of the noise in the obtained second image will be very different from the noise distribution in the original image, and the degree of noise pollution of the second image is greater.
[0089] Therefore, if when performing the operation of denoising the noisy sample image through the denoising model to be trained, the obtained first noise data is accurate and reasonable enough, that is, it can accurately extract the noise values in the area that was originally noise in the original image, then the degree of noise pollution of the obtained second image data should not be much different from the degree of noise pollution of the original image, that is, the distribution gap between the two is smaller.
[0090] Therefore, after obtaining the distribution gap between the noisy sample image and each of the second images, the denoising model is continuously trained based on this distribution gap. After the distribution gap converges, the training of the denoising model is completed. Once the training of the denoising model is completed, the trained denoising model can accurately extract the noise in the noisy image, that is, accurately find the distribution area of the noise in the noisy image. Among them, the distribution gap is the mean square error between the noisy sample image and each of the second images. As an implementation manner, the convergence of the distribution gap can be that the distribution difference is less than a specified gap, and a continuous plurality of distribution gaps are all less than the specified gap, and there is no obvious downward trend.
[0091] As an implementation manner, when determining the distribution gap between the noisy sample image and each of the second images, in order to compare the two groups of images on the same scale, scale change processing needs to be performed on them. Specifically, the implementation manner of using the distribution gap between the noisy sample image and each of the second images to train the denoising model to be trained can be to perform scale transformation operations on the noisy sample image and multiple second images; obtain the distribution gap between the scaled noisy sample image and each of the second images; and train the denoising model to be trained based on each of the distribution gaps.
[0092] By performing scale transformation operations on the noisy sample image and multiple second images, the scales of the noisy sample image and multiple second images are transformed into the same scale space. For example, it can be achieved through linear transformation or neural network simulation. Obtain the distribution gap between the scaled noisy sample image and each scaled second image, so as to obtain a plurality of distribution gaps, and then train the denoising model to be trained based on each of the distribution gaps.
[0093] S550: Obtain a target image to be processed, where the target image contains noise to be processed.
[0094] S560: Based on the pre-trained denoising model, determine the distribution area of the noise to be processed in the target image as the first area.
[0095] S570: Remove the noise to be processed in the first area according to the image in the second area of the target image, and obtain the denoised target image.
[0096] Among them, the first area includes a plurality of first sub-areas, and the second area includes a plurality of second sub-areas. Then the implementation manner of S570 can be as Figure 7 shown. S570 may include: S571 and S572.
[0097] S571: Determine the second sub-region corresponding to each first sub-region based on the distance between each first sub-region and each second sub-region.
[0098] As an implementation, both the first sub-region and the second sub-region are composed of a specified number of pixels, and the shapes of the first sub-region and the second sub-region are the same. Specifically, both the first sub-region and the second sub-region can be rectangles. The first sub-region is provided with a first position point, and the second sub-region is provided with a second position point. The distance between the first sub-region and the second sub-region can be the distance between the first position point of the first sub-region and the second position point of the second sub-region, that is, the distance between the first coordinate point of the first position point and the second coordinate of the second position point. Among them, the first position point can be the position of a certain pixel within the first sub-region, and then the first coordinate is the pixel coordinate of this pixel.
[0099] As an implementation, the first position point can be the vertex of the first sub-region. In some embodiments, the origin of the pixel coordinates of the image to be processed is located at the upper left corner of the image to be processed. Then, the first position point can be the vertex within the first sub-region that is closest to the origin of the pixel coordinates of the image to be processed. For example, the first position point can be the vertex at the upper left corner of the first sub-region. Similarly, the second position point of the second sub-region can also be the vertex within the second sub-region that is closest to the origin of the pixel coordinates of the image to be processed.
[0100] Thus, determine the distance between the coordinate of the first position point within the first sub-region and the coordinate of the second position point within the second sub-region as the distance between the first sub-region and the second sub-region, so as to determine the distance between each first sub-region and each second sub-region.
[0101] As an implementation, the implementation of determining the second sub-region corresponding to each first sub-region based on the distance between each first sub-region and each second sub-region can be to use the second sub-region corresponding to the shortest distance among the multiple distances corresponding to each first sub-region as the second sub-region corresponding to this first sub-region.
[0102] Specifically, a sub-region is determined from multiple first sub-regions in sequence as the target sub-region, the distances between the target sub-region and each second sub-region are determined, and the second sub-region with the shortest distance to the target sub-region is found as the second sub-region corresponding to the target sub-region. In some embodiments, the distances between the target sub-region and the alternative first sub-regions corresponding thereto are determined, where the images corresponding to the alternative first sub-regions have been modified. For example, the images corresponding to the alternative first sub-regions have been modified by the images within the second sub-regions. If the target sub-region has corresponding alternative first sub-regions, the distances between the target sub-region and the alternative first sub-regions are determined, and the alternative first sub-region with the shortest distance to the target sub-region is found as the first sub-region to be confirmed. Then, the designated sub-region is determined from the second sub-region corresponding to the target sub-region and the first sub-region to be confirmed, and the designated sub-region is used as the second sub-region corresponding to the finally determined target sub-region.
[0103] Specifically, the implementation manner of determining the designated sub-region from the second sub-region corresponding to the target sub-region and the first sub-region to be confirmed may be as follows: the distance between the second sub-region corresponding to the target sub-region and the target sub-region is determined and denoted as the first distance, the distance between the first sub-region to be confirmed corresponding to the target sub-region and the target sub-region is determined and denoted as the second distance, the minimum distance is determined from the first distance and the second distance, and the sub-region corresponding to the minimum distance is used as the designated sub-region. If the first distance and the second distance are the same, a sub-region may be randomly determined as the designated sub-region.
[0104] As Figure 8 shown, Figure 8 In the figure, the gray area is the first sub-region, the white area is the second sub-region, the position point P1 is the first position point of the first sub-region, the position point P2 is the second position point of the second sub-region. The distances between the first sub-region 701 and the second sub-regions 702 and 703 are the same, and the distances between the first sub-region 701 and the second sub-regions 702 and 703 are the shortest. That is, the second sub-regions with the shortest distance to the first sub-region 701 are the second sub-region 702 and the second sub-region 703. Then, one sub-region may be randomly selected from the second sub-region 702 and the second sub-region 703 as the second sub-region corresponding to the first sub-region 701.
[0105] S572: Modify the image within the first sub-region corresponding to the second sub-region according to the image within each second sub-region to obtain the denoised target image.
[0106] As an implementation manner, the implementation manner of modifying the image in the first sub-region corresponding to each second sub-region according to the image in each second sub-region may be to modify the image in the first sub-region to the image in the second sub-region corresponding to the first sub-region, that is, replace the image in the first sub-region corresponding to each second sub-region according to the image in each second sub-region. As Figure 9 shown, if the second sub-region corresponding to the first sub-region 701 is the second sub-region 703, the image of the first sub-region 701 is replaced with the image in the second sub-region 703.
[0107] As an implementation manner, after the image in the first sub-region has been replaced with the image of the second sub-region corresponding to each first sub-region, the pixel data of the first sub-region after replacement can be superimposed on the first image to form a target image, that is, the denoised image.
[0108] In the above process of determining the second sub-region corresponding to each first sub-region based on the distance between each first sub-region and each second sub-region, modifying the image in the first sub-region corresponding to the second sub-region according to the image in each second sub-region to obtain the denoised target image, the second sub-region closest to each first sub-region can be determined by convolution, and the pixel data in the first sub-region can be replaced with the pixel data in the second sub-region closest to the first sub-region. In addition, considering that the spatial scales of the images of multiple second sub-regions are irregular and cannot be directly convolved. The entire image to be processed can be abstracted into a graph, where the second sub-region corresponding to each first sub-region is used as a node of the graph (graph node), and the distance between the first sub-region and the second sub-region is used as an edge of the graph (graph edge), so as to determine the second sub-region closest to each first sub-region by convolution, and replace the pixel data in the first sub-region with the pixel data in the second sub-region closest to the first sub-region.
[0109] Please refer to Figure 10 , Figure 10 shows an image processing method provided in an embodiment of the present application. The execution subject of this method may be the above-mentioned server or the above-mentioned user terminal. In the embodiment of the present application, this method is applied to a user terminal, and the user terminal may be a head-mounted display device that can implement Augmented Reality (AR) or Virtual Reality (VR). Specifically, this method includes: S1001 to S1004.
[0110] S1001: Obtain the image rendered by the renderer as the target image to be processed.
[0111] After decoding the video data, the decoded video data is sent to the renderer. After the renderer renders and composes the decoded video data, it is displayed on the display screen. Among them, the renderer can be SurfaceFlinger. SurfaceFlinger is an independent Service that receives the Surfaces of all Windows as inputs, calculates the positions of each Surface in the final composite image according to parameters such as ZOrder, transparency, size, and position, and then hands it over to HWComposer or OpenGL to generate the final display Buffer, and then displays it on a specific display device.
[0112] The renderer requires a large amount of time and computing power to render high-quality (noise-free) pictures, and it is difficult to directly apply it to real-time rendering scenarios (such as games, VR, etc.). Because, if the renderer takes too much time to render high-quality images, it will cause the images displayed on the display device to freeze. For example, a head-mounted display device that can render VR scenarios will cause users to feel dizzy due to the freezing or high latency of the displayed images, reducing the user experience. And if high-quality images are abandoned and low-quality images are used, it will greatly reduce the user's visual experience.
[0113] Therefore, the renderer can first render a lower-quality image, denoted as the noise to be processed. This noise to be processed still has a lot of noise. After subsequent steps, a target image is obtained, and then the target image is displayed. Thus, not only can high-quality images be displayed, but also the excessive time consumption of the renderer for rendering noise-free images can be reduced, so as to achieve the purpose of real-time rendering of real game scenes and VR scenes.
[0114] Specifically, the renderer can use the first strategy to generate a lower-quality image, that is, the noise to be processed. After the noise to be processed is denoised by the image processing method of the embodiments of the present application, the denoised target image is then displayed. The renderer can also use the second strategy to generate a high-quality image, that is, a noise-free image.
[0115] Specifically, the first strategy or the second strategy can be selected according to the real-time nature of the image to be displayed. Specifically, the real-time nature of the image to be displayed can correspond to the application program that requests to display the image, that is, determine the application program corresponding to the image to be displayed, and determine the real-time nature of the image to be displayed according to this application program.
[0116] As an implementation manner, determine the identifier of the application corresponding to the image to be displayed, and then determine the real-time level of the image to be displayed according to the identifier of the application. Specifically, determine the identifier of the application that sends the playback request of the image to be displayed, and determine the type of the application corresponding to the identifier of the application. As an implementation manner, the real-time levels can be preset, and different categories of applications correspond to different real-time levels. For example, the real-time level of the game category is J1, the real-time level of the video application is J2, and the real-time level of the audio application is J3. Among them, the level of J1 is the highest, and then, J2 and J3 decrease in turn.
[0117] Judge whether the real-time level corresponding to the image to be displayed meets the preset level. If it meets the preset level, the renderer uses the first strategy to render the image; otherwise, the renderer uses the second strategy to render the image. Among them, the preset level is the preset real-time level, which can be set by the user according to requirements. For example, the preset level is J2 and below. Then, if the real-time level corresponding to the image to be displayed is J3, the real-time level of the image to be displayed meets the preset level. That is to say, for the image to be displayed with relatively high real-time requirements, the second strategy can be not executed, but a low-quality noisy image is first generated, and then the image is denoised according to the method of the embodiment of the present application, so as to avoid excessive time consumption for rendering high-quality images and affect the user experience.
[0118] S1002: Obtain a target image to be processed, where the target image contains noise to be processed.
[0119] S1003: Based on a pre-trained denoising model, determine the distribution area of the noise to be processed in the target image as the first area.
[0120] S1004: Remove the noise to be processed in the first area according to the image in the second area in the target image to obtain the denoised target image.
[0121] Therefore, the renderer can first render a lower-quality image, denoted as the noise to be processed, which still has a lot of noise. Then, through subsequent steps, the target image is obtained, and the target image is displayed. Thus, not only can a high-quality image be displayed, but also the excessive time consumption of the renderer for rendering a noise-free image can be reduced, so as to achieve the purpose of real-time rendering of real game scenes and VR scenes.
[0122] Please refer to Figure 11 , which shows a structural block diagram of an image processing device 1100 provided by an embodiment of the present application. The device may include: an acquisition unit 1101, a determination unit 1102, and a processing unit 1103.
[0123] An acquisition unit 1101, configured to acquire a target image to be processed, where the target image contains noise to be processed.
[0124] A determination unit 1102, configured to determine, based on a pre-trained denoising model, a distribution area of the noise to be processed in the target image as a first area, where the denoising model is trained based on a noisy sample image, and the noise contained in the noisy sample image is the noise collected by an image acquisition device when acquiring an image.
[0125] A processing unit 1103, configured to remove the noise to be processed in the first area according to the image in a second area in the target image, to obtain the denoised target image, where the second area is the area outside the first area in the target image.
[0126] Further, the processing unit 1103 is further configured to determine, based on the distance between each first sub-area and each second sub-area, a second sub-area corresponding to each first sub-area; modify the image in the first sub-area corresponding to the second sub-area according to the image in each second sub-area, to obtain the denoised target image.
[0127] Further, the processing unit 1103 is further configured to use, as the second sub-area corresponding to the first sub-area, the second sub-area corresponding to the shortest distance among the multiple distances corresponding to each first sub-area.
[0128] Further, the image processing device 1000 further includes a training unit, configured to perform denoising processing on the noisy sample image through a denoising model to be trained, to obtain first noise data and a first image, where the first noise data includes multiple noise values and pixel coordinates corresponding to each noise value; obtain multiple second noise data based on at least some of the noise values of the first noise data and the pixel coordinates corresponding to at least some of the noise values; obtain multiple second images from the multiple second noise data and the first image; and train the denoising model to be trained based on the distribution gap between the noisy sample image and each second image, to obtain a trained denoising model, where the distribution gap obtained by the trained denoising model is less than a specified value.
[0129] Further, the training unit is further configured to superimpose each second noise data on the first image, to obtain multiple second images.
[0130] Further, the training unit is further configured to perform multiple pixel position swapping operations on the first noise data to obtain multiple second noise data, where the multiple pixel position swapping operations are not all the same, and the pixel position swapping operation is to swap the pixel coordinates corresponding to at least some of the multiple noise values in the first noise data.
[0131] Further, the training unit is further configured to perform a scale transformation operation on the noisy sample image and each of the multiple second images; obtain the distribution gap between the scale-transformed noisy sample image and each of the second images; and train the denoising model to be trained based on each of the distribution gaps. Wherein, the distribution gap is the mean square error between the noisy sample image and each of the second images.
[0132] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0133] In several embodiments provided in the present application, the coupling between modules may be electrical, mechanical, or other forms of coupling.
[0134] In addition, in each embodiment of the present application, each functional module may be integrated in a processing module, may exist separately physically for each module, or two or more modules may be integrated in one module. The above integrated modules may be implemented in the form of hardware or in the form of software functional modules.
[0135] Please refer to Figure 12 , which shows a structural block diagram of an electronic device provided by an embodiment of the present application. The electronic device 100 may be the above-mentioned user terminal and server. The user terminal may be an image acquisition device (for example, a camera, a video camera, etc.), may also be a terminal equipped with an image acquisition device, for example, a mobile terminal, etc., or may also be a device with functions of rendering and displaying images, for example, a head-mounted display device capable of implementing AR or VR, a mobile terminal, a computer device, etc.
[0136] As an implementation manner, if the electronic device is a camera, since the electronic device can implement the above-mentioned image processing method, even if the lens quality of the camera is not high or the exposure is poor, through the above method, the disordered high-brightness noise generated by the poor-exposure lens can be removed to a certain extent. Therefore, even when the hardware of the camera is not good, the above method can still be used to make the image collected by the camera exceed the image quality level that can be obtained by the hardware parameters of the camera, thus reducing the requirement for the camera hardware.
[0137] As an implementation, the electronic device is a head-mounted display device. A renderer is provided in the device, and the renderer is used to render images and display them on the screen. By using the method of the embodiment of the present application, the renderer can first render a lower-quality image, denoted as the noise to be processed, which still has a lot of noise. After subsequent steps, a target image is obtained, and then the target image is displayed. Thus, not only can high-quality images be displayed, but also the excessive time consumption of the renderer for rendering noise-free images can be reduced, so that the purpose of real-time rendering of real game scenes and VR scenes can be achieved. Therefore, compared with obtaining high-quality pictures at a greater cost, obtaining high-quality pictures through this method is faster and has a lower cost.
[0138] The electronic device 100 in the present application may include one or more of the following components: a processor 110, a memory 120, and one or more application programs. One or more application programs may be stored in the memory 120 and configured to be executed by one or more processors 110. One or more programs are configured to execute the method described in the foregoing method embodiments.
[0139] The processor 110 may include one or more processing cores. The processor 110 connects various parts within the entire electronic device 100 through various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 120, and by calling data stored in the memory 120, the processor 110 executes various functions of the electronic device 100 and processes data. Optionally, the processor 110 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 110 may integrate one or several combinations of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, and application programs, etc.; the GPU is responsible for the rendering and drawing of the display content; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor 110 and may be implemented separately through a communication chip.
[0140] The memory 120 may include a Random Access Memory (RAM), and may also include a Read-Only Memory. The memory 120 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the following various method embodiments, etc. The data storage area may also store data created during the use of the terminal 100 (such as a phone book, audio and video data), etc.
[0141] Please refer to Figure 13 , which shows a structural block diagram of a computer-readable storage medium provided by an embodiment of the present application. Program code is stored in the computer-readable storage medium 1300, and the program code can be called by a processor to execute the method described in the above method embodiments.
[0142] The computer-readable storage medium 1300 can be an electronic memory such as a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), an EPROM, a hard disk, or a ROM. Optionally, the computer-readable storage medium 1300 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 1300 has a storage space for the program code 1310 that executes any method step in the above method. These program codes can be read out from or written into one or more computer program products. The program code 1310 can be compressed in an appropriate form, for example.
[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. An image processing method, characterized in that, Including: Obtain a target image to be processed, where the target image contains noise to be processed; Based on a pre-trained denoising model, determine the distribution area of the noise to be processed in the target image as a first area. Among them, the denoising model is trained based on noisy sample images, and the noise contained in the noisy sample images is the noise collected by an image acquisition device when acquiring images; According to the image in the second area of the target image, remove the noise to be processed in the first area to obtain the denoised target image, where the second area is the area outside the first area in the target image; Before determining the first area of the noise to be processed in the target image based on the pre-trained denoising model, the method further includes: Denoise the noisy sample image through a denoising model to be trained to obtain first noise data and a first image. The first noise data includes multiple noise values and the pixel coordinates corresponding to each noise value; Based on at least some of the noise values of the first noise data and the pixel coordinates corresponding to at least some of the noise values, obtain multiple second noise data; Obtain multiple second images from the multiple second noise data and the first image; Train the denoising model to be trained based on the distribution gap between the noisy sample image and each second image to obtain a trained denoising model, where the distribution gap obtained by the trained denoising model is less than a specified value.
2. The method according to claim 1, wherein The obtaining multiple second images from the multiple second noise data and the first image includes: Superimpose each second noise data on the first image to obtain multiple second images.
3. The method according to claim 1, wherein Both the first noise data and the multiple second noise data correspond to the noise area of the noisy sample image. The obtaining multiple second noise data based on at least some of the noise values of the first noise data and the pixel coordinates corresponding to at least some of the noise values includes: Perform multiple pixel position swapping operations on the first noise data to obtain multiple second noise data. Among them, the multiple pixel position swapping operations are not all the same. The pixel position swapping operation is to swap the pixel coordinates corresponding to at least some of the multiple noise values in the first noise data.
4. The method according to any one of claims 1 to 3, characterized in that Training the denoising model to be trained based on the distribution gap between the noisy sample image and each second image includes: Perform a scale transformation operation on both the noisy sample image and the multiple second images; Obtain the distribution gap between the scaled noisy sample image and each second image; Train the denoising model to be trained based on each distribution gap.
5. The method according to claim 4, characterized in that, The distribution gap is the mean square error between the noisy sample image and each second image.
6. The method according to claim 1, wherein The first area includes multiple first sub-areas, and the second area includes multiple second sub-areas. The removing the noise to be processed in the first area according to the image in the second area of the target image to obtain the denoised target image includes: Based on the distance between each first sub-area and each second sub-area, determine the second sub-area corresponding to each first sub-area; Modify the image in the first sub-region corresponding to each second sub-region according to the image in the second sub-region to obtain the denoised target image.
7. The method according to claim 6, characterized in that, Determine the second sub-region corresponding to each first sub-region based on the distance between each first sub-region and each second sub-region, including: Use the second sub-region corresponding to the shortest distance among the multiple distances corresponding to each first sub-region as the second sub-region corresponding to this first sub-region.
8. The method according to claim 1, wherein The target image to be processed is an image rendered by a renderer.
9. An image processing apparatus, characterized in that, It includes: An acquisition unit for acquiring a target image to be processed, where the target image contains noise to be processed; A determination unit for determining the distribution area of the noise to be processed in the target image as the first area based on a pre-trained denoising model. Among them, the denoising model is trained based on a noisy sample image, and the noise contained in the noisy sample image is the noise collected by an image acquisition device when acquiring an image; A processing unit for removing the noise to be processed in the first area according to the image in the second area of the target image to obtain the denoised target image, where the second area is the area outside the first area in the target image; The image processing device further includes a training unit for: Perform denoising processing on the noisy sample image through a denoising model to be trained to obtain first noise data and a first image. The first noise data includes multiple noise values and the pixel coordinates corresponding to each noise value; Obtain multiple second noise data based on at least some of the noise values of the first noise data and the pixel coordinates corresponding to at least some of the noise values; Obtain multiple second images from the multiple second noise data and the first image; Train the denoising model to be trained based on the distribution gap between the noisy sample image and each second image to obtain a trained denoising model, where the distribution gap obtained by the trained denoising model is less than a specified value.
10. The device according to claim 9, characterized in that, The training unit is further used for: superimposing each second noise data on the first image to obtain multiple second images.
11. The device according to claim 9, characterized in that, Both the first noise data and the multiple second noise data correspond to the noise area of the noisy sample image; the training unit is further used for: performing multiple pixel position swapping operations on the first noise data to obtain multiple second noise data, where the multiple pixel position swapping operations are not all the same, and the pixel position swapping operation is to swap the pixel coordinates corresponding to at least some of the multiple noise values in the first noise data.
12. The device according to any one of claims 9-11, characterized in that, The training unit is further used for: Perform a scale transformation operation on both the noisy sample image and multiple second images; Obtain the distribution gap between the scale-transformed noisy sample image and each second image; Train the denoising model to be trained based on each distribution gap.
13. The device according to claim 12, characterized in that, The distribution gap is the mean square error between the noisy sample image and each second image.
14. The device according to claim 9, wherein The first region includes a plurality of first sub-regions, and the second region includes a plurality of second sub-regions. The processing unit is further configured to: determine a second sub-region corresponding to each first sub-region based on the distance between each first sub-region and each second sub-region; Modify the image within the first sub-region corresponding to the second sub-region according to the image within each second sub-region, so as to obtain the denoised target image.
15. The device according to claim 14, characterized in that, The processing unit is further configured to: use the second sub-region corresponding to the shortest distance among the multiple distances corresponding to each first sub-region as the second sub-region corresponding to the first sub-region.
16. The device according to claim 9, characterized in that, The target image to be processed is an image rendered by a renderer.
17. An electronic device, characterized in that, Comprising: One or more processors; A memory; One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the method according to any one of claims 1-8.
18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program code executable by a processor, and when the program code is executed by the processor, the processor executes the method according to any one of claims 1-8.
19. A computer program product, characterized in that, The computer program product includes program code, and when the program code is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Video denoising method and device
CN104010114A
Image processing method and apparatus, and electronic device
CN108737750A
Medical image denoising method and device
CN110443758A