Image processing method and device, equipment and storage medium
By performing joint denoising and demosaic processing on the current frame RAW image in image processing, using the time domain information of the front and back frame images, the problem of poor image denoising and demosaic effects in the prior art is solved, and higher image quality and time domain stability are achieved.
Patent Information
- Application Number
- CN202510127562.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art is difficult to effectively denoise and demosaic in image processing, resulting in blurred image details, distorted color, and poor time domain stability.
By acquiring the current frame RAW image to be processed and performing joint denoising and demosaic processing based on the associated image, the time domain information of the front and back frame images can be used to improve the display quality and signal-to-noise ratio of the image.
It improves the display quality of RAW images, enhances the signal-to-noise ratio and time-domain stability of the output images, and improves the detailed expression and color reduction of the image.
Smart Images

Figure CN119967299A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to computer technology, and relate to but are not limited to an image processing method and apparatus, device, and storage medium. Background Art
[0002] In the current electronic equipment field, especially in digital cameras, smart phones, tablet computers and various image acquisition and processing devices, image quality is one of the important indicators to measure the performance of the equipment. As users' requirements for image clarity, color reproduction and detail expression increase, electronic equipment faces unprecedented challenges in image processing.
[0003] When electronic devices collect images, due to the limitations of sensor technology and optical systems, the images obtained are often accompanied by noise and mosaic phenomena. Image noise will cause image details to be blurred and colors to be distorted, while the mosaic phenomenon will cause the image to show obvious color block distribution before processing.
[0004] Therefore, image denoising and demosaicing are often the key stages in the image signal processing process. Summary of the invention
[0005] In view of this, the image processing method, apparatus, device, and storage medium provided in the embodiments of the present application can improve the display quality of RAW images and improve the signal-to-noise ratio and time domain stability of the output images. The image processing method, apparatus, device, and storage medium provided in the embodiments of the present application are implemented as follows:
[0006] In a first aspect, an embodiment of the present application provides an image processing method, comprising:
[0007] Acquire an image to be processed, where the image to be processed is a RAW image of the current frame;
[0008] According to the associated image and the image to be processed, the image to be processed is jointly denoised and demosaiced to obtain a current frame output image. When the current frame RAW image is the first frame image, the associated image is a preset image. When the current frame RAW image is not the first frame image, the associated image is the previous frame output image. The previous frame output image is an image obtained after the joint denoising and demosaicing processing.
[0009] In a second aspect, an embodiment of the present application provides an image processing device, including:
[0010] An acquisition module is used to acquire an image to be processed, where the image to be processed is a RAW image of the current frame;
[0011] A processing module is used to perform joint denoising and demosaicing on the image to be processed according to the associated image and the image to be processed to obtain a current frame output image. When the current frame RAW image is the first frame image, the associated image is a preset image. When the current frame RAW image is not the first frame image, the associated image is a previous frame output image, and the previous frame output image is an image obtained after the joint denoising and demosaicing processing.
[0012] In a third aspect, an embodiment of the present application provides a computer device, including a memory and a processor, wherein the memory stores a computer program that can be executed on the processor, and when the processor executes the program, the method described in the embodiment of the present application is implemented.
[0013] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method provided in the embodiment of the present application.
[0014] The image processing method, apparatus, computer device and computer-readable storage medium provided in the embodiments of the present application, after acquiring the current frame RAW image to be processed, jointly denoise and demosaic the current frame RAW image based on the associated image and the current frame RAW image to obtain the current frame output image, and when the current frame RAW image is the first frame image, the associated image is the preset image, and when the current frame RAW image is not the first frame image, the associated image is the previous frame output image obtained after the joint denoise and demosaic processing.
[0015] In this way, when the current frame RAW image is jointly denoised and demosaiced, not only the information of the current frame RAW image itself is considered, but also the time domain information between the previous frame image associated with it is considered. Based on the time domain information between the previous and next frame images, the display quality of the RAW image can be improved, and the signal-to-noise ratio and time domain stability of the output image can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and are used together with the specification to illustrate the technical solution of the present application.
[0017] Figure 1 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;
[0018] Figure 2 A schematic diagram of an implementation flow of an image processing method provided in an embodiment of the present application;
[0019] Figure 3A schematic diagram of the implementation flow of another image processing method provided in an embodiment of the present application;
[0020] Figure 4 A schematic diagram of an implementation flow of another image processing method provided in an embodiment of the present application;
[0021] Figure 5 A schematic diagram of the structure of a pre-trained target network provided in an embodiment of the present application;
[0022] Figure 6 A schematic diagram of the effect of a color correction process provided by an embodiment of the present application;
[0023] Figure 7 A schematic diagram of a target network training implementation process provided in an embodiment of the present application;
[0024] Figure 8 A schematic diagram of the effect of an output image sequence set provided by an embodiment of the present application;
[0025] Fig. 9 A schematic diagram of an implementation flow of obtaining an output image sequence set provided in an embodiment of the present application;
[0026] Fig.10 A schematic diagram of the implementation process of obtaining a time domain training data set provided in an embodiment of the present application;
[0027] Fig.11 A structural schematic diagram of a degradation simulation process provided in an embodiment of the present application;
[0028] Fig.12 A schematic diagram of a training process of a global alignment module provided in an embodiment of the present application;
[0029] Fig.13 A schematic diagram of a training process of a local alignment module provided in an embodiment of the present application;
[0030] Fig.14 A schematic diagram of a training process of a processing module provided in an embodiment of the present application;
[0031] Fig.15 A schematic diagram of the structure of an image processing device provided in an embodiment of the present application;
[0032] Fig.16 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0033] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the specific technical solution of the present application will be further described in detail below in conjunction with the drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application, but are not used to limit the scope of the present application.
[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0035] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0036] It should be pointed out that the terms "first\second\third" involved in the embodiments of the present application are used to distinguish similar or different objects, and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0037] Before introducing the relevant technical solutions of the present application, the terms involved in the embodiments of the present application are first explained.
[0038] Image Signal Processor (ISP): ISP is a hardware or processing unit dedicated to processing image signals. It is widely used in the signal conversion process between image sensors (such as CMOS or CCD sensors) and display devices. The core function of ISP is to extract and optimize image information from the raw data output by the sensor, and finally output an image that can be displayed, stored or further processed. It is usually integrated in the entire link of image processing, involving a variety of image processing algorithms, covering all aspects from signal acquisition to final output.
[0039] Demosaic (DM): Demosaic technology is also called color filter array interpolation (CFA interpolation) or color reconstruction. The purpose of demosaicing is to recover the missing pixel values from the RAW format image to form a standard RGB image. In a digital camera, the image sensor samples the original image in Bayer format through the color filter array (CFA), and each pixel only records one color information (red, green or blue). The demosaicing process is to restore the complete RGB value of each pixel through an interpolation algorithm based on this incomplete information.
[0040] RAW domain noise reduction: RAW domain noise reduction is an important part of image processing, especially in the ISP pipeline. RAW domain noise reduction refers to the noise reduction of RAW format image data in the early stage of image signal processing, that is, before the image data is converted to common formats such as JPEG or PNG. RAW data is the original data obtained directly from the image sensor, usually in Bayer arrangement, that is, each pixel records only one color component (red, green, blue or a combination thereof).
[0041] Joint Denoising and Demosaicing (JDD): Denoising and demosaicing are treated as a whole task, aiming to reduce noise and restore missing pixel values at the same time. This method is usually based on a deep learning model, which learns the ability of denoising and demosaicing at the same time by training the model. This method can more effectively utilize image information, reduce information loss, and improve the visual effect of the image.
[0042] Ground Truth (GT) for network training: The ground truth label refers to the correct answer or true state of each sample in the dataset. In the context of machine learning and deep learning, it represents the actual situation or expected result of each piece of data in actual application. These labels are usually provided by human experts and provide a standard reference system for the model to judge whether its prediction results are accurate.
[0043] Color Conversion Matrix (CCM): The color correction matrix is a 3x3 or 3x4 matrix used to perform color correction on the original RGB image output by the image sensor. Its main purpose is to adjust the RGB values captured by the sensor to make them closer to the real colors perceived by the human eye or the response values of the standard color space.
[0044] Color Space Transforming (CST): CST refers to the process of mapping color data in one color model or representation to another color model. A color space is essentially a three-dimensional coordinate system where each point represents a certain color in an image. Common color spaces include RGB, HSV / HSL, Lab, YCbCr, etc. Different color spaces have different advantages in different processing tasks, so color space transformation is widely used in image processing, computer vision, image compression, video coding and other fields.
[0045] Traditional RAW domain denoising and demosaicing are two independent modules in the RAW domain processing of the ISP pipeline. Most existing ISPs place the RAW domain denoising module before the demosaicing module. RAW domain denoising will inevitably lose some details, resulting in the subsequent demosaicing being unable to obtain this lost information, reducing the image resolution. In order to avoid interference between the denoising module and the demosaicing module, a deep learning method is usually used to complete the denoising and demosaicing tasks together, namely the joint denoising and demosaicing (JDD) method.
[0046] Regarding the joint denoising and demosaicing methods in the related technologies, some methods directly explicitly model the joint denoising and demosaicing processing tasks into the denoising part and the demosaicing part, first design the denoising network and the demosaicing network respectively, and then train the two parts of the neural network in series or in parallel as the final joint denoising and demosaicing network; another part of the method adopts implicit end-to-end modeling, and directly uses one network to solve the denoising and demosaicing problems; and another part of the method introduces priors into the neural network, for example, using black and white images as guide images to guide the joint denoising and demosaicing of color images, or introducing distribution learning into the joint denoising and demosaicing processing method, using the variational method to derive a closed solution, and guide the neural network training process.
[0047] The joint denoising and demosaicing methods in related technologies are single-frame input and output, and do not consider using time domain information for further detail enhancement. The time domain information of the previous and next frames in the video is highly correlated. Incorporating time domain information into the denoising framework can improve the signal-to-noise ratio and time domain stability of the video image.
[0048] Although there is a time domain noise reduction method in the ISP of the related technology, this method simply aligns the previous and next frames and then fuses them. When the object moves significantly, ghosting and motion blur are likely to occur.
[0049] In view of this, an embodiment of the present application provides an image processing method, which is applied to electronic devices. The electronic devices may include but are not limited to mobile phones, wearable devices (such as smart watches, smart bracelets, smart glasses, etc.), tablet computers, laptops, vehicle terminals, PCs (Personal Computers), etc.
[0050] The functions implemented by the method can be implemented by calling program codes by a processor in the electronic device. Of course, the program codes can be stored in a computer storage medium. It can be seen that the electronic device at least includes a processor and a storage medium.
[0051] Figure 1 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0052] For example, Figure 1 As shown, the electronic device 10 may include a processor 101 , an external memory interface 102 , an internal memory 103 and a universal serial bus (USB) interface 104 .
[0053] It is to be understood that the structure illustrated in the embodiment of the present application does not constitute a specific limitation on the electronic device 10. In other embodiments of the present application, the electronic device 10 may include more or fewer components than shown in the figure, or combine some components, or split some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0054] The processor 101 may include one or more processing units, for example: the processor 101 may include an application processor (application processor, AP), a modem processor, a graphics processor (graphics processing unit, GPU), an image signal processor (image signal processor, ISP), a controller, a video codec, a digital signal processor (digital signal processor, DSP), a baseband processor, and / or a neural network processing unit (neural network processing unit, NPU), etc. Among them, different processing units can be independent devices or integrated into one or more processors. Exemplarily, the processor 101 can be a smart terminal CPU, such as a Snapdragon series processor, etc. In some embodiments, the processor 101 may include one or more interfaces. The interface may include an inter integrated circuit (I2C) interface, an inter integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0055] The external memory interface 102 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 10. The external memory card communicates with the processor 101 through the external memory interface 102 to implement a data storage function, such as storing music, video and other files in the external memory card.
[0056] The internal memory 103 can be used to store computer executable program codes, which include instructions. The internal memory 103 can include a program storage area and a data storage area. The program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area can store data created during the use of the electronic device 10 (such as audio data, a phone book, etc.), etc.
[0057] In addition, the electronic device involved in the embodiments of the present application may also be installed with an operating system, and application programs may be installed and run on the operating system, which is not limited in the embodiments of the present application.
[0058] Figure 2 The following is a schematic diagram of the implementation process of the image processing method provided in the embodiment of the present application. Figure 2 As shown, the method may include the following steps 201 to 202:
[0059] Step 201, obtaining an image to be processed, where the image to be processed is a RAW image of the current frame.
[0060] It should be noted that when an electronic device takes a photo, the image sensor captures light and converts it into electrical signals. These electrical signals are then converted into digital signals and saved in RAW format. RAW images contain the original pixel information captured from the image sensor without any image processing or color encoding.
[0061] In order to improve the image quality and color accuracy of the current frame RAW image, the current frame RAW image may be subjected to joint denoising and demosaicing processing.
[0062] Step 202, based on the associated image and the image to be processed, the image to be processed is jointly denoised and demosaiced to obtain a current frame output image. When the current frame RAW image is the first frame image, the associated image is a preset image. When the current frame RAW image is not the first frame image, the associated image is the previous frame output image, and the previous frame output image is an image obtained after the joint denoising and demosaicing processing.
[0063] In the embodiments of the present application, there is no limitation on the type of output image, for example, in some embodiments, the output image may be an RGB image, wherein each pixel in the RGB image contains three components, R, G, and B, so the amount of data for each pixel is relatively large.
[0064] Alternatively, in some other embodiments, the output image may be a YUV image, wherein the YUV image achieves downsampling of the chrominance component by separating the luminance and chrominance components, thereby reducing the amount of data.
[0065] In this way, when the image is subsequently processed based on ISP, since the size and bit width of the RGB image are larger than the size and bit width of the YUV image, when the output image is a YUV image, the bandwidth and power consumption required by the electronic device for subsequent image processing can be effectively reduced.
[0066] In the embodiment of the present application, when the joint denoising and demosaicing processing is performed on the current frame RAW image, the processing is performed based on the current frame RAW image's own information and the information of the associated images.
[0067] When the current RAW image is the first image, the associated image may be a preset image. Here, the preset image may be different based on the type of the associated image.
[0068] For example, when the associated image type is an RGB image, the preset image may be an image in which the R, G, and B pixel values are all assigned a value of 0; when the associated image type is a YUV image, the preset image may be an image in which the brightness Y is 0, and the hue U and saturation V are also 0.
[0069] Of course, the pixel values in the preset image may also be other values, which is not limited to this.
[0070] When the current RAW image is not the first image, the associated image is the previous output image obtained by performing joint denoising and demosaicing on the previous RAW image according to the previous output image and the previous RAW image. That is, the associated image used to process the RAW image here is the output image obtained after the previous processing. After this cycle, joint denoising and demosaicing can be performed on each RAW image to obtain the output image corresponding to each RAW image.
[0071] In an embodiment of the present application, after the current frame RAW image to be processed is acquired, the current frame RAW image is jointly denoised and demosaiced based on the associated image and the current frame RAW image to obtain the current frame output image, and when the current frame RAW image is the first frame image, the associated image is a preset image, and when the current frame RAW image is not the first frame image, the associated image is the previous frame output image obtained after the joint denoising and demosaicing processing.
[0072] In this way, when the current frame RAW image is jointly denoised and demosaiced, not only the information of the current frame RAW image itself is considered, but also the time domain information between the previous frame image associated with it is considered. Based on the time domain information between the previous and next frame images, the display quality of the RAW image can be improved, and the signal-to-noise ratio and time domain stability of the output image can be improved.
[0073] Figure 3 The following is a schematic diagram of the implementation process of the image processing method provided in the embodiment of the present application. Figure 3 As shown, the method may include the following steps 301 to 302:
[0074] Step 301, obtaining an image to be processed, where the image to be processed is a RAW image of the current frame.
[0075] Here, the method of executing step 301 is the same as the method of executing step 201 in the above embodiment, and will not be repeated here.
[0076] Step 302, according to the associated image, the image to be processed and the pre-trained target network, perform joint denoising and demosaicing processing on the image to be processed to obtain the current frame output image.
[0077] Among them, when the current frame RAW image is the first frame image, the associated image is the preset image. When the current frame RAW image is not the first frame image, the associated image is the previous frame output image, and the previous frame output image is an image obtained after joint denoising and demosaicing processing.
[0078] In addition, the pre-trained target network includes an alignment module and a processing module. The alignment module is used to perform time domain alignment processing on the associated image and the image to be processed to obtain a target aligned image. The processing module is used to perform joint denoising and demosaicing processing on the image to be processed through the target aligned image to obtain the current frame output image.
[0079] In an embodiment of the present application, the alignment module and the processing module in a pre-trained target network are connected. After the image to be processed is acquired, the image to be processed and the associated image can be input into the pre-trained target network, so as to first perform time domain alignment processing on the associated image and the image to be processed through the alignment module in the pre-trained target network to obtain a target aligned image; then, the target aligned image and the image to be processed (the current frame RAW image) are input into the processing module in the pre-trained target network, so that the processing module performs joint denoising and demosaicing processing on the image to be processed (the current frame RAW image) through the target aligned image to obtain the current frame output image.
[0080] After the above cycle, the joint denoising and demosaicing processing of each frame of RAW image can be realized to obtain the output image corresponding to each frame of RAW image.
[0081] In an embodiment of the present application, after the current frame RAW image to be processed is obtained, the associated image and the current frame RAW image are input into a pre-trained target network to perform joint denoising and demosaicing on the current frame RAW image to obtain the current frame output image.
[0082] In this way, when the current frame RAW image is jointly denoised and demosaiced, the current frame RAW image and the associated image are aligned in the time domain by using the information of the two frames of images in the time domain through the pre-trained target network, and then the current frame RAW image is jointly denoised and demosaiced, which can improve the display quality of the RAW image and improve the signal-to-noise ratio and time domain stability of the output image.
[0083] Figure 4 The following is a schematic diagram of the implementation process of the image processing method provided in the embodiment of the present application. Figure 4 As shown, the method may include the following steps 401 to 404:
[0084] Step 401, obtaining an image to be processed, where the image to be processed is a RAW image of the current frame.
[0085] Here, the method of executing step 401 is the same as the method of executing step 201 in the above embodiment, and will not be repeated here.
[0086] Step 402: Perform global alignment processing on the associated image and the image to be processed by using the pre-trained global alignment module of the target network to obtain a globally aligned image.
[0087] Among them, when the current frame RAW image is the first frame image, the associated image is the preset image. When the current frame RAW image is not the first frame image, the associated image is the previous frame output image, and the previous frame output image is an image obtained after joint denoising and demosaicing processing.
[0088] In an embodiment of the present application, the pre-trained target network includes an alignment module and a processing module, the alignment module includes a global alignment module and a local alignment module, the global alignment module is connected to the local alignment module, and the local alignment module is connected to the processing module.
[0089] Figure 5 A schematic diagram of the structure of a pre-trained target network is given.
[0090] like Figure 5 As shown, after the current frame RAW image is acquired, the current frame RAW image and the associated images can be input into the global alignment module of the pre-trained target network, so that the associated images associated with the current frame RAW image and the current frame RAW image can be globally aligned through the global alignment module to obtain a globally aligned image.
[0091] Thus, in image processing, due to the movement of the electronic device or the change of the shooting angle, there may be an obvious displacement between the two acquired frames. Through the global alignment module, the current frame RAW image can be aligned in position to the associated image (i.e., the image obtained by the pre-trained target network through the previous frame RAW image and the previous frame output image, and the previous frame RAW image is jointly denoised and demosaiced), so that the obtained global alignment image includes the time domain information of the current frame RAW image and the associated image for subsequent processing and analysis.
[0092] It should be noted that in the process of digital signal processing performed by electronic devices, de-mosaicing of images will transform the images from the RAW domain to the RGB domain, and from the RGB domain to the YUV domain, it also includes color correction processing, gamma processing and color space conversion processing, etc. However, in practice, it is found that the neural network model can directly learn the gamma processing method and color space conversion processing method from the data, but cannot learn the color correction processing.
[0093] Based on this, in some embodiments, such as Figure 5 As shown, the pre-trained target network may further include a pre-processing module, and the pre-processing module is connected to the global alignment module.
[0094] In this way, before globally aligning the associated image and the image to be processed to obtain the globally aligned image, the image to be processed may be firstly subjected to color correction through a preprocessing module to obtain a corrected image to be processed; and then globally aligning the associated image and the corrected image to be processed may be subjected to global alignment through a global alignment module to obtain a globally aligned image.
[0095] In some embodiments, Figure 6 As shown, a schematic diagram of the effect of color correction processing is given.
[0096] like Figure 6 As shown, in order to implement color correction processing on the image to be processed and obtain the corrected image to be processed, the initial pixel value corresponding to each pixel in the image to be processed can be determined first; according to the preset color correction matrix, the initial pixel value corresponding to each pixel is corrected to obtain the corrected image to be processed, and the pixel value of each pixel in the corrected image to be processed is the corrected pixel value after correction.
[0097] That is to say, when performing color correction (CCM) processing on each pixel in the image to be processed, the initial pixel value of each pixel in the image to be processed (current frame RAW image) can be obtained first. For example, the initial pixel value of the current frame RAW image may include R, G or B.
[0098] In this way, after the initial pixel value of each pixel in the current frame RAW image is obtained, the initial pixel value corresponding to each pixel can be corrected according to a preset color correction matrix to obtain a corrected image to be processed.
[0099] like Figure 6 As shown, in some embodiments, when the initial pixel value of a pixel point in the current frame RAW image is R or B, the two adjacent G channels can be averaged through a preset color correction matrix to calculate the corrected pixel value corresponding to the initial pixel value, and the preset color correction matrix can be calculated by ISP; when the initial pixel value of a pixel point in the current frame RAW image is G, the adjacent B pixel points and R pixel points are used to calculate the corrected pixel value corresponding to the initial pixel value.
[0100] In some embodiments, after obtaining the corrected image to be processed, the associated image and the corrected image to be processed can be input into a global alignment module to obtain a projection transformation matrix; then, the associated image and the corrected image to be processed are aligned through the projection transformation matrix to obtain a globally aligned image.
[0101] That is, the associated image (i.e., the image obtained by the pre-trained target network after joint denoising and demosaicing of the previous frame RAW image and the previous frame output image) and the corrected image to be processed can be input into the global alignment module together, so that the global alignment module can regress the projection transformation matrix through training and learning. Among them, the projection transformation is a geometric transformation that projects a two-dimensional image onto another two-dimensional plane, and the projection transformation matrix is a 3x3 matrix, which describes the mapping relationship from the associated image to the current frame RAW image. Subsequently, the projection transformation matrix is used to project the associated image and the corrected image to be processed, so as to realize the alignment of the two frames of images in spatial position and obtain a globally aligned image.
[0102] Here, the global alignment module can reduce the displacement between the associated image and the current frame RAW image, so that the associated image and the corrected image to be processed are aligned in spatial position.
[0103] Step 403: Perform local alignment processing on the global aligned image and the image to be processed by using the pre-trained local alignment module of the target network to obtain a local aligned image.
[0104] It is understandable that when an electronic device captures an image, there may be local changes such as rotation, scaling, and distortion in the image. In order to better adapt to local changes in the image, local alignment processing can be performed to improve the accuracy and robustness of the aligned image.
[0105] Based on this, in the embodiment of the present application, after globally aligning the associated image and the current frame RAW image to obtain the globally aligned image, the globally aligned image and the image to be processed can be further locally aligned to obtain the locally aligned image. Here, since the two frames of images have been preliminarily aligned by the global alignment module, the local alignment module mainly focuses on the compensation of local displacement.
[0106] In some embodiments, in order to locally align the global alignment image and the image to be processed to obtain the locally aligned image, the global alignment image and the image to be processed can be input into a local alignment module in a pre-trained target network, and the local alignment module can regress the global alignment image and the image to be processed to obtain an optical flow map through training and learning. The optical flow map is a two-dimensional vector field, in which each vector represents the movement direction and speed of the corresponding pixel point in the image.
[0107] Through the optical flow map, the motion area of the global alignment image is aligned with the motion area of the image to be processed to obtain a local alignment image. The alignment process moves the pixel points in the global alignment image to the corresponding position in the image to be processed based on the motion vector of each pixel point in the optical flow map.
[0108] By implementing this embodiment, it can be ensured that the local displacement in the global alignment image is accurately compensated in the image to be processed (the current frame RAW image).
[0109] In the embodiment of the present application, when determining the optical flow map, the optical flow algorithm used may be an optical flow estimation algorithm based on dense sampling (Dense Inverse Search, DIS). The DIS optical flow algorithm is mainly used to estimate the motion vector of the pixel in the image. Using dense sampling, the optical flow is estimated for each pixel in the image, so that a global dense optical flow field can be obtained.
[0110] Step 404 , using the processing module of the pre-trained target network, jointly denoise and demosaic the local aligned image and the image to be processed to obtain an output image of the current frame.
[0111] Understandably, in some scenarios, the associated image and the corrected image to be processed (the current frame RAW image processed by the preprocessing module) have almost no identical content. In this case, even if the associated image and the corrected image to be processed are globally aligned and locally aligned, the locally aligned image obtained is not conducive to the joint denoising and demosaicing of the image to be processed (the current frame RAW image).
[0112] Based on this, Figure 5 As shown, in some embodiments, the processing module may also include a reliability assessment module and a reconstruction module, and the reliability assessment module is connected to the local alignment module.
[0113] In some embodiments, the image similarity between the local aligned image and the image to be processed can be determined through a reliability assessment module; and the image to be processed is jointly denoised and demosaiced according to the image similarity and the local aligned image through a reconstruction module to obtain a current frame output image.
[0114] In some embodiments, the local aligned image and the image to be processed may be first input into a reliability evaluation module, and the image similarity between the local aligned image and the image to be processed may be calculated by the reliability evaluation module.
[0115] Here, there is no limitation on the type of presentation result of image similarity. For example, image similarity can be presented in the form of a reliability image, the weight of which is between 0 and 1, and the weight is used to characterize the image similarity between the local alignment image and the image to be processed. The larger the weight, the greater the image similarity between the local alignment image and the image to be processed.
[0116] Alternatively, the image similarity can also be directly presented in a numerical form such as percentage. For example, the image similarity can be directly output as 80%, which indicates that the similarity between the local aligned image and the image to be processed is large, and there are more similar contents. The local aligned image plays a guiding role in the reconstruction of the image to be processed.
[0117] In some embodiments, Figure 5As shown, the reliability assessment module is also connected to the reconstruction module, so that after determining the image similarity between the local aligned image and the image to be processed, the reliability image (a form used to characterize the image similarity between the local aligned image and the image to be processed), the local aligned image and the image to be processed can be input into the reconstruction module together, so that the reconstruction module performs joint denoising and demosaicing on the image to be processed according to the image similarity and the local aligned image to obtain the current frame output image.
[0118] Among them, the greater the image similarity, the more information in the local aligned image is used after the processed image is jointly denoised and demosaiced, and the smaller the image similarity is, the less information in the local aligned image is used after the processed image is jointly denoised and demosaiced, thereby avoiding the situation where the reconstruction effect of the image to be processed is affected due to too little similar content in the associated image and the image to be processed.
[0119] By implementing this embodiment, the reliability assessment module intercepts the extreme situation that the associated image and the corrected image to be processed (the current frame RAW image after being processed by the preprocessing module) have almost no identical content, as well as the situation that the global alignment processing and the local alignment processing are inaccurate, thereby avoiding the reconstruction effect of the image to be processed (the current frame RAW image) deteriorated by the local aligned image.
[0120] In an embodiment of the present application, the current frame RAW image is obtained and used as the image to be processed. The global alignment module of the pre-trained target network is used to perform global alignment processing on the associated image and the image to be processed to obtain a globally aligned image; then, the local alignment module of the pre-trained target network is used to perform local alignment processing on the global aligned image and the image to be processed to obtain a locally aligned image; finally, the processing module of the pre-trained target network is used to perform joint denoising and demosaicing processing on the locally aligned image and the image to be processed to obtain the current frame output image. In this way, when the current frame RAW image is subjected to joint denoising and demosaicing processing, the current frame RAW image and the associated image are subjected to time domain alignment processing by using the information of the two frames of images in the time domain through the pre-trained target network, and then the current frame RAW image is subjected to joint denoising and demosaicing processing, which can improve the display quality of the RAW image and improve the signal-to-noise ratio and time domain stability of the output image.
[0121] Understandably, the target network needs to be trained with the data set first, and can be put into use only after the training is completed to improve the joint denoising and demosaicing effects of the image.
[0122] Based on this, in some embodiments, before acquiring the image to be processed, the target network may be trained to obtain a trained target network, and then the RAW image may be denoised and demosaiced through the trained target network.
[0123] In the embodiment of the present application, there is no limitation on the training method of the target network.
[0124] For example, in some embodiments, to implement training of the target network, the following steps 701 to 702 may be performed:
[0125] Step 701: Acquire a time domain training data set, where the time domain training data set includes a plurality of image pairs, and each image pair includes a noisy RAW image and a noise-free output image.
[0126] It is understandable that the training dataset is the basis of image training. It contains a large number of image samples for training models. These samples cover different scenes, objects and features, providing rich visual information for the model. By learning and analyzing these samples, the model can gradually grasp the laws and features in the image, thus having the corresponding functions.
[0127] Based on this, in an embodiment of the present application, before training the target network, a time domain training data set that can be used to train the target network can be obtained. Here, the time domain training data set may include multiple image pairs, each of which includes a noisy RAW image and a noise-free output image, wherein the noise-free output image can be understood as a true value image, which is a control image of the noisy RAW image.
[0128] In the embodiment of the present application, there is no limitation on the method of obtaining the time domain training data set.
[0129] For example, in some embodiments, to obtain a time domain training data set, an output image sequence set may be first obtained, the output image sequence set including a plurality of continuous frame output images; and then a degradation simulation process is performed on the output image sequence set to obtain a time domain training data set.
[0130] Here, there is also no limitation on the method for obtaining the output image sequence set.
[0131] For example, in some embodiments, the output image sequence set is an image sequence captured in a scene where the moving speed is greater than a speed threshold. Here, there is no limitation on the value of the speed threshold. For example, when the speed threshold is a large value, the scene where the moving speed is greater than the speed threshold can be a scene shot by a high-speed camera.
[0132] Figure 8 A schematic diagram of the effect of outputting a set of image sequences is given.
[0133] like Figure 8 As shown, the output image sequence set can be a sequence set composed of multiple images collected by a high-speed camera (i.e., the actual collected data set). It is understandable that when ordinary cameras record scenes with fast-moving objects, motion blur will occur, affecting the quality of the output image sequence set. Existing high-speed cameras can capture hundreds of thousands of images per second, which is enough to cope with high-speed motion scenes, such as players competing on the court, pets running, fireworks, etc. Therefore, using a high-speed camera to focus on collecting a batch of image sequences or videos of high-speed motion scenes to obtain an output image sequence set can effectively improve the effect of the target network in sudden large displacement scenes.
[0134] In some other embodiments, the set of output image sequences is a sequence of images including different motion characteristics.
[0135] In the embodiment of the present application, based on the different types of the output image sequence set, the corresponding acquisition method is also different.
[0136] For example, in some embodiments, when the output image sequence set is a sequence of images including different motion characteristics, the output image sequence set may be obtained by executing the following steps 901 to 902:
[0137] Step 901: Acquire an initial image, where the initial image is an RGB image with a resolution greater than a resolution threshold.
[0138] It can be understood that in order to better train the target network, training images with better quality are required. Therefore, the initial image obtained in the embodiment of the present application may be a high-resolution image, that is, the resolution of the initial image is greater than the resolution threshold.
[0139] Step 902, simulate motion processing is performed on the initial image to obtain at least one set of output image sequence sets, where the simulated motion processing includes any one or more of replication processing, random cropping processing, projection transformation processing, optical flow processing and hybrid processing.
[0140] In an embodiment of the present application, after acquiring the initial image, a series of simulated motion processing can be performed on the initial image to obtain multiple groups of initial image sequence sets under different types of simulated motion processing, thereby expanding the number of images and obtaining a high-quality output image sequence set.
[0141] In some embodiments, when the simulated motion processing is a copy processing, the initial image may be repeated several times. Here, the initial image may be repeated several times to obtain an RGB sequence image in a static scene, and this series of RGB sequence images connected in time sequence is used as a set of output image sequence sets. Figure 8 As shown, based on this method, a simulation data set 1 can be obtained.
[0142] In some embodiments, when the simulated motion processing is a random cropping process, after determining the size of the image to be cropped, a plurality of cropped images can be acquired based on the position of a corner of the image to be cropped in the initial image.
[0143] Here, we take the upper left corner of the image to be cropped as an example. The pixel position of the upper left corner of the cropped image can be modeled by random walk, that is, given the initial position of the upper left corner pixel of the cropped image, a unique row and column is randomly generated each time to obtain the new position of the upper left corner pixel of the cropped image, thereby obtaining multiple cropped images. Finally, a time sequence is assigned to the obtained multiple cropped images to obtain a set of output image sequence sets. Figure 8 As shown, based on this method, a simulation data set 2 can be obtained.
[0144] In some embodiments, when the simulated motion processing is a projection transformation processing, since the projection transformation can be used to simulate the transformation of the shooting angle, the standard projection transformation has 8 parameters, and four pairs of coordinates are required to determine the parameters of the projection transformation. Based on this, the coordinates of the four corners of the initial image can be modeled by random walks, and a continuous projection transformation parameter matrix can be obtained. By using these continuous matrices to perform projection transformation on the image, a continuous projection transformation image sequence can be obtained, and a set of output image sequence sets can be obtained. Figure 8 As shown, based on this method, a simulation data set 3 can be obtained.
[0145] In some embodiments, when the simulated motion processing is optical flow processing, since the optical flow describes the displacement of the projection point of a spatial point in the imaging plane, the initial image can be first divided into several grids, which is usually achieved by specifying the number of rows and columns of the grid, thereby dividing the image into multiple small areas, and the vertices of each small area will be used to generate the optical flow; then, the optical flow vector can be randomly generated at the vertex of each grid. The optical flow vector represents the direction and speed of movement of the pixel in the image from the current frame to the next frame. These vectors can be two-dimensional, containing components in the horizontal and vertical directions. With the optical flow of the grid vertices, the interpolation method (such as bilinear interpolation, bicubic interpolation, etc.) can be used to calculate the optical flow of other pixels in the image. In this way, the entire image has an optical flow field, that is, each pixel has a corresponding optical flow vector.
[0146] Secondly, after obtaining the current frame image (initial image) and the optical flow field, the next frame image (the next frame image corresponding to the initial image) can be calculated by reverse mapping or forward mapping. Reverse mapping is the mapping from the next frame to the current frame. The position of each pixel in the current frame in the next frame is found through the optical flow vector, and the pixel value is sampled from the current frame. After determining the next frame image corresponding to the initial image, it can be added to the output image sequence set. By looping the above method, multiple frames of images can be obtained, thereby obtaining a set of output image sequence sets. For example Figure 8 As shown, based on this method, a simulation data set 4 can be obtained.
[0147] In some embodiments, when the simulated motion processing is a hybrid processing, after obtaining the simulation data sets 1 to 4, the above simulation data sets may be hybrid processed to obtain an output image sequence set. Figure 8 As shown, based on this method, a simulation data set 5 can be obtained.
[0148] There is no limitation on the mixing method. For example, some images in each simulation data set may be randomly selected and mixed to form an output image sequence set; or at least two simulation data sets may be randomly selected and mixed to form an output image sequence set; or all images in some simulation data sets and some images in some simulation data sets may be randomly selected and mixed to form an output image data set.
[0149] In the embodiment of the present application, there is no limitation on the image type in the output image sequence set, for example, it can be a YUV image. After the initial image is subjected to simulated motion processing, the image obtained is an RGB image, so the RGB image can be further converted to a YUV image.
[0150] By implementing this embodiment, on the one hand, by obtaining multiple groups of initial image sequence sets under different types of simulated motion processing, the number of images can be expanded to obtain a high-quality output image sequence set; on the other hand, by introducing the acquisition of fast-motion images in high-speed scenes, the defect of motion blur caused by ordinary imaging equipment when acquiring high-speed motion scenes is compensated, thereby ensuring the quality of the output image data set.
[0151] It can be understood that the images in the output image sequence set contain rich brightness and chromaticity information. Through data degradation and simulation, various degradation conditions that images may encounter in the real world can be simulated, such as noise, blur, illumination changes, etc. These degradation conditions are more common in actual application scenarios. Therefore, using the degraded and simulated time domain training data set to train the target network can make the target network better adapt to the complexity of the real world.
[0152] Based on this, in some embodiments, after obtaining the output image sequence set, in order to implement degradation simulation processing on the output image sequence set to obtain a time domain training data set, the following steps 1001 to 1005 may be performed:
[0153] Step 1001: randomly generate image signal processing parameters, and perform inverse transformation on the image signal processing parameters to obtain inverse transformation parameters.
[0154] Image signal processing parameters refer to a series of parameters used by the image signal processor when processing images, which have a crucial impact on the final presentation of the image. Image signal processing parameters usually include color correction matrix (CCM), automatic white balance (AWB) parameters, automatic exposure control (AE) parameters, etc., which work together on the original signal output by the image sensor to produce high-quality image output.
[0155] Among them, the CCM parameter is used to correct the color deviation of the image to ensure the accurate color reproduction of the image. It converts the original color data output by the image sensor into a color that is closer to the human eye's perception through a series of mathematical operations.
[0156] The AWB parameter is used to adjust the color temperature of the image to ensure that the image maintains natural colors under different lighting conditions. It automatically adjusts the color temperature of the image by analyzing the color information in the image to make it closer to the real color environment.
[0157] The AWB parameter is used to adjust the color temperature of the image to ensure that the image maintains natural colors under different lighting conditions.
[0158] In some embodiments, the image signal processing parameters may also include sharpness adjustment parameters, noise reduction processing parameters, etc. These parameters act together on the image signal processor to generate high-quality image output.
[0159] Fig.11 A structural schematic diagram of degradation simulation processing is given.
[0160] like Fig.11 As shown, in an embodiment of the present application, after obtaining the output image sequence set, in order to realize the degradation and simulation processing of the output image data set, the image signal processing parameters can be randomly generated first, and then the image signal processing parameters are inversely transformed to obtain inverse transformation parameters.
[0161] Step 1002, inputting a target output image and inverse transformation parameters in the output image sequence set into an inverse image simulation processor to obtain a noise-free RAW image, wherein the target output image is any image in the output image sequence set.
[0162] Here, any image in the output image sequence set can be randomly selected and used as the target output image, and input into the inverse image simulation processor together with the inverse transformation parameters to obtain a noise-free RAW image. Repeating the above steps for each image in the output image sequence set can obtain a noise-free RAW image corresponding to each image in the output image sequence set.
[0163] In the embodiment of the present application, the inverse image simulation processor may be a script for data degradation, or may be software, or may be in a neural network.
[0164] Step 1003 , according to a preset calibration noise model, the noise-free RAW image is subjected to noise adding processing to obtain a noisy RAW image.
[0165] In an embodiment of the present application, the calibration noise model may be acquired in advance. When acquiring the preset calibration noise model, the noise calibration may be performed on the initial RAW image directly output from the sensor to obtain the preset calibration noise model.
[0166] After the preset calibration noise model is obtained, the noise-free RAW image obtained in step 1002 may be subjected to noise addition processing based on the preset calibration noise model to obtain a noisy RAW image corresponding to each image in the output image sequence set.
[0167] Step 1004: input the noisy RAW image and the inverse transformation parameters into an image simulation processor to obtain a noise-free output image.
[0168] Here, after acquiring a noisy RAW image, the noisy RAW image and inverse transformation parameters can be input into an image simulation processor, so that the noisy RAW image can be processed by the image simulation processor according to the inverse transformation parameters, thereby obtaining a noise-free output image corresponding to each noisy RAW image.
[0169] Step 1005, repeat the above steps for each output image in the output image sequence set to obtain a time domain training data set.
[0170] Here, the steps from step 1001 to step 1004 are repeatedly performed for each output image in the output image sequence set, and a noisy RAW image and a noise-free output image corresponding to each output image can be obtained. The noisy RAW image and the noise-free output image corresponding to each output image constitute an image pair, and multiple image pairs can constitute a time domain training data set.
[0171] When implementing this embodiment, when performing degradation and simulation processing on an image, the output image sequence set used includes a combination of multiple simulated motion modes, which effectively simulates the motion modes in real scenes, can obtain unlimited training data based on limited actual data, and provides a sufficient number of high-quality training data sets to support the training of the target network.
[0172] Step 702: Train the target network using the time domain training data set to obtain a pre-trained target network.
[0173] Here, training the target network can be training the parameters of each module in the target network. In the related technology, when performing time domain noise reduction, the network parameters of the model are numerous and cumbersome, and it is not easy to call out a set of network parameters that are suitable for each scene. Therefore, the mobile terminal urgently needs an intelligent joint denoising and demosaicing network structure that can avoid ghosting, improve time domain consistency and video details.
[0174] And in the process of model training, the loss function plays a vital role. The loss function is a function that measures the difference between the model prediction and the actual result, and guides the optimization of the model parameters. During the training process, it is necessary to select a loss function that is suitable for the current task. Model training based on the loss function is an iterative process, which requires continuous adjustment and optimization of model parameters to improve prediction performance.
[0175] Based on this, in some embodiments, in order to ensure training efficiency and effect, when training the target network, each target submodule in the target network can be trained separately through the time domain training data set to obtain multiple trained target submodules, and the target submodules include a global alignment module, a local alignment module and a processing module. Here, the processing module may include a reliability assessment module and a reconstruction module.
[0176] In the embodiment of the present application, when any target sub-module is trained, the weights of other target sub-modules are kept unchanged.
[0177] For example, in some embodiments, when the target submodule is a global alignment module, in order to train the global alignment module through a time domain training data set to obtain a trained global alignment module, the following method may be performed:
[0178] The global alignment module of the target network is trained using the first training data set and the first loss function to obtain a trained global alignment module.
[0179] Fig.12 A schematic diagram of the training process of a global alignment module is given.
[0180] like Fig.12As shown, the first training data set includes output images with global information in the output image sequence set, such as the above-obtained simulation data set 1 (the output image sequence set obtained when the simulated motion processing is a copy processing), simulation data set 2 (the output image sequence set obtained when the simulated motion processing is a random cropping processing) and simulation data set 3 (the output image sequence set obtained when the simulated motion processing is a projection transformation processing).
[0181] The first loss function may include a first loss sub-function and a second loss sub-function. The first loss sub-function is determined by the transformation parameters of the projection transformation matrix obtained after the associated image and the image to be processed are processed by the global alignment module and the transformation parameters of the preset projection transformation matrix. The second loss sub-function is determined by the difference between the global aligned image and the processed image to be processed.
[0182] Here, the transformation parameters of the preset projection transformation matrix may be true values of the projection transformation parameters, and the first loss sub-function may be determined by calculating the difference between the transformation parameters of the projection transformation matrix and the true values of the projection transformation parameters.
[0183] like Fig.12 As shown, the processed image to be processed refers to the image to be processed after preprocessing (color correction processing), demosaicing processing, gamma processing and color space conversion processing. By calculating the difference between it and the global alignment image, the second loss sub-function can be obtained.
[0184] Among them, the demosaicing module can be a demosaicing network or a demosaicing algorithm in related technologies, and the gamma processing and color space conversion processing can be implemented using the processing methods in related technologies, which will not be repeated here.
[0185] In some embodiments, after obtaining the first loss sub-function and the second loss sub-function, a weighted average of the first loss sub-function and the second loss sub-function is taken as the first loss function.
[0186] Here, when training the global alignment module, the training parameters of the local alignment module and the processing module may be kept unchanged to reduce the training difficulty and improve the training efficiency.
[0187] In some embodiments, when the target submodule is a local alignment module, in order to train the local alignment module through the time domain training data set to obtain a trained local alignment module, the following method may be performed:
[0188] The local alignment module of the target network is trained using the second training data set and the second loss function to obtain a trained local alignment module.
[0189] Fig.13A schematic diagram of the training process of a local alignment module is given.
[0190] like Fig.13 As shown, the second training data set includes output images with motion information in the output image sequence set, such as the above-obtained simulation data set 4 (the output image sequence set obtained when the simulated motion processing is optical flow processing) and the simulation data set 5 (the output image sequence set obtained when the simulated motion processing is hybrid processing).
[0191] The second loss function includes a third loss sub-function, a fourth loss sub-function and a fifth loss sub-function. The third loss sub-function is determined by the difference between the local aligned image and the processed image to be processed. The fourth loss sub-function is determined by the difference between the optical flow image obtained after the associated image and the image to be processed are processed by the local alignment module and the preset optical flow image. The fifth loss sub-function is an optical flow field smoothing function.
[0192] like Fig.13 As shown in FIG. 1 , the processed image to be processed refers to the image to be processed after preprocessing (color correction processing), de-mosaicing, gamma processing and color space conversion. By calculating the difference between it and the local alignment image, the third loss sub-function can be obtained.
[0193] Here, the preset optical flow image may be a true value of the optical flow image, and the fourth loss sub-function may be determined by calculating the difference between the optical flow image and the true value of the optical flow image.
[0194] The fifth loss sub-function may be taken as the 2-norm of the optical flow image.
[0195] In some embodiments, after obtaining the third loss sub-function, the fourth loss sub-function and the fifth loss sub-function, a weighted average of the third loss sub-function, the fourth loss sub-function and the fifth loss sub-function may be taken as the second loss function.
[0196] Here, when training the local alignment module, the training parameters of the global alignment module and the processing module may be kept unchanged to reduce the training difficulty and improve the training efficiency.
[0197] In some embodiments, when the target submodule is a processing module, in order to train the processing module through the time domain training data set to obtain a trained local alignment module, the following method may be performed:
[0198] The processing module of the target network is trained by using the third training data set and the third loss function to obtain a trained processing module. The processing module includes a reliability evaluation module and a reconstruction module.
[0199] Fig.14A schematic diagram of the training process of a processing module is given.
[0200] like Fig.14 As shown, the third training data set is an output image sequence set, which may be the above-obtained simulation data set 5 (the output image sequence set obtained when the simulated motion processing is a hybrid processing) and the actual acquired data set (i.e., the output image sequence set acquired by a high-speed camera).
[0201] The third loss function includes a sixth loss sub-function and a seventh loss sub-function. The sixth loss sub-function is determined by the difference between the current frame output image obtained after the processing module performs joint denoising and demosaicing on the local aligned image and the image to be processed and the preset current frame output image. The seventh loss sub-function is determined based on image similarity, the current frame output image and the associated image.
[0202] Here, the preset current frame output image may be the output true value image, and the sixth loss sub-function may be determined by calculating the difference between the current frame output image and the output true value image.
[0203] In some embodiments, when the seventh loss sub-function is determined according to the image similarity, the current frame output image and the associated image, it can be implemented by the following formula 1:
[0204]
[0205] Among them, W, X prv and X JDD They are all one-dimensional vectors, representing image similarity, associated images, and JDD current frame RAW images, respectively. N represents the number of elements in the vector, and M represents the total number of samples.
[0206] In some embodiments, after the sixth loss sub-function and the seventh loss sub-function are obtained, a weighted average of the sixth loss sub-function and the seventh loss sub-function may be taken as the third loss function.
[0207] Here, when training the processing module, the training parameters of the global alignment module and the local alignment module may be kept unchanged to reduce the difficulty of training and improve the training efficiency.
[0208] In some embodiments, the learning rate may be reduced, the training parameters of all modules may be fine-tuned, and the final training may be completed to obtain a pre-trained target network.
[0209] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.
[0210] First, in the embodiments of the present application, the time-domain JDD data set is mainly produced by data degradation and simulation.
[0211] In the embodiment of the present application, since the input of the target network is a RAW image and the output is a YUV image, it is first necessary to obtain a batch of high-quality YUV sequence sets for data degradation and simulation.
[0212] like Figure 8 As shown, the methods of obtaining a YUV sequence set (ie, an output image sequence set) are divided into two categories.
[0213] The first category: obtaining a YUV sequence based on a single high-resolution image.
[0214] (1) By repeating a single high-resolution image several times, the RGB sequence of the static scene can be obtained, which is then converted into a YUV sequence (i.e., one of the output image sequence sets).
[0215] (2) After the size of the cropped image block is given, the cropped image block can be obtained by simply knowing the position of the upper left corner of the cropped image block in the original image. The pixel position of the upper left corner of the cropped image block is modeled as a random walk, that is, given the initial position of the upper left corner pixel of the cropped image, the row and column displacements are randomly generated each time to obtain the new position of the upper left corner pixel, and then the cropped image is obtained. Multiple frames of continuous cropped images are put together to form a continuously moving RGB sequence, which is then converted into a YUV sequence (i.e., one of the output image sequence sets).
[0216] (3) Projection transformation can be used to simulate the change of shooting angle. The standard projection transformation has 8 parameters, and four pairs of coordinates are required to determine the parameters of the projection transformation. By modeling the coordinates of the four corners of the image using random walks, a continuous projection transformation parameter matrix can be obtained. Using these continuous matrices to perform projection transformation on the image, a continuous projection transformation RGB sequence can be obtained, which can then be converted into a YUV sequence (i.e., one of the output image sequence sets).
[0217] (4) In a video image sequence, optical flow describes the displacement of the projection point of a spatial point in the imaging plane. In this scheme, optical flow is mainly used to solve the problem of local motion. The specific approach is: first divide the image into several grids and randomly generate the optical flow of the grid vertices; then, use the interpolation method to obtain the optical flow of the entire image; finally, based on the optical flow, calculate the next frame image from the current frame image to obtain the RGB sequence, and then obtain the YUV sequence (i.e., one of the output image sequence sets).
[0218] (5) Randomly combine the YUV sequence sets generated in the above four methods (1) to (4) to obtain a mixed motion YUV sequence (i.e., one of the output image sequence sets).
[0219] The second category: using high-speed cameras to capture YUV images (i.e., one of the output image sequence sets). When ordinary cameras record scenes where some objects are moving quickly, motion blur will occur, affecting the quality of the YUV sequence. Existing high-speed cameras can capture hundreds of thousands of images per second, which is enough to cope with high-speed motion scenes, such as players competing on the field, pets running, fireworks, etc. Therefore, a high-speed camera is used to focus on collecting a batch of image sequences or videos of high-speed motion scenes to improve the effect of the target network in sudden large displacement scenes.
[0220] After producing a sufficient number of high-quality YUV sequence sets, these YUV sequence sets can be used for data degradation and simulation to obtain the training data set of the target network. The steps are as follows:
[0221] (1) Collect the RAW image directly output by the sensor for noise calibration and obtain the calibrated noise model.
[0222] (2) Randomly generate a set of ISP meta parameters, such as CCM parameters, AWB parameters, etc., perform inverse transformation on these parameters, input the inverse transformed parameters and the YUV image together into the inverse ISP simulator, and degrade to obtain a clean RAW image (i.e., a noise-free RAW image). The inverse ISP simulator can be a script, software, or neural network for data degradation.
[0223] (3) Use the calibrated noise model to add noise to the clean RAW image to obtain a noisy RAW image. The noisy RAW image and the ISP meta parameters obtained in step (2) are sent to the ISP simulator to obtain a noisy RAW image of the target network input.
[0224] (4) The clean RAW image and the ISP meta parameters obtained in step (2) are input into the ISP simulator to obtain a clean YUV image (noise-free YUV image). The noise-free YUV image is the ground truth for training the JDD network.
[0225] (5) Repeat steps (2) to (4) to obtain a data set for training the target network.
[0226] The target network provided in the embodiment of the present application can be divided into five parts: a preprocessing module, a global alignment module, a local alignment module, a reliability assessment module and a reconstruction module.
[0227] The pre-processing module is used to perform CCM processing on the RAW image. In the ISP architecture, the demosaicing module transforms the image from the RAW domain to the RGB domain, and the process from the RGB domain to the YUV domain also includes color correction processing, gamma processing, and color space conversion processing.
[0228] It is found in practice that the neural network model can directly learn the gamma processing method and the color space conversion processing method from the data, but cannot learn the color correction processing. Based on this, in an embodiment of the present application, a RAW domain CCM algorithm is designed for preprocessing the RAW image.
[0229] The RAW domain CCM algorithm directly performs CCM transformation on the RAW image, and the CCM transformation matrix is calculated and output by the ISP. The RAW domain CCM method is similar to the RGB domain CCM transformation method, which also uses the CCM matrix and the values of the RGB three channels to calculate the transformed RGB three-channel values. That is, if the current pixel point of the RAW image is R or B, the two adjacent G channels are averaged to calculate the pixel value after CCM transformation; if the current pixel point of the RAW image is G, the adjacent B and R pixels are used to calculate the pixel value after CCM transformation.
[0230] The global alignment module is used to deal with global rigid displacements, such as rotation, translation, and change of shooting angle of view of the imaging device. The global motion is modeled as a projection transformation. The YUV image of the previous frame and the RAW image after CCM transformation of the current frame are input to the global motion estimation network, and the projection transformation matrix is regressed (i.e., the output projection transformation matrix). Then, the YUV image of the previous frame is globally aligned to the RAW image of the current frame using the projection transformation matrix.
[0231] The local alignment module is used to deal with local displacement, such as people walking, leaves swaying, etc. The local motion is modeled as optical flow, and the global aligned YUV image of the previous frame and the RAW image of the current frame after CCM transformation are input to the local motion estimation network to regress the optical flow image. Then, the optical flow image is used to locally align the globally aligned YUV image of the previous frame to the RAW image of the current frame.
[0232] Here, the optical flow algorithm may be a DIS optical flow algorithm.
[0233] The reliability assessment module is used to evaluate whether the global alignment image and the local alignment image are conducive to the reconstruction of the current frame RAW image. In some extreme scenarios, the associated image acquired in the previous frame may have almost no identical content with the current frame RAW image. In this case, even if the global alignment and local alignment are performed, the associated image acquired in the previous frame is not conducive to the joint denoising and demosaicing of the current frame RAW image.
[0234] The reliability assessment module is used to intercept such extreme situations, as well as situations where global alignment and local motion alignment are inaccurate, to prevent the associated image acquired in the previous frame from degrading the joint denoising and demosaicing processing effect of the current frame RAW image. The reliability assessment network inputs the YUV image (associated image) of the previous frame after global and local alignment and the RAW image after CCM transformation of the current frame, and outputs a reliability image (i.e., obtains image similarity). The weight of the reliability image is between 0 and 1.
[0235] The reconstruction module consists of a joint denoising and demosaicing network. The input of the joint denoising and demosaicing network includes the YUV image of the previous frame (associated image) after global and local alignment, the RAW image of the current frame after CCM transformation, and the reliability image, and the output is the YUV image of the current frame.
[0236] The network structure proposed in this scheme is relatively complex. In order to ensure the training efficiency and effect, the entire training process is divided into four stages.
[0237] First stage of training: Use Figure 8 The simulation data sets 1, 2, and 3 shown are used to train the global motion estimation module and freeze the weights of other modules. The training loss function is divided into two parts. The first loss sub-function calculates the difference between the projection transformation parameters regressed by the global motion estimation network and the projection transformation parameters GT; the second loss sub-function calculates the difference between the YUV image of the last frame of global alignment and the YUV image of the current frame RAW image after CCM, DM, gamma, and CST transformation.
[0238] The weighted average of the first loss sub-function and the second loss sub-function is taken as the final loss function (i.e., the first loss function) of the first stage of training. The DM module can be a DM network or a traditional DM algorithm, and the gamma and CST transformations are implemented using traditional methods.
[0239] Second stage of training: Use Figure 8 The simulation data 4, simulation data 5 and actual acquisition data shown are used to train the local motion estimation module, and the weights of other modules are frozen. The loss function is a weighted average of three parts. The third loss sub-function is the difference between the YUV image domain of the previous frame after local motion alignment and the YUV image of the current frame RAW image after CCM, DM, gamma and CST transformation; the fourth loss sub-function is the difference between the optical flow image regressed by the local motion estimation network and the true value of the optical flow image; the fifth loss sub-function is the optical flow field smoothing function, where the 2-norm of the optical flow image is taken as the loss function.
[0240] Phase 3 training: Use Figure 8The simulation data 5 and the actual acquisition data shown are used to train the reliability estimation module and the reconstruction module, and the previous global motion estimation module and the local motion estimation module are frozen. The loss function is a weighted combination of two parts: the sixth loss sub-function calculates the difference between the current frame YUV image reconstructed by the target network and the YUV GT image; the seventh loss sub-function is calculated based on the reliability weight map (image similarity), the current frame YUV image reconstructed by the target network, and the aligned previous frame YUV map (associated image), as shown in the above formula 1.
[0241] The fourth stage of training: reduce the learning rate, fine-tune the training parameters of all modules, complete the final training, and obtain the final target network.
[0242] The embodiment of the present application provides a method for producing time domain simulation data. The introduction of a high-speed camera to capture fast motion scenes makes up for the defect that ordinary imaging equipment will produce motion blur when capturing high-speed motion scenes, and ensures the quality of degraded source data. During degradation, a combination of multiple motion modes is considered, effectively simulating the motion mode in the real scene, and it is possible to obtain unlimited training data based on limited actual acquisition data, providing a sufficient number of high-quality training data sets to support the training target network.
[0243] The embodiment of the present application provides a method for color correction processing of RAW images, which makes up for the disadvantage that the data-driven neural network cannot learn CCM transformation. The method can also be used in other similar problems.
[0244] The embodiment of the present application provides a network structure and training method for joint denoising and demosaicing in the time domain. The network structure fully considers the modeling of global motion and local motion, and evaluates the reliability of the alignment of the previous and next frames to avoid degrading the reconstruction result of the current frame due to adding the previous frame information. It improves the temporal consistency of the video and avoids ghosting.
[0245] In the embodiment of the present application, the output image sequence is obtained based on high-resolution image simulation motion, which not only reduces the workload of data collection, but also improves the temporal consistency of the output image sequence.
[0246] In an embodiment of the present application, a high-speed camera is used to capture images or video sequences of fast-motion scenes for the production of a time domain data set, thereby ensuring the data quality of the fast-motion scenes.
[0247] In the embodiment of the present application, data is produced by using degradation and simulation methods, which can obtain unlimited training data based on a limited image sequence to support the training of a neural network.
[0248] In the embodiment of the present application, the proposed target network structure fully considers the modeling of global motion and local motion, and evaluates the reliability of the alignment of the previous and next frames to avoid deteriorating the reconstruction result of the current frame due to the addition of the previous frame information.
[0249] Using the image processing method provided in the embodiment of the present application in video can not only obtain better resolution than traditional de-mosaic time domain noise reduction, but also lower the threshold for time domain noise reduction debugging, avoid ghosting due to inaccurate alignment of previous and next frames, and improve time domain consistency.
[0250] The image processing method provided in the embodiment of the present application includes two parts: the preparation of a time domain training data set and the design and training of a joint denoising and demosaicing network structure. When preparing the time domain data, a dynamic output image sequence is first obtained based on the collected high-resolution photos and image sequences or videos captured by a high-speed camera, and then the output image sequence is used to perform inverse ISP degradation and ISP simulation to obtain a time domain training data set for the target network.
[0251] In the time domain network structure design and training part, the motion between the previous and next frames of the video is divided into global motion and local motion, the previous frame image is aligned to the current frame image, and then the reliability evaluation network is used to score the reliability of the aligned image. Finally, the reliability map, the aligned previous frame image (associated image) and the current frame EAW image are input into the processing module together to obtain a clean output image.
[0252] In some embodiments, 3DGS technology can be used when obtaining an output image sequence set. Based on a 2D image, 3DGS technology is used to model a 3D scene, simulate the movement of the three-dimensional world, and project a two-dimensional YUV sequence. Compared with the method of simulating movement proposed in the embodiment of the present application, the output image sequence set obtained in this way will be more realistic and more consistent with the real scene. Degradation and simulation of the output image sequence set obtained by this means can further improve the processing effect of the target network.
[0253] In some embodiments, when obtaining the output image sequence set, the AIGC technology may also be considered. Combining the human body prior model with the AIGC technology to generate a real person-related video for degradation can not only improve the performance of the person scene, but also reduce the cost of video acquisition.
[0254] In some embodiments, a network may be used to first perform deblurring on the RAW image, and then input into the target network provided in the embodiment of the present application for de-noising and de-mosaicing, which can solve the problem of motion blur in the RAW image itself.
[0255] The method proposed in the embodiment of the present application can also be extended to other time domain methods, such as time domain noise reduction in the RAW domain, time domain noise reduction in the RGB domain, etc.
[0256] In some embodiments, multiple frames of images may be input into the target network, and multiple previous and next frame images may be used to improve the temporal expression.
[0257] It should be understood that, although the steps in the above-mentioned flowcharts are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above-mentioned flowcharts may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.
[0258] Based on the foregoing embodiments, an embodiment of the present application provides an image processing device, which includes the modules included and the units included in the modules, which can be implemented by a processor; of course, it can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP) or a field programmable gate array (FPGA), etc.
[0259] Fig.15 A schematic diagram of the structure of an image processing device provided in an embodiment of the present application is shown in FIG. Fig.15 As shown, the device 1500 includes an acquisition module 1501 and a processing module 1502, wherein:
[0260] An acquisition module 1501 is used to acquire an image to be processed, where the image to be processed is a RAW image of a current frame;
[0261] The processing module 1502 is used to perform joint denoising and demosaicing on the image to be processed according to the associated image and the image to be processed to obtain a current frame output image. When the current frame RAW image is the first frame image, the associated image is a preset image. When the current frame RAW image is not the first frame image, the associated image is a previous frame output image, and the previous frame output image is an image obtained after the joint denoising and demosaicing processing.
[0262] In some embodiments, the output image is a YUV image.
[0263] In some embodiments, the processing module 1502 is specifically configured to perform joint denoising and demosaicing on the image to be processed according to the associated image, the image to be processed and a pre-trained target network to obtain a current frame output image;
[0264] Among them, the pre-trained target network includes an alignment module and a processing module. The alignment module is used to perform time domain alignment processing on the associated image and the image to be processed to obtain a target aligned image. The processing module is used to perform joint denoising and demosaicing processing on the image to be processed through the target aligned image to obtain the current frame output image.
[0265] In some embodiments, the alignment module includes a global alignment module and a local alignment module, the global alignment module is connected to the local alignment module, and the local alignment module is connected to the processing module. The processing module 1502 is specifically used to perform global alignment processing on the associated image and the image to be processed through the global alignment module to obtain a globally aligned image; to perform local alignment processing on the global aligned image and the image to be processed through the local alignment module to obtain a locally aligned image; and to perform joint denoising and demosaicing processing on the locally aligned image and the image to be processed through the processing module to obtain the current frame output image.
[0266] In some embodiments, the processing module includes a reliability assessment module and a reconstruction module, and the reliability assessment module is connected to the local alignment module. The processing module 1502 is specifically used to determine the image similarity between the local aligned image and the image to be processed through the reliability assessment module; through the reconstruction module, the image to be processed is jointly denoised and demosaiced according to the image similarity and the local aligned image to obtain the current frame output image.
[0267] In some embodiments, the pre-trained target network further includes a pre-processing module, and the pre-processing module is connected to the global alignment module;
[0268] The processing module 1502 is further specifically used to perform color correction processing on the image to be processed through the preprocessing module to obtain a corrected image to be processed; and to perform global alignment processing on the associated image and the corrected image to be processed through the global alignment module to obtain the globally aligned image.
[0269] In some embodiments, the processing module 1502 is further specifically used to determine the initial pixel value corresponding to each pixel point in the image to be processed; according to a preset color correction matrix, the initial pixel value corresponding to each pixel point is corrected to obtain the corrected image to be processed, and the pixel value of each pixel point in the corrected image to be processed is the corrected pixel value.
[0270] In some embodiments, the processing module 1502 is further specifically used to input the associated image and the corrected image to be processed into the global alignment module to obtain a projection transformation matrix; through the projection transformation matrix, the associated image and the corrected image to be processed are positionally aligned to obtain the globally aligned image.
[0271] In some embodiments, the processing module 1502 is further specifically used to input the globally aligned image and the image to be processed into the local alignment module to obtain an optical flow map; through the optical flow map, the moving area of the globally aligned image is aligned with the moving area of the image to be processed to obtain the local aligned image.
[0272] In some embodiments, the apparatus further comprises a training module;
[0273] The acquisition module 1501 is further used to acquire a time domain training data set, wherein the time domain training data set includes a plurality of image pairs, each of which includes a noisy RAW image and a noise-free output image;
[0274] The training module is used to train the target network through the time domain training data set to obtain the pre-trained target network.
[0275] In some embodiments, the acquisition module 1501 is further specifically used to acquire an output image sequence set, wherein the output image sequence set includes a plurality of continuous frame output images; and perform degradation simulation processing on the output image sequence set to obtain the time domain training data set.
[0276] In some embodiments, the output image sequence set is an image sequence acquired in a scene where the moving speed is greater than a speed threshold, or the output image sequence set is an image sequence including different motion characteristics.
[0277] In some embodiments, the acquisition module 1501 is specifically used to acquire an initial image, which is an RGB image with a resolution greater than a resolution threshold; simulate motion processing is performed on the initial image to obtain at least one set of output image sequence sets, and the simulated motion processing includes any one or more of copy processing, random cropping processing, projection transformation processing, optical flow processing and mixing processing.
[0278] In some embodiments, the acquisition module 1501 is specifically used to randomly generate image signal processing parameters, and perform inverse transformation processing on the image signal processing parameters to obtain inverse transformation parameters; input the target output image in the output image sequence set and the inverse transformation parameters to an inverse image simulation processor to obtain a noise-free RAW image, wherein the target output image is any image in the output image sequence set; according to a preset calibration noise model, the noise-free RAW image is subjected to noise addition processing to obtain the noisy RAW image; the noisy RAW image and the inverse transformation parameters are input to an image simulation processor to obtain the noise-free output image; and the above steps are repeated for each output image in the output image sequence set to obtain the time domain training data set.
[0279] In some embodiments, the training module is specifically used to train each target sub-module in the target network through the time domain training data set to obtain multiple trained target sub-modules. The target sub-modules include a global alignment module, a local alignment module and a processing module. When training any of the target sub-modules, the weights of other target sub-modules are kept unchanged.
[0280] In some embodiments, the training module is specifically used to train the global alignment module of the target network through a first training data set and a first loss function to obtain a trained global alignment module, wherein the first training data set includes an output image with global information in the output image sequence set, and the first loss function includes a first loss sub-function and a second loss sub-function, wherein the first loss sub-function is determined by the transformation parameters of the projection transformation matrix obtained after the global alignment module processes the associated image and the image to be processed and the transformation parameters of the preset projection transformation matrix, and the second loss sub-function is determined by the difference between the global alignment image and the processed image to be processed.
[0281] In some embodiments, the training module is specifically used to train the local alignment module of the target network through a second training data set and a second loss function to obtain a trained local alignment module, wherein the second training data set includes an output image with motion information in the output image sequence set, and the second loss function includes a third loss sub-function, a fourth loss sub-function and a fifth loss sub-function, wherein the third loss sub-function is determined by the difference between the local alignment image and the processed image to be processed, the fourth loss sub-function is determined by the difference between the optical flow image obtained after the associated image and the image to be processed are processed by the local alignment module and a preset optical flow image, and the fifth loss sub-function is an optical flow field smoothing function.
[0282] In some embodiments, the training module is specifically used to train the processing module of the target network through a third training data set and a third loss function to obtain a trained processing module, wherein the third training data set is the output image sequence set, and the third loss function includes a sixth loss sub-function and a seventh loss sub-function. The sixth loss sub-function is determined by the difference between a current frame output image obtained after the processing module performs joint denoising and demosaicing on the local aligned image and the image to be processed and a preset current frame output image, and the seventh loss sub-function is determined based on the image similarity, the current frame output image and the associated image.
[0283] The description of the above device embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. For technical details not disclosed in the device embodiment of the present application, please refer to the description of the method embodiment of the present application for understanding.
[0284] It should be noted that in the embodiments of this application Fig.15 The division of modules in the image processing device shown is schematic and is only a logical function division. There may be other division methods in actual implementation. In addition, each functional unit in each embodiment of the present application may be integrated into a processing unit, or may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit. It may also be implemented in the form of a combination of software and hardware.
[0285] It should be noted that in the embodiment of the present application, if the above method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application can be essentially or partly embodied in the form of a software product that contributes to the relevant technology. The computer software product is stored in a storage medium, including several instructions to enable an electronic device to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific combination of hardware and software.
[0286] The embodiment of the present application provides a computer device, which may be a server, and its internal structure diagram may be as follows: Fig.16As shown. The computer device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the above method is implemented.
[0287] An embodiment of the present application provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the method provided in the above embodiment are implemented.
[0288] An embodiment of the present application provides a computer program product including instructions, which, when executed on a computer, enables the computer to execute the steps of the method provided in the above method embodiment.
[0289] Those skilled in the art will understand that Fig.16 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0290] In one embodiment, the image processing apparatus provided by the present application can be implemented in the form of a computer program. The computer program can be Fig.16 The computer device shown in the figure is run. The memory of the computer device can store various program modules constituting the above-mentioned device. The computer program composed of various program modules enables the processor to execute the steps in the method of each embodiment of the present application described in this specification.
[0291] It should be noted here that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments, and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium, storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.
[0292] It should be understood that "one embodiment" or "an embodiment" or "some embodiments" mentioned throughout the specification means that specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in one embodiment" or "in some embodiments" appearing throughout the specification may not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the size of the sequence number of the above-mentioned processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application. The above-mentioned sequence numbers of the embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments. The above description of each embodiment tends to emphasize the differences between the various embodiments, and the same or similar aspects can be referenced to each other. For the sake of brevity, this article will not repeat them.
[0293] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there may be three relationships. For example, object A and / or object B can represent three situations: object A exists alone, object A and object B exist at the same time, and object B exists alone.
[0294] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0295] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The embodiments described above are only schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation, such as: multiple modules or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or modules can be electrical, mechanical or other forms.
[0296] The modules described above as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules; they may be located in one place or distributed on multiple network units; some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment.
[0297] In addition, all functional modules in the embodiments of the present application may be integrated into one processing unit, or each module may be a separate unit, or two or more modules may be integrated into one unit; the above-mentioned integrated modules may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0298] A person skilled in the art can understand that all or part of the steps of implementing the above method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, etc., various media that can store program codes.
[0299] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application can essentially or in other words, the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling an electronic device to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0300] The methods disclosed in several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0301] The features disclosed in several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0302] The features disclosed in several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0303] The above is only an implementation method of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. An image processing method, characterized in that: The method comprises: Acquire an image to be processed, where the image to be processed is a RAW image of the current frame; According to the associated image and the image to be processed, the image to be processed is jointly denoised and demosaiced to obtain a current frame output image. When the current frame RAW image is the first frame image, the associated image is a preset image. When the current frame RAW image is not the first frame image, the associated image is the previous frame output image. The previous frame output image is an image obtained after the joint denoising and demosaicing processing.
2. The method according to claim 1, characterized in that The output image is a YUV image.
3. The method according to claim 1 or 2, characterized in that: The step of performing denoising and demosaicing on the image to be processed according to the associated image and the image to be processed to obtain a current frame output image includes: According to the associated image, the image to be processed and a pre-trained target network, performing joint denoising and demosaicing processing on the image to be processed to obtain a current frame output image; Among them, the pre-trained target network includes an alignment module and a processing module. The alignment module is used to perform time domain alignment processing on the associated image and the image to be processed to obtain a target aligned image. The processing module is used to perform joint denoising and demosaicing processing on the image to be processed through the target aligned image to obtain the current frame output image.
4. The method according to claim 3, characterized in that The alignment module includes a global alignment module and a local alignment module, the global alignment module is connected to the local alignment module, the local alignment module is connected to the processing module, and the image to be processed is subjected to joint denoising and demosaicing processing according to the associated image, the image to be processed and the pre-trained target network to obtain a current frame output image, including: By means of the global alignment module, the associated image and the image to be processed are globally aligned to obtain a globally aligned image; By means of the local alignment module, the global alignment image and the image to be processed are locally aligned to obtain a local alignment image; The processing module performs joint denoising and demosaicing processing on the local aligned image and the image to be processed to obtain the current frame output image.
5. The method according to claim 4, characterized in that The processing module includes a reliability evaluation module and a reconstruction module, wherein the reliability evaluation module is connected to the local alignment module, and the local alignment image and the image to be processed are jointly denoised and demosaiced by the processing module to obtain the current frame output image, including: Determining the image similarity between the local aligned image and the image to be processed by the reliability evaluation module; The reconstruction module performs joint denoising and demosaicing processing on the image to be processed according to the image similarity and the local alignment image to obtain the current frame output image.
6. The method according to any one of claims 2 to 5, characterized in that: The pre-trained target network further includes a pre-processing module, the pre-processing module is connected to the global alignment module, and the global alignment processing is performed on the associated image and the image to be processed through the global alignment module to obtain a global aligned image, including: By means of the preprocessing module, color correction processing is performed on the image to be processed to obtain a corrected image to be processed; The global alignment module performs global alignment processing on the associated image and the corrected image to be processed to obtain the global aligned image.
7. The method according to claim 6, characterized in that The step of performing color correction on the image to be processed to obtain a corrected image to be processed includes: Determine an initial pixel value corresponding to each pixel in the image to be processed; According to a preset color correction matrix, the initial pixel value corresponding to each pixel point is corrected to obtain the corrected image to be processed, and the pixel value of each pixel point in the corrected image to be processed is the corrected pixel value after correction.
8. The method according to claim 6, characterized in that The globally aligning module performs global alignment processing on the associated image and the corrected image to be processed to obtain the globally aligned image, including: Inputting the associated image and the corrected image to be processed into the global alignment module to obtain a projection transformation matrix; The associated image and the corrected image to be processed are subjected to position alignment processing through the projection transformation matrix to obtain the global alignment image.
9. The method according to claim 4, characterized in that The locally aligning the global aligned image and the image to be processed is performed locally by the local alignment module to obtain a locally aligned image, including: Inputting the global alignment image and the image to be processed into the local alignment module to obtain an optical flow map; The motion region of the global alignment image is aligned with the motion region of the image to be processed by using the optical flow map to obtain the local alignment image.
10. The method according to any one of claims 1 to 9, characterized in that: Before acquiring the image to be processed, the method further includes: Acquire a time domain training data set, wherein the time domain training data set includes a plurality of image pairs, each of the image pairs includes a noisy RAW image and a noise-free output image; The target network is trained using the time domain training data set to obtain the pre-trained target network.
11. The method according to claim 10, characterized in that The step of obtaining a time domain training data set includes: Acquire an output image sequence set, wherein the output image sequence set includes a plurality of continuous frame output images; A degradation simulation process is performed on the output image sequence set to obtain the time domain training data set.
12. The method according to claim 11, characterized in that The output image sequence set is an image sequence acquired in a scene where the moving speed is greater than a speed threshold, or the output image sequence set is an image sequence including different motion characteristics.
13. The method according to claim 12, characterized in that The step of obtaining an output image sequence set comprises: Acquire an initial image, where the initial image is an RGB image with a resolution greater than a resolution threshold; The initial image is subjected to simulated motion processing to obtain at least one set of output image sequence sets, wherein the simulated motion processing includes any one or more of replication processing, random cropping processing, projection transformation processing, optical flow processing and mixing processing.
14. The method according to claim 11, characterized in that The step of performing degradation simulation processing on the output image sequence set to obtain the time domain training data set includes: Randomly generating image signal processing parameters, and performing inverse transformation processing on the image signal processing parameters to obtain inverse transformation parameters; Inputting the target output image in the output image sequence set and the inverse transformation parameters into an inverse image simulation processor to obtain a noise-free RAW image, wherein the target output image is any image in the output image sequence set; According to a preset calibration noise model, the noise-free RAW image is subjected to noise processing to obtain the noisy RAW image; Inputting the noisy RAW image and the inverse transformation parameters into an image simulation processor to obtain the noise-free output image; The above steps are repeated for each output image in the output image sequence set to obtain the time domain training data set.
15. The method according to claim 10, characterized in that The step of training the target network by using the time domain training data set to obtain the pre-trained target network includes: Through the time domain training data set, each target sub-module in the target network is trained respectively to obtain multiple trained target sub-modules, wherein the target sub-modules include a global alignment module, a local alignment module and a processing module. When any of the target sub-modules is trained, the weights of other target sub-modules are kept unchanged.
16. The method according to claim 15, characterized in that In the case where the target submodule is a global alignment module, the global alignment module is trained by using the time domain training data set to obtain a trained global alignment module, including: The global alignment module of the target network is trained by a first training data set and a first loss function to obtain a trained global alignment module, wherein the first training data set includes an output image with global information in the output image sequence set, and the first loss function includes a first loss sub-function and a second loss sub-function, wherein the first loss sub-function is determined by transformation parameters of a projection transformation matrix obtained after the associated image and the image to be processed are processed by the global alignment module and transformation parameters of a preset projection transformation matrix, and the second loss sub-function is determined by the difference between the global alignment image and the processed image to be processed.
17. The method according to claim 15, characterized in that In the case where the target submodule is a local alignment module, the local alignment module is trained by using the time domain training data set to obtain a trained local alignment module, including: The local alignment module of the target network is trained by a second training data set and a second loss function to obtain a trained local alignment module, wherein the second training data set includes an output image with motion information in the output image sequence set, and the second loss function includes a third loss sub-function, a fourth loss sub-function and a fifth loss sub-function, wherein the third loss sub-function is determined by the difference between the local alignment image and the processed image to be processed, the fourth loss sub-function is determined by the difference between an optical flow image obtained after the associated image and the image to be processed are processed by the local alignment module and a preset optical flow image, and the fifth loss sub-function is an optical flow field smoothing function.
18. The method according to claim 15, characterized in that In the case where the target submodule is a processing module, the processing module is trained using the time domain training data set to obtain a trained local alignment module, including: The processing module of the target network is trained by a third training data set and a third loss function to obtain a trained processing module, wherein the third training data set is the output image sequence set, and the third loss function includes a sixth loss sub-function and a seventh loss sub-function, wherein the sixth loss sub-function is determined by the difference between a current frame output image obtained after the processing module performs joint denoising and demosaicing on the local aligned image and the image to be processed and a preset current frame output image, and the seventh loss sub-function is determined based on the image similarity, the current frame output image and the associated image.
19. An image processing device, characterized in that: include: An acquisition module is used to acquire an image to be processed, where the image to be processed is a RAW image of the current frame; A processing module is used to perform joint denoising and demosaicing on the image to be processed according to the associated image and the image to be processed to obtain a current frame output image. When the current frame RAW image is the first frame image, the associated image is a preset image. When the current frame RAW image is not the first frame image, the associated image is a previous frame output image, and the previous frame output image is an image obtained after the joint denoising and demosaicing processing.
20. A computer device comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, wherein: When the processor executes the program, the steps of the method according to any one of claims 1 to 18 are implemented.
21. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 18 is implemented.