Image signal processing method, electronic equipment and computer readable medium
By using AI models for noise reduction and demosaic processing in ISP, combined with ISO sensitivity data and LSC data, the problem of poor coupling between AI models and ISP in the prior art is solved, and efficient data interaction and powerful adjustable controllable capabilities are achieved.
Patent Information
- Application Number
- CN202311663818.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-01
- Publication Date
- 2025-06-10
AI Technical Summary
The existing AI models are poorly coupled with image signal processing (ISP), have large data interaction overhead, and have poor adjustable controllability based on local and global.
An AI model is used to perform noise reduction (NR) and demosaic (DMC) processing in ISP, and data interaction is realized through a data transmission. The input data includes ISO sensitivity data and lens shadow correction (LSC) data to achieve global and local intensity control.
The deep fusion of AI models and ISPs is achieved, reducing data interaction overhead and improving the tunable and controllable capabilities based on local and global.
Smart Images

Figure CN120128810A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of image signal processing, and particularly to a method for image signal processing, an electronic device, and a computer-readable medium. Background Art
[0002] The raw data directly collected by an image acquisition device (such as a camera, a mobile phone, etc.) through a photosensitive device (Sensor) is the raw data of an image, that is, a Bell domain (RAW domain, or Bayer domain) image.
[0003] The quality of the RAW domain image is poor, so it needs to be processed by image signal processing (ISP, Image Signal Processing) for image quality enhancement (distortion correction, noise removal, detail enhancement, color correction, tone enhancement, etc.) before it can be used as a normal image.
[0004] Digital signal processing can be implemented through an artificial intelligence (AI, Artificial Intelligence) model. However, the existing AI models are not well coupled with ISP, resulting in large data interaction overhead and poor adjustable and controllable capabilities based on local and global aspects. Summary of the Invention
[0005] The present disclosure provides a method for image signal processing, an electronic device, and a computer-readable medium.
[0006] In a first aspect, an embodiment of the present disclosure provides a method for image signal processing, which includes noise reduction and demosaicing processing, and the noise reduction and demosaicing processing includes:
[0007] Inputting the raw data into a preset image processing model; the raw data includes a RAW domain image, ISO sensitivity data of the International Organization for Standardization, and lens shading correction (LSC) data, and the image processing model is an artificial intelligence (AI) model for simultaneously performing noise reduction processing and demosaicing processing on the RAW domain image;
[0008] Obtaining a result image output by the image processing model; the resolution of the result image is greater than the resolution of the RAW domain image.
[0009] In a second aspect, an embodiment of the present disclosure provides an electronic device, which includes a memory and a processor; the memory stores a computer program executable by the processor, and when the computer program is executed by the processor, it implements any one of the methods for image signal processing according to the embodiments of the present disclosure.
[0010] In a third aspect, an embodiment of the present disclosure provides a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements any one of the methods for image signal processing according to the embodiments of the present disclosure.
[0011] In the embodiments of the present disclosure, an AI model is used for NR and DMC in ISP. Since NR and DMC are highly complex and non-linear, they are most suitable to be performed by an AI model. Moreover, NR and DMC are performed simultaneously in one AI model, so only one data transfer is required between the chip (such as an ISP chip) and the device running the AI model (such as a GPU), and the data interaction overhead is low. At the same time, the data input to the AI model also includes ISO sensitivity data and LSC data. Through the ISO sensitivity data, the processing results at the required ISO sensitivity can be obtained to achieve global intensity control, while through the LSC data, the noise level difference caused by lens shading can be eliminated to achieve local intensity control.
[0012] Thus, the embodiments of the present disclosure achieve a deep integration of the AI model and ISP, with low data interaction overhead and strong adjustable and controllable capabilities based on both local and global aspects. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In the drawings of the embodiments of the present disclosure:
[0014] Figure 1 is a flowchart of a method for image signal processing provided by an embodiment of the present disclosure;
[0015] Figure 2 is a block diagram of the composition of an electronic device provided by an embodiment of the present disclosure;
[0016] Figure 3 is a block diagram of the composition of a computer-readable medium provided by an embodiment of the present disclosure;
[0017] Figure 4 is a flowchart of the model training process in another method for image signal processing provided by an embodiment of the present disclosure;
[0018] Figure 5 is a schematic diagram of the input and output of an image processing model in another method for image signal processing provided by an embodiment of the present disclosure;
[0019] Figure 6 is a schematic diagram of the input and output of an image processing model in another method for image signal processing provided by an embodiment of the present disclosure;
[0020] Figure 7 is a schematic diagram of the input and output of an image processing model in another method for image signal processing provided by an embodiment of the present disclosure;
[0021] Figure 8 is a schematic diagram of the input and output of an image processing model in another method for image signal processing provided by an embodiment of the present disclosure;
[0022] Figure 9Schematic diagram of the input and output of an image processing model in another image signal processing method provided by an embodiment of the present disclosure;
[0023] Figure 10 Flow chart of a method for image information processing in the related art;
[0024] Figure 11 Flow chart of another image signal processing method provided by an embodiment of the present disclosure;
[0025] Figure 12 Flow chart of another image signal processing method provided by an embodiment of the present disclosure;
[0026] Figure 13 Flow chart of another image signal processing method provided by an embodiment of the present disclosure. Detailed implementation manners
[0027] To enable those skilled in the art to better understand the technical solutions of the present disclosure, the image signal processing methods, electronic devices, and computer-readable media provided by the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0028] The present disclosure will be described more fully hereinafter with reference to the accompanying drawings. However, the illustrated embodiments may be embodied in different forms and the present disclosure should not be construed as limited to the embodiments set forth hereinafter. On the contrary, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art.
[0029] The accompanying drawings of the embodiments of the present disclosure are used to provide a further understanding of the embodiments of the present disclosure, and constitute a part of the specification, and are used to explain the present disclosure together with the detailed embodiments, and do not constitute a limitation to the present disclosure. By describing the detailed embodiments with reference to the accompanying drawings, the above and other features and advantages will become more obvious to those skilled in the art.
[0030] The present disclosure may be described with reference to the plan views and / or cross-sectional views by means of the ideal schematic diagrams of the present disclosure. Therefore, the example illustrations may be modified according to the manufacturing technology and / or tolerances.
[0031] Without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other.
[0032] The terms used in this disclosure are only for describing specific embodiments and are not intended to limit this disclosure. As used in this disclosure, the term "and / or" includes any and all combinations of one or more of the associated listed items. As used in this disclosure, the singular forms "a" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. As used in this disclosure, the terms "comprising", "made of", specify the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their groups.
[0033] Unless otherwise defined, all terms (including technical and scientific terms) used in this disclosure have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in common dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning unless this disclosure clearly so defines.
[0034] This disclosure is not limited to the embodiments shown in the drawings, but includes modifications to the configurations formed based on the manufacturing process. Therefore, the regions illustrated in the drawings have schematic properties, and the shapes of the regions shown in the figures illustrate the specific shapes of the regions of the elements, but are not intended to be restrictive.
[0035] The raw data of an image directly collected by an image acquisition device (such as a camera, a mobile phone, etc.) through a photosensitive device is the raw image data, that is, the RAW domain image, and the RAW domain image needs to be subjected to ISP image quality enhancement.
[0036] ISP can be carried out in a chip (such as an ISP chip) in the form of an ISP pipeline (ISP-pipeline), referring to Figure 10 , the ISP pipeline includes a plurality of sequentially performed processes. The specific processes include but are not limited to high dynamic range rendering (HDR, High Dynamic Range) processing, noise reduction (NR, Noise Reduction) processing, black level correction (BL, Black Level Correction) processing, digital gain (Dgain, Digital Gain) processing, lens shading correction (LSC, Lens Shading Correction) processing, auto white balance correction (AWB, Auto White Balance) processing, demosaicing (DMC) processing, etc.
[0037] With the enhancement of the computing power of professional processing devices such as Graphics Processing Units (GPUs) and Neural Processing Units (NPUs), and the development of AI technology, image processing using AI models is no longer limited to high-level tasks such as tracking, detection, recognition, and classification, but can also be used for low-level tasks in ISP, etc.
[0038] However, in some related technologies, there are many problems in using AI models for ISP.
[0039] For example, ISP is performed in a chip (such as an ISP chip), but the AI model needs to run on a dedicated device such as a GPU. Therefore, when using an AI model for ISP, a large amount of data needs to be transmitted between the chip and the GPU, resulting in a large data interaction overhead.
[0040] Again, the AI model has the property of being a black box and is difficult to directly control, resulting in poor adjustable and controllable capabilities based on its local and global aspects.
[0041] Therefore, in related technologies, the coupling between the AI model and ISP is poor, the data interaction overhead is large, and the adjustable and controllable capabilities based on local and global aspects are poor.
[0042] In a first aspect, embodiments of the present disclosure provide a method for Image Signal Processing (ISP).
[0043] Embodiments of the present disclosure are used to perform ISP on the original RAW domain image collected by an image acquisition device to enhance its image quality and obtain a directly usable normal image.
[0044] The ISP method of the embodiments of the present disclosure includes noise reduction and demosaicing processing; referring to Figure 1 , the above noise reduction and demosaicing processing includes:
[0045] S101: Input the original data into a preset image processing model.
[0046] Among them, the original data includes the RAW domain image, International Organization for Standardization (ISO) sensitivity data, and LSC data, and the image processing model is an AI model used to simultaneously perform noise reduction processing and demosaicing processing on the RAW domain image.
[0047] S102: Obtain the result image output by the image processing model.
[0048] Among them, the resolution of the result image is greater than the resolution of the RAW domain image.
[0049] The ISP in the embodiments of the present disclosure includes noise reduction (NR) and demosaicing (DMC) processes, and both processes are implemented by an AI model, that is, the noise reduction and demosaicing (AI NRDMC) process.
[0050] In the noise reduction and demosaicing process, the original data including the RAW domain image is first input into a pre-set image processing model, and the image processing model will automatically output the result image, that is, the image obtained after performing NR and DMC on the RAW domain image.
[0051] Among them, the RAW domain image is the original light intensity data directly sensed by the photosensitive device of the camera, and the photosensitive device can be divided into different colors. For example, the photosensitive device can adopt the color filter array (CFA, Color Filter Array) mode, such as the RGB (red-green-blue) mode, and its variants RGGB, BGGR, GRBG, GBRG, etc., or any mode that can represent image color information such as RCCB (red-transmissive-transmissive blue), RYYCY (red-yellow-yellow-transmissive-yellow), etc.
[0052] Among them, the DMC process restores the "mosaic" to multiple pixels, so it will improve the resolution of the image, and thus the resolution of the result image must be higher than that of the RAW domain image.
[0053] For example, referring to Figure 5 , the dimension of the RAW domain image input to the avatar processing model can be h*w*1, where h represents the height, w represents the width, and 1 represents a single channel, that is, the size of the RAW domain image is h*w, and the number of channels is 1. The colors of the pixels at different positions in the RAW domain image are different; while the dimension of the result image generated by the image processing model can be h*w*3, where the number of channels is 3, such as the channels of red, green, and blue respectively. Thus, the result image is an RGB domain image (RGB domain), and the resolution is 3 times that of the RAW domain image.
[0054] In the embodiments of the present disclosure, in addition to the RAW domain image, the original data input to the image processing model also includes ISO sensitivity data and LSC data, that is, the ISO sensitivity data and LSC data are also introduced as reference factors in the processing of the image processing model.
[0055] Among them, the ISO sensitivity data is defined according to the ISO standard and represents the data of the sensitivity of the camera (such as the photosensitive device or film) to light. The ISO sensitivity data of the entire image is usually the same value; thus, the ISO sensitivity data is equivalent to the global intensity indication of the image processing model.
[0056] Among them, since the camera lens is not an ideal device, it will cause the brightness (such as gray level) at different positions of the image to be different under uniformly incident light (usually the brightness is the highest at the center of the image and gradually decreases towards the periphery), that is, there is lens shading. Therefore, different degrees of compensation gain are required for different positions of the image, that is, LSC is performed. When the lens shading is different, the noise level caused by the shading in the image is also different. Therefore, the NR process also needs to be based on the LSC data (that is, the data used for LSC compensation, such as the data in the lookup table LUT), so that the LSC data can also be input into the image processing model to indicate the noise conditions of different local parts of the image.
[0057] In the embodiments of the present disclosure, an AI model is used for NR and DMC in ISP. Since NR and DMC are highly complex and non-linear, they are most suitable to be performed by an AI model. Moreover, NR and DMC are performed simultaneously in one AI model, so only one data transfer is required between the chip (such as an ISP chip) and the device running the AI model (such as a GPU), and the data interaction overhead is low. At the same time, the data input into the AI model also includes ISO sensitivity data and LSC data. Through the ISO sensitivity data, the processing results at the required ISO sensitivity can be obtained to achieve global intensity control, and through the LSC data, the noise level difference caused by lens shading can be eliminated to achieve local intensity control.
[0058] Thus, the embodiments of the present disclosure achieve a deep integration of the AI model and ISP, with small data interaction overhead and strong adjustable and controllable capabilities based on both local and global aspects.
[0059] In some embodiments, the raw data further includes: position encoding (PE, Position Encoding) data; the PE data represents the intensity information of the positions of the pixels in the RAW domain image.
[0060] As a way of the embodiments of the present disclosure, the raw data input into the image processing model may further include PE data, which is obtained according to the spatial position (such as two-dimensional coordinates) of each pixel, so that the image processing model can perceive the spatial position of the currently processed pixel or image patch, and then perform different NR and DMC according to the differences at different spatial positions (such as the noise level difference caused by lens shading) to achieve better processing effects.
[0061] In some embodiments, the PE data is calculated from the RAW domain image according to a preset encoding algorithm;
[0062] Or,
[0063] The PE data is preset.
[0064] As a mode of the embodiments of the present disclosure, it may be to calculate the applicable PE data according to the RAW domain image to be processed through a predetermined encoding algorithm, and then input it together with the RAW domain image into an image processing model for processing.
[0065] For example, the above encoding algorithm may be an AI model obtained by pre-training or an algorithm set by humans.
[0066] Alternatively, as another mode of the embodiments of the present disclosure, the PE data may also be pre-determined fixed data, such as preferred data determined according to experience.
[0067] In some embodiments, the image processing model includes at least one of the following:
[0068] U-Net model;
[0069] Transformer-based model.
[0070] As a mode of the embodiments of the present disclosure, the specific form of the image processing model adopted may specifically be in the form of a U-Net model or a Transformer-based model.
[0071] Among them, the U-Net model is a type of convolutional neural network (CNN, Convolutional Neural Networks). Referring to Figure 6 , the U-Net model includes a plurality of sequentially connected downsampling blocks (Blocks), and after the lowest-level downsampling block, a plurality of sequentially connected upsampling blocks are connected, and the output of each downsampling block is also connected to the upsampling block of the same level.
[0072] The Transformer-based model is a model established according to the attention mechanism and adopting an encoding-decoding structure, and it can also use CNN as the network backbone.
[0073] It should be understood that the specific form of the image processing model used in the embodiments of the present disclosure is not limited to the above examples, and it may also be other CNN-based models, or other models with an encoding-decoding structure, or any other AI model that can implement NR and DMC.
[0074] It should be understood that the original data should all be input into the image processing model in a manner matching the form of the image processing model.
[0075] For example, for the U-net model, different information in the original data can be processed into images of the same size, and then the images of different information are concatenated (contact) at the channel scale as a multi-channel input image with the same size, and this input image is then actually input into the U-net model.
[0076] For example, referring to Figure 6 , the dimension of the RAW domain image can be h*w*1; the PE data can be processed into an image (PE Map) with a dimension of h*w*2, where the number of channels is 2 because the spatial position needs to be determined by 2 coordinate values (such as row coordinate, column coordinate); the ISO sensitivity data and LSC data can be processed into images (ISO Map and LSC Map) with a dimension of h*w*1 respectively. All pixel values in the ISO Map are the same ISO sensitivity value, and each pixel value in the LSC Map is the compensation value at the corresponding position in the lookup table (LUT). Thus, the above original data can be concatenated into an input image with a dimension of h*w*5 and then input into the U-net model, while the dimension of the resulting image is still h*w*3.
[0077] Again, for example, referring to Figure 7 , the RAW domain image can be split into multiple channels according to color. For example, the RAW domain image in the RGGB color mode is split into 4 channels, that is, the dimension of the RAW domain image becomes h / 2*w / 2*4. Correspondingly, the dimensions of the PE Map, ISO Map, and LSC Map become h / 2*w / 2*2, h / 2*w / 2*1, and h / 2*w / 2*1 respectively. After their concatenation, an input image with a dimension of h / 2*w / 2*8 is obtained.
[0078] Again, for example, referring to Figure 8 , other information (PE Map, ISO Map, LSC Map) in the original data except the RAW domain image is also input into each sampling block (Block) of the U-net model respectively. It should be understood that the size of other data input into each sampling block should match the requirements of that sampling block. For example, different downsamplings can be performed on the PE Map, ISO Map, and LSC Map to reduce their sizes.
[0079] Again, for example, referring to Figure 9 , when the RAW domain image is multiple consecutive frames in a video, it can be that the information (Multi-frame info) of multiple frames of RAW domain images is input into the image processing model, which can be continuous input of multiple frames of RAW domain images, or the concatenation of multiple frames of RAW domain images as an image input.
[0080] In some embodiments, the method of the present disclosure embodiments further includes at least one of the following processes: HDR process; BLC process; Dgain process; LSC process; AWB process.
[0081] In some embodiments, the noise reduction and demosaicing process is performed after the HDR process; after the noise reduction and demosaicing process, it further includes at least one of the following processes: BLC process; Dgain process; AWB process; LSC process.
[0082] Referring to Figure 11 , in the IPS process of the present disclosure embodiments, in addition to the above noise reduction and demosaicing (AI NRDMC) process performed by the AI model, other processing steps may also be included, such as HDR process, BLC process, Dgain process, LSC process, AWB process, etc.
[0083] It should be understood that the above LSC process is a process for eliminating lens shadows in an image, which requires the use of data in a lookup table (LUT) (LSC data); although the image processing model also uses LSC data, its role is to perform better NR processing and DMC processing based on the LSC data, rather than directly performing the LSC process in the image processing model.
[0084] Further, referring to Figure 11 , the noise reduction and demosaicing process can be after the HDR process and before other processing steps, that is, in the ISP pipeline, the noise reduction and demosaicing process can be directly performed after the HDR process, and then other processing is performed.
[0085] It should be understood that the ISP process of the present disclosure embodiments is not limited to the above examples.
[0086] For example, referring to Figure 12 , in the ISP process, the noise reduction and demosaicing process can also be located at other positions.
[0087] Again, for example, referring to Figure 13 , the BLC process and the AWB process are also embedded in the image processing model as a "preprocessing module" (the results of the BLC process and the AWB process are both represented by signed numbers).
[0088] Again, for example, it can also be that only a part of the above-listed processes are included in the ISP process (the noise reduction and demosaicing process must be present).
[0089] Again, for example, it can also be that other processes not listed above are included in the ISP process, such as dead pixel removal process, tone mapping (TM) process, etc.
[0090] In some embodiments, referring to Figure 4, before denoising and demosaicing processing, it further includes training an image processing model; training the image processing model includes:
[0091] A101. Determine training sample pairs.
[0092] Among them, the training sample pair includes corresponding training original data and training standard images. The training original data includes training RAW domain images, training ISO sensitivity data, and training LSC data.
[0093] A102. Input the training original data into the image processing model to obtain the training result image output by the image processing model.
[0094] A103. Determine the loss function according to the training standard image and the training result image, and adjust the image processing model according to the loss function.
[0095] As a way of the embodiment of the present disclosure, it can be pre-
[0096] Train the image processing model to improve the performance of the image processing model.
[0097] Among them, each training uses a training sample pair. The training sample pair includes the training original data for input into the image processing model, and the "ideal result (training standard image)" that should theoretically be generated according to the training original data when the performance of the image processing model is good.
[0098] After actually inputting the training original data into the image processing model, the image processing model will generate a training result image. And according to the training result image and the training standard image, the loss function (Loss) can be determined. This loss function represents the gap between the current actual processing result of the image processing model and the ideal result. Therefore, through optimization algorithms such as the gradient descent method, the parameters (such as the element values of the convolution kernel, the weight values of the connections, etc.) and the structure of the image processing model can be adjusted in the direction of optimizing the loss function to optimize the image processing model.
[0099] It should be understood that the training of the image processing model can be carried out multiple times through different training sample pairs until the training end condition is met. Among them, the specific form of the training end condition is diverse. For example, it can be that the image processing model reaches a certain performance (such as signal-to-noise ratio PSNR, maximum similarity SSIM) threshold, or the performance of the image processing model converges, or a certain number of training times is reached, etc.
[0100] In some embodiments, referring to Figure 4 , determining the training sample pair (A101) includes:
[0101] A1011. Obtain a high-definition and low-noise image, determine the high-definition and low-noise image as the training standard image, and determine the training ISO sensitivity data and training LSC data.
[0102] A1012. Add noise to the high-definition and low-noise image according to the noise function corresponding to the training ISO sensitivity data to obtain a high-definition image.
[0103] A1013. Downsample the high-definition image to obtain the training RAW domain image.
[0104] As a way of the embodiments of the present disclosure, it may be directly obtaining a high-definition (resolution), low-noise image (high-definition and low-noise image) as the training standard image, and setting the training ISO sensitivity data and training LSC data required in the training.
[0105] Meanwhile, obtain a noise function, which characterizes the noise that should be in the image when collecting images with the target image acquisition device (i.e., the device that actually collects the RAW domain image) at different ISO sensitivities; thus, according to the selected training ISO sensitivity data, the corresponding noise can be "added" to the high-definition and low-noise image through the noise function to obtain an image with noise (high-definition image), and the noise in the high-definition image is the noise that should be in the RAW domain image directly collected by the target image acquisition device under the training ISO sensitivity data.
[0106] Continue to downsample (reduce the resolution) the high-definition image, then an image with a normal noise level and a lower resolution (training RAW domain image) can be obtained, which is equivalent to the RAW domain image directly collected by the target image acquisition device. Therefore, it corresponds to the high-definition and low-noise image. That is, ideally, the image processing model should be able to generate a high-definition and low-noise image based on the training RAW domain image.
[0107] In some embodiments, adding noise to the high-definition and low-noise image according to the noise function corresponding to the training ISO sensitivity data to obtain a high-definition image (A1012) includes:
[0108] A10121. Determine the noise compensation at each position of the high-definition and low-noise image according to the training LSC data.
[0109] A10122. Add noise to the high-definition and low-noise image according to the noise function and the noise compensation to obtain a high-definition image.
[0110] As a way of the embodiments of the present disclosure, when adding noise to the high-definition and low-noise image, the noise at each position of the image can be further compensated according to the set LSC data (training LSC data) so that the finally obtained noise is more in line with the actual noise when affected by the lens shadow.
[0111] Example:
[0112] Next, an exemplary introduction to a specific ISP method of the present disclosure embodiment will be given.
[0113] The ISP method of this example may include two parts: model (image processing model) training and model inference (model usage).
[0114] Among them, the model training part may include the following steps:
[0115] (1) Construct a dataset of RGB images (RGB domain) with multi-scenes, high definition, and low noise (a collection of multiple high-definition and low-noise images)
[0116] Among them, an RGB image dataset can be collected using a high-resolution, high-definition, and low-noise image acquisition device.
[0117] Furthermore, multiple images of the same scene can be collected, and after aligning the multiple images, they can be fused into one image to improve the clarity of the image and reduce the noise level therein.
[0118] Among them, the RGB image dataset can contain images under various illuminance, color temperature, and detailed scenes, such as high-noon high-brightness scenes, dusk medium-brightness scenes, night low-illuminance scenes, high-texture scenes, low-texture flat scenes, and other various scenes.
[0119] (2) Calibrate the noise parameters of the target image acquisition device at each ISO sensitivity
[0120] Among them, the target image acquisition device is used to collect images at different ISO sensitivities, and a professional device is used to calibrate the noise parameters (noise profile) therein.
[0121] For example, it can be calibrated that the noise of the target image acquisition device at a certain ISO sensitivity follows a Poisson-Gaussian distribution, that is:
[0122]
[0123] Among them, x noisy represents the image containing noise, x * represents the expected value of the noise-free image, Possion(.) and Gaussian(.) respectively represent the Poisson distribution and the Gaussian distribution, so (k, σ 2 ) represents the noise coefficient at this ISO sensitivity.
[0124] (3) Construct a noisy image set (a collection of multiple high-definition images)
[0125] As before, the noise coefficient at a certain ISO sensitivity is (k, σ 2), so that multiple different noise intensity levels can be further set, such as there are 3 noise intensity levels of low, mid, and high, and the corresponding 3 noise variances var low , var m id , var hi gh are respectively:
[0126] var low = α low,0 ·k·x intensity + α low,1 ·σ 2 ;
[0127] var mid = α mid,0 ·k·x intensity + α mid,1 ·σ 2 ;
[0128] var high = α high,0 ·k·x intensity + α high,1 ·σ 2 ;
[0129] Among them, α low,0 , α mid,0 , α high,0 , α low,1 , α mid,1 , α high,1 respectively represent the intensity coefficients that can be set under each noise intensity level. For example, [α low,0 , α mid,0 , α high,0 =[0.9, 1.0, 1.1], [α low,1 , α mid,1 , α high,1 =[0.9, 1.0, 1.1]; x intensity represents the value of a pixel (such as grayscale). For example, for an 8-bit image, x intensity =[0, 255].
[0130] Thus, the above is equivalent to determining the noise function.
[0131] Furthermore, the LSC data of the target image acquisition device is a known look-up table (LUT). The resolution of this look-up table is less than the resolution of the image. Therefore, methods such as bi-linear interpolation and bi-cubic interpolation can be selected to complete the data therein to obtain the training LSC data. Among them, each value of the LSC data represents the intensity of the LSC data used for the image at the corresponding spatial position index (i, j).
[0132] For example, the intensity of the LSC data defining the center position of the image is lsc base , then the noise intensity compensation (noise compensation determined according to the training LSC data) at different spatial positions of the image can be obtained as follows:
[0133] var low,lsc = var low ·(lsc i,j / lsc base );
[0134] var mid,lsc = var mid ·(lsc i,j / lsc base );
[0135] var high,lsc = var high ·(lsc i,j / lsc base );
[0136] Furthermore, noise can be added to the RGB image dataset based on the above noise variance and noise intensity compensation, that is:
[0137] y low,intensity = x low,intensity + Gaussian(0, var low,lsc );
[0138] y mid,intensity = x mid,intensity + Gaussian(0, var mid,lsc );
[0139] y high,intensity = x high,intensity + Gaussian(0, var high,lsc );
[0140] Among them, y low,intensity , y mid,intensity , y high,intensity respectively represent the noise-added images (high-definition images) obtained at a certain noise intensity level; Gaussian(0, var) represents a Gaussian distribution with a mean of 0 and a variance of var.
[0141] Thus, for the same image, the noise intensity and variance of the noise-added images obtained at different noise intensity levels are different. And for the same noise-added image, the noise intensity and variance at different positions are also different.
[0142] (4) Downsample the noise-added image slices to the RAW domain to obtain training RAW domain images and form training sample pairs
[0143] Downsample the noisy RGB image data to make it enter the RAW domain, that is:
[0144] y raw,mid,intensity = DS(y low,intensity );
[0145] y raw,mid,intensity = DS(y mid,intensity );
[0146] y raw,high,intensity = DS(y high,intensity );
[0147] where DS(.) represents the downsampling function, and y raw,low,intensity , y raw,mid,intensity , y raw,high,intensity respectively represent the training RAW domain images obtained after downsampling the noisy images for different noise intensity levels. The directly acquired images x intensity corresponding to them are used as the training standard images to obtain training sample pairs. For example, for the low level, there is a training sample pair [y raw,low,intensity , x intensity .
[0148] (5) Train the U-net model or Transformer-based model (image processing model)
[0149] Use the training sample pairs obtained above to train by: inputting the training raw data into the image processing model - determining the loss function based on the obtained training result image and the training standard image - adjusting the image processing model according to the loss function until the above training end condition is met; save the U-net model or Transformer-based model at this time as the image processing model for subsequent use.
[0150] For the PE data (PE map) to be used subsequently, it can also be pre-computed at this time. For example, it is mapped according to the spatial positions [i, j] of each pixel in the RAW domain image through the function (encoding algorithm) func([i, j]).
[0151] where the function func(.) can be a pre-set function or an AI model obtained through training.
[0152] And the obtained PE map can be at the block level, that is, multiple adjacent pixels belonging to a block in the PE map have the same value.
[0153] Among them, the model inference part may include the following steps:
[0154] (6) Obtain the input of the image processing model
[0155] Use the RAW domain image collected by the target image acquisition device or the RAW domain image processed by HDR as the input of the image processing model, where the pixel at the position index [i, j] is defined as p i,j , where i = 0, 1,..., h, j = 0, 1,..., w, and h and w represent the height and width of the image respectively.
[0156] (7) Perform BLC and AWB preprocessing
[0157] Perform BLC processing and AWB processing respectively through the preprocessing unit of the image processing model.
[0158] Among them, for pixel p i,j , the BLC is as follows:
[0159] p blc,i,j = p i,j - blc offset ;
[0160] Among them, p blc,i,j is represented by a signed number, and the blc offset compensation value can compensate different values for different color channels.
[0161] In the AWB processing, the AWB compensation coefficient can be defined as [gain R , gain Gr , gain Gb , gain B , then there is:
[0162]
[0163] Among them, p awb,i,j is represented by a signed number, and pixel ∈ represents the position where pixel belongs to the corresponding color of the CFA mode.
[0164] As before, the data after BLC processing and AWB processing can be retained as signed numbers to improve the processing ability of NR under low illumination.
[0165] (8) The image processing model performs NR and DMC
[0166] Define the function of the image processing model as forward(.), and the output (result image) is model output , then there is:
[0167] model output = clip(forward(P awb,ISO,LSC,PE),0,wl aiNrDmc );
[0168] Where P awb represents a patch composed of p awb,i,j . ISO, LSC, and PE represent ISO Map, LSCMap, and PE Map respectively, where the boldface words represent images (or matrices), clip(.) represents a truncation function, and wl aiNrDmc represents the maximum value allowed for the output of the image processing model.
[0169] (9) Perform Dgain processing and LSC processing
[0170] Continue to perform Dgain processing and LSC processing on the resulting image. The gain for Dgain processing is:
[0171] gain digital ;
[0172] Therefore, the Dgain processing for a 3-channel RGB image can be expressed as:
[0173] p dgain,i,j = clip(model output,i,j × gain digital , 0, wl dgain );
[0174] where model output,i,j represents the value of the model output (resulting image) at the [i, j] position, and wl dgain represents the maximum output value allowed by the Dgain module.
[0175] The gain for LSC processing is:
[0176] [gain lsc,R,i,j , gain lsc,Gr,i,j , gain lsc,Gb,i,j , gain lsc,B,i,j ;
[0177] where gain lsc,.,i,j represents the LSC gain value corresponding to the color pixel at the [i, j] spatial position;
[0178] Therefore, the LSC processing for a 3-channel RGB image can be expressed as:
[0179]
[0180] where p lsc,R,i,j , p lsc,G,i,j , p lsc,B,i,jrespectively represent the values of the corresponding color channels after LSC processing, wl lsc represents the maximum output value allowed by LSC processing.
[0181] Thus, p lsc,R,i,j 、p lsc,G,i,j 、p lsc,B,i,j are the final processing results of the IPS.
[0182] In a second aspect, referring to Figure 2 , an embodiment of the present disclosure provides an electronic device, which includes a memory and a processor; the memory stores a computer program executable by the processor, and when the computer program is executed by the processor, it implements any one of the image signal processing methods of the embodiments of the present disclosure.
[0183] In a third aspect, referring to Figure 3 , an embodiment of the present disclosure provides a computer-readable medium, on which a computer program is stored, and when the computer program is executed by the processor, it implements any one of the image signal processing methods of the embodiments of the present disclosure.
[0184] Among them, the processor is a device with data processing capabilities, which includes but is not limited to a central processing unit (CPU), etc.; the memory is a device with data storage capabilities, which includes but is not limited to a random access memory (RAM, more specifically such as SDRAM, DDR, etc.), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory (FLASH); the I / O interface (read / write interface) is connected between the processor and the memory and can realize the information interaction between the memory and the processor, and it includes but is not limited to a data bus (Bus), etc.
[0185] Those of ordinary skill in the art can understand that all or some of the steps, systems, and functional modules / units in the devices disclosed above can be implemented as software, firmware, hardware, and their appropriate combinations.
[0186] In a hardware implementation, the division between the functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component can have multiple functions, or a function or step can be executed by several physical components in cooperation.
[0187] Some physical components or all physical components may be implemented as software executed by a processor, such as a central processing unit (CPU), a digital signal processor or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-temporary medium) and a communication medium (or temporary medium). As known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules or other data). Computer storage media include, but are not limited to, random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory (FLASH) or other disk storage; read-only compact disk (CD-ROM), digital versatile disk (DVD) or other optical disk storage; magnetic cassettes, magnetic tapes, disk storage or other magnetic storage; any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically embodies computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0188] The present disclosure has disclosed example embodiments, and although specific terms are employed, they are used and should be interpreted only in a general illustrative sense and not for limiting purposes. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly stated, features, characteristics, and / or elements described in conjunction with a particular embodiment may be used alone or in combination with features, characteristics, and / or elements described in conjunction with other embodiments. Therefore, those skilled in the art will appreciate that various changes in form and detail may be made without departing from the scope of the present disclosure as set forth in the appended claims.
Claims
1. A method for image signal processing, wherein, it includes noise reduction and demosaicing processing, and the noise reduction and demosaicing processing includes: inputting the original data into a preset image processing model; the original data includes a Bayer RAW domain image, ISO sensitivity data of the International Organization for Standardization, and lens shading correction (LSC) data, and the image processing model is an artificial intelligence (AI) model for simultaneously performing noise reduction processing and demosaicing processing on the RAW domain image; obtaining the result image output by the image processing model; the resolution of the result image is greater than the resolution of the RAW domain image.
2. The method according to claim 1, wherein, the original data further includes: position encoding (PE) data; the PE data represents the intensity information of the positions of the pixels in the RAW domain image.
3. The method according to claim 2, wherein, the PE data is calculated from the RAW domain image according to a preset encoding algorithm; or, the PE data is preset.
4. The method according to claim 1, wherein, the image processing model includes at least one of the following: U-Net model; Transformer-based model.
5. The method according to claim 1, wherein, it further includes at least one of the following processes: high dynamic range (HDR) rendering processing; black level correction (BLC) processing; digital gain (Dgain) processing; LSC processing; auto white balance (AWB) correction processing.
6. The method according to claim 1, wherein, it further includes HDR processing; the noise reduction and demosaicing processing is performed after the HDR processing; after the noise reduction and demosaicing processing, it further includes at least one of the following processes: BLC processing; Dgain processing; AWB processing; LSC processing.
7. The method according to claim 1, wherein, before the noise reduction and demosaicing processing, it further includes training the image processing model; the training of the image processing model includes: determining a training sample pair; the training sample pair includes corresponding training original data and a training standard image, and the training original data includes a training RAW domain image, training ISO sensitivity data, and training LSC data; inputting the training original data into the image processing model to obtain the training result image output by the image processing model; determining a loss function according to the training standard image and the training result image, and adjusting the image processing model according to the loss function.
8. The method according to claim 7, wherein, the determining of the training sample pair includes: obtaining a high-definition and low-noise image, determining the high-definition and low-noise image as the training standard image, and determining the training ISO sensitivity data and the training LSC data; adding noise to the high-definition and low-noise image according to the noise function corresponding to the training ISO sensitivity data to obtain a high-definition image; downsampling the high-definition image to obtain the training RAW domain image.
9. The method according to claim 8, wherein, Adding noise to the high-definition low-noise image according to the noise function corresponding to the trained ISO sensitivity data to obtain a high-definition image includes: Determining the noise compensation at each position of the high-definition low-noise image according to the trained LSC data; Adding noise to the high-definition low-noise image according to the noise function and the noise compensation to obtain the high-definition image.
10. An electronic device, comprising a memory and a processor; the memory stores a computer program executable by the processor, and when the computer program is executed by the processor, it implements the method for image signal processing according to any one of claims 1 to 9.
11. A computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the method for image signal processing according to any one of claims 1 to 9.