Image signal processing method, electronic device and computer-readable storage medium
By simultaneously performing noise reduction and demosaic processing in the image processing model, the problem of poor coupling between the AI model and the ISP is solved, and efficient data interaction and powerful adjustable controllable capabilities are achieved.
Patent Information
- Application Number
- PCT/CN2024/123629
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-01
- Filing Date
- 2024-10-09
- Publication Date
- 2025-06-05
AI Technical Summary
The poor coupling of existing AI models with image signal processing (ISP) results in large data interaction overhead and poor adjustable controllability based on local and global.
An image processing model is adopted, which can perform noise reduction and demosaic processing simultaneously. By inputting the original data into the preset AI model for processing, the resolution of the output result image is greater than the resolution of the original RAW domain image.
The deep integration of AI models and ISPs is achieved, reducing data interaction overhead, and enhancing the tunable and controllable capabilities based on local and global.
Smart Images

Figure CN2024123629_05062025_PF_FP_ABST
Abstract
Description
Image signal processing method, electronic device, and computer-readable storage medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to Chinese patent application No. 202311663818.7 filed on December 1, 2023, the contents of which are incorporated herein by reference in their entirety. Technical Field
[0003] The present disclosure relates to the technical field of image signal processing, and in particular to an image signal processing method, an electronic device, and a computer-readable storage medium. Background Art
[0004] Image acquisition devices (such as cameras, mobile phones, etc.) directly capture the original image data through the photosensitive device (Sensor), that is, the Bayer domain (RAW domain, or Bayer domain) image.
[0005] The quality of RAW domain images is poor, so they need to undergo image signal processing (ISP) to enhance the image quality (distortion correction, noise removal, detail enhancement, color correction, tone enhancement, etc.) before they can be used as normal images.
[0006] Digital signal processing can be implemented through artificial intelligence (AI) models, but existing AI models are poorly coupled with ISPs, have high data interaction overhead, and have poor local and global adjustability and controllability.
[0007] Summary of the Invention
[0008] In a first aspect, an embodiment of the present disclosure provides a method for image signal processing, which includes noise reduction and demosaicing processing, wherein the noise reduction and demosaicing processing includes: inputting raw data into a preset image processing model; the raw data includes a RAW domain image, International Standards Organization (ISO) sensitivity data, and Lens Shading Correction (LSC) data, and the image processing model is an artificial intelligence (AI) model for simultaneously performing noise reduction and demosaicing on the RAW domain image; obtaining a result image output by the image processing model; the resolution of the result image is greater than the resolution of the RAW domain image.
[0009] In a second aspect, an embodiment of the present disclosure provides an electronic device comprising a memory and a processor; the memory stores a computer program that can be executed by the processor, and the computer program is executed by the processor, so that the processor implements any one of the image signal processing methods of the embodiments of the present disclosure.
[0010] In a third aspect, an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor so that the processor implements any one of the image signal processing methods of the embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] FIG1 is a flowchart of a method for image signal processing provided by an embodiment of the present disclosure;
[0012] FIG2 is a block diagram of an electronic device according to an embodiment of the present disclosure;
[0013] FIG3 is a block diagram of a computer-readable medium according to an embodiment of the present disclosure;
[0014] FIG4 is a flowchart of a model training process in another image signal processing method provided by an embodiment of the present disclosure;
[0015] FIG5 is a schematic diagram of input and output of an image processing model in another image signal processing method provided by an embodiment of the present disclosure;
[0016] FIG6 is a schematic diagram of input and output of an image processing model in another image signal processing method provided by an embodiment of the present disclosure;
[0017] FIG7 is a schematic diagram of input and output of an image processing model in another image signal processing method provided by an embodiment of the present disclosure;
[0018] FIG8 is a schematic diagram of input and output of an image processing model in another image signal processing method provided by an embodiment of the present disclosure;
[0019] FIG9 is a schematic diagram of input and output of an image processing model in another image signal processing method provided by an embodiment of the present disclosure;
[0020] FIG10 is a flow chart of a method for image information processing in the related art;
[0021] FIG11 is a flow chart of another method for image signal processing provided by an embodiment of the present disclosure;
[0022] FIG12 is a flow chart of another method for image signal processing provided by an embodiment of the present disclosure; and
[0023] FIG13 is a flow chart of another method for image signal processing provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0024] To enable those skilled in the art to better understand the technical solutions of the present disclosure, the image signal processing method, electronic device, and computer-readable medium provided by the embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.
[0025] The present disclosure will be described more fully hereinafter with reference to the accompanying drawings, but the illustrated embodiments may be embodied in different forms, and the present disclosure should not be construed as limited to the embodiments set forth below. These embodiments are provided to make the present disclosure more thorough and complete and to enable those skilled in the art to fully understand the scope of the present disclosure.
[0026] The accompanying drawings of the embodiments of the present disclosure are used to provide a further understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the detailed embodiments, they are used to explain the present disclosure and do not constitute a limitation of the present disclosure. The above and other features and advantages will become more apparent to those skilled in the art by describing the detailed embodiments with reference to the accompanying drawings.
[0027] The present disclosure may be described with reference to plan views and / or cross-sectional views by way of ideal schematic views of the present disclosure. Therefore, the exemplary illustrations may be modified according to manufacturing techniques and / or tolerances.
[0028] In the absence of conflict, the various embodiments of the present disclosure and the various features therein may be combined with each other.
[0029] The terms used in this disclosure are only used to describe specific embodiments and are not intended to limit the disclosure. As used in this disclosure, the term "and / or" includes any and all combinations of one or more related enumerated items. As used in this disclosure, the singular forms "a" and "the" are also intended to include plural forms, unless the context clearly indicates otherwise. As used in this disclosure, the terms "comprising" and "made of" specify the presence of the features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof.
[0030] Unless otherwise defined, all terms (including technical and scientific terms) used in this disclosure have the same meanings as those commonly understood by those skilled in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning unless expressly defined in this disclosure.
[0031] The present disclosure is not limited to the embodiments shown in the drawings, but includes modifications of the configurations formed based on the manufacturing process. Therefore, the regions illustrated in the drawings have schematic properties, and the shapes of the regions shown in the drawings illustrate the specific shapes of the regions of the elements, but are not intended to be limiting.
[0032] Image acquisition devices (such as cameras, mobile phones, etc.) directly capture the original image data through photosensitive devices, that is, RAW domain images, and RAW domain images need to be enhanced by ISP image quality.
[0033] ISP can be performed in the form of an ISP pipeline in a chip (such as an ISP chip). Referring to Figure 10, the ISP pipeline includes multiple processes performed in sequence, including but not limited to high dynamic range rendering (HDR) processing, noise reduction (NR) processing, black level correction (BL) processing, digital gain (Dgain) processing, lens shading correction (LSC) processing, automatic white balance correction (AWB) processing, demosaicing (DMC) processing, etc.
[0034] With the increase in computing power of specialized processing devices such as graphics processing units (GPUs) and neural processing units (NPUs) and the development of AI technology, image processing using AI models is no longer limited to high-level tasks such as tracking, detection, recognition, and classification, but can also be used for low-level tasks in ISP.
[0035] However, in some related technologies, there are many problems with using AI models for ISP.
[0036] For example, ISP is performed in chips (such as ISP chips), but AI models need to run in specialized devices such as GPUs. Therefore, when using AI models for ISP, a large amount of data must be transmitted between the chip and the GPU, resulting in high data exchange overhead.
[0037] For example, AI models have black box properties and are difficult to control directly, resulting in poor local and global adjustability and controllability.
[0038] As a result, the AI model and ISP in related technologies are poorly coupled, the data interaction overhead is high, and the local and global adjustability and controllability are poor.
[0039] In a first aspect, an embodiment of the present disclosure provides an image signal processing (ISP) method.
[0040] The embodiment of the present disclosure is used to perform ISP on the original RAW domain image captured by an image capture device to enhance its image quality and obtain a directly usable normal image.
[0041] The ISP method of the embodiment of the present disclosure includes noise reduction and demosaicing processing; referring to FIG. 1 , the noise reduction and demosaicing processing includes the following steps S101 to S102 .
[0042] In step S101 , raw data is input into a preset image processing model.
[0043] In some embodiments, the original data includes RAW domain images, International Organization for Standardization (ISO) sensitivity data, and LSC data, and the image processing model is an AI model for simultaneously performing noise reduction and demosaicing on the RAW domain images.
[0044] In step S102, the result image output by the image processing model is obtained.
[0045] In some embodiments, the resolution of the resulting image is greater than the resolution of the RAW domain image.
[0046] The ISP of the embodiment of the present disclosure includes noise reduction (NR) and demosaicing (DMC) processing, and the two processes are implemented through an AI model, namely, noise reduction and demosaicing (AI NRDMC) processing.
[0047] In the noise reduction and demosaicing process, the original data including the RAW domain image is first input into a preset image processing model. The image processing model will automatically output the result image, that is, the image obtained after NR and DMC are performed on the RAW domain image.
[0048] In some embodiments, a RAW image is raw light intensity data directly sensed by a camera's photosensitive device, where the photosensitive device can be classified into different colors. For example, the photosensitive device can use a color filter array (CFA) mode, such as RGB (red-green-blue) mode, and its variants such as RGGB, BGGR, GRBG, and GBRG, or any other mode capable of representing image color information, such as RCCB (red-transmitted-transmitted blue) and RYYCY (red-yellow-yellow-transmitted-yellow).
[0049] In some embodiments, the DMC process restores the “mosaic” to multiple pixels, so it will increase the resolution of the image, so that the resolution of the resulting image is necessarily higher than that of the RAW domain image.
[0050] For example, referring to Figure 5, the dimensions of the RAW domain image (RAW domain) input to the image processing model can be h*w*1, where h represents height, w represents width, and 1 represents a single channel. That is, the size of the RAW domain image is h*w, the number of channels is 1, and the colors of pixels at different positions in the RAW domain image are different; while the dimensions of the result image generated by the image processing model can be h*w*3, where the number of channels is 3, such as red, green, and blue channels respectively, so that the result image is an RGB domain image (RGB domain), and the resolution is 3 times the resolution of the RAW domain image.
[0051] In the embodiment of the present disclosure, the original data input into the image processing model includes not only the RAW domain image but also ISO sensitivity data and LSC data, that is, ISO sensitivity data and LSC data are also introduced as reference factors in the processing of the image processing model.
[0052] In some embodiments, ISO sensitivity data is defined according to ISO standards and is data that characterizes the sensitivity of a camera (such as a photosensitive device or film) to light. The ISO sensitivity data of the entire image is usually the same value; therefore, the ISO sensitivity data is equivalent to a global strength indication of the image processing model.
[0053] In some embodiments, because the camera lens is not an ideal device, the brightness (e.g., grayscale) of different image locations under uniform incident light varies (typically, brightness is highest at the center of the image and gradually decreases toward the periphery). This is due to the presence of lens shading. Therefore, different image locations require different levels of compensation gain, i.e., LSC. Different lens shading results in different levels of noise in the image caused by the shading. Therefore, NR processing also needs to be based on LSC data (i.e., data used for LSC compensation, such as data in a lookup table (LUT)). Therefore, LSC data can be input into the image processing model to indicate the noise level in different parts of the image.
[0054] The disclosed embodiment uses an AI model to perform NR and DMC in ISP. Since NR and DMC are highly complex and nonlinear, they are most suitable for being performed by an AI model. Moreover, NR and DMC are performed simultaneously in one AI model, so only one data transmission is required between the chip (such as the ISP chip) and the device running the AI model (such as the GPU), and the data interaction overhead is low. At the same time, the data input to the AI model also includes ISO sensitivity data and LSC data. The ISO sensitivity data can be used to obtain the processing results under the required ISO sensitivity, thereby realizing global intensity control, while the LSC data can be used to eliminate the noise level differences caused by lens shading, thereby realizing local intensity control.
[0055] Therefore, the embodiments of the present disclosure achieve deep integration of AI models and ISP, with low data interaction overhead and strong local and global adjustable and controllable capabilities.
[0056] In some embodiments, the raw data further includes: Position Encoding (PE) data; the PE data represents intensity information of the position of a pixel in the RAW domain image.
[0057] As one embodiment of the present disclosure, the original data input into the image processing model may also include PE data, which is derived based on the spatial position (such as two-dimensional coordinates) of each pixel (pixel) and is used to enable the image processing model to perceive the spatial position of the currently processed pixel or image block (patch), and then perform different NR and DMC according to the differences at different spatial positions (such as the difference in noise levels caused by lens shading) to achieve better processing effects.
[0058] In some embodiments, the PE data is calculated from the RAW domain image according to a preset encoding algorithm; or, the PE data is preset.
[0059] As one method of an embodiment of the present disclosure, applicable PE data may be calculated based on the RAW domain image to be processed using a predetermined encoding algorithm, and then input into the image processing model together with the RAW domain image for processing.
[0060] For example, the above encoding algorithm can be a pre-trained AI model or a manually set algorithm.
[0061] Alternatively, as another embodiment of the present disclosure, the PE data may also be predetermined fixed data, such as preferred data determined based on experience.
[0062] In some embodiments, the image processing model includes at least one of the following: a U-Net model; a transformer-based model.
[0063] As one embodiment of the present disclosure, the image processing model used may specifically be in the form of a U-Net model or a Transformer-based model.
[0064] In some embodiments, the U-Net model is a type of convolutional neural network (CNN). Referring to Figure 6, the U-Net model includes multiple downsampling blocks connected in sequence. The lowest-level downsampling block is followed by multiple upsampling blocks connected in sequence, and the output of each downsampling block is also connected to the upsampling block at the same level.
[0065] The Transformer-based model is a model based on the attention mechanism and adopts an encoding-decoding structure. It can also use CNN as the network backbone.
[0066] It should be understood that the specific form of the image processing model used in the embodiments of the present disclosure is not limited to the above examples. It can also be other CNN-based models, or other models with encoding-decoding structures, or any other AI models that can achieve NR and DMC.
[0067] It should be understood that the original data should be input into the image processing model in a manner that matches the form of the image processing model.
[0068] For example, for the U-net model, the different information in the original data can be processed into images of the same size, and then the images with different information are used as input and spliced (contacted) at the channel scale into an input image with unchanged size and multiple channels, and the input image is actually input into the U-net model.
[0069] For example, referring to Figure 6, the dimensions of a RAW domain image can be h*w*1; PE data can be processed into an image (PE Map) of h*w*2 dimensions, where the number of channels is 2 because the spatial position is determined by two coordinate values (such as row coordinates and column coordinates); ISO sensitivity data and LSC data can be processed into images of h*w*1 dimensions (ISO Map and LSC Map), respectively. The values of all pixels in the ISO Map are the same ISO sensitivity value, and the value of each pixel in the LSC Map is the compensation value of the corresponding position in the lookup table (LUT). Thus, the above raw data can be spliced into an input image of h*w*5 dimensions and then input into the U-net model, while the dimensions of the resulting image are still h*w*3.
[0070] For another example, referring to Figure 7, the RAW domain image can be split into multiple channels according to color. For example, the RAW domain image in RGGB color mode is split into 4 channels, that is, the dimension of the RAW domain image becomes h / 2*w / 2*4. Correspondingly, the dimensions of PE Map, ISO Map, and LSC Map become h / 2*w / 2*2, h / 2*w / 2*1, and h / 2*w / 2*1, respectively. After splicing, the input image with a dimension of h / 2*w / 2*8 is obtained.
[0071] For another example, referring to Figure 8 , in addition to the RAW domain image, other information in the original data (PE Map, ISO Map, LSC Map) is also simultaneously input into each sampling block of the U-net model. It should be understood that the size of the other data input to each sampling block should match the requirements of the sampling block. For example, the PE Map, ISO Map, and LSC Map can be downsampled differently to reduce their size.
[0072] For another example, referring to FIG9 , when the RAW domain images are multiple consecutive frames in a video, the information of the multiple RAW domain images (Multi-frame info) can be input into the image processing model. The multiple RAW domain images can be input continuously, or the multiple RAW domain images can be stitched together and input as one image.
[0073] In some embodiments, the method of the embodiment of the present disclosure further includes at least one of the following processing: HDR processing; Black Level Correction (BLC) processing; Dgain processing; LSC processing; AWB processing.
[0074] In some embodiments, the noise reduction and demosaicing process is performed after the HDR process; after the noise reduction and demosaicing process, at least one of the following processes is also included: BLC process; Dgain process; AWB process; LSC process.
[0075] 11 , in the ISP process of the embodiment of the present disclosure, in addition to the noise reduction and demosaicing (AI NRDMC) processing performed by the AI model mentioned above, other processing steps may also be included, such as HDR processing, BLC processing, Dgain processing, LSC processing, AWB processing, etc.
[0076] It should be understood that the above LSC processing is a process for eliminating lens shading in the image, which requires the use of lookup table (LUT) data (LSC data); and although the image processing model also uses LSC data, its function is to better perform NR processing and DMC processing based on the LSC data, rather than directly performing LSC processing in the image processing model.
[0077] Further, referring to Figure 11, the noise reduction and demosaicing processing can be located after the HDR processing and before other processing steps. That is, in the ISP pipeline, the noise reduction and demosaicing processing can be performed directly after the HDR processing, and then other processing can be performed.
[0078] It should be understood that the ISP process of the embodiments of the present disclosure is not limited to the above examples.
[0079] For example, referring to FIG. 12 , during the ISP process, the noise reduction and demosaicing processing may also be located at other positions.
[0080] For another example, referring to FIG13 , BLC processing and AWB processing are also embedded into the image processing model as “pre-processing modules” (the results of BLC processing and AWB processing are both represented by signed numbers).
[0081] For another example, the ISP process may include only part of the above-listed processing (the noise reduction and demosaicing processing is necessary).
[0082] For another example, the ISP process may also include other processing not listed above, such as bad pixel removal processing, tone mapping (TM) processing, etc.
[0083] In some embodiments, referring to FIG. 4 , before the noise reduction and demosaicing process, an image processing model is further included; and the image processing model training includes the following steps A101 to A103 .
[0084] In step A101 , training sample pairs are determined.
[0085] In some embodiments, the training sample pairs include corresponding training raw data and training standard images, and the training raw data include training RAW domain images, training ISO sensitivity data, and training LSC data.
[0086] In step A102, the original training data is input into the image processing model to obtain the training result image output by the image processing model.
[0087] In step A103, a loss function is determined based on the training standard image and the training result image, and the image processing model is adjusted based on the loss function.
[0088] As one method of an embodiment of the present disclosure, the image processing model may be trained in advance to improve the performance of the image processing model.
[0089] In some embodiments, each training uses a training sample pair, which includes the original training data for input to the image processing model, and the "ideal result" (training standard image) that should theoretically be produced based on the original training data when the image processing model performs well.
[0090] After the original training data is actually input into the image processing model, the image processing model will generate a training result image, and the loss function (Loss) can be determined based on the training result image and the training standard image. The loss function represents the gap between the current actual processing result of the image processing model and the ideal result. Therefore, the parameters of the image processing model (such as the element value of the convolution kernel, the weight value of the connection, etc.) and the structure can be adjusted in the direction of optimizing the loss function through optimization algorithms such as the gradient descent method to optimize the image processing model.
[0091] It should be understood that the image processing model can be trained multiple times using different training sample pairs until a training termination condition is met. The specific forms of the training termination condition are various, such as when the image processing model reaches a certain performance threshold (e.g., signal-to-noise ratio (PSNR) or maximum similarity (SSIM), when the image processing model's performance converges, or when a certain number of training cycles are reached.
[0092] In some embodiments, referring to FIG. 4 , determining training sample pairs (step A101 ) includes the following steps A1011 to A1013 .
[0093] In step A1011 , a high-definition low-noise image is acquired, the high-definition low-noise image is determined to be a training standard image, and training ISO sensitivity data and training LSC data are determined.
[0094] In step A1012 , noise is added to the high-definition low-noise image according to the noise function corresponding to the training ISO sensitivity data to obtain a high-definition image.
[0095] In step A1013 , the high-definition image is downsampled to obtain a training RAW domain image.
[0096] As one method of an embodiment of the present disclosure, it is possible to directly obtain a high-definition (resolution), low-noise image (high-definition low-noise image) as a training standard image, and set the training ISO sensitivity data and training LSC data required for training.
[0097] At the same time, a noise function is obtained, which represents the noise that should be present in the image when the target image acquisition device (i.e., the device that actually acquires the RAW domain image) is used to acquire the image at different ISO sensitivities. Therefore, according to the selected training ISO sensitivity data, the corresponding noise can be "added" to the high-definition low-noise image through the noise function to obtain an image with noise (a high-definition image), and the noise in the high-definition image is the noise that should be present in the RAW domain image directly acquired by the target image acquisition device under the training ISO sensitivity data.
[0098] By continuing to downsample the high-definition image, an image with a normal noise level and lower resolution (training RAW domain image) can be obtained, which is equivalent to the RAW domain image directly acquired by the target image acquisition device. Therefore, it corresponds to the high-definition, low-noise image. That is, ideally, the image processing model should be able to generate a high-definition, low-noise image based on the training RAW domain image.
[0099] In some embodiments, adding noise to the high-definition low-noise image according to the noise function corresponding to the training ISO sensitivity data to obtain a high-definition image (step A1012) includes the following steps A10121 to A10122.
[0100] In step A10121, noise compensation at each position of the high-definition low-noise image is determined based on the training LSC data.
[0101] In step A10122, noise is added to the high-definition low-noise image according to the noise function and noise compensation to obtain a high-definition image.
[0102] As one method of an embodiment of the present disclosure, when adding noise to a high-definition low-noise image, the noise at each position of the image can be further compensated according to the set LSC data (training LSC data) so that the final noise is more consistent with the actual noise when affected by lens shading.
[0103] Example:
[0104] The following is an exemplary introduction to a specific ISP method according to an embodiment of the present disclosure.
[0105] The ISP method of this example may include two parts: model (image processing model) training and model inference (model usage).
[0106] In some embodiments, the model training portion may include the following steps (1) to (9).
[0107] (1) Construct a multi-scene, high-definition, low-noise RGB image (RGB domain) dataset (a collection of multiple high-definition and low-noise images).
[0108] In some embodiments, a high-resolution, high-definition, low-noise image acquisition device may be used to acquire the RGB image dataset.
[0109] Furthermore, multiple images of the same scene can be collected and aligned and fused into one image to improve the clarity of the image and reduce the noise level therein.
[0110] In some embodiments, the RGB image dataset may include images in a variety of illumination, color temperature, and detail scenes, such as high-brightness scenes at noon, medium-brightness scenes at dusk, low-illumination scenes at night, high-texture scenes, low-texture flat scenes, and other scenes.
[0111] (2) Calibrate the noise parameters of the target image acquisition device at various ISO sensitivities.
[0112] In some embodiments, a target image acquisition device is used to capture images at different ISO sensitivities, and professional equipment is used to calibrate noise profiles therein.
[0113] For example, the noise of the target image acquisition device at a certain ISO sensitivity can be calibrated to obey the Poisson-Gaussian distribution, that is:
[0114] Among them, x noisy represents an image with noise, x * represents the expected value of the noise-free image, Possion(.) and Gaussian(.) represent Poisson distribution and Gaussian distribution respectively, so (k,σ 2 ) represents the noise coefficient at that ISO sensitivity.
[0115] (3) Construct a noisy image set (a collection of multiple high-definition images).
[0116] As before, the noise coefficient at a certain ISO sensitivity is (k,σ 2 ), so that a variety of noise intensity levels can be further set, such as low, medium, and high, and the corresponding three noise variances var low var mid var high They are: var low =α low,0 ·k·x intensity +α low,1 ·σ 2 ; var mid =α mid,0 ·k·x intensity +α mid,1 ·σ 2 ; varhigh =α high,0 ·k·x intensity +α high,1 ·σ 2 ;
[0117] Among them, α low,0 , α mid,0 , α high,0 and α low,1 , α mid,1 , α high,1 Respectively represent the intensity coefficients that can be set at each noise intensity level, such as [α low,0 ,α mid,0 ,α high,0 ]=[0.9,1.0,1.1],[α low,1 ,α mid,1 ,α high,1 ]=[0.9,1.0,1.1];x intensity Indicates the value of a pixel (such as grayscale), such as x for 8-bit images intensity =[0,255].
[0118] Therefore, the above is equivalent to determining the noise function.
[0119] Furthermore, the LSC data of the target image acquisition device is a known lookup table (LUT), and the resolution of the lookup table is smaller than the resolution of the image. Therefore, bilinear interpolation, bicubic interpolation and other methods can be used to complete the data therein to obtain training LSC data, where each value of the LSC data represents the intensity of the LSC data used by the image of the corresponding spatial position index (i, j).
[0120] For example, the intensity of the LSC data at the center of the image is defined as lsc base , then the noise intensity compensation at different spatial positions of the image (noise compensation determined based on training LSC data) is: var low,lsc =var low ·(lsc i,j / lsc base ); var mid,lsc =var mid ·(lsc i,j / lsc base ); var high,lsc =var high ·(lsc i,j / lsc base );
[0121] Furthermore, noise can be added to the RGB image dataset based on the above noise variance and noise intensity compensation, that is: ylow,intensity=xlow,intensity+Gaussian(0,var low,lsc ); ymid,intensity=xmid,intensity+Gaussian(0,var mid,lsc ); yhigh,intensity=xhigh,intensity+Gaussian(0,var high,lsc );
[0122] Among them, ylow,intensity, ymid,intensity, and yhigh,intensity represent the noise-added images (high-definition images) obtained under a certain noise intensity level; Gaussian (0 ,var ) represents a Gaussian distribution with mean 0 and variance var.
[0123] Therefore, for the same image, the noise intensity and variance of the noisy images obtained at different noise intensity levels are different. And for the same noisy image, the noise intensity and variance at different positions are also different.
[0124] (4) Downsample the noisy image slices to the RAW domain to obtain training RAW domain images and form training sample pairs.
[0125] Downsample the noisy RGB image data to enter the RAW domain, that is: yraw,low,intensity=DS(ylow,intensity); yraw,mid,intensity=DS(ymid,intensity); yraw,high,intensity=DS(yhigh,intensity);
[0126] Where DS(.) represents the downsampling function from RGB domain to RAW domain, yraw,low,intensity, yraw,mid,intensity, and yraw,high,intensity represent the training RAW domain images obtained by downsampling the noisy images at different noise intensity levels, and the corresponding directly acquired images x are intensity As the training standard image, we get the training sample pairs, such as the training sample pairs [yraw, low, intensity, x intensity ].
[0127] (5) Train U-net model or Transformer-based model (image processing model).
[0128] Using the training sample pairs obtained above, training is performed by: inputting the original training data into the image processing model - determining the loss function based on the obtained training result images and training standard images - adjusting the image processing model according to the loss function until the above training end conditions are met; and saving the current U-net model or Transformer-based model as the image processing model for subsequent use.
[0129] The PE data (PE map) to be used subsequently can also be pre-calculated at this time, for example, by mapping it through the function (encoding algorithm) func([i,j]) based on the spatial position [i,j] of each pixel in the RAW domain image.
[0130] Among them, the function func(.) can be a pre-set function or a trained AI model.
[0131] The obtained PE map can be block-level, that is, multiple adjacent pixels belonging to a block in the PE map have the same value.
[0132] In some embodiments, the model inference portion may include the following steps (6) to (9).
[0133] (6) Obtain input for the image processing model.
[0134] The RAW domain image captured by the target image acquisition device or the RAW domain image processed by HDR is used as the input of the image processing model, where the pixel (pixel) at the position index [i, j] is defined as p i,j , where i = 0, 1, ..., h, j = 0, 1, ..., w, h and w represent the height and width of the image respectively.
[0135] (7) Perform BLC and AWB pretreatment.
[0136] The BLC processing and AWB processing are performed respectively through the pre-processing unit of the image processing model.
[0137] In some embodiments, pixelp i,j The BLCs are: blc,i,j =p i,j -blc offset ;
[0138] Among them, p blc,i,j Use signed numbers, and blc offset The compensation value can be used to compensate different values for different color channels. In AWB processing, the AWB compensation coefficient can be defined as [gainR ,gain Gr ,gain G b, gain B ], then:
[0139] Among them, p awb,i,j Using signed numbers, pixel∈ indicates that the pixel belongs to the position of the corresponding color of the CFA pattern.
[0140] As before, the data after BLC processing and AWB processing can be retained as signed numbers to improve the NR processing capability under low illumination.
[0141] (8) Image processing model performs NR and DMC.
[0142] The function that defines the image processing model is forward(.), and the output (result image) is model output , then we have: model output =clip(forward(P awb ,ISO,LSC,PE),0,wl aiNrDmc );
[0143] Among them, P awb Indicated by p awb,i,j The image block (patch) is composed of ISO, LSC, and PE, respectively. The bold words represent images (or matrices), and clip (.) represents the truncation function. aiNrDmc The maximum value allowed for output by the image processing model.
[0144] (9) Perform Dgain treatment and LSC treatment.
[0145] Continue to perform Dgain processing and LSC processing on the result image, where the gain of Dgain processing is: gain digital ;
[0146] Therefore, the Dgain processing for the 3-channel RGB image can be expressed as: dgain,i,j =clip(modeloutput,i,j×gain digital ,0,wl dgain );
[0147] Among them, modeloutput,i,j represents the value of the model output (result image) at the [i,j] position, wl dgain Indicates the maximum output value allowed by the Dgain module.
[0148] The gain of LSC processing is: lsc,R,i,j ,gainlsc,Gr,i,j,gainlsc,Gb,i,j,gain lsc,B,i,j ];
[0149] Among them, gain lsc,.,i,j Represents the LSC gain value of the corresponding color pixel at the [i, j] spatial position;
[0150] Therefore, the LSC processing for 3-channel RGB images can be expressed as:
[0151] Among them, p lsc,R,i,j 、p lsc,G,i,j 、p lsc,B,i,j Respectively represent the values of the corresponding color channels after LSC processing, wl lsc Indicates the maximum output value allowed by LSC processing.
[0152] Therefore, p lsc,R,i,j 、p lsc,G,i,j 、p lsc,B,i,j This is the final processing result of the ISP.
[0153] In a second aspect, referring to FIG2 , an embodiment of the present disclosure provides an electronic device comprising a memory and a processor; the memory stores a computer program that can be executed by the processor, and the computer program is executed by the processor, so that the processor implements any one of the image signal processing methods of the embodiments of the present disclosure.
[0154] In a third aspect, referring to FIG3 , an embodiment of the present disclosure provides a computer-readable storage medium on which a computer program is stored. The computer program is executed by a processor, so that the processor implements any one of the image signal processing methods of the embodiments of the present disclosure.
[0155] In some embodiments, the processor is a device with data processing capabilities, including but not limited to a central processing unit (CPU); the memory is a device with data storage capabilities, including but not limited to random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), and flash memory (FLASH); the I / O interface (read-write interface) is connected between the processor and the memory, and can realize information interaction between the memory and the processor, including but not limited to a data bus (Bus), etc.
[0156] Those skilled in the art will appreciate that all or some of the steps, systems, and functional modules / units in the apparatus disclosed above may be implemented as software, firmware, hardware, or a suitable combination thereof.
[0157] In hardware implementations, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, one physical component may have multiple functions, or one function or step may be performed by several physical components in cooperation.
[0158] Some or all of the physical components may be implemented as software executed by a processor (such as a central processing unit (CPU), a digital signal processor, or a microprocessor), or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or temporary media). As is well known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory (FLASH) or other disk storage; compact disc (CD-ROM), digital versatile disc (DVD) or other optical disc storage; magnetic cassettes, tapes, disk storage or other magnetic storage; any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, as is well known to those skilled in the art, communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0159] The present disclosure has disclosed example embodiments, and although specific terms are employed, they are used and should be interpreted only in a general illustrative sense and not for purposes of limitation. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly indicated, features, characteristics, and / or elements described in conjunction with a particular embodiment may be used alone or in combination with features, characteristics, and / or elements described in conjunction with other embodiments. Therefore, it will be understood by those skilled in the art that various changes in form and detail may be made without departing from the scope of the present disclosure as set forth in the appended claims.
Claims
1. A method for image signal processing, comprising noise reduction and demosaicing processing, wherein the noise reduction and demosaicing processing comprises: Inputting raw data into a preset image processing model; the raw data includes a Bayer RAW domain image, ISO sensitivity data of the International Organization for Standardization, and lens shading correction LSC data, and the image processing model is an artificial intelligence AI model for simultaneously performing noise reduction and demosaicing on the RAW domain image; and A result image output by the image processing model is obtained; the resolution of the result image is greater than the resolution of the RAW domain image.
2. The method according to claim 1, wherein: The original data also includes: Position-encoded PE data; the PE data represents intensity information of the position of the pixel in the RAW domain image.
3. The method according to claim 2, wherein: The PE data is calculated from the RAW domain image according to a preset encoding algorithm; or, The PE data is preset.
4. The method according to claim 1, wherein: The image processing model includes at least one of the following: U-Net model; and Transformer-based model based on transformer.
5. The method according to claim 1, wherein: Also includes at least one of the following treatments: High dynamic range rendering HDR processing; Black level correction BLC processing; Digital gain Dgain processing; LSC treatment; and Automatic white balance correction AWB processing.
6. The method according to claim 1, wherein: It also includes HDR processing; the noise reduction and demosaicing processing is performed after the HDR processing; after the noise reduction and demosaicing processing, it also includes at least one of the following processing: BLC treatment; Dgain processing; AWB processing; and LSC treatment.
7. The method according to claim 1, wherein: Before the noise reduction and demosaicing process, the image processing model is further trained; the training of the image processing model includes: Determine a training sample pair; the training sample pair includes corresponding training original data and training standard images, The training raw data includes training RAW domain images, training ISO sensitivity data, and training LSC data; Inputting the original training data into the image processing model to obtain a training result image output by the image processing model; and A loss function is determined according to the training standard image and the training result image, and the image processing model is adjusted according to the loss function.
8. The method according to claim 7, wherein: Determining the training sample pairs comprises: Acquire a high-definition low-noise image, determine that the high-definition low-noise image is the training standard image, and determine the training ISO sensitivity data and the training LSC data; adding noise to the high-definition low-noise image according to the noise function corresponding to the training ISO sensitivity data to obtain a high-definition image; and The high-definition image is down-sampled to obtain the training RAW domain image.
9. The method according to claim 8, wherein: The adding noise to the high-definition low-noise image according to the noise function corresponding to the training ISO sensitivity data to obtain a high-definition image comprises: Determining noise compensation at each position of the high-definition low-noise image according to the training LSC data; and Noise is added to the high-definition low-noise image according to the noise function and the noise compensation to obtain the high-definition image.
10. An electronic device comprising a memory and a processor; the memory stores a computer program executable by the processor, the computer program is executed by the processor, so that the processor implements the method for image signal processing according to any one of claims 1 to 9.
11. A computer-readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor so that the processor implements the image signal processing method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Image processing method and device, electronic equipment and computer readable storage medium
CN113689335A
Network model training method, image processing method and related equipment
CN113850367A
Image processing method and device, electronic equipment and computer readable storage medium
CN115063333A
Image processing method and device
CN116739948A
Image processing apparatus and its control method, computer program and computer readable storage medium
JP2008042682A