Image signal processing method and corresponding storage medium
By dividing the image signal processing process into two parts—traditional processing and artificial intelligence model-based processing—the problem of large parameter libraries and difficult debugging in existing technologies is solved, resulting in a significant improvement in image quality, especially in the optimization of peak signal-to-noise ratio, structural similarity index, and chroma channel.
Patent Information
- Application Number
- PCT/CN2024/116766
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-25
- Filing Date
- 2024-09-04
- Publication Date
- 2026-01-02
AI Technical Summary
Existing image signal processing methods face problems such as large parameter libraries, difficulty in debugging, and long development cycles, making it difficult to meet the special image quality needs of different users.
The image signal processing process is divided into two parts. One part, which is related to the non-shared parameters of the shooting device, adopts traditional methods. The other part is processed using an artificial intelligence-based image signal processing model, which uses a machine learning model to process the image after the first processing.
It significantly improves image quality, especially in peak signal-to-noise ratio, structural similarity index and chroma channel performance, meeting the special needs of different users in different scenarios.
Smart Images

Figure CN2024116766_02012026_PF_FP_ABST
Abstract
Description
Image signal processing method and corresponding storage medium TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, in particular to an image signal processing method and a storage medium. BACKGROUND
[0002] An optical image of an object in nature is projected onto the surface of an image sensor with a color filter array through a camera lens, is first converted into an analog electrical signal through photoelectric conversion, is then converted into a digital image signal, i.e., a Bayer format image, through an analog to digital converter (A / D), and is then sent to a digital signal processing chip (DSP) for image signal processing (ISP) to form an RGB format image.
[0003] General ISP processing can include black level compensation, lens shading correction, bad pixel correction, color interpolation, Bayer noise removal, white balance correction, color correction, Gamma correction, color space conversion, color noise removal and edge enhancement in the YUV color space, color and contrast enhancement, etc. With the increasing complexity of image processing scenarios and the increasing special requirements for image quality, general ISP faces challenges such as a large parameter library, difficulty in debugging, and a long development cycle.
[0004] SUMMARY
[0005] The embodiments of the present application provide an image signal processing method and a storage medium, which improve the quality of the processed image.
[0006] In an aspect, the embodiments of the present application provide an image signal processing method, comprising:
[0007] obtaining an original shooting format image;
[0008] performing first image signal processing on the original shooting format image to obtain a first processed image; the first image signal processing is processing related to non-shared parameters of a shooting device;
[0009] performing second image signal processing on the first processed image according to a preset image signal processing model to obtain a display format image; the image signal processing model is used for the second image signal processing on the first processed image, and the second image signal processing is other image signal processing than the first image signal processing.
[0010] In another aspect, the present application provides a computer readable storage medium storing a plurality of computer programs, which are adapted to be loaded and executed by a processor to implement the image signal processing method according to the first aspect of the present application.
[0011] In the method of the present application, the ISP processing on the original shooting format image is divided into two parts, one part of the first image signal processing related to the non-shared parameters of the shooting device is performed by using the conventional method, and the other part of the second image signal processing is performed by using the artificial intelligence based method, i.e. by using the preset image signal processing model. By comparing the quality of the display format images obtained by using the method of the present application and the conventional ISP processing, it can be found that the quality of the display format image obtained by using the method of the present application can be greatly improved, especially in the peak signal-to-noise ratio, structural similarity index and chroma channel. BRIEF DESCRIPTION OF DRAWINGS
[0012] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0013] Fig. 1 is a schematic diagram of an image signal processing method according to an embodiment of the present application;
[0014] Fig. 2 is a flowchart of an image signal processing method according to an embodiment of the present application;
[0015] Fig. 3 is a flowchart of a method for training an image signal processing model according to an embodiment of the present application;
[0016] Fig. 4 is a schematic diagram of an initial image signal processing model according to an embodiment of the present application;
[0017] Fig. 5 is a schematic diagram of a VGG16 network according to an embodiment of the present application;
[0018] Fig. 6 is a schematic diagram of a calculation of a generative adversarial loss function according to an embodiment of the present application;
[0019] Fig. 7 is a flowchart of an image signal processing method according to a specific application embodiment of the present application;
[0020] Fig. 8 is a flowchart of a method for training an image signal processing model according to a specific application embodiment of the present application;
[0021] Fig. 9 is a schematic diagram of the logical structure of a server according to an embodiment of the present application. DETAILED DESCRIPTION
[0022] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work belong to the protection scope of the present application.
[0023] The terms "first", "second", "third", "fourth" and the like (if any) in the description, claims, and above drawings of the present application are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or device including a series of steps or units does not necessarily have to include those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to the process, method, product, or device.
[0024] The embodiments of the present application provide an image signal processing method, which can be mainly applied to the process of obtaining a display format image by processing an original shooting format image after the image is shot by a shooting device, as shown in FIG. 1, the image signal processing system can perform the following image signal processing:
[0025] obtaining an original shooting format image; performing first image signal processing on the original shooting format image to obtain a first processed image; the first image signal processing is processing related to non-shared parameters of the shooting device; performing reprocessing on the first processed image according to a preset image signal processing model to obtain a display format image, the image signal processing model is used to perform second image signal processing on the first processed image, and the second image signal processing is other image signal processing than the first image signal processing.
[0026] Among them, the image signal processing system can be applied to a terminal device or a server, and the terminal device can include but is not limited to the following electronic devices that need to perform image processing: mobile phones, computers, smart voice interaction devices, smart home appliances, vehicle-mounted terminals, etc.
[0027] The preset image signal processing model is a machine learning model based on artificial intelligence, which can be obtained by a certain method, wherein the artificial intelligence (AI) is a theory, method, technology and application system for perceiving environment, acquiring knowledge and using knowledge to obtain the best results by using a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence. In other words, artificial intelligence is a comprehensive technology of computer science, which attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that the machine has the functions of perception, reasoning and decision-making.
[0028] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, machine learning and deep learning.
[0029] Machine learning (ML) is a multi-disciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It is a subject that studies how computers simulate or implement human learning behavior to acquire new knowledge or skills, and reorganize existing knowledge structure to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent, and its applications are widespread in various fields of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and teaching learning.
[0030] In this way, the quality of the display format image obtained by the method of the embodiment and the display format image obtained by the traditional ISP processing are compared, and the quality of the display format image obtained by the method of the embodiment can be greatly improved, especially in peak signal-to-noise ratio, structural similarity index and chroma channel.
[0031] An embodiment of the present application provides an image signal processing method, which is mainly a method executed by an image signal processing system, and a flowchart is shown in FIG. 2, comprising:
[0032] Step 101, obtaining an original shooting format image.
[0033] It can be understood that when the photographing apparatus photographs an image, an optical image generated by an object passing through a camera lens is projected onto a surface of an image sensor using a color filter array, is first converted into an analog electrical signal through photoelectric conversion, and is then converted into a digital image signal through an analog-to-digital converter, that is, a Bayer format image. In this embodiment, the Bayer format image can be taken as a raw photographing format image. In order to meet special scene requirements and special quality requirements of different users for a photographed image, image signal processing, that is, ISP processing, needs to be performed on the raw photographing format image to obtain a display format image, such as a joint photographic experts group (JPEG) image.
[0034] Specifically, the image signal processing performed on the raw photographing format image can include but is not limited to the following processing: black level compensation, lens shading correction, bad pixel correction, color interpolation (demosaic), Bayer noise removal, auto white balance (AWB) correction, color correction matrix (CCM), Gamma correction, demosaicing, color space conversion (conversion from RGB to YUV), removal of color noise and edge enhancement in the YUV color space, color and contrast enhancement, and automatic exposure control.
[0035] Step 102, first image signal processing is performed on the raw photographing format image to obtain a first processed image, where the first image signal processing is processing related to non-shared parameters of the photographing apparatus.
[0036] The first image signal processing herein mainly refers to processing related to non-shared parameters of the photographing apparatus, where the non-shared parameters refer to parameters that are not shared (that is, the same) by all photographing apparatuses, and the non-shared parameters of different photographing apparatuses are different. The photographing apparatus herein refers to an apparatus used to photograph the raw photographing format image. For example, black level correction, normalization, automatic white balance, color correction, demosaicing, and the like mainly involve basic color correction, color space conversion, and color detail restoration, and have high requirements for accuracy and stability, and the parameters based on these processes of different photographing apparatuses are different, and are related to non-shared parameters of the photographing apparatus. In this embodiment, the first image signal processing is performed on the raw photographing format image through a traditional method, which can provide accurate and stable processing results, resulting from long-term accumulated experience, strict mathematical models, and optimized algorithms.
[0037] Step 103, re-processing the first processed image according to the preset image signal processing model to obtain a display format image, the image signal processing model being used for second image signal processing of the first processed image, the second image signal processing being other image signal processing than the first image signal processing.
[0038] Here, the image signal processing model is an artificial intelligence-based machine learning model, which can be trained by a certain training method, and the running program of the image signal processing model is set in the image signal processing system in advance. When the image signal processing process of the embodiment is initiated, the image signal processing model can be directly called to re-process the first processed image.
[0039] Here, the second image signal processing is other image signal processing than the first image signal processing, which is generally related to the shared parameters of the shooting device. Here, the shared parameters refer to parameters that can be shared (i.e. the same) by all shooting devices, such as parameters for color restoration, detail texture recovery, brightness adjustment, contrast adjustment, and noise removal. These processes include color, texture, brightness, contrast, noise, and other complex data features, and different shooting devices can have the same parameters based on these processes.
[0040] It can be seen that in the method of the embodiment, the ISP processing of the original shooting format image by the image signal processing system is divided, a part of the first image signal processing related to the non-shared parameters of the shooting device is processed by a traditional method, and another part of the second image signal processing is processed by an artificial intelligence-based method, i.e. by using a preset image signal processing model. The quality of the display format image obtained by using the method of the embodiment and the traditional ISP processing is compared, and the quality of the display format image obtained by using the method of the embodiment can be greatly improved, especially in terms of peak signal-to-noise ratio, structural similarity index, and chroma channel.
[0041] In a specific embodiment, the training of the image signal processing model used in step 103 can be implemented according to the following supervised training method, as shown in the flowchart of FIG. 3, which includes:
[0042] Step 201, determining an initial image signal processing model.
[0043] It can be understood that the image signal processing system determines the initial model of image signal processing when determining the initial model of image signal processing, and determines the multi-layer structure included in the initial model of image signal processing and the initial value of the parameters in each layer structure. Among them, the parameters of the initial model of image signal processing refer to the fixed parameters used in the calculation process of each layer structure in the initial model of image signal processing, which do not need to be assigned at any time, such as parameter scale, weight value, user vector length and the like.
[0044] Specifically, the initial model of image signal processing is specifically used for performing second image signal processing on the first processed image (obtained after the first image signal processing) and outputting a display format image. In the specific implementation process, as shown in FIG. 3, the initial model of image signal processing can include an image signal processing network based on multi-level wavelet transform of convolutional networks for biomedical image segmentation (U-Net), and can include a multi-layer U-Net (taking a two-layer U-Net as an example in FIG. 3), each layer of the U-Net can include a wavelet transform down-sampling layer, a wavelet transform up-sampling layer, a convolution and activation layer and a residual module, and in the multi-layer U-Net, any layer of the U-Net is nested in another layer of the U-Net.
[0045] Among them, the residual module can be specifically composed of a residual channel attention module (RCAB), which can minimize the loss of information in these layers.
[0046] Step 202, determining a training sample, the training sample including a plurality of groups of sample images, each group of sample images including an original photographed sample image and a corresponding display sample image.
[0047] Specifically, when determining the training sample, the corresponding display sample image can be obtained after performing traditional ISP processing on each original photographed sample image in the shooting device; or the corresponding display sample image can be obtained after processing the original photographed sample image by using an image editing program (such as PS software, etc.).
[0048] Step 203, adjusting the initial model of image signal processing according to the initial model of image signal processing and the training sample, to obtain the above-mentioned preset image signal processing model.
[0049] Specifically, the image signal processing system can first perform the above-mentioned first image signal processing on each original shooting sample image to obtain a first processed image, and then perform second image signal processing on the first processed image through the image signal processing initial model to obtain a corresponding display format image; then, according to each display format image obtained by the image signal processing initial model and the corresponding display sample image in the training sample, a loss function related to the image signal processing initial model is calculated; and then, according to the calculated loss function, the parameter value in the image signal processing initial model is adjusted.
[0050] Further, the training process of the image signal processing model is to minimize the value of the above-mentioned loss function, and the training process is to continuously optimize the parameter value of the parameter in the image signal processing initial model determined in the above-mentioned step 201 through a series of mathematical optimization methods such as back propagation derivation and gradient descent, and to minimize the calculated value of the above-mentioned loss function.
[0051] It should be noted that the above-mentioned steps 202 to 203 are a one-time adjustment of the parameter value in the image signal processing initial model through each display format image obtained by the image signal processing initial model, and in actual application, the above-mentioned steps 202 to 203 need to be executed repeatedly until the adjustment of the parameter value meets a certain stop condition.
[0052] Therefore, after the image signal processing system executes the above-mentioned embodiment steps 201 to 203, it also needs to judge whether the current adjustment of the parameter value meets the preset stop condition, and when it meets, the process ends, and the parameter value of the image signal processing initial model adjusted in the above-mentioned step 203 is taken as the parameter value of the image signal processing model finally trained; when it does not meet, the image signal processing initial model after adjusting the parameter value is returned to execute the above-mentioned steps 202 to 203, that is, a new batch of training samples is used to adjust the parameter value in the image signal processing initial model. The preset stop condition includes but is not limited to any one of the following conditions: the difference between the current adjusted parameter value and the last adjusted parameter value is less than a threshold value, that is, the adjusted parameter value converges; and the number of adjustments of the parameter value is equal to the preset number, etc.
[0053] The loss function calculated by the image signal processing system in the process of adjusting the parameter value of the image signal processing initial model can be used to indicate the difference between each display format image obtained by the image signal processing initial model and the actual display sample image with better quality based on the corresponding original shooting sample image (i.e. the display sample image in the training sample), such as cross-entropy loss function, etc.
[0054] In a specific embodiment, the image signal processing system can first calculate one or more loss functions, not limited to, such as the pixel loss function Loss
[0055] (1) pixel loss function Loss L1
[0056] The pixel loss function L1 is a loss function that measures the difference between the display format image obtained by the image signal processing initial model and the corresponding display sample image in the training sample at the pixel level. The pixel loss function helps to improve the stability and detail retention capability of the image signal processing initial model for images, and through the characteristics of noise resistance, maintaining image details, sparsity, etc., the generated display format image is more consistent with human eye perception, clearer and more stable. For an original shooting sample image, the calculation method of the pixel loss function is the average of the absolute difference between the display format image obtained by the image signal processing initial model and the corresponding display sample image y i , as shown in the following formula 1, where n is the number of original shooting sample images in the training sample:
[0057] (2) perceptual loss function Loss vgg
[0058] The perceptual loss function, also known as the mean square error loss function (MSE), is a commonly used index to measure the difference between two numerical values. In this embodiment, it is the mean square value of the difference between the high-order features of the display format image obtained by the image signal processing initial model and the high-order features of the corresponding display sample image in the training sample. It can be specifically represented by the following formula 2:
[0059] In this embodiment, X and Y in the above formula 2 can be the high-order features feature y and of the display format image obtained by the image signal processing initial model and the corresponding display sample image in the training sample, respectively. Specifically, a pre-trained Visual Geometry Group (VGG) 16 network can be used to extract features from the display format image obtained by the image signal processing initial model and the corresponding display sample image in the training sample, respectively.
[0060] Among them, as shown in FIG. 5, the VGG16 network is a deep neural network that can extract high-order features, which can include 16 hidden layers (13 convolutional layers and 3 fully connected layers), and the perception loss function based on the VGG16 network helps the image signal processing initial model to learn the detailed information of the image.
[0061] (3) Structure loss function Loss ssim
[0062] The structure loss function is also called a structural similarity index (SSIM) loss function, which is an index for measuring the similarity between two images. In this embodiment, the two images are specifically the display format image obtained by the image signal processing initial model and the corresponding display sample image y in the training sample. The SSIM of the two images comprehensively considers multiple aspects of information such as brightness, contrast and structure, and the value range of the SSIM is usually between [0, 1], the closer to 1, the more similar the two images are. Through the structure loss function, the image signal processing initial model can capture the structure and texture in the image, thereby generating a high-quality image that more comprehensively considers the features of the image. Among them, the SSIM of the two images can be represented by the following public
[0063] formula 3:
[0064] Among them, C1=(K1L) 2 , C2=(K2L) 2 , L=2 B -1, u y , respectively represent the mean of the two images, respectively represent the variance of the two images, represents the covariance of the two images, C1 and C2 are stability coefficients, K1 and K2 are both default parameters 0.01 and 0.03, and B is the bit depth of the image. Therefore, the structure loss function includes the SSIM between the display format image obtained by the image signal processing initial model and the corresponding display sample image y in the training sample, which can be specifically represented by the following formula 4: Loss ssim =1-SSIM (4)
[0065] Further, in order to make the trained image signal processing model more accurate in processing image signals, after adjusting the parameter values in the image signal processing initial model based on the first loss function calculated in steps 201 to 203, a preliminary image signal processing initial model is obtained, and in other embodiments, the parameter values in the preliminary image signal processing initial model can be further fine-tuned. The fine-tuning method is similar to the method of adjusting the parameter values in the image signal processing initial model based on the first loss function described above, except that in the fine-tuning process, one or more loss functions are first calculated, and a second loss function is calculated as a loss function related to the image signal processing initial model based on these loss functions. For example, the second loss function is the weighted sum of these loss functions, or the second loss function is the weighted sum of these loss functions and the first loss function described above. Specifically, these loss functions can include:
[0066] (1) Chroma channel loss function Loss UV
[0067] The chroma channel (UV) loss function is a measure of the difference between the generated image and the target image in the UV color space. Compared with the RGB channel, the UV channel is more sensitive to color changes, which can more effectively affect the perception of color by the human eye. In this embodiment, the UV loss function is used to measure the difference between the display format image obtained by the image signal processing initial model and the corresponding display sample image in the training sample. The introduction of the UV loss function can enable the image signal processing model to learn and adjust colors more carefully, thereby reducing the perceived color error.
[0068] Specifically, a Gaussian blur operator G(x) with a mean of 0 and a variance σ of 20 is used to blur the display format image obtained by the image signal processing initial model and the display sample image in the training sample, respectively, as shown in the following formula 5: and G(y) i ; Finally, the loss function between the UV channels of the display format image obtained by the image signal processing initial model is calculated, as shown in the following formula 6:
[0069] (2) Adversarial loss function Loss GAN
[0070] The adversarial loss function L1 is based on the loss function of a generative adversarial network (GAN), which is a deep learning framework composed of a generator and a discriminator. The generator and the discriminator are trained in an adversarial manner in the GAN network.
[0071] In this embodiment, the generator is specifically the image signal processing initial model, and the purpose is to make the display format image generated by the image signal processing initial model approach the display sample image in the training sample, and the discriminator discriminates the display format image as the display sample image. The purpose of the discriminator is to accurately discriminate whether the input image is a display sample image or a display format image generated by the generator. Usually, a scalar real = 1 is used to represent that the discriminator is very confident that the input image is a display sample image, and a scalar fake = 0 is used to represent that the discriminator is very confident that the input image is not a display sample image.
[0072] Specifically, the selection of the discriminator in this embodiment mainly adopts a PatchGAN structure. The PatchGAN discriminator discriminates a local area of an image, and its output is an N*N matrix X. Each element X[i][j] of the matrix X is also called a local area (Patch), which represents the discrimination output of the discriminator on a local receptive field of the input image. Finally, the mean value of each Patch is taken as the discrimination result output of the discriminator PatchGAN on the input image. Such local area judgment helps the image signal processing initial model to better generate realistic image local details and textures in the adversarial learning process, and has great significance for improving the overall quality and authenticity of the display format image obtained by the image signal processing model.
[0073] In the training process of this embodiment, the generator (i.e. the image signal processing initial model) and the discriminator are alternately trained, so that they are mutually antagonistic and learn from each other, and finally the display format image obtained by the image signal processing initial model approaches the display sample image in the training sample, and the discriminator cannot accurately discriminate the difference between the display format image and the display sample image.
[0074] Specifically, as shown in FIG. 6, the adversarial training of the generator and the discriminator, in this process, the display format image generated by the image signal processing initial model is first input into the discriminator for discrimination, and the probability of discriminating as a display sample image is output The adversarial loss function Loss is calculated through the probability GAN , the initial model of image signal processing (i.e., the generator) is updated by optimization algorithms such as back propagation and gradient descent. The probability that the display format image obtained by the initial model of image signal processing is identified as the corresponding display sample image by the discriminator is included in the adversarial loss function, which aims to minimize the probability that the display format image is identified as a non-display sample image by the discriminator, and is used to guide the generator to obtain a display format image that approximates the display sample image. Specifically, it can be represented by the following formulas 7 and 8, where real is 1, indicating that the discriminator is very confident that the display format image is a display sample image:
[0075] Further, after completing the parameter update of the generator, the display format image obtained by the initial model of image signal processing and the display sample image y in the training sample are input into the discriminator to obtain the corresponding probabilities and P y , respectively, and the loss function L2 is calculated with real and fake, respectively, to obtain two loss functions loss real and loss fake , and finally the average is taken to obtain the adversarial loss function Loss PatchGAN of the discriminator, which aims to minimize the probability that the display sample image is incorrectly identified as a non-display sample image and the probability that the display format image is incorrectly identified as a display sample image. Specifically, it can be represented by the following formulas 9 to 11: loss real =L2(PatchGAN(y),real) (9)
[0076] After calculating the adversarial loss function of the discriminator, the parameters of the discriminator are updated by optimization algorithms such as back propagation and gradient descent. The steps of updating the above initial model of image signal processing (i.e., the generator) and the discriminator are alternated, which helps to balance the learning process of the generator and the discriminator and prevents one of them from being too powerful, causing unstable training. At the same time, such a strategy also enables the generator and the discriminator to learn from each other and ultimately achieve the goal of generating images that approximate good quality images.
[0077] The following is a specific application example to illustrate the image signal processing method in the embodiment of the present application, as shown in FIG. 7, which can specifically include the following process:
[0078] Step 301, obtaining an original shooting format image shot by a shooting device, which can be a Bayer format image.
[0079] Step 302, black level correction and normalization processing is performed on the original shooting format image to obtain a normalized image, specifically:
[0080] The sensors and hardware configurations of different shooting devices can have large differences, resulting in different black level and white level values. In the embodiment, traditional black level correction and normalization processing is adopted. Specifically, first, the data type of the original shooting format image Image unit16 is converted from uint16 to float32 to ensure that no overflow occurs during numerical calculation, then each pixel of the converted image is subtracted by the black level value, and values less than 0 are truncated to 0, i.e. the BLC process; finally, the corrected image value is divided by the difference between the white level white_level and the black level black_level, and the dynamic range of the image is normalized to [0, 1] to obtain the normalized image. This makes subsequent numerical operations easier and also makes the image present appropriate contrast and details under different lighting conditions. Specifically, it can be represented by the following formula 12:
[0081] Step 303, de-mosaicing processing is performed on the normalized image to obtain a de-mosaiced image. Specifically:
[0082] The original shooting format, i.e. the Bayer format image, is a color filter array (Color Filter Array) image that supports different format arrangements. Each pixel contains only red, green, and blue (R, G, B) monochromatic light, not complete RGB information, so a de-mosaicing algorithm is needed to convert the single-channel image into a complete RGB color image to obtain a de-mosaiced image. De-mosaicing mainly interpolates the original shooting format image into a complete RGB image. Taking the bilinear interpolation algorithm as an example, it interpolates in the horizontal and vertical directions, and estimates the missing color channel by weighted average of surrounding pixels. Specifically, interpolation in the horizontal and vertical directions can be performed in the following formula 13 to estimate the G channel pixel:
[0083] As shown in the following formula 14, interpolation is performed in the diagonal direction to estimate the R and B channel pixels:
[0084] As shown in the following formula 15, the interpolated G, R, and B channel values are used to obtain a complete RGB image:
[0085] Because the Bayer format is not fixed, there are many different formats, and different cameras can choose different Bayer formats. For each Bayer format, the demosaicing process involves complex image features, nonlinear relationships, etc. The traditional demosaicing process has unique advantages because it can directly restore complete RGB information through interpolation, supports different Bayer formats, has good anti-aliasing effect, can preserve image details and edges, and accurately restores color information.
[0086] Step 304, white balance processing is performed on the demosaiced image to obtain a white balance processed image. Specifically:
[0087] Cameras usually use red gain (R Gain ), green gain (G Gain ), and blue gain (B Gain ) to adjust the white balance so that the white point in the image looks neutral under various lighting conditions. Specifically, for the original RGB value of each pixel, multiply the corresponding color gain to achieve the white balance adjustment effect. Because different camera systems use different white balance algorithms, they have different color gain parameters. In this embodiment, by directly reading the color gain parameters inside the camera, the R, G, and B values of each pixel in the demosaiced image are multiplied by the corresponding color gain to obtain the white balance processed image Image wb .
[0088] Step 305, color correction is performed on the white balance processed image to obtain a corrected image, which is the first processed image mentioned above. Specifically:
[0089] The camera color space is designed and implemented by the camera manufacturer to represent and process the color information captured by the camera. Different camera manufacturers may use different sensors, image processing flows, and color models, so their camera color spaces will be different. To ensure color consistency of images on different hardware devices, color correction is needed to map the image from the camera color space to the standard color space. The camera color correction matrix is obtained by calibration and calibration by the camera manufacturer. By multiplying the original color vector of the image by the color correction matrix, the image can be mapped from the camera color space to the standard color space, so that the image appears in standard colors on different devices.
[0090] Because different cameras have different color correction matrices, in this embodiment, the white balance processed image Image wb is multiplied by the color correction matrix M to obtain the corrected image Image corrected, which is shown in the following formula 16:
[0091] Step 306, the above-mentioned corrected image is processed again according to the preset image signal processing model to obtain a display format image.
[0092] It should be noted that, in the training process, the preset image signal processing model can first complete the first image signal processing of the original sample image according to the traditional method, which can provide more accurate reference for the training of the image signal processing model, so that the image signal processing model can better adapt to the characteristics and hardware differences of different cameras, and improve the generalization ability of the model to different cameras. Specifically, as shown in FIG. 8, the image signal processing model can be trained by the following steps:
[0093] Step 401, determining the initial values of the layers and parameters in the initial image signal processing model, specifically, the structure of the initial image signal processing model can be as shown in FIG. 3, which will not be repeated here.
[0094] Step 402, determining the training sample, the training sample includes multiple groups of sample images, and each group of sample images includes an original sample image and a corresponding display sample image.
[0095] Step 403, performing the first image signal processing (as described above in steps 302 to 305) on each original sample image to obtain a first processed image, and then performing the second image signal processing on the first processed image through the initial image signal processing model to obtain a corresponding display format image.
[0096] Step 404, calculating the pixel loss function Loss L1 , the perceptual loss function Loss vgg , and the structure loss function Loss ssim according to each display format image obtained by the initial image signal processing model and the corresponding display sample image in the training sample, and calculating the first loss function according to the weighted sum of the pixel loss function, the perceptual loss function and the structure loss function, which can be shown in the following formula 17: Loss pretrain =loss L1 +loss vgg +loss ssim *0.15 (17)
[0097] The pixel loss function focuses on the accurate matching at the bottom pixel level, and the pixel loss function is an important part of ensuring the quality of the generated image, and the perceptual loss function considers the high-level perceptual features of the image, so that the generated image is more consistent with the human eye perception. In the research of the second image signal processing based on deep learning, setting their weights to 1 is a common practice, which comprehensively considers information at different levels to ensure that the generated image achieves a good balance in bottom and high-level features.
[0098] The structure loss function further improves the structural integrity, and the overall structure of the image involves the relationship between different elements, such as the relative position between objects, the coherence of textures, etc. In this embodiment, on the basis of maintaining pixel-level and perceptual similarity, the structure loss function is introduced to participate in training, which further pushes the generated image to better maintain structural similarity with the display sample image as a whole. During the training process, it is found that the value of the structure loss function is large (the order of magnitude is 1e-1), in order to avoid the training process from excessively focusing on the structure loss function and ignoring the pixel loss function and the perceptual loss function (both of which are 1e-2), the weight of the structure loss function is adjusted to 0.15, so that its numerical magnitude is consistent with the other two, which helps to ensure that the model considers detail accuracy, human eye perceptual similarity and overall structural similarity when generating images.
[0099] In step 405, the initial value of the parameter in the image signal processing initial model is adjusted according to the first loss function, and the purpose of the adjustment is to minimize the first loss function.
[0100] In step 406, it is judged whether the adjustment of the parameter value in the image signal processing initial model satisfies the preset stop condition, if yes, the parameter value of the signal processing initial model adjusted in the above step 405 is taken as the parameter value in the trained image signal processing model; if not, the image signal processing initial model after adjusting the parameter value is returned to execute the above step 402, that is, a batch of training samples are replaced, and the parameter value in the image signal processing initial model is adjusted according to the replaced training samples.
[0101] It should be noted that the image signal processing model can be trained by the above steps 401 to 406, and the image signal processing model can be tested on the test set, and the peak signal to noise ratio (PSNR), the structural similarity coefficient (SSIM) and the UV value calculated are 23.49, 0.8865 and 0.0125 respectively, that is, there is still a color error between the display format image obtained by the trained image signal processing model and the display sample image.
[0102] Thus in one specific embodiment, the trained image signal processing model needs to be further fine-tuned, specifically, the image signal processing model can be fine-tuned according to the steps of steps 402 to 406, in this process, a second loss function can be introduced, specifically, a chroma channel loss function, that is, when fine-tuning the parameter values in the image signal processing model according to the first loss function in step 405, a loss function can be calculated according to formula 18 as follows, and the parameter values in the trained image signal processing model are adjusted according to the calculated loss function: Loss finetrain = Loss pretrain + a x loss UV (18)
[0103] Wherein, a represents the weight value of the chroma channel loss function, since the chroma channel loss function and the pixel loss function, the perceptual loss function are in the same order of magnitude, in order to make the image signal processing model more focused on color restoration and reduce color error in the fine-tuning stage, a relatively large weight value is allocated to the chroma loss function in this embodiment. When the weight value of the chroma channel loss function is 1, 4, 10, 15, and 20, respectively, the results on the test set are shown in Table 1 below, and it is found that when the weight value is 10, not only can the UV color error be further reduced, but also the PSNR and SSIM values can be further improved.
[0104] Table 1
[0105] In another specific embodiment, after training the image signal processing model through the above steps 401 to 406, in order to further promote the image signal processing model to obtain more real display format images and have certain stylization effect. In this embodiment, an adversarial loss function Loss GAN is further introduced in the fine-tuning stage. Since the training of GAN network is unstable, it is a reasonable choice to set a small weight value of the adversarial loss function, especially in the fine-tuning stage. Because in the fine-tuning stage, the main goal is to keep the accuracy of the image restoration and authenticity, and not to emphasize the authenticity generated by the GAN network too much.
[0106] In this case, when the trained image signal processing model is further fine-tuned, specifically, the image signal processing model can be fine-tuned according to the steps of steps 402 to 406, and in this process, another second loss function, specifically, an adversarial loss function, can be introduced, that is, when the parameter values in the image signal processing model are fine-tuned according to the first loss function in step 405, a loss function can be calculated according to the following formula 19, and the parameter values in the trained image signal processing model are adjusted according to the calculated loss function: Loss finetrain =Loss pretrain +α×loss UV +β×loss GAN (19)
[0107] In the fine-tuning stage, it is found that the numerical magnitude of the adversarial loss function is 1e-1, in order to make its numerical magnitude less than that of other loss functions, in the embodiment, the weight values of the adversarial loss function are respectively set to 0.0001, 0.001, 0.01 and 0.1. The PSNR, SSIM and UV values calculated on the test set are shown in Table 2 below, and it is found that when the weight value is set to 0.0001, the adversarial loss function has less influence on the image signal processing model (i.e., the second row of data is close to the first row), and when the weight value is set to 0.001, not only can the UV color error be further reduced, but also the PSNR and SSIM values can be further improved.
[0108] Table 2
[0109] The server provided by the embodiment of the present application has a structure diagram as shown in FIG. 9, and the server can have great differences due to different configurations or performances, and can include one or more central processing units (CPU) 20 (for example, one or more processors) and a memory 21, one or more storage media 22 (for example, one or more mass storage devices) storing application programs 221 or data 222. The memory 21 and the storage medium 22 can be temporary storage or persistent storage. The programs stored in the storage medium 22 can include one or more modules (not shown in the figure), and each module can include a series of instruction operations in the server. Further, the central processing unit 20 can be configured to communicate with the storage medium 22 and execute a series of instruction operations in the storage medium 22 on the server.
[0110] Specifically, the application 221 stored in the storage medium 22 comprises an application of image signal processing, and the application is adapted to be loaded and executed by the processor to perform the image signal processing method as described above, which will not be repeated here. Further, the central processor 20 can be configured to communicate with the storage medium 22, and perform a series of operations corresponding to the application of image signal processing stored in the storage medium 22 on the server.
[0111] The server can further comprise one or more power supplies 23, one or more wired or wireless network interfaces 24, one or more input / output interfaces 25, and / or one or more operating systems 223, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0112] The steps performed by the image signal processing system in the above method embodiments can be based on the structure of the server shown in FIG. 9.
[0113] Further, the embodiments of the present application also provide a computer readable storage medium storing a plurality of computer programs, and the computer programs are adapted to be loaded and executed by the processor to perform the image signal processing method performed by the image signal processing system as described above.
[0114] Those skilled in the art can understand that all or part of the steps in the above-mentioned embodiments can be instructed by a program to relevant hardware, and the program can be stored in a computer readable storage medium, which can include read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.
[0115] The above describes in detail the image signal processing method, system, storage medium and server provided by the embodiments of the present application. The principles and implementation manners of the present application are described by applying specific examples in this paper. The above embodiment descriptions are only used to help understand the method of the present application and its core idea. Meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation manners and application ranges will be changed. In summary, the content of the present description should not be understood as a limitation of the present application.
Claims
1. An image signal processing method comprising: obtaining an original capture format image; performing first image signal processing on the original capture format image to obtain a first-processed image; the first image signal processing is a processing related to a non-shared parameter of a capturing device; wherein the first image signal processing comprises black level correction and normalization, white balance processing, demosaicing and color correction; performing second image signal processing on the first-processed image according to a preset image signal processing model to obtain a display format image, the image signal processing model being used to perform the second image signal processing on the first-processed image, the second image signal processing being other image signal processing than the first image signal processing; the second image signal processing comprising color restoration, detail texture recovery, brightness adjustment, contrast adjustment and noise reduction; wherein the preset image signal processing model is obtained by: determining an image signal processing initial model; determining training samples, the training samples comprising a plurality of groups of sample images, each group of sample images comprising an original capture sample image and a corresponding display sample image thereof; adjusting the image signal processing initial model according to the image signal processing initial model and the training samples to obtain the preset image signal processing model.
2. The image signal processing method of claim 1, wherein the step of determining the image signal processing initial model comprises: determining the image signal processing initial model to comprise a plurality of layers of image segmentation convolution networks, each layer of the image segmentation convolution networks comprising a wavelet transform down-sampling layer, a wavelet transform up-sampling layer, a convolution and activation layer and a residual module; wherein any layer of the image segmentation convolution networks is nested in another layer of the image segmentation convolution networks.
3. The image signal processing method of claim 2, wherein the step of adjusting the image signal processing initial model according to the image signal processing initial model and the training samples comprises: performing the first image signal processing on each of the original capture sample images to obtain a first-processed image, and then performing the second image signal processing on the first-processed image by the image signal processing initial model to obtain a corresponding display format image; calculating a loss function related to the image signal processing initial model according to each of the display format images obtained by the image signal processing initial model and the corresponding display sample images in the training samples; adjusting parameter values in the image signal processing initial model according to the loss function.
4. The image signal processing method of claim 3, wherein the step of calculating the loss function related to the image signal processing initial model according to each of the display format images obtained by the image signal processing initial model and the corresponding display sample images in the training samples comprises: calculating a pixel loss function, the pixel loss function comprising an average value of absolute differences between the display format image obtained by the image signal processing initial model and the corresponding display sample image in the training samples. calculating a perceptual loss function, the perceptual loss function comprising a mean square value of a difference between high-level features of the display format image obtained by the image signal processing initial model and high-level features of the corresponding display sample image in the training sample; calculating a structural loss function, the structural loss function comprising a structural similarity index between the display format image obtained by the image signal processing initial model and the corresponding display sample image in the training sample; calculating a first loss function as the loss function related to the image signal processing initial model according to the pixel loss function, the perceptual loss function and the structural loss function.
5. The image signal processing method of claim 4, wherein the method further comprises: calculating a chroma channel loss function, the chroma channel loss function being used to indicate a difference between chroma channels of the display format image obtained by the image signal processing initial model and the corresponding display sample image in the training sample; calculating an adversarial loss function, the adversarial loss function comprising a probability that the display format image obtained by the image signal processing initial model is judged as the corresponding display sample image by a discriminator; calculating a second loss function according to the chroma channel loss function and the adversarial loss function; fine-tuning parameter values in the image signal processing initial model according to the second loss function to obtain the preset image signal processing model.
6. An image signal processing method, comprising: obtaining an original shooting format image; performing first image signal processing on the original shooting format image to obtain a first-processed image; the first image signal processing is processing related to non-shared parameters of a shooting device; performing second image signal processing on the first-processed image according to a preset image signal processing model to obtain a display format image, the image signal processing model being used to perform the second image signal processing on the first-processed image, the second image signal processing being image signal processing other than the first image signal processing.
7. The image signal processing method of claim 6, wherein the first image signal processing comprises: black level correction and normalization, white balance processing, demosaicing and color correction; the second image signal processing comprises color restoration, detail texture restoration, brightness adjustment, contrast adjustment and noise reduction.
8. The image signal processing method of claim 6, wherein the method further comprises: determining an image signal processing initial model; determining training samples, the training samples comprising a plurality of groups of sample images, each group of sample images comprising an original shooting sample image and a corresponding display sample image thereof; adjusting the image signal processing initial model according to the image signal processing initial model and the training samples to obtain the preset image signal processing model.
9. The image signal processing method of claim 8, wherein the step of determining the image signal processing initial model specifically comprises: determining that the image signal processing initial model comprises a plurality of layers of image segmentation convolutional networks, each layer of the image segmentation convolutional networks comprising a wavelet transform down-sampling layer, a wavelet transform up-sampling layer, a convolution and activation layer and a residual module. Any one of the multi-layer image segmentation convolutional networks is nested in another image segmentation convolutional network. 10.The image signal processing method of claim 9, wherein the adjusting the image signal processing initial model according to the image signal processing initial model and the training samples comprises: obtaining a first processed image by performing the first image signal processing on each of the original shooting sample images; obtaining a corresponding display format image by performing the second image signal processing on the first processed image according to the image signal processing initial model; calculating a loss function related to the image signal processing initial model according to each of the display format images obtained by the image signal processing initial model and a corresponding display sample image in the training samples; and adjusting a parameter value in the image signal processing initial model according to the loss function. 11.The image signal processing method of claim 10, wherein the calculating the loss function related to the image signal processing initial model according to each of the display format images obtained by the image signal processing initial model and a corresponding display sample image in the training samples comprises: calculating a pixel loss function, wherein the pixel loss function comprises an average value of an absolute difference between the display format image obtained by the image signal processing initial model and the corresponding display sample image in the training samples; calculating a perception loss function, wherein the perception loss function comprises a mean square value of a difference between a high-order feature of the display format image obtained by the image signal processing initial model and a high-order feature of the corresponding display sample image in the training samples; and calculating a structure loss function, wherein the structure loss function comprises a structural similarity index between the display format image obtained by the image signal processing initial model and the corresponding display sample image in the training samples; and calculating a first loss function as the loss function related to the image signal processing initial model according to the pixel loss function, the perception loss function and the structure loss function. 12.The image signal processing method of claim 11, wherein the method further comprises: calculating a chroma channel loss function, wherein the chroma channel loss function is used to indicate a difference between a chroma channel of the display format image obtained by the image signal processing initial model and a chroma channel of the corresponding display sample image in the training samples; calculating an adversarial loss function, wherein the adversarial loss function comprises a probability that the display format image obtained by the image signal processing initial model is identified as the corresponding display sample image by a discriminator; calculating a second loss function according to the chroma channel loss function and the adversarial loss function; and fine-tuning the parameter value in the image signal processing initial model according to the second loss function to obtain the preset image signal processing model. 13.A computer readable storage medium, wherein the computer readable storage medium stores a plurality of computer programs, and the computer programs are adapted to be loaded and executed by a processor to implement the image signal processing method of claim 1.
Citation Information
Patent Citations
Model training and parameter prediction method and device, electronic equipment and storage medium
CN116187396A
Image quality improvement method based on layer-by-layer training pyramid network
CN117058062A
Image signal processor, image processing system, and operating method of image signal processor
US20200221027A1
Method and device for joint denoising and demosaicing using neural network
US20220164926A1