Image processing method and device and electronic equipment

By generating a personalized color correction matrix based on image scene recognition and color temperature and brightness parameters, the problem of poor color correction effect in existing technologies is solved, and better matching of image color with scene is achieved.

CN120953145APending Publication Date: 2025-11-14VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511070945.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

In existing technologies, color correction methods based on brightness and color temperature parameters are difficult to adapt to various types of images, resulting in poor color correction effects.

Method used

By performing scene recognition on the image and combining color temperature, brightness, and scene category to generate a personalized color correction matrix, color correction processing is performed.

Benefits of technology

Optimize the color temperature and brightness of the image to improve the color correction effect, making the optimized image colors more compatible with the scene.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953145A_ABST
    Figure CN120953145A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method and device and electronic equipment, and belongs to the technical field of image processing. The image processing method comprises the following steps: performing scene recognition on the first image to obtain a scene category of the first image; generating a first color correction matrix of the first image based on the first image, the color temperature parameter of the first image, the brightness parameter of the first image and the scene category; and performing color correction processing on the first image based on the first color correction matrix to obtain a second image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, specifically to an image processing method, apparatus, and electronic device. Background Technology

[0002] In related technologies, color correction is often required for images during image processing. Common color correction methods typically involve determining a Color Correction Matrix (CCM) based on the luminance (lux) and color temperature (cct) parameters of the image to be processed. Then, color correction is performed on the image based on the determined CCM. However, due to the diverse types of images, it is difficult to guarantee good correction results for all image types using the aforementioned color correction methods. Summary of the Invention

[0003] This application provides an image processing method, apparatus, and electronic device. Since the first color correction matrix is ​​determined based on the first image, the color temperature parameters of the first image, the brightness parameters of the first image, and the scene category, the color temperature and brightness of the first image can be optimized during the color correction process based on the first color correction matrix. This also makes the color of the optimized second image more compatible with the image scene in the second image, thereby optimizing the color correction effect of the second image.

[0004] Firstly, an image processing method is provided, including:

[0005] Scene recognition is performed on the first image to obtain the scene category of the first image;

[0006] Based on the first image, the color temperature parameter of the first image, the brightness parameter of the first image, and the scene category, a first color correction matrix for the first image is generated.

[0007] Based on the first color correction matrix, the first image is subjected to color correction processing to obtain the second image.

[0008] Secondly, an image processing apparatus is provided, comprising:

[0009] The recognition module is used to perform scene recognition on the first image to obtain the scene category of the first image;

[0010] The generation module is used to generate a first color correction matrix for the first image based on the first image, the color temperature parameters of the first image, the brightness parameters of the first image, and the scene category.

[0011] The correction module is used to perform color correction processing on the first image based on the first color correction matrix to obtain the second image.

[0012] Thirdly, an electronic device is provided, comprising a processor and a memory, wherein the memory stores a program or instructions executable on the processor, the program or instructions, when executed by the processor, perform the steps of the method described in the first aspect.

[0013] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0014] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the steps of the method described in the first aspect.

[0015] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the steps of the method described in the first aspect.

[0016] In this embodiment, the determined first color correction matrix is ​​based on the first image, the color temperature parameter of the first image, the brightness parameter of the first image, and the scene category of the first image. That is, in the process of determining the color correction matrix, in addition to considering the color temperature and brightness of the first image, the scene category of the first image is also considered. Therefore, in the process of color correction of the first image based on the first color correction matrix, the color temperature and brightness of the first image can be optimized, and the optimized color of the second image can be better matched with the image scene in the second image, thereby optimizing the color correction effect of the second image. Attached Figure Description

[0017] Figure 1 This is a schematic flowchart of an image processing method provided by some embodiments of this application;

[0018] Figure 2 This is a schematic flowchart of an image processing method provided by some embodiments of this application;

[0019] Figure 3 This is a schematic flowchart of an image processing method provided by some embodiments of this application;

[0020] Figure 4 This is a flowchart illustrating the process of human visual sensory annotation of an image to be processed, provided in some embodiments of this application;

[0021] Figure 5 This is a flowchart illustrating the image processing procedure of the first sub-network provided in some embodiments of this application;

[0022] Figure 6 This is a flowchart illustrating the image processing procedure of the first sub-network provided in some embodiments of this application;

[0023] Figure 7 This is a flowchart illustrating the image processing procedure performed by the second sub-network according to some embodiments of this application;

[0024] Figure 8 This is a flowchart illustrating the image processing procedure performed by the third sub-network according to some embodiments of this application;

[0025] Figure 9 This is a schematic flowchart illustrating the image processing process of the initial generation network provided in some embodiments of this application;

[0026] Figure 10 This is a flowchart illustrating the process of fusing the input map input3 and the feature map feature4 according to some embodiments of this application;

[0027] Figure 11 This is a schematic diagram of the structure of an image processing apparatus provided in some embodiments of this application;

[0028] Figure 12 Schematic diagrams of the structure of electronic devices provided for some embodiments of this application;

[0029] Figure 13 A schematic diagram of the hardware structure of an electronic device provided for some embodiments of this application. Detailed Implementation

[0030] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0031] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, not limited in number; for example, the first object can be one or more. Furthermore, "or" in this application indicates at least one of the connected objects. For example, the scope of protection for "A or B" covers at least three scenarios: Scenario 1: including A but not B; Scenario 2: including B but not A; Scenario 3: including both A and B. In addition, the terms "A and / or B," "at least one of A and B," and "at least one of A or B" also cover at least the above three scenarios. The character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0032] The terminology used in the implementation section of this application is only for explaining specific embodiments of this application and is not intended to limit this application. The terminology involved in the embodiments of this application is explained below.

[0033] Scene category can refer to the scene type of an image, and the scene category can be preset as needed. For example, in some embodiments of this application, the scene category can be divided into the following seven categories: 0 represents a shopping mall scene, 1 represents an outdoor sunlight scene, 2 represents an indoor office scene, 3 represents a sunset scene, 4 represents a complex light source scene, 5 represents a green plant scene, and 6 represents other scenes. As another example, in some embodiments of this application, the scene category is divided into the following two types: face scenes and non-face scenes.

[0034] Color temperature parameter refers to the color temperature of an image, and its unit is Kelvin (K).

[0035] The brightness parameter can refer to the brightness of an image, and its unit is lux (Lux).

[0036] The Color Correction Matrix (CCM) is a matrix used to calibrate the RGB values ​​of an image.

[0037] Color correction processing refers to the process of calibrating the RGB values ​​of an image according to CCM.

[0038] The generative network can be a pre-trained network model that can be used to generate a corresponding color correction matrix based on the input image.

[0039] The image processing method, apparatus, and electronic device provided in this application will be described in detail below with reference to the accompanying drawings and through some embodiments and application scenarios.

[0040] Please see Figure 1 , Figure 1 This is a flowchart illustrating an image processing method provided in an embodiment of this application. The image processing method includes the following steps:

[0041] Step 101: Perform scene recognition on the first image to obtain the scene category of the first image.

[0042] The first image can be an image that requires color correction in various scenarios, such as an image acquired during shooting or video recording. In some embodiments of this application, the first image is the initial (raw) image output by the demosaic module of the image signal processing (ISP) link in an electronic device.

[0043] The scene recognition of the first image described above can be performed based on a pre-trained network model. Accordingly, the scene categories can be pre-set as needed. For example, in some embodiments of this application, the scene categories can be divided into the following seven categories: 0 represents a shopping mall scene, 1 represents an outdoor sunny scene, 2 represents an indoor office scene, 3 represents a sunset scene, 4 represents a complex light source scene, 5 represents a green plant scene, and 6 represents other scenes. Thus, the scene category obtained from the scene recognition of the first image can refer to a value from 0 to 6.

[0044] Step 102: Based on the first image, the color temperature (cct) parameter of the first image, the brightness (lux) parameter of the second image, and the scene category, generate a first color correction matrix for the first image.

[0045] The cct parameter described above is used to characterize the color temperature of the first image. The lux parameter described above is used to characterize the brightness of the first image.

[0046] It is understood that the aforementioned first color correction matrix can be simultaneously related to the color temperature parameter, the brightness parameter, and the scene category. That is, different color parameter ranges can correspond to different first color correction matrices, different brightness parameter ranges can correspond to different first color correction matrices, and different scene categories can correspond to different first color correction matrices.

[0047] Step 103: Based on the first color correction matrix, perform color correction processing on the first image to obtain the second image.

[0048] The second image can be obtained by performing color correction processing on the first image using the determined first color correction matrix, based on the existing CCM module in the electronic device.

[0049] In this embodiment, since the determined first color correction matrix is ​​determined based on the first image, the color temperature parameter of the first image, the brightness parameter of the first image, and the scene category of the first image, in the process of determining the color correction matrix, in addition to considering the color temperature and brightness of the first image, the scene category of the first image is also considered. Therefore, in the process of color correction of the first image based on the first color correction matrix, the color temperature and brightness of the first image can be optimized, and the optimized color of the second image can be better matched with the image scene in the second image, thereby optimizing the color correction effect of the second image.

[0050] Optionally, generating a first color correction matrix for the first image based on the first image, the color temperature parameter of the first image, the brightness parameter of the first image, and the scene category includes:

[0051] Among the S generation networks, the generation network that matches the color temperature parameter, the brightness parameter, and the scene category is determined; wherein, different generation networks correspond to different parameter sets, the parameter set includes a color temperature range, a brightness range, and a scene category, and S is an integer greater than 1;

[0052] The first image is input into the generation network that matches the color temperature parameter, the brightness parameter, and the scene category to generate a first color correction matrix.

[0053] The aforementioned S generative networks can be pre-trained. It is understood that the generative networks can be networks trained based on images whose image parameters belong to their corresponding parameter groups. Thus, during training, the generative networks learn the correspondence between the image parameters of the input image and the corresponding correction matrix, and optimize the network parameters. This allows the trained generative networks to determine the corresponding color correction matrix based on the optimized network parameters when they subsequently receive images whose image parameters belong to their corresponding parameter groups, thereby realizing the generation process of the first color correction matrix.

[0054] It is understood that the color temperature parameters, brightness parameters, and scene category of the first image belong to the parameter group corresponding to the generative network that matches the color temperature parameters, brightness parameters, and scene category.

[0055] In this embodiment, by pre-obtaining S generation networks, and with different parameter sets corresponding to different generation networks, it can be ensured that each generation network can generate a more suitable first color correction matrix when generating a first color correction matrix for an image whose corresponding parameter set matches. Based on this, in this embodiment, by determining the generation network that matches the color temperature parameter, the brightness parameter, and the scene category among the S generation networks, and inputting the first image into the generation network that matches the color temperature parameter, the brightness parameter, and the scene category to generate the first color correction matrix, it is beneficial to improve the matching degree between the determined first color correction matrix and the first image. Thus, color correction of the first image based on the first color correction matrix is ​​beneficial to improving the color accuracy in the corrected second image.

[0056] Optionally, before determining the generation network that matches the color temperature parameter, the brightness parameter, and the scene category among the S generation networks, the method further includes:

[0057] The range of values ​​for the color temperature parameter is divided to obtain M color temperature ranges;

[0058] The range of values ​​for the brightness parameter is divided into N brightness ranges;

[0059] Determine T candidate scene categories;

[0060] Based on the M color temperature ranges, the N brightness ranges, and the T candidate scene categories, S parameter groups are obtained;

[0061] Wherein, the T candidate scene categories include the scene category of the first image, and M, N and T are all integers greater than 1; S = M × N × T, and the S parameter groups correspond one-to-one with the S generator networks.

[0062] The method of dividing the color temperature parameter and the number of color temperature ranges can be set as needed. For example, in some embodiments of this application, the value of M is 5, and the five color temperature ranges are: below 2570, 3350 to 2640, 3450 to 4600, 4700 to 6500, and above 6500. As another example, in some embodiments of this application, the value of M is 6, and the six color temperature ranges are: below 2570, 3350 to 2640, 3450 to 4600, 4700 to 5500, 5500 to 6500, and above 6500.

[0063] The method of dividing the brightness range and the number of brightness ranges can be set as needed. For example, in some embodiments of this application, the value of N is 3, and the three brightness ranges are 0 to 180, 190 to 385, and above 410. As another example, in some embodiments of this application, the value of N is 4, and the four brightness ranges are 0 to 180, 190 to 290, 290 to 385, and above 410.

[0064] The value of T can be set as needed; for example, T can be 5, 6, or other numbers.

[0065] In some embodiments of this application, the value of M is 5, the value of N is 3, and the value of T is 6. In this case, S = M × N × T = 90. In other embodiments of this application, the value of M is 5, the value of N is 3, and the value of T is 5. In this case, S = M × N × T = 75.

[0066] The above-mentioned method of obtaining S parameter groups based on the M color temperature ranges, the N brightness ranges, and the T candidate scene categories can be described as follows: the M color temperature ranges, the N brightness ranges, and the T candidate scene categories are arranged and combined, with each combination including a color temperature range, a brightness range, and a candidate scene category. In this way, M×N×T combinations can be obtained, and each combination is used as a parameter group, thereby obtaining S parameter groups.

[0067] In this embodiment, the range of values ​​for the color temperature parameter is divided into M color temperature ranges, and the range of values ​​for the brightness parameter is divided into N brightness ranges, thus determining T candidate scene categories. Then, based on the M color temperature ranges, the N brightness ranges, and the T candidate scene categories, S parameter groups are obtained, thereby realizing the generation process of S parameter groups. In this way, the images of each parameter group can be determined according to the parameter groups, and the corresponding generative networks can be trained using the determined images, thereby realizing the training process of S generative networks.

[0068] Optionally, the generating network includes a first sub-network, a second sub-network, and a third sub-network, wherein the first sub-network is used to identify differences between different images, the second sub-network is used to generate a color correction matrix, and the third sub-network is used to identify the scene category of the image;

[0069] The step of inputting the first image into the generation network that matches the color temperature parameter, the brightness parameter, and the scene category to generate a first color correction matrix includes:

[0070] The first image is subjected to feature extraction through the third sub-network to obtain a first feature map;

[0071] The second feature map is obtained by extracting features from the first image through the second sub-network.

[0072] The second feature map is fused with the first image to obtain the third feature map;

[0073] The third feature map is input into the first sub-network to perform sensory difference recognition and output sensory difference information. The sensory difference information is used to characterize the sensory difference between the third feature map and the fourth feature map. The fourth feature map is the feature map of the image that matches the image content of the first image in the training data corresponding to the generator network. Alternatively, if the first image is a frame in video data, the fourth feature map is the feature map of the previous frame of the first image in the video data.

[0074] The first feature map, the second feature map, the third feature map, and the sensory difference information are input into the fully connected layer of the second sub-network, and the first color correction matrix is ​​output.

[0075] The first feature map can be the input of the fully connected layer of the third sub-network, that is... Figure 9 Feature map 7 in the illustrated embodiment. Please refer to [link / reference]. Figure 9 The second feature map can be the feature map output by the convolutional layer 10 of the second sub-network, that is... Figure 9 Feature map 4 in the illustrated embodiment. The third feature map is... Figure 9 Feature map 5 in the illustrated embodiment. The aforementioned sensory difference information is... Figure 9 Output1 in the illustrated embodiment. The first color correction matrix mentioned above is... Figure 9 Output2 in the illustrated embodiment.

[0076] In some embodiments of this application, the aforementioned sensory difference information may be sensory difference information between the first image identified by the first sub-network and historical images from the training phase, wherein the historical images may be target images in the training data whose image content and scene are similar to the first image. In other embodiments of this application, the first image is a frame in video data, and the fourth feature map is a feature map of the previous frame of the first image in the video data.

[0077] In this embodiment, by inputting the first feature map, the second feature map, the third feature map, and the sensory difference information into the fully connected layer of the second sub-network, the first color correction matrix output by the second sub-network is obtained. This helps to improve the matching degree between the generated first color correction matrix and the first image, thereby improving the color correction effect of the second image.

[0078] Optionally, before determining the generation network that matches the color temperature parameter, the brightness parameter, and the scene category among the S generation networks, the method further includes:

[0079] The first initial network, the second initial network, and the third initial network are trained respectively to obtain the first sub-network, the second sub-network, and the third sub-network;

[0080] Based on the first, second, and third sub-networks obtained through training, an initial generative network is constructed.

[0081] The initial generator network is trained to obtain S generator networks.

[0082] Understandably, the first, second, and third sub-networks can be pre-trained individually based on the training data. Then, the model parameters from these trained sub-networks are used to update the individual sub-networks in the initial generator network. This ensures the initial generator network has good initial performance, thereby reducing the number of iterations required to train multiple generator networks subsequently.

[0083] The above-mentioned training of the initial generator network to obtain the S generator networks can refer to: training the initial generator network based on S types of training data respectively to obtain the S generator networks, wherein the S generator networks correspond one-to-one with the S types of training data, and the generator network is obtained by training the initial generator network based on the corresponding type of training data.

[0084] In this embodiment, the first sub-network, the second sub-network, and the third sub-network are obtained through pre-training, and an initial generator network is constructed based on the trained first sub-network, the second sub-network, and the third sub-network. In this way, it can be ensured that the initial generator network has good initial performance, thereby reducing the number of iterations required in the subsequent training of S generator networks.

[0085] Optionally, the first sub-network is a network obtained by training the first initial network based on the first training dataset, wherein the first training dataset includes multiple first training data, and the first training data includes an image to be processed, a reference image, and difference annotation information of the two images;

[0086] The second sub-network is a network trained on the second initial network based on the second training dataset, wherein the second training dataset includes multiple second training data, and the second training data includes an image to be processed and a second color correction matrix;

[0087] The third sub-network is a network trained based on a third training dataset, wherein the third training dataset includes multiple third training data sets, and the third training data set includes an image to be processed and a scene category label.

[0088] The number of first training data points included in the first training dataset can be set as needed. For example, in some embodiments of this application, the number of first training data points included in the first training dataset can be greater than 100, or greater than 1000. Correspondingly, the number of second training data points included in the second training dataset can be set as needed. For example, in some embodiments of this application, the number of second training data points included in the second training dataset can be greater than 200, or greater than 1000. The number of third training data points included in the third training dataset can be set as needed. For example, in some embodiments of this application, the number of third training data points included in the third training dataset can be greater than 500, or greater than 2000.

[0089] Please see Figure 4 The first sub-network is constructed to determine the color difference between the image to be processed and the reference image and represent it in numerical form. While quantifying the color difference between the two images, the subjective perception of the human eye is introduced. By using the light perception of the human eye to annotate the difference between the image to be processed and the reference image, the annotated difference information can represent the human eye difference between the image to be processed and the reference image, so that the generated difference annotation information is closer to the difference perceived by the human eye.

[0090] The image processing flow of the first sub-network mentioned above is as follows:

[0091] Step 2021: Input the image to be processed as input1 and the reference image as input2 into the first sub-network. The size of the image to be processed is 256*256*3, and the size of the reference image is 256*256*3.

[0092] Step 2022: Subtract the corresponding channels of input1 and input2 to get the size of feature1, which is 256*256*3. This step is to obtain the difference between the two images so that the final difference value can be better simulated later.

[0093] Step 2023: Concatenate input1, input2 and feature map feature1 to obtain feature map feature2; feed feature map feature2 into convolutional layer 1 to output feature map a1, and feed feature map a1 into the ReLU activation function layer to obtain feature map a2.

[0094] Specifically, feature1, input1, and input2 are concatenated, meaning feature1 is directly appended with input1 and input2 to form feature2, which has a size of 256*256*9. Feature2 is then fed into the first convolutional layer 1, with a layer size of 1×1×9 and 20 kernels, outputting a feature map a1 with a size of 256×256×20. The resulting feature data a1 is then fed into a ReLU activation function layer to process the feature map output from the first convolutional layer, i.e., performing a ReLU function operation on the feature map to obtain feature map a2. The ReLU function is:

[0095]

[0096] Where x is the feature map and y is the output after passing through the ReLU activation function layer.

[0097] Step 2024: Feed feature map a2 into the convolution channel of a multi-layer convolutional layer combination, and output the corresponding feature map a3.

[0098] The result of step 3 is then fed into a second convolutional + ReLU layer. The second convolutional layer has a kernel size of 1*1*20, a stride of 1, and 20 kernels, resulting in a feature map size of 256*256*20. The output of the second convolutional + ReLU layer is then fed into a third convolutional + ReLU layer. The third convolutional layer has a kernel size of 4*4*20 and 10 kernels, resulting in a feature map size of 64*64*10. The output of the third convolutional + ReLU layer is then fed into a fourth convolutional + ReLU layer. The fourth convolutional layer has a kernel size of 1*1*10, a stride of 1, and 10 kernels, resulting in a feature map size of 64*64*10. The output of the fourth convolutional layer + ReLU layer is then fed into the fifth convolutional layer + ReLU layer. The fifth convolutional layer has a kernel size of 4*4*10, a stride of 1, and 8 kernels, resulting in a feature map size of 16*16*8. The output of the fifth convolutional layer + ReLU layer is then fed into the sixth convolutional layer + ReLU layer. The sixth convolutional layer has a kernel size of 1*1*8, a stride of 1, and 10 kernels. The output of the sixth convolutional + ReLU layer is fed into the seventh convolutional + ReLU layer. The kernel size of the seventh convolutional layer is 1*1*6, the stride is 1, and the number of kernels is 6, resulting in a feature map size of 16*16*6. The output of the seventh convolutional + ReLU layer is then fed into the eighth convolutional + ReLU layer. The kernel size of the eighth convolutional layer is 4*4*6, the stride is 1, and the number of kernels is 6, resulting in a feature map size of 4*4*6. The output of the +ReLU layer is then fed into the ninth convolutional +ReLU layer. The kernel size of the ninth convolutional layer is 4*4*6, the stride is 1, and the number of kernels is 3, resulting in a feature map of size 4*4*3. The output of the ninth convolutional +ReLU layer is then fed into the tenth convolutional +ReLU layer. The kernel size of the tenth convolutional layer is 1*1*3, the stride is 1, and the number of kernels is 3, resulting in a feature map of size 4*4*3. The feature map of size 4*4*3 output by the tenth convolutional +ReLU layer is the aforementioned feature map a3.

[0099] Step 2025: Upsample the feature map a3 obtained in step 2024 to obtain feature map feature3, which has a size of 10*10*3.

[0100] Step 2026: Feed the feature map feature3 into the fully connected layer to obtain the output value output1. The input image size of the fully connected layer is required to be 10*10*3, and the dimension of the output value of the fully connected layer is 1. output1 is the final difference value, i.e., the sensory difference information.

[0101] Step 203: Construct the second sub-network, wherein step 203 may be a step performed before the above step "training the first initial network, the second initial network and the third initial network respectively to obtain the first sub-network, the second sub-network and the third sub-network".

[0102] After step 202, the construction of the second sub-network will use the first sub-network constructed in the previous layer. The first sub-network is to better fit different images after taking into account the differences in human visual perception. The second sub-network is to obtain a CCM matrix that is closer to the reference image through network training. The first sub-network is introduced at the same time as the construction of the second sub-network in order to take into account the changes in human visual perception in the network construction, which is more conducive to network convergence and iteration.

[0103] The image processing flow of the second sub-network mentioned above is as follows:

[0104] Step 2031: Input the image to be processed as input3 into the second sub-network, where the size of the image to be processed is 256*256*3.

[0105] Step 2032: Feed the input map input3 into a convolutional layer and a ReLU activation function layer to obtain the output feature map b1.

[0106] The input image (input3) is fed into the first convolutional layer (layer 1), which has a size of 1×1×3 and 9 kernels, outputting a 256×256×9 feature map. This feature map is then processed by a ReLU activation function layer, which applies the ReLU function to the feature map output from the first convolutional layer. The ReLU function is as follows:

[0107]

[0108] Step 2033: Feed feature map b1 into multiple convolution channels to obtain output feature map b2.

[0109] The result b1 from the previous step is fed into the second convolutional + ReLU layer. The kernel size of the second convolutional layer is 1*1*9, the stride is 1, and the number of kernels is 9, resulting in a feature map size of 256*256*9. The output of the second convolutional + ReLU layer is then fed into the third convolutional + ReLU layer. The kernel size of the third convolutional layer is 4*4*20, and the number of kernels is 10, resulting in a feature map size of 64*64*10. The output of the third convolutional + ReLU layer is then fed into the fourth convolutional + ReLU layer. The fourth convolutional layer (Layer 4 + ReLU) has a kernel size of 1*1*10, a stride of 1, and 10 kernels, resulting in a feature map of size 64*64*10. The output of this layer is then fed into a fifth convolutional layer (Layer 5 + ReLU), with a kernel size of 4*4*10, a stride of 1, and 8 kernels, resulting in a feature map of size 16*16*8. Finally, the output of this layer is fed into a sixth convolutional layer (Layer 6 + ReLU), with the kernel size of 1*1*10, a stride of 1, and 8 kernels, resulting in a feature map of size 16*16*8. The kernel size is 1*1*8, the stride is 1, and the number of kernels is 6, resulting in a feature map size of 16*16*6. The output of the sixth convolutional + ReLU layer is then fed into the seventh convolutional + ReLU layer, which also has a kernel size of 1*1*6, a stride of 1, and 6 kernels, resulting in another feature map size of 16*16*6. Finally, the output of the seventh convolutional + ReLU layer is fed into the eighth convolutional + ReLU layer, which has a kernel size of 4*4*6, a stride of 1, and 6 kernels. The number of kernels is 6, resulting in feature maps of size 4*4*6. The output of the eighth convolutional + ReLU layer is then fed into the ninth convolutional + ReLU layer. The kernel size of the ninth convolutional layer is 4*4*6, the stride is 1, and the number of kernels is 3, resulting in feature maps of size 4*4*3. The output of the ninth convolutional + ReLU layer is then fed into the tenth convolutional + ReLU layer. The kernel size of the tenth convolutional layer is 1*1*3, the stride is 1, and the number of kernels is 3, resulting in feature map b2 of size 4*4*3.

[0110] Step 2034: Feed feature map b2 into the convolutional layer, and then perform downsampling on the output to obtain feature map feature4.

[0111] In this process, the final feature map b2 from step 2033 is input into the eleventh convolutional layer for convolution processing to obtain a feature map of size 4*4*1. Then, the feature map of size 4*4*1 is downsampled to obtain a feature map feature4 of size 3*3*1. The kernel size of the eleventh convolutional layer is 1*1*3, the stride is 1, and the number of kernels is 1.

[0112] Step 2035: Fuse the input map input3 with the feature map feature4 to obtain the feature map feature5.

[0113] In this process, feature4 is multiplied by input3 to simulate the image after the initial CCM matrix is ​​generated. The multiplication method is to multiply each column of the feature3*3 matrix with the value of the corresponding channel of input3 and then add them together to obtain the feature map feature5.

[0114] For example, see Figure 10 The values ​​of each row and column of feature map feature4 are denoted as f1-9; input1 is 256*256*3, with the three channels corresponding to the r, g, and b channels respectively; multiplying them together yields a 256*256*3 feature map. Figure 5 The first channel of the feature map, corresponding to f1, f2, and f3, is multiplied by the values ​​of the corresponding channels r, g, and b, and then summed. The second channel of the feature map, corresponding to f4, f5, and f6, is multiplied by the values ​​of the corresponding channels r, g, and b, and then summed. The third channel of the feature map, corresponding to f7, f8, and f9, is multiplied by the values ​​of the corresponding channels r, g, and b, and then summed.

[0115] Step 2036: Input the feature map feature5 and the 256*256*3 RGB image input2 into the first sub-network constructed in the previous step to obtain the output output1, which has a size of 1*1.

[0116] Step 2037: Upsample feature map 4 to a 10*10*1 feature map; downsample feature map 5 to a 10*10*3 feature map; upsample output 1 to a 10*10*1 feature map; and concatenate the three downsampled feature maps to obtain the feature map. Figure 10 *10*5.

[0117] Step 2038: Obtain the features from step 2037 Figure 10 The 10*5 kernel size is fed into the twelfth convolutional layer for convolution processing to obtain a feature map of size 10*10*3. Then, the 10*10*3 feature map is fed into the fully connected layer to obtain a feature map of size 9*1*1 output by the fully connected layer. The kernel size of the twelfth convolutional layer is 1*1*5, the stride is 1, and the number of kernels is 3.

[0118] Step 2039: Resize the 9*1*1 features obtained in step 2038 to 3*3 data to obtain the final output result output2.

[0119] Step 204: Construct the third sub-network, wherein step 204 may be a step performed before the above step "training the first initial network, the second initial network and the third initial network respectively to obtain the first sub-network, the second sub-network and the third sub-network".

[0120] The image processing flow of the third sub-network mentioned above is as follows:

[0121] Step 2041: The image to be processed is used as input4 and concatenated with lux and cct information to obtain feature map feature6.

[0122] The image to be processed has a size of 256*256*3, and lux and cct are the brightness and color temperature parameters of the image to be processed. The lux and cct are upsampled to 256*256*1 data, and then concat with input4 to obtain feature6256*256*5.

[0123] Step 2042: Feed the feature map feature6 into the multi-layer convolution channel to obtain the output feature map c1.

[0124] The network structure consists of 10 groups, with feature6 fed into a convolutional layer + ReLU layer for training. The first group of convolutional layers has a kernel size of 1×1×5 and 10 kernels, with an output feature map of 256*256*10. The second group of convolutional layers has a kernel size of 1×1×10 and 20 kernels, with an output feature map of 256*256*20. The third to eighth groups of convolutional layers have a kernel size of 1×1×20 and 8 kernels, with an output feature map of 256*256*8. The ninth and tenth groups of convolutional layers have a kernel size of 1×1×8 and 3 kernels, with an output feature map of 10×10×3.

[0125] Step 2043: The feature map c1 obtained in the previous step is fed into a convolutional layer. The convolutional layer has a 1×1×1 kernel and a kernel count of 1, and outputs the feature map. Figure 10 *10*1, obtain the output features Figure 7 ;feature Figure 7 This is scene map information.

[0126] Step 2044: Add features Figure 7 The output is fed into a fully connected layer, resulting in a 1*1 size output value, output3, which is recorded as the scene label value.

[0127] Step 205: Construct the initial generator network. Specifically, the initial generator network can be constructed based on the first sub-network, the second sub-network, and the third sub-network obtained from training.

[0128] The steps above constructed the subnetworks required for the initial generator network. Next, we will construct the initial generator network. Please refer to [link / reference]. Figure 9 The initial generation network is built on the basis of the second sub-network. The difference between the two is that scene detection features are introduced during the network construction process.

[0129] Step 2051: First, the second and third sub-networks need to be constructed; based on the second sub-network, steps 2052 and 2053 need to be modified to construct the final initial generator network.

[0130] Step 2052: Sub-step 2036 within step 203 of the second sub-network, input2RGB input image and feature map feature5 of the first sub-network are sent in. This is modified to only send feature map feature5 to the first sub-network, that is, remove the step of sending in the input2 target image, and keep the rest unchanged.

[0131] Step 2053: Sub-step 2037 within step 203 of the second sub-network, the original three features (feature4, feature5, and output1) are concatenated and replaced with features4, feature5, output1, and the features obtained from sub-step 2043 of step 204 of the third sub-network. Figure 7 The four parts are spliced ​​together, and the spliced ​​feature map is fed into a fully connected layer to obtain the feature map.

[0132] In the second sub-network, during the downsampling or upsampling concatenation process of features4, features5, and output1 after the downsampling or upsampling of the second sub-network (sub-step 2037), the feature map feature7 obtained from the sub-step 2043 of the third sub-network is added. That is, the data size fed into the convolutional layer 12 changes from the original 10*10*5 to 10*10*6, and then the convolutional layer 12 obtains a feature map with a size of 10*10*3. Then, this feature map is fed into the fully connected layer to obtain the 9*1*1 feature output by the fully connected layer. The convolutional kernel size of the convolutional layer 12 is 1*1*6, the stride is 1, and the number of convolutional kernels is 3.

[0133] Step 2054: Resize the 9*1*1 feature obtained in the previous step to 3*3 to obtain the final result output2.

[0134] Step 206: Training the network, including: training the first initial network, the second initial network, and the third initial network respectively to obtain the first sub-network, the second sub-network, and the third sub-network; constructing an initial generator network based on the trained first sub-network, the second sub-network, and the third sub-network; and training the initial generator network to obtain S generator networks.

[0135] After the initial network is built, the model needs to be trained first to obtain the internal parameters of the network. The trained network can be directly applied, that is, given input, the corresponding output result is obtained.

[0136] Before training the initial generator network, the first sub-network, the second sub-network, and the third sub-network need to be trained first. Then, the training parameters of these sub-networks are replaced with the corresponding parameters of the initial generator network, and then the initial generator network is trained.

[0137] Network training loss function:

[0138] The gradients of each parameter in the network are calculated using backpropagation, and the network parameters are updated using stochastic gradient descent until the target loss function of the subnetwork converges numerically. The network parameters are then saved, and the training process of the entire network ends.

[0139] The first subnetwork utilizes the mean squared error loss function;

[0140] The objective loss function of the first sub-network is the mean squared error loss function, specifically as follows:

[0141]

[0142] xi represents the image output by the network, and T represents the error value corresponding to the ground truth annotation.

[0143] The second sub-network and the initial generator network use the mean squared error loss function, which is still the same as the above formula;

[0144] xi represents the network output result, and T represents the ground truth annotation result.

[0145] The loss function for the third subnetwork utilizes the cross-entropy loss function.

[0146] In this implementation, the first sub-network, the second sub-network, and the third sub-network are pre-trained based on the training dataset. This allows the training parameters of these sub-networks to replace the corresponding parameters of the initial generated network, thereby improving the convergence speed of the training process of the initial generated network.

[0147] Optionally, training the initial generator network to obtain S generator networks includes:

[0148] Obtain S fourth training datasets that correspond one-to-one with the S generator networks, wherein each fourth training dataset includes multiple fourth training data sets, each fourth training data set includes a to-be-processed image, a reference image, and a second color correction matrix, and each of the S fourth training datasets corresponds one-to-one with the S parameter sets, wherein the image parameters of all to-be-processed images in the same fourth training dataset belong to the corresponding parameter set, wherein the image parameters include color temperature parameters, brightness parameters, and scene category;

[0149] The initial generator network is trained based on the S fourth training datasets to obtain S generator networks.

[0150] Before training the initial generator network, scene classification can be performed on all fourth training data based on the third sub-network. Specifically, the images to be processed in the fourth training data can be input into the third sub-network for scene classification, and all fourth training data with the same scene category can be stored in the same fourth training dataset, thus obtaining S fourth training datasets.

[0151] It is understood that the training process of the initial generator network described above is only an example of this application. In fact, the training process of each of the S generator networks described above can only be obtained by training according to the training method of the generator network described above.

[0152] Step 207: Apply the S trained generative networks to the ISP module of the electronic device.

[0153] Steps 201 to 206 above yield a trained generative network, meaning that a relatively well-matched CCM matrix can be obtained by inputting the image to be processed. Compared to manually adjusting the CCM matrix, since the generative network considers scene information and human visual perception, the image after color correction based on the CCM matrix generated by the generative network has a better visual effect, and the color of the corrected image matches the image scene in the image better. At the same time, since no manual adjustment is required, the CCM matrix can be automatically generated.

[0154] In some embodiments of this application, M is 5, N is 3, and T is 6. In this case, S = M × N × T = 90, meaning 90 generative networks can be trained. For example, when using a generative network trained on a dataset of lux partition 1, cct partition 1, and scene detection category 1, the ISP can select the corresponding trained CCM network based on the current category and then output the optimal CCM parameters.

[0155] The specific CCM parameter file format is as follows:

[0156] / *Color Correction* /

[0157] .outdoorStart=190, / / outdoorStart and End represent the lux partition of the outdoor scene. When the lux is less than 180, it is used as lux segment 1;

[0158] .outdoorEnd=180, / / lux190-385 serves as the second lux partition;

[0159] .lowlightStart=385, / /

[0160] .lowlightEnd=410, / / lux greater than 410 is used as the third lux partition;

[0161] .day=6500, / / Color temperature distinction parameter, color temperature greater than 6500 is CCT partition 1;

[0162] .dayfStart=4600, / / Color temperature 6500 to 4700 is CCT partition 2

[0163] .dayfEnd=4700,

[0164] .faStart=3350, / / Color temperature 4600-3450 is CCT partition 3

[0165] .faEnd=3450,

[0166] .ahStart=2570, / / Color temperature 3350-2640 is CCT zone 4

[0167] .ahEnd=2640, / / Less than color temperature 2570 is CCT zone 5

[0168] The lux, cct, and scene detection segments can be determined according to project requirements. Currently, the lux segment has 3 segments, the cct segment has 5 segments, and the scene detection segment has 6 segments. Each lux segment corresponds to 5 cct segments, and each cct segment corresponds to 6 types of scene detection. The final segment will have corresponding CCM parameter values. A total of 3×5×6=90 CCM parameters are needed. Manually debugging these parameters would be a huge task. However, using a trained neural network can save manpower and obtain more accurate CCM parameters, improving the accuracy of the final image color and making the image closer to human vision.

[0169] The overall process is as follows: Specific steps are as follows Figure 2As shown, after the Demosaic module in the ISP link, the output image to be processed is sent to the third sub-network of scene classification to obtain scene category labels; then, based on the scene category labels, the current Lux, and CCT, the corresponding trained generator network is selected; after selecting the corresponding generator network, the image to be processed is input into the generator network to obtain the CCM parameters output by the generator network, and the CCM parameters are sent to the CCM module for CCM processing, and then the subsequent ISP link processing can be performed; this ISP link belongs to the preview stream and the image capture stream link.

[0170] In this embodiment, by obtaining S fourth training datasets that correspond one-to-one with the S generator networks, and training the initial generator network based on the S fourth training datasets respectively, S generator networks can be obtained. In this way, the training process of the multiple generator networks can be realized.

[0171] It should be noted that the image processing method provided in this application embodiment has at least the following beneficial effects:

[0172] In some embodiments of this application, AI technology is used to assist in calibration, which can obtain a more standardized CCM matrix that is adapted to the real scene. The obtained CCM matrix is ​​more in line with human visual perception and can reduce the amount of manual adjustment and increase color accuracy.

[0173] In this embodiment, a first sub-network is constructed to calculate the color difference between two images, so as to better calculate the color difference between images. The network uses human visual perception to mark the color difference between the image to be modified and the target image as gt, and the image to be modified and the target image as input, and human visual perception is introduced as a guide in calculating the image difference.

[0174] In this embodiment, a second sub-network is constructed. After outputting scene numbers, the graphs of different scenes are regularized. Graphs with the same scene and the same label (trigger) are fed into the first sub-network for training to obtain the corresponding CCM matrix adapted to each lux, cct, and scene.

[0175] The CCM matrix trained by the network can be directly applied, reducing the cost of manual parameter tuning. Furthermore, it combines traditional constraints with the AI ​​network to learn the final CCM parameters. The overall process is as follows: Figure 2 As shown.

[0176] By incorporating human visual perception differences and scene category information into the network construction, more suitable CCM matrices can be adapted for different images, reducing the cost of manual debugging. This part of the network can be used alone to optimize the initial parameters, or it can be combined with the scene classification network to output the corresponding CCM matrix as a CCM module in the ISP link.

[0177] The network training process incorporates image scene information and image features related to previous modules. By combining scene features with information about the image itself, more suitable CCM parameters are adapted for the ISP path. Furthermore, the image processing method described above can also be used to adjust the skin tone of people in images.

[0178] The image processing method provided in this application, while considering differences in human visual perception, incorporates scene information. This improves the visual appeal of the optimized second image and makes the colors of the optimized second image more closely match the scene within it.

[0179] The image processing method provided in this application can be executed by an image processing device. This application uses an image processing device executing the image processing method as an example to illustrate the image processing device provided in this application.

[0180] Please see Figure 11 , Figure 11 This is a schematic diagram of the structure of an image processing apparatus 1100 provided in an embodiment of this application. The image processing apparatus 1100 includes:

[0181] The recognition module 1101 is used to perform scene recognition on the first image to obtain the scene category of the first image;

[0182] The generation module 1102 is used to generate a first color correction matrix for the first image based on the first image, the color temperature parameter of the first image, the brightness parameter of the first image, and the scene category.

[0183] The correction module 1103 is used to perform color correction processing on the first image based on the first color correction matrix to obtain the second image.

[0184] Optionally, the generation module 1102 is specifically used for:

[0185] Among the S generation networks, the generation network that matches the color temperature parameter, the brightness parameter, and the scene category is determined; wherein, different generation networks correspond to different parameter sets, the parameter set includes a color temperature range, a brightness range, and a scene category, and S is an integer greater than 1;

[0186] The first image is input into the generation network that matches the color temperature parameter, the brightness parameter, and the scene category to generate a first color correction matrix.

[0187] Optionally, the device further includes:

[0188] The division module is used to divide the value range of the color temperature parameter to obtain M color temperature ranges;

[0189] The division module is also used to divide the value range of the brightness parameter to obtain N brightness ranges;

[0190] The determination module is used to determine T candidate scene categories;

[0191] The determining module is further configured to obtain S parameter groups based on the M color temperature ranges, the N brightness ranges, and the T candidate scene categories;

[0192] Wherein, the T candidate scene categories include the scene category of the first image, and M, N and T are all integers greater than 1; S = M × N × T, and the S parameter groups correspond one-to-one with the S generator networks.

[0193] Optionally, the generating network includes a first sub-network, a second sub-network, and a third sub-network, wherein the first sub-network is used to identify differences between different images, the second sub-network is used to generate a color correction matrix, and the third sub-network is used to identify the scene category of the image;

[0194] The generation module 1102 is specifically used for:

[0195] The first image is subjected to feature extraction through the third sub-network to obtain a first feature map;

[0196] The second feature map is obtained by extracting features from the first image through the second sub-network.

[0197] The second feature map is fused with the first image to obtain the third feature map;

[0198] The third feature map is input into the first sub-network to perform sensory difference recognition and output sensory difference information. The sensory difference information is used to characterize the sensory difference between the third feature map and the fourth feature map. The fourth feature map is the feature map of the image that matches the image content of the first image in the training data corresponding to the generator network. Alternatively, if the first image is a frame in video data, the fourth feature map is the feature map of the previous frame of the first image in the video data.

[0199] The first feature map, the second feature map, the third feature map, and the sensory difference information are input into the fully connected layer of the second sub-network, and the first color correction matrix is ​​output.

[0200] Optionally, the device further includes:

[0201] The training module is used to train the first initial network, the second initial network, and the third initial network respectively to obtain the first sub-network, the second sub-network, and the third sub-network;

[0202] A construction module is used to construct an initial generator network based on the first, second, and third sub-networks obtained through training.

[0203] The training module is also used to train the initial generator network to obtain S generator networks.

[0204] Optionally, the first sub-network is a network obtained by training the first initial network based on the first training dataset, wherein the first training dataset includes multiple first training data, and the first training data includes an image to be processed, a reference image, and difference annotation information of the two images;

[0205] The second sub-network is a network trained on the second initial network based on the second training dataset, wherein the second training dataset includes multiple second training data, and the second training data includes an image to be processed and a second color correction matrix;

[0206] The third sub-network is a network trained based on a third training dataset, wherein the third training dataset includes multiple third training data sets, and the third training data set includes an image to be processed and a scene category label.

[0207] Optionally, the training module is specifically used for:

[0208] Obtain S fourth training datasets that correspond one-to-one with the S generator networks, wherein each fourth training dataset includes multiple fourth training data sets, each fourth training data set includes a to-be-processed image, a reference image, and a second color correction matrix, and each of the S fourth training datasets corresponds one-to-one with the S parameter sets, wherein the image parameters of all to-be-processed images in the same fourth training dataset belong to the corresponding parameter set, wherein the image parameters include color temperature parameters, brightness parameters, and scene category;

[0209] The initial generator network is trained based on the S fourth training datasets to obtain S generator networks.

[0210] In this embodiment, since the determined first color correction matrix is ​​determined based on the first image, the color temperature parameter of the first image, the brightness parameter of the first image, and the scene category of the first image, in the process of determining the color correction matrix, in addition to considering the color temperature and brightness of the first image, the scene category of the first image is also considered. Therefore, in the process of color correction of the first image based on the first color correction matrix, the color temperature and brightness of the first image can be optimized, and the optimized color of the second image can be better matched with the image scene in the second image, thereby optimizing the color correction effect of the second image.

[0211] The image processing device 1100 in this embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This embodiment does not specifically limit the specific type of device.

[0212] The image processing device 1100 in this embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this embodiment does not specifically limit its use.

[0213] The image processing device 1100 provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method embodiments achieve the same technical effect, and will not be described again here to avoid repetition.

[0214] In some embodiments, such as Figure 12As shown, this application embodiment also provides an electronic device 1200, including a processor 1201, a memory 1202, and a program or instructions stored in the memory 1202 and executable on the processor 1201. When the program or instructions are executed by the processor 1201, they implement the various processes of the above-described image processing method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.

[0215] Figure 13 A schematic diagram of the hardware structure of an electronic device according to an embodiment of this application.

[0216] The electronic device 1300 includes, but is not limited to, components such as: radio frequency unit 1301, network module 1302, audio output unit 1303, input unit 1304, sensor 1305, display unit 1306, user input unit 1307, interface unit 1308, memory 1309, and processor 1310.

[0217] The processor 1310 is used to perform scene recognition on the first image to obtain the scene category of the first image;

[0218] The processor 1310 is configured to generate a first color correction matrix for the first image based on the first image, the color temperature parameters of the first image, the brightness parameters of the first image, and the scene category.

[0219] The processor 1310 is used to perform color correction processing on the first image based on the first color correction matrix to obtain a second image.

[0220] Optionally, the processor 1310 is configured to determine, among S generation networks, a generation network that matches the color temperature parameter, the brightness parameter, and the scene category; wherein different generation networks correspond to different parameter sets, the parameter set including a color temperature range, a brightness range, and a scene category, and S is an integer greater than 1;

[0221] The processor 1310 is used to input the first image into the generation network that matches the color temperature parameter, the brightness parameter and the scene category, and generate a first color correction matrix.

[0222] Optionally, the processor 1310 is used to divide the range of values ​​for the color temperature parameter to obtain M color temperature ranges;

[0223] The processor 1310 is used to divide the value range of the brightness parameter to obtain N brightness ranges;

[0224] The processor 1310 is used to determine T candidate scene categories;

[0225] The processor 1310 is used to obtain S parameter groups based on the M color temperature ranges, the N brightness ranges, and the T candidate scene categories;

[0226] Wherein, the T candidate scene categories include the scene category of the first image, and M, N and T are all integers greater than 1; S = M × N × T, and the S parameter groups correspond one-to-one with the S generator networks.

[0227] Optionally, the generating network includes a first sub-network, a second sub-network, and a third sub-network, wherein the first sub-network is used to identify differences between different images, the second sub-network is used to generate a color correction matrix, and the third sub-network is used to identify the scene category of the image;

[0228] The processor 1310 is used to extract features from the first image through the third sub-network to obtain a first feature map;

[0229] The processor 1310 is used to extract features from the first image through the second sub-network to obtain a second feature map;

[0230] The processor 1310 is used to perform image fusion between the second feature map and the first image to obtain a third feature map;

[0231] The processor 1310 is configured to input the third feature map into the first sub-network, perform sensory difference recognition, and output sensory difference information. The sensory difference information is used to characterize the sensory difference between the third feature map and the fourth feature map. The fourth feature map is a feature map of an image in the training data corresponding to the generator network that matches the image content of the first image. Alternatively, if the first image is a frame in video data, the fourth feature map is a feature map of the previous frame of the first image in the video data.

[0232] The processor 1310 is used to input the first feature map, the second feature map, the third feature map and the sensory difference information into the fully connected layer of the second sub-network and output the first color correction matrix.

[0233] Optionally, the processor 1310 is used to train the first initial network, the second initial network, and the third initial network respectively to obtain the first sub-network, the second sub-network, and the third sub-network;

[0234] The processor 1310 is used to construct an initial generator network based on the first sub-network, the second sub-network, and the third sub-network obtained through training.

[0235] The processor 1310 is used to train the initial generator network to obtain S generator networks.

[0236] Optionally, the first sub-network is a network obtained by training the first initial network based on the first training dataset, wherein the first training dataset includes multiple first training data, and the first training data includes an image to be processed, a reference image, and difference annotation information of the two images;

[0237] The second sub-network is a network trained on the second initial network based on the second training dataset, wherein the second training dataset includes multiple second training data, and the second training data includes an image to be processed and a second color correction matrix;

[0238] The third sub-network is a network trained based on a third training dataset, wherein the third training dataset includes multiple third training data sets, and the third training data set includes an image to be processed and a scene category label.

[0239] Optionally, the processor 1310 is configured to acquire S fourth training datasets that correspond one-to-one with the S generator networks, wherein the fourth training dataset includes multiple fourth training data, each of which includes a to-be-processed image, a reference image, and a second color correction matrix. The S fourth training datasets correspond one-to-one with the S parameter groups, and the image parameters of all to-be-processed images in the same fourth training dataset belong to the corresponding parameter group. The image parameters include color temperature parameters, brightness parameters, and scene category.

[0240] The processor 1310 is used to train the initial generator network based on the S fourth training datasets respectively, to obtain S generator networks.

[0241] Those skilled in the art will understand that the electronic device 1300 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1310 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 13 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0242] It should be understood that, in this embodiment, the input unit 1304 may include a graphics processing unit (GPU) 13041 and a microphone 13042. The GPU 13041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1306 may include a display panel 13061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1307 includes a touch panel 13071 and other input devices 13072. The touch panel 13071 is also called a touch screen. The touch panel 13071 may include a touch detection device and a touch controller. Other input devices 13072 may include, but are not limited to, a physical keyboard, function keys (such as volume control buttons, power buttons, etc.), a trackball, a mouse, and a joystick, which will not be described in detail here.

[0243] The memory 1309 can be used to store software programs and various data. The memory 1309 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1309 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1309 in this embodiment includes, but is not limited to, these and any other suitable types of memory.

[0244] Processor 1310 may include one or more processing units; optionally, processor 1310 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 1310.

[0245] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described image processing method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.

[0246] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0247] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described image processing method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0248] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0249] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0250] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0251] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An image processing method, characterized in that, include: Scene recognition is performed on the first image to obtain the scene category of the first image; Based on the first image, the color temperature parameter of the first image, the brightness parameter of the first image, and the scene category, a first color correction matrix for the first image is generated. Based on the first color correction matrix, the first image is subjected to color correction processing to obtain the second image.

2. The image processing method according to claim 1, characterized in that, The step of generating a first color correction matrix for the first image based on the first image, the color temperature parameters of the first image, the brightness parameters of the first image, and the scene category includes: Among the S generation networks, the generation network that matches the color temperature parameter, the brightness parameter, and the scene category is determined; wherein, different generation networks correspond to different parameter sets, the parameter set includes a color temperature range, a brightness range, and a scene category, and S is an integer greater than 1; The first image is input into the generation network that matches the color temperature parameter, the brightness parameter, and the scene category to generate a first color correction matrix.

3. The image processing method according to claim 2, characterized in that, Before determining the generation network that matches the color temperature parameter, the brightness parameter, and the scene category among the S generation networks, the method further includes: The range of values ​​for the color temperature parameter is divided to obtain M color temperature ranges; The range of values ​​for the brightness parameter is divided into N brightness ranges; Determine T candidate scene categories; Based on the M color temperature ranges, the N brightness ranges, and the T candidate scene categories, S parameter groups are obtained; Wherein, the T candidate scene categories include the scene category of the first image, and M, N and T are all integers greater than 1; S = M × N × T, and the S parameter groups correspond one-to-one with the S generator networks.

4. The image processing method according to claim 3, characterized in that, The generating network includes a first sub-network, a second sub-network, and a third sub-network. The first sub-network is used to identify differences between different images, the second sub-network is used to generate a color correction matrix, and the third sub-network is used to identify the scene category of the image. The step of inputting the first image into the generation network that matches the color temperature parameter, the brightness parameter, and the scene category to generate a first color correction matrix includes: The first image is subjected to feature extraction through the third sub-network to obtain a first feature map; The second feature map is obtained by extracting features from the first image through the second sub-network. The second feature map is fused with the first image to obtain the third feature map; The third feature map is input into the first sub-network to perform sensory difference recognition and output sensory difference information. The sensory difference information is used to characterize the sensory difference between the third feature map and the fourth feature map. The fourth feature map is the feature map of the image that matches the image content of the first image in the training data corresponding to the generator network. Alternatively, if the first image is a frame in video data, the fourth feature map is the feature map of the previous frame of the first image in the video data. The first feature map, the second feature map, the third feature map, and the sensory difference information are input into the fully connected layer of the second sub-network, and the first color correction matrix is ​​output.

5. The image processing method according to claim 4, characterized in that, Before determining the generation network that matches the color temperature parameter, the brightness parameter, and the scene category among the S generation networks, the method further includes: The first initial network, the second initial network, and the third initial network are trained respectively to obtain the first sub-network, the second sub-network, and the third sub-network; Based on the first, second, and third sub-networks obtained through training, an initial generative network is constructed. The initial generator network is trained to obtain S generator networks.

6. The image processing method according to claim 5, characterized in that, The first sub-network is a network obtained by training the first initial network based on the first training dataset. The first training dataset includes multiple first training data, which includes an image to be processed, a reference image, and difference annotation information between the two images. The second sub-network is a network trained on the second initial network based on the second training dataset, wherein the second training dataset includes multiple second training data, and the second training data includes an image to be processed and a second color correction matrix; The third sub-network is a network trained based on a third training dataset, wherein the third training dataset includes multiple third training data sets, and the third training data set includes an image to be processed and a scene category label.

7. The image processing method according to claim 5, characterized in that, The process of training the initial generator network to obtain S generator networks includes: Obtain S fourth training datasets that correspond one-to-one with the S generator networks, wherein each fourth training dataset includes multiple fourth training data sets, each fourth training data set includes a to-be-processed image, a reference image, and a second color correction matrix, and each of the S fourth training datasets corresponds one-to-one with the S parameter sets, wherein the image parameters of all to-be-processed images in the same fourth training dataset belong to the corresponding parameter set, wherein the image parameters include color temperature parameters, brightness parameters, and scene category; The initial generator network is trained based on the S fourth training datasets to obtain S generator networks.

8. An image processing apparatus, characterized in that, include: The recognition module is used to perform scene recognition on the first image to obtain the scene category of the first image; The generation module is used to generate a first color correction matrix for the first image based on the first image, the color temperature parameters of the first image, the brightness parameters of the first image, and the scene category. The correction module is used to perform color correction processing on the first image based on the first color correction matrix to obtain the second image.

9. The image processing apparatus according to claim 8, characterized in that, The generation module is specifically used for: Among the S generation networks, the generation network that matches the color temperature parameter, the brightness parameter, and the scene category is determined; wherein, different generation networks correspond to different parameter sets, the parameter set includes a color temperature range, a brightness range, and a scene category, and S is an integer greater than 1; The first image is input into the generation network that matches the color temperature parameter, the brightness parameter, and the scene category to generate a first color correction matrix.

10. The image processing apparatus according to claim 9, characterized in that, The device further includes: The division module is used to divide the value range of the color temperature parameter to obtain M color temperature ranges; The division module is also used to divide the value range of the brightness parameter to obtain N brightness ranges; The determination module is used to determine T candidate scene categories; The determining module is further configured to obtain S parameter groups based on the M color temperature ranges, the N brightness ranges, and the T candidate scene categories; Wherein, the T candidate scene categories include the scene category of the first image, and M, N and T are all integers greater than 1; S = M × N × T, and the S parameter groups correspond one-to-one with the S generator networks.

11. The image processing apparatus according to claim 10, characterized in that, The generating network includes a first sub-network, a second sub-network, and a third sub-network. The first sub-network is used to identify differences between different images, the second sub-network is used to generate a color correction matrix, and the third sub-network is used to identify the scene category of the image. The generation module is specifically used for: The first image is subjected to feature extraction through the third sub-network to obtain a first feature map; The second feature map is obtained by extracting features from the first image through the second sub-network. The second feature map is fused with the first image to obtain the third feature map; The third feature map is input into the first sub-network to perform sensory difference recognition and output sensory difference information. The sensory difference information is used to characterize the sensory difference between the third feature map and the fourth feature map. The fourth feature map is the feature map of the image that matches the image content of the first image in the training data corresponding to the generator network. Alternatively, if the first image is a frame in video data, the fourth feature map is the feature map of the previous frame of the first image in the video data. The first feature map, the second feature map, the third feature map, and the sensory difference information are input into the fully connected layer of the second sub-network, and the first color correction matrix is ​​output.

12. The image processing apparatus according to claim 11, characterized in that, The device further includes: The training module is used to train the first initial network, the second initial network, and the third initial network respectively to obtain the first sub-network, the second sub-network, and the third sub-network; A construction module is used to construct an initial generator network based on the first, second, and third sub-networks obtained through training. The training module is also used to train the initial generator network to obtain S generator networks.

13. The image processing apparatus according to claim 12, characterized in that, The first sub-network is a network obtained by training the first initial network based on the first training dataset. The first training dataset includes multiple first training data, which includes an image to be processed, a reference image, and difference annotation information between the two images. The second sub-network is a network trained on the second initial network based on the second training dataset, wherein the second training dataset includes multiple second training data, and the second training data includes an image to be processed and a second color correction matrix; The third sub-network is a network trained based on a third training dataset, wherein the third training dataset includes multiple third training data sets, and the third training data set includes an image to be processed and a scene category label.

14. The image processing apparatus according to claim 12, characterized in that, The training module is specifically used for: Obtain S fourth training datasets that correspond one-to-one with the S generator networks, wherein each fourth training dataset includes multiple fourth training data sets, each fourth training data set includes a to-be-processed image, a reference image, and a second color correction matrix, and each of the S fourth training datasets corresponds one-to-one with the S parameter sets, wherein the image parameters of all to-be-processed images in the same fourth training dataset belong to the corresponding parameter set, wherein the image parameters include color temperature parameters, brightness parameters, and scene category; The initial generator network is trained based on the S fourth training datasets to obtain S generator networks.

15. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a program or instructions that can run on the processor, and the program or instructions, when executed by the processor, implement the steps of the image processing method as described in any one of claims 1-7.