Aesthetically Guided Automatic White Balance Method and Device
Through the aesthetic-guided reinforcement learning model, the correction strategy of pixel-by-pixel optimization of image is solved, and the existing technology has inaccurate lighting estimation in complex environments is achieved, achieving efficient and accurate automatic white balance effect.
Patent Information
- Application Number
- CN202510286634.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-03-12
AI Technical Summary
The existing automatic white balance technology is difficult to achieve accurate lighting estimation when facing solid color scenes or images in complex multi-light environments, which affects the processing effect of white balance.
Using an aesthetic-guided reinforcement learning model, the images are input into the reinforcement learning model by obtaining the image to be corrected and the instance segmentation information, and the correction strategy is optimized using the reward function, and the correction processing is performed pixel by pixel, and the cycle optimization is optimized until the preset number is reached.
Accurate color offset correction of images in any environment is achieved, the efficiency and accuracy of automatic white balance is improved, and the correction process takes into account the coordination between the local and the overall, achieving color fidelity effect.
Smart Images

Figure CN119788978B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer science and technology, and particularly to an aesthetic-guided automatic white balance method and device. Background Art
[0002] Automatic white balance is a key image processing technology, whose main goal is to eliminate the influence of environmental light sources on images and make the originally white areas present a true white effect.
[0003] Currently, the processing of automatic white balance for images mainly relies on a data-driven approach. Data-driven usually extracts the illumination features in the image through a convolutional neural network or other deep learning models for color cast correction.
[0004] However, the data-driven approach highly depends on the diversity and quality of the training dataset. Therefore, when facing images in a pure color scene or a complex multi-light source environment, it is difficult to achieve accurate illumination estimation, and thus the processing effect of white balance will be greatly reduced. Therefore, providing a method that can accurately correct the color cast of images in any environment is an urgent problem to be solved in this field. Summary of the Invention
[0005] The embodiments of this application provide an aesthetic-guided automatic white balance method and device, which can improve the efficiency and accuracy of the automatic white balance solution.
[0006] In a first aspect, the embodiments of this application provide an aesthetic-guided automatic white balance method, which is applied to a computing device and includes:
[0007] Obtain the image to be corrected and the instance segmentation information corresponding to the image to be corrected;
[0008] Input the image to be corrected into a reinforcement learning model to obtain a reward value and a correction strategy for each pixel point of the image to be corrected;
[0009] Optimize the reinforcement learning model based on the reward value to obtain an optimized reinforcement learning model;
[0010] Based on the correction strategy for each pixel point of the image to be corrected and the instance segmentation information, perform correction processing on each pixel point of the image to be corrected to obtain a corrected image;
[0011] Execute the loop to input the corrected image into the reinforcement learning model to obtain a reward value and a correction strategy for each pixel of the corrected image, optimize the reinforcement learning model based on the reward value to obtain an optimized reinforcement learning model, and correct each pixel of the corrected image based on the correction strategy and instance segmentation information until the number of loop executions reaches a preset number to obtain the final image after final correction.
[0012] In a possible implementation, the reinforcement learning model includes: a state module, an agent module, and a reward function module, where the reward function module includes: a guidance reward function, an aesthetic quality reward function, and an artifact loss function;
[0013] Accordingly, the step of inputting the image to be corrected into the reinforcement learning model to obtain a reward value and a correction strategy for each pixel of the image to be corrected includes:
[0014] Input the image to be corrected into the reinforcement learning model, and perform light estimation on the image to be corrected based on the state module to obtain a preprocessing strategy;
[0015] The agent module obtains a correction strategy based on the preprocessing strategy and the historical correction strategy;
[0016] Based on the reward function module, obtain a first reward value corresponding to the guidance reward function, a second reward value corresponding to the aesthetic quality reward function, and a third reward value corresponding to the artifact loss function.
[0017] In a possible implementation, the reinforcement learning model further includes: an environment module;
[0018] Accordingly, the step of performing light estimation on the image to be corrected based on the state module to obtain a preprocessing strategy includes:
[0019] Obtain the light information of the image to be corrected;
[0020] Based on the environment module, perform light estimation and global correction on the light information of the image to be corrected through a diagonal matrix to obtain a corrected image;
[0021] Determine the corrected image as the preprocessing result.
[0022] In a possible implementation, the step of optimizing the reinforcement learning model based on the reward value to obtain an optimized reinforcement learning model includes:
[0023] Sum the first reward value, the second reward value, and the third reward value to obtain a total reward value;
[0024] Discount the total reward value based on a preset discount function to obtain a discounted reward value;
[0025] Based on the discounted reward value, correct the parameters of the agent module in the reinforcement learning model to obtain an optimized reinforcement learning model.
[0026] In a possible implementation, the step of correcting each pixel point of the image to be corrected based on the correction strategy and the instance segmentation information of each pixel point of the image to be corrected to obtain a corrected image includes:
[0027] Determine a plurality of target regions corresponding to the image to be corrected based on the instance segmentation information;
[0028] Among the correction strategies corresponding to the pixel points included in each target region, determine the first correction strategy with the highest proportion, and determine the first correction strategy as the target correction strategy corresponding to the target region;
[0029] Execute the corresponding target correction strategy for each pixel point in the target region to obtain a corrected image.
[0030] In a possible implementation, before inputting the image to be corrected into the reinforcement learning model, the method further includes:
[0031] Obtain training data, where the training data includes multiple groups of training images and corresponding standard images;
[0032] Train based on the multiple groups of training images and a preset pixel reinforcement learning framework corresponding to the standard images to obtain the reinforcement learning model.
[0033] In a possible implementation, the step of training based on the multiple groups of training images and a preset pixel reinforcement learning framework corresponding to the standard images to obtain the reinforcement learning model includes:
[0034] For each group of training images and corresponding standard images, input the training images and corresponding standard images into the preset pixel reinforcement learning framework to obtain a reward value and a correction strategy for each pixel point of the training image;
[0035] Optimize the pixel reinforcement learning framework based on the reward value to obtain an optimized pixel reinforcement learning framework;
[0036] Correct each pixel point of the training image based on the correction strategy to obtain a corrected image;
[0037] Execute the loop to input the corrected image into the pixel reinforcement learning framework to obtain a reward value and a correction strategy for each pixel point of the corrected image, optimize the pixel reinforcement learning framework based on the reward value to obtain an optimized pixel reinforcement learning framework, and correct each pixel point of the corrected image based on the correction strategy until the number of loop executions reaches a preset number;
[0038] Update the parameters of the pixel reinforcement learning framework based on all the reward values corresponding to each set of training images and the corresponding standard images until the reward values converge to obtain the reinforcement learning model.
[0039] In a second aspect, the present application provides an aesthetics-guided automatic white balance device, which is applied to a computing device and includes:
[0040] An acquisition module, configured to acquire a to-be-corrected image and instance segmentation information corresponding to the to-be-corrected image;
[0041] A processing module, configured to input the to-be-corrected image into the reinforcement learning model to obtain a reward value and a correction strategy for each pixel point of the to-be-corrected image;
[0042] The processing module is further configured to optimize the reinforcement learning model based on the reward value to obtain an optimized reinforcement learning model;
[0043] The processing module is further configured to perform correction processing on each pixel point of the to-be-corrected image based on the correction strategy for each pixel point of the to-be-corrected image and the instance segmentation information to obtain a corrected image;
[0044] The processing module is further configured to execute in a loop the step of inputting the corrected image into the reinforcement learning model to obtain a reward value and a correction strategy for each pixel point of the corrected image, optimizing the reinforcement learning model based on the reward value to obtain an optimized reinforcement learning model, and correcting each pixel point of the corrected image based on the correction strategy and the instance segmentation information until the number of loop executions reaches a preset number to obtain a final image after final correction is completed.
[0045] In a third aspect, the present application provides a computer device, including: a memory, a processor;
[0046] The memory stores computer execution instructions;
[0047] The processor executes the computer execution instructions stored in the memory, so that the processor executes the method as described above.
[0048] Fourthly, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the method described above.
[0049] Fifthly, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method described above.
[0050] The aesthetic-guided automatic white balance method and device provided by the embodiments of the present application first obtain a to-be-corrected image and instance segmentation information corresponding to the to-be-corrected image, input the to-be-corrected image into a reinforcement learning model to obtain a reward value and a correction strategy for each pixel point of the to-be-corrected image, optimize the reinforcement learning model based on the reward value to obtain an optimized reinforcement learning model, and correct each pixel point of the to-be-corrected image based on the correction strategy and the instance segmentation information to obtain a corrected image. Then, the above steps are cycled until the number of loop executions reaches a preset number to obtain a final image after final correction. The present application corrects each pixel of the to-be-corrected image multiple times through a reinforcement learning model to achieve color balance, and performs correction with reference to the instance segmentation information, which can ensure that the coordination between the local and the whole is taken into account during the correction process to achieve the effect of color fidelity. Description of the Drawings
[0051] The drawings here are incorporated into the description and form a part of this description, showing embodiments consistent with the present application, and are used together with the description to explain the principles of the present application.
[0052] Figure 1 It is a schematic diagram of the scenario provided by the present application;
[0053] Figure 2 It is a flowchart of the aesthetic-guided automatic white balance method provided by the present application Figure 1 ;
[0054] Figure 3 It is a schematic diagram of instance segmentation for example;
[0055] Figure 4 It is a flowchart of the aesthetic-guided automatic white balance method provided by the present application Figure 2 ;
[0056] Figure 5 It is a flowchart of the aesthetic-guided automatic white balance method provided by the present application Figure 3 ;
[0057] Figure 6 It is a flowchart of the aesthetic-guided automatic white balance method provided by the present application Figure 4 ;
[0058] Figure 7 Schematic structural diagram of an example reinforcement learning model;
[0059] Figure 8 Flow schematic of the aesthetic-guided automatic white balance method provided by this application Figure 5 ;
[0060] Figure 9 Schematic comparison diagram of image correction processing between this application and the prior art for example;
[0061] Figure 10 Flow schematic of the aesthetic-guided automatic white balance method provided by this application Figure 6 ;
[0062] Figure 11 Flow schematic of the aesthetic-guided automatic white balance method provided by this application Figure 7 ;
[0063] Figure 12 Schematic structural diagram of the aesthetic-guided automatic white balance device provided by this application;
[0064] Figure 13 Schematic structural diagram of the electronic device provided by this application.
[0065] Through the above-mentioned drawings, specific embodiments of this application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of this application in any way, but to illustrate the concept of this application to those skilled in the art by referring to specific embodiments. Detailed Description of the Embodiments
[0066] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. On the contrary, they are merely examples of devices and methods consistent with some aspects of this application as detailed in the appended claims.
[0067] First, the terms involved in this application are explained:
[0068] Automatic white balance: It is an image processing technology widely used in devices such as digital cameras, camcorders, and mobile phones. By estimating the scene illumination color during shooting, it automatically adjusts the color temperature of the photo or video to ensure that the restoration of all colors in the image is as close as possible to the true colors seen by the human eye under the corresponding illumination conditions.
[0069] Instance segmentation: Identifying the exact pixel boundaries of each individual object instance in an image.
[0070] GT Image: Ground Truth, which is used as the reference true annotation image or data in computer vision or aesthetic-guided automatic white balance tasks to test the performance of algorithms. In this invention, it specifically refers to the image after expert color correction.
[0071] Solid Color Scene: A scene mainly composed of single or a small number of simple color areas, lacking complex textures and detailed variations, such as a blue sky, grassland, wall, etc.
[0072] Color Aesthetics: A discipline that studies the impact of color combinations, contrasts, and harmonies on visual perception and emotional transmission, involving the application principles and aesthetic values of colors in art, design, and daily life.
[0073] Pixel Reinforcement Learning: A multi-agent reinforcement learning method in which each pixel of an image is regarded as an independent agent. They learn decision-making strategies from local information through interaction, cooperation, or competition to jointly complete the overall task.
[0074] Action Space: In reinforcement learning, it defines the set of behaviors or decisions that an agent can execute, which describes all possible actions that an agent can choose at each time step.
[0075] Automatic white balance is a key image processing technology. Its main goal is to eliminate the influence of environmental light sources on images, so that originally white areas present a true white effect. Currently, it mainly relies on a data-driven approach to process automatic white balance for images. Data-driven usually extracts the illumination features in the image through a convolutional neural network or other deep learning models for color cast correction. However, the data-driven approach highly depends on the diversity and quality of the training dataset. Therefore, when facing images in a solid color or complex multi-light source environment, it is difficult to achieve accurate illumination estimation, and thus the processing effect of white balance will be greatly reduced. Therefore, providing a method that can accurately correct the color cast of images in any environment is an urgent problem to be solved in this field.
[0076] Figure 1 For the scene schematic diagram provided in this application, the image to be corrected and the corresponding instance segmentation information can be input into the reinforcement learning model, so that the reinforcement learning module discounts the obtained correction strategy according to the obtained reward value to obtain the correction strategy for each pixel point of the image to be corrected. Based on the correction strategy and the instance segmentation information, each pixel point of the image to be corrected is corrected to obtain the corrected image, and then the above steps are looped until the number of loop executions reaches the preset number to obtain the final image after the final correction is completed.
[0077] The following uses specific embodiments to elaborate in detail on the technical solution of the present application and how the technical solution of the present application solves the above technical problems. The following several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0078] Figure 2 Flow schematic of the aesthetic-guided automatic white balance method provided by the present application Figure 1 , the method is applied to a computing device, such as Figure 2 shown, including:
[0079] S201. Obtain the image to be corrected and the instance segmentation information corresponding to the image to be corrected.
[0080] The execution subject of the present application is a computing device, which can be a computer or a server providing services.
[0081] Combined with a scenario example, the instance segmentation (Segmentation, abbreviated as SEG) information refers to multiple regions in the image to be corrected. Figure 3 For the schematic diagram of instance segmentation as an example, as Figure 3 shown, the picture can be divided into four regions, namely the region where the person is located, the region where the pet is located, the sky region, and the lawn region.
[0082] S202. Input the image to be corrected into the reinforcement learning model to obtain a reward value and a correction strategy for each pixel of the image to be corrected.
[0083] Combined with a scenario example, the initial value correction strategy includes the correction strategy for each pixel in the image to be corrected. The correction strategy includes color temperature adjustment processing, etc. The reward value represents the optimization degree of the correction strategy for each pixel in the image to be corrected.
[0084] S203. Optimize the reinforcement learning model based on the reward value to obtain an optimized reinforcement learning model.
[0085] Combined with a scenario example, the parameters in the reinforcement learning model can be discounted based on a preset discount reward coefficient, that is, the parameters in the reinforcement learning model are corrected to optimize the reinforcement learning model.
[0086] S204. Correct each pixel of the image to be corrected based on the correction strategy for each pixel of the image to be corrected and the instance segmentation information to obtain a corrected image.
[0087] Combined with the scenario example, according to the correction strategy and instance segmentation information of each pixel point obtained above, the correction strategy corresponding to each region is obtained, and the correction strategy corresponding to each pixel in each region is the same.
[0088] S205. Loop to input the corrected image into the reinforcement learning model to obtain a reward value and a correction strategy for each pixel point of the corrected image, optimize the reinforcement learning model based on the reward value to obtain an optimized reinforcement learning model, and correct each pixel point of the corrected image based on the correction strategy and instance segmentation information, until the number of loop executions reaches a preset number to obtain the final image after final correction.
[0089] Combined with the scenario example, after obtaining the corrected image above, the corrected image can be input into the reinforcement learning model again. Since the reinforcement learning model input again has been optimized by the reward value obtained last time, a new corresponding correction strategy for this time can be obtained, and the corrected image is corrected again based on the correction strategy corresponding to this time, and the above steps are looped until the number of loops reaches a preset number. For example, after 10 loop corrections, the final image after final correction is obtained.
[0090] Based on the method provided in this example, each pixel of the image to be corrected can be corrected multiple times through the reinforcement learning model to achieve color balance, and the correction is performed with reference to the instance segmentation information, which can ensure that the coordination between the local and the whole is taken into account during the correction process to achieve the effect of color fidelity.
[0091] Optionally, Figure 4 is the process schematic of the aesthetic-guided automatic white balance method provided in this application Figure 2 , the reinforcement learning model includes: a state module, an agent module, and a reward function module, where the reward function module includes: a guidance reward function, an aesthetic quality reward function, and an artifact loss function;
[0092] Correspondingly, as Figure 4 shown, S202 includes:
[0093] S401. Input the image to be corrected into the reinforcement learning model, and perform light estimation on the image to be corrected based on the state module to obtain a preprocessing strategy.
[0094] Combined with the scenario example, the reinforcement learning model includes a state module, an agent module, and a reward function module. After inputting the image to be corrected into the reinforcement learning model, the state module is used to perform light estimation on the image to be corrected to obtain a preprocessing strategy for the image to be corrected.
[0095] S402. The agent module obtains the correction strategy based on the preprocessing strategy and the historical correction strategy.
[0096] Combined with the scenario example, if this correction is the first correction, the policy network of the agent module will give the correction strategy based on the preprocessing strategy given by the state module. If this correction is not the first correction, the correction strategies given in each previous correction will be used as the historical correction strategies, and the policy network of the agent module will give the current correction strategy based on the preprocessing strategy given by the state module and the historical correction strategies given in each previous correction.
[0097] S403. Based on the reward function module, obtain the first reward value corresponding to the guidance reward function, the second reward value corresponding to the aesthetic quality reward function, and the third reward value corresponding to the artifact loss function.
[0098] Combined with the scenario example, the preprocessing strategy is mainly used to correct the color cast of the image to be corrected. After that, it is necessary to further introduce an aesthetic guidance model to simulate manual adjustment of the color temperature by humans to optimize the color of the image, and assist with appropriate other color adjustment actions to enhance the color aesthetic performance of the corrected image. The main goal of the subsequent adjustment is to improve the aesthetic score of the image while minimizing the color cast as much as possible. The overall objective function to be achieved is as follows:
[0099]
[0100] where / / Lc - L / / 2 represents the distance between the estimated illumination and the true illumination, A(Lc) represents the aesthetic score of the image, and α and β are used to adjust the emphasis ratio of color fidelity and color aesthetics. Specifically, the above process can be realized by the reward values given by the guidance reward function, the aesthetic quality reward function, and the artifact loss function of the reward function module. The guidance reward function, the aesthetic quality reward function, and the artifact loss function of the reward function module are used to optimize and guide the policy network of the agent module. Specifically, the guidance reward function is defined as the change in the distance between the image to be corrected and the preset target GT map, and learns the expert aesthetics and color fidelity information in the preset expert color correction map. It can be expressed by the following formula:
[0101] Guidance reward function:
[0102] where I target represents the target GT map, s t represents the state before correction, and s (t+1) represents the state after correction.
[0103] The aesthetic quality reward function is defined as the change in the aesthetic feeling before and after the correction of the current image, and is used to optimize the aesthetic feeling before and after the correction of the image, and correct the image in the direction of better aesthetic evaluation. It can be expressed by the following formula:
[0104] Aesthetic quality reward function:
[0105] where Aes score (s t ) represents the evaluation of the image by the aesthetic model.
[0106] The artifact loss function is used to calculate the difference in gradient performance between the corrected image and the GT target image, and by suppressing this difference, the generation of artifacts in the image is reduced. It can be expressed by the following formula:
[0107] Artifact loss function:
[0108] where grad c (V i,j ) is used to calculate the gradient information of the image.
[0109] Based on the above guidance reward function, the aesthetic quality reward function, and the artifact loss function, the corresponding first reward value, second reward value, and third reward value are obtained respectively.
[0110] Based on the method provided in this example, the purpose of obtaining the preprocessing strategy, correction strategy, and reward value can be achieved.
[0111] Optionally, Figure 5 is a schematic flow of the aesthetic-guided automatic white balance method provided by this application Figure 3 , and the reinforcement learning model further includes: an environment module;
[0112] Correspondingly, as Figure 5 shown, S401 includes:
[0113] S501. Obtain the illumination information of the image to be corrected.
[0114] Combined with the scenario example, for the image I to be corrected, the illumination information can be defined as {LR, LG, LB}.
[0115] S502. Based on the environment module, perform illumination estimation and global correction on the illumination information of the image to be corrected through a diagonal matrix to obtain the corrected image.
[0116] Combined with the scenario example, in the environment module, the color cast of the image to be corrected can be globally corrected through a diagonal matrix, and the corresponding formula is as follows:
[0117]
[0118] where Iwb is the corrected image.
[0119] S503. Determine the corrected image as the preprocessing result.
[0120] Combine with the scenario example and determine the Iwb as the preprocessing result.
[0121] Based on the method provided in this example, preprocessing of the image to be corrected can be achieved.
[0122] Optionally, Figure 6 The flow chart of the aesthetic-guided automatic white balance method provided by this application Figure 4 , such as Figure 6 shown, S203 includes:
[0123] S601, Sum the first reward value, the second reward value and the third reward value to obtain the total reward value.
[0124] Combine with the scenario example, sum the obtained first reward value, the second reward value and the third reward value. For example, if the first reward value is 40, the second reward value is 30, and the third reward value is 30, then the obtained total reward value is 100.
[0125] S602, Based on a preset discount function, perform discount processing on the total reward value to obtain a discounted reward value.
[0126] Combine with the scenario example, the following is an example of the discount function:
[0127]
[0128] where R(t) represents the reward for correction at step t, π represents the policy for correction at step t, and γ represents the discount reward coefficient.
[0129] If the obtained reward value is 100 and the discount reward coefficient is 0.9, the finally discounted reward value is 90.
[0130] S603, Based on the discounted reward value, correct the parameters of the agent module in the reinforcement learning model to obtain an optimized reinforcement learning model.
[0131] Combine with the scenario example, based on the obtained discounted reward value, superimpose the reward value on the correction policy to optimize the parameters of the agent module in the reinforcement learning model. Specifically, the policy network and value network in the agent module can be optimized to obtain the correction policy.
[0132] Specifically, based on the correction policy, map the correction policy and the actions in the action space, and apply the adjustment to the image. The specific adjustment formula is as follows:
[0133]
[0134] Where I(t + 1) represents the corrected image, and xi represents the pixel position corresponding to the adjustment of the "Action" strategy.
[0135] Figure 7 It is a schematic structural diagram of an example reinforcement learning model. As Figure 7 shown, the image to be corrected is input into the reinforcement learning model. Based on the state module, light estimation is performed to obtain a preprocessing strategy. Based on the agent module, a correction strategy is obtained. The correction strategy includes the adjustment actions corresponding to each pixel point. However, each pixel point can only have one adjustment action during one correction process. The adjustment actions specifically include: increasing the color temperature and decreasing the color temperature. Based on the reward function module, the total reward value is obtained. Specifically, the reward function module includes an artifact loss function, an aesthetic quality reward function, and a guidance reward function. Then, after discounting the total reward value, the agent module of the reinforcement learning model is optimized. Then, based on the environment module, color cast correction is performed on the image to be corrected based on the diagonal matrix correction of the preprocessing strategy, and the adjustment actions corresponding to each pixel point in the image to be corrected are adjusted based on the correction strategy. In this way, it is looped a preset number of times, for example, looped 10 times. The correction strategy corresponding to each loop includes the adjustment actions corresponding to each pixel point.
[0136] Based on the method provided in this example, a correction strategy for the image to be corrected can be obtained, and the correction process of the image to be corrected can be completed based on the correction strategy.
[0137] Figure 8 It is a flowchart of the aesthetic-guided automatic white balance method provided by this application Figure 5 As Figure 8 shown, S204 includes:
[0138] S801. Determine multiple target regions corresponding to the image to be corrected based on the instance segmentation information.
[0139] Combined with the scene example, different regions in the instance segmentation information are determined as different target regions.
[0140] S802. Determine the first correction strategy with the highest proportion in the correction strategies corresponding to the pixel points included in each target region, and determine the first correction strategy as the target correction strategy corresponding to the target region.
[0141] Combined with scene examples, based on the correction strategy, determine the adjustment actions corresponding to the pixel points included in each target area, and determine the proportion of each type of adjustment action. For example, taking the target area where a person is located as an example, if the target area where a person is located contains 100 pixel points, among which the pixel points subjected to color temperature increase processing account for 30%, and the pixel points subjected to color temperature decrease processing account for 70%, so the color temperature decrease processing with a higher proportion can be determined as the first correction strategy, and the color temperature adjustment processing can be determined as the target correction strategy corresponding to the target area where the person is located. Similarly, in this way, determine the target correction strategy corresponding to each target area.
[0142] S803. Execute the corresponding target correction strategy for each pixel point in the target area to obtain a corrected image.
[0143] Combined with scene examples, based on the target correction strategy corresponding to each target area, execute the corresponding target correction strategy for each pixel point in the target area. For example, if the target correction strategy corresponding to the target area where a person is located is color temperature adjustment processing, then execute color temperature adjustment processing for all 100 pixel points in the target area where the person is located.
[0144] Based on the method provided in this example, referring to the instance segmentation information to correct the image to be corrected can ensure the coordination between the local and the whole during the correction process.
[0145] Figure 9 The comparison schematic diagram of the present application for the example and the prior art to correct the image is as Figure 9 shown. Taking the same image to be corrected as the input image, based on the mainstream white balance of the prior art, the obtained output image can achieve color fidelity. Based on the method provided in this embodiment, by correcting the color deviation of the image to be corrected in the way of diagonal matrix correction, color fidelity can be achieved, and by optimizing the reward value of the reward function module, the aesthetic feeling of the image to be corrected can be improved to take into account the goal of color aesthetics.
[0146] Optionally, Figure 10 is the flow schematic of the aesthetic-guided automatic white balance method provided by the present application Figure 6 , before S202, as Figure 10 shown, the method further includes:
[0147] S1001. Obtain training data, where the training data includes multiple groups of training images and corresponding standard images.
[0148] Combined with scene examples, obtain paired images before and after manual color adjustment to obtain a data set {I1, I2, , In}, generate paired labels {L1, L2, , Ln}, use the instance segmentation algorithm to obtain the instance segmentation information {M1, M2, , Mn} for each element in the dataset. Combine the paired dataset and label data, and then divide them into a training set and a test set according to a ratio of 8:2. Use the training set as the training data, where each element includes a pair of images. The image before color adjustment can be determined as the training image, and the image after manual color adjustment can be determined as the standard image. The image before manual color adjustment in the test set can be used as the image to be corrected mentioned above.
[0149] S1002. Train based on the pixel reinforcement learning framework preset for the multi-group training images and the corresponding standard images to obtain the reinforcement learning model.
[0150] Combine the scenario examples, combine the above Figure 7 , and based on the training data obtained above, train the preset pixel reinforcement learning framework until each reward function in the reward function module converges to complete the training of the pixel reinforcement learning framework. Based on the method provided in this example, the trained reinforcement learning model can be obtained.
[0151] Optionally, Figure 11 is the flowchart of the aesthetic-guided automatic white balance method provided by this application Figure 7 , as Figure 11 shown, S1002 includes:
[0152] S1101. For each group of training images and the corresponding standard images, input the training images and the corresponding standard images into the preset pixel reinforcement learning framework to obtain the reward value and the correction strategy for each pixel point of the training images.
[0153] Combine the scenario examples, combine Figure 7 , output the training images and standard images in the training set to the preset pixel reinforcement learning framework, estimate the illumination of the training images based on the state module to obtain the preprocessing strategy for the training images. If this is the first processing of the training images, the agent module can obtain the adjustment strategy based on the preprocessing strategy for the training images. If this is not the first processing of the training images, the agent module can obtain the current adjustment strategy based on the preprocessing strategy for the training images and the previous historical adjustment strategies.
[0154] S1102. Optimize the pixel reinforcement learning framework based on the reward value to obtain the optimized pixel reinforcement learning framework.
[0155] Combined with the scenario example, the reward function module obtains the total reward value based on the artifact loss function, the aesthetic quality reward function, and the guidance reward function, and discounts the total reward value based on a preset discount function to obtain the discounted reward value. The parameters in the pixel reinforcement learning framework are optimized based on the discounted reward value.
[0156] S1103. Correct each pixel point of the training image based on the correction strategy to obtain the corrected image.
[0157] Combined with the scenario example, in the environment module, based on the preprocessing strategy, color bias correction is performed on the image to be corrected for the training image through diagonal matrix correction, and the adjustment action corresponding to each pixel point is executed based on the correction strategy to obtain the corresponding corrected image this time.
[0158] S1104. Loop to input the corrected image into the pixel reinforcement learning framework to obtain the reward value and the correction strategy for each pixel point of the corrected image, optimize the pixel reinforcement learning framework based on the reward value to obtain the optimized pixel reinforcement learning framework, and correct each pixel point of the corrected image based on the correction strategy, until the number of loop executions reaches the preset number.
[0159] Combined with the scenario example, based on the above process, the training image is repeatedly executed a preset number of times, for example, looped 10 times, and the discounted reward values corresponding to the 10 loops can be obtained.
[0160] S1105. Update the parameters of the pixel reinforcement learning framework based on all the reward values corresponding to each group of training images and the corresponding standard images until the reward value converges to obtain the reinforcement learning model.
[0161] Combined with the scenario example, the 10 discounted reward values are superimposed to obtain the total discounted reward value, and the parameters of the pixel reinforcement learning framework are updated based on the obtained total discounted reward value to complete the training of the pixel reinforcement learning framework based on this training image.
[0162] For each training image in the training set, repeat the steps of S1201 - S1205 above to complete the training of the pixel reinforcement learning framework for each training image, until in the training process, the artifact loss function, the aesthetic quality reward function, and the guidance reward function all reach the convergence effect, and then the trained reinforcement learning model is obtained.
[0163] In this application, each pixel of the image to be corrected is corrected multiple times through a reinforcement learning model to achieve color balance, and the correction is carried out with reference to instance segmentation information, which can ensure that the coordination between the local and the whole is taken into account during the correction process to achieve the effect of color fidelity.
[0164] Figure 12 FIG. 4 is a schematic structural diagram of an aesthetic-guided automatic white balance device provided by this application. The device is applied to a computing device, such as Figure 12 shown, and includes:
[0165] An acquisition module 121, configured to acquire the image to be corrected and the instance segmentation information corresponding to the image to be corrected;
[0166] A processing module 122, configured to input the image to be corrected into a reinforcement learning model to obtain a reward value and a correction strategy for each pixel point of the image to be corrected;
[0167] The processing module 122 is further configured to optimize the reinforcement learning model based on the reward value to obtain an optimized reinforcement learning model;
[0168] The processing module 122 is further configured to perform correction processing on each pixel point of the image to be corrected based on the correction strategy for each pixel point of the image to be corrected and the instance segmentation information to obtain a corrected image;
[0169] The processing module 122 is further configured to repeatedly execute the steps of inputting the corrected image into the reinforcement learning model to obtain a reward value and a correction strategy for each pixel point of the corrected image, optimizing the reinforcement learning model based on the reward value to obtain an optimized reinforcement learning model, and performing correction on each pixel point of the corrected image based on the correction strategy and the instance segmentation information until the number of repeated executions reaches a preset number to obtain a final image after the final correction is completed.
[0170] The aesthetic-guided automatic white balance device provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.
[0171] Figure 13 FIG. 5 is a schematic structural diagram of an electronic device provided by this application. As Figure 13 shown, the electronic device 50 provided in this embodiment includes: at least one processor 501 and a memory 502. Optionally, the device 50 further includes a communication component 503. Among them, the processor 501, the memory 502, and the communication component 503 are connected through a bus.
[0172] In a specific implementation process, at least one processor 501 executes computer-executable instructions stored in a memory 502, so that at least one processor 501 executes the above-mentioned method.
[0173] For the specific implementation process of the processor 501, reference can be made to the above method embodiment, and its implementation principle and technical effect are similar, so they will not be elaborated here in this embodiment.
[0174] In the above embodiment, it should be understood that the processor may be a central processing unit (English: Central Processing Unit, abbreviated as: CPU), or may also be other general-purpose processors, digital signal processors (English: Digital Signal Processor, abbreviated as: DSP), application specific integrated circuits (English: Application Specific Integrated Circuit, abbreviated as: ASIC), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed and completed by a hardware processor, or can be executed and completed by a combination of hardware and software modules in the processor.
[0175] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.
[0176] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.
[0177] This application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the above-mentioned method is implemented.
[0178] This application also provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the processor executes the computer-executable instructions, the above-mentioned method is implemented.
[0179] The above-readable storage medium may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk. The readable storage medium may be any available medium accessible by a general-purpose or special-purpose computer.
[0180] An exemplary readable storage medium is coupled to the processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium may also be an integral part of the processor. The processor and the readable storage medium may be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium may also exist as discrete components in a device.
[0181] The division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other may be indirect couplings or communication connections through some interfaces, devices, or units, and may be in electrical, mechanical, or other forms.
[0182] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0183] In addition, in each embodiment of the present invention, the functional units may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit.
[0184] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, etc., various media that can store program codes.
[0185] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When this program is executed, it executes the steps including the above method embodiments; and the aforementioned storage medium includes: ROMs, RAMs, magnetic disks, or optical discs, etc., various media that can store program codes.
[0186] Finally, it should be noted that: After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily think of other implementation manners of the present invention. The present invention is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include common general knowledge or conventional technical means in the technical field not disclosed in the present invention. It is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.
Claims
1. An aesthetically guided automatic white balance method, characterized in that The method is applied to a computing device, comprising: Acquire the image to be corrected and instance segmentation information corresponding to the image to be corrected; The image to be corrected is input into a reinforcement learning model to obtain a reward value and a correction strategy for each pixel of the image to be corrected, wherein the reinforcement learning model includes: a state module, an agent module and a reward function module, wherein the reward function module includes: a guidance reward function, an aesthetic quality reward function and an artifact loss function; accordingly, the image to be corrected is input into the reinforcement learning model to obtain a reward value and a correction strategy for each pixel of the image to be corrected, comprising: inputting the image to be corrected into the reinforcement learning model, and performing illumination estimation on the image to be corrected based on the state module, to obtain a preprocessing strategy; the agent module obtains a correction strategy based on the preprocessing strategy and the historical correction strategy; based on the reward function module, obtains a first reward value corresponding to the guidance reward function, a second reward value corresponding to the aesthetic quality reward function and a third reward value corresponding to the artifact loss function; the guidance reward function is the distance change between the image to be corrected and the preset target GT map; the aesthetic quality reward function is the aesthetic change before and after the current image is corrected, which is used to optimize the aesthetics before and after the image is corrected; the artifact loss function is used to calculate the gradient performance difference between the corrected image and the target GT map; summing the first reward value, the second reward value and the third reward value to obtain a total reward value; Based on a preset discount function, the total reward value is discounted to obtain a discounted reward value; based on the discounted reward value, the parameters of the agent module in the reinforcement learning model are modified to obtain an optimized reinforcement learning model; Determine a plurality of target regions corresponding to the image to be corrected based on the instance segmentation information; Determine a first correction strategy with the highest proportion among the correction strategies corresponding to the pixels included in each target area, and determine the first correction strategy as the target correction strategy corresponding to the target area; execute the corresponding target correction strategy on each pixel in the target area to obtain a corrected image; The corrected image is input into the reinforcement learning model in a loop to obtain a reward value and a correction strategy for each pixel of the corrected image, the reinforcement learning model is optimized based on the reward value to obtain an optimized reinforcement learning model, and each pixel of the corrected image is corrected based on the correction strategy and instance segmentation information, until the number of loop executions reaches a preset number, so as to obtain a final image after the final correction is completed.
2. The method according to claim 1, characterized in that The reinforcement learning model also includes: environment module; Accordingly, the step of performing illumination estimation on the image to be corrected based on the state module to obtain a preprocessing strategy includes: Acquiring illumination information of the image to be corrected; Based on the environment module, illumination information of the image to be corrected is estimated and globally corrected by using a diagonal matrix to obtain a corrected correction image; The correction image is determined as the preprocessing result.
3. The method according to claim 1 or 2, characterized in that: Before inputting the image to be corrected into the reinforcement learning model, the method further includes: Acquiring training data, wherein the training data includes multiple sets of training images and corresponding standard images; Training is performed based on a preset pixel reinforcement learning framework of the multiple groups of training images and corresponding standard images to obtain the reinforcement learning model.
4. The method according to claim 3, characterized in that The pixel reinforcement learning framework preset based on the multiple sets of training images and corresponding standard images is trained to obtain the reinforcement learning model, including: For each set of training images and corresponding standard images, the training images and the corresponding standard images are input into a preset pixel reinforcement learning framework to obtain a reward value and a correction strategy for each pixel point of the training image; Optimizing the pixel reinforcement learning framework based on the reward value to obtain an optimized pixel reinforcement learning framework; Correcting each pixel of the training image based on the correction strategy to obtain a corrected image; The corrected image is input into the pixel reinforcement learning framework in a loop to obtain a reward value and a correction strategy for each pixel of the corrected image, the pixel reinforcement learning framework is optimized based on the reward value to obtain an optimized pixel reinforcement learning framework, and each pixel of the corrected image is corrected based on the correction strategy, until the number of loop executions reaches a preset number; The parameters of the pixel reinforcement learning framework are updated based on all reward values corresponding to each group of training images and the corresponding standard images until the reward values converge to obtain the reinforcement learning model.
5. An aesthetically guided automatic white balance device, characterized in that: The device is applied to a computing device, comprising: An acquisition module, used for acquiring an image to be corrected and instance segmentation information corresponding to the image to be corrected; A processing module, used for inputting the image to be corrected into a reinforcement learning model to obtain a reward value and a correction strategy for each pixel of the image to be corrected; The processing module is further used to optimize the reinforcement learning model based on the reward value to obtain an optimized reinforcement learning model; The processing module is further used to perform correction processing on each pixel of the image to be corrected based on the correction strategy for each pixel of the image to be corrected and the instance segmentation information, so as to obtain a corrected image; The processing module is further used to cyclically execute the steps of inputting the corrected image into the reinforcement learning model to obtain a reward value and a correction strategy for each pixel of the corrected image, optimizing the reinforcement learning model based on the reward value to obtain an optimized reinforcement learning model, and correcting each pixel of the corrected image based on the correction strategy and instance segmentation information, until the number of loop executions reaches a preset number of times, so as to obtain a final image after the final correction is completed; The reinforcement learning model includes: a state module, an agent module and a reward function module, wherein the reward function module includes: a guidance reward function, an aesthetic quality reward function and an artifact loss function. Accordingly, the processing module is specifically used to: input the image to be corrected into the reinforcement learning model, and perform illumination estimation on the image to be corrected based on the state module to obtain a preprocessing strategy; the agent module obtains a correction strategy based on the preprocessing strategy and the historical correction strategy; based on the reward function module, obtain a first reward value corresponding to the guidance reward function, a second reward value corresponding to the aesthetic quality reward function and a third reward value corresponding to the artifact loss function; the guidance reward function is the distance change between the image to be corrected and the preset target GT map; the aesthetic quality reward function is the aesthetic change of the current image before and after correction, which is used to optimize the aesthetics before and after image correction; the artifact loss function is used to calculate the gradient performance difference between the corrected image and the target GT map; The processing module is further specifically used to sum the first reward value, the second reward value and the third reward value to obtain a total reward value; based on a preset discount function, discount the total reward value to obtain a discounted reward value; based on the discounted reward value, modify the parameters of the agent module in the reinforcement learning model to obtain an optimized reinforcement learning model; The processing module is also specifically used to determine multiple target areas corresponding to the image to be corrected based on the instance segmentation information; determine the first correction strategy with the highest proportion among the correction strategies corresponding to the pixels contained in each target area, and determine the first correction strategy as the target correction strategy corresponding to the target area; execute the corresponding target correction strategy for each pixel in the target area to obtain a corrected image.
6. A computer device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 4 when executed by a processor.
Citation Information
Patent Citations
Underwater image enhancement method and device for reinforcement learning parameter optimization, and medium
CN115423724A
Image white balance adjustment method and device, electronic equipment and storage medium
CN117221741A