AWB system and method for complex light environment

Through feature decoupling, reweighting and light-sensitive feature comparison learning modules, the feature selection and processing problems in the multi-light color constancy algorithm are solved, and the efficient white balance effect is achieved in complex lighting environments, improving the accuracy of light estimation and model lightweighting.

CN120375010APending Publication Date: 2025-07-25BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510537682.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing multi-light color constancy algorithm has shortcomings in feature selection, supervision and feature processing, resulting in low accuracy of lighting estimation and high computational complexity, making it difficult to effectively deal with color shift phenomena in complex lighting environments.

Method used

The feature decoupling module, feature reweighting module, efficient and lightweight extraction shallow feature module and a comparison learning module based on light-sensitive feature Loguv are used to decouple and recombinate chromaticity features and color context features respectively. The complexity of the model is reduced through depth separation technology, and the comparison learning is used to improve the light estimation effect using light-sensitive features.

Benefits of technology

It realizes a more accurate white balance effect under multi-light source conditions, reduces model complexity, improves the accuracy and efficiency of light estimation, and is suitable for color constancy tasks in complex lighting environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375010A_ABST
    Figure CN120375010A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of color constancy, and discloses an AWB system and method for a complex light environment. Through a comparative learning module based on the illumination sensitive feature Loguv, the essential difference between single light and multiple light when the illumination condition changes is revealed by using the illumination sensitive feature Loguv, and through comparative learning, the illumination estimation effect is improved from the illumination dimension. Through a feature decoupling module and a feature reweighting module, illumination estimation is carried out by using chromaticity features and color context features at the same time. The high-efficiency lightweight shallow feature extraction module solves the problem of high-redundancy local representation when VIT-AWB captures shallow features, reduces the complexity of the model by using a depth separable technology, and further meets the lightweight demand of a multi-light color constancy network. The invention further provides an AWB method, the color constancy of the color cast image under the multi-light-source condition can be effectively achieved, and the more accurate white balance effect is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of color constancy in computer vision, and particularly relates to an AWB system and method for complex illumination environments. Background Art

[0002] Computing color constancy is an important part of the image processing pipeline. It mimics the color constancy ability of the human visual system (even when the scene illumination changes, the human visual system can perceive the color of an object as constant). This process is also called white balance and is the first step in the camera image processing pipeline. In other words, the purpose of computing color constancy is to endow photographic equipment with the ability of color constancy and eliminate the color cast phenomenon in images caused by ambient light sources. To achieve the function of color constancy and eliminate the color cast in images, many color constancy methods have been proposed.

[0003] Currently, most of the research on color constancy algorithms focuses on single light source conditions, that is, it is assumed that there is only one uniform light source in the scene. However, the premise assumptions of these methods are too ideal. In real life, multi-light scenarios are more common, and there are often multiple light sources with different quantities, ranges, positions, and intensities in the scene, which usually produce shadows and cause mutual reflection between objects. This makes multi-light color constancy more challenging than single uniform illumination scenarios.

[0004] Secondly, there are problems in feature selection and processing in the current research on multi-light color constancy algorithms. That is: (1) In the current multi-light color constancy field, most feature selections directly use the original features of the image or transform the original features into another color space as input. However, the original features of the image include chromaticity features, context features, texture features, frequency domain features, etc., and not all features are beneficial to light estimation. (2) In the feature selection under model supervision, features unrelated to light are used for supervision loss. Due to the coupling problem between light and the surface reflectivity of objects, these will limit the accuracy of light estimation. (3) In feature processing, when the selected features are input into VIT-AWB for shallow feature extraction, there are problems of high redundancy and high computational complexity.

[0005] In summary, existing multi-light color constancy algorithms still need to be improved and perfected. Summary of the Invention

[0006] Aiming at the deficiencies in the prior art, the present invention proposes an AWB system and method for complex illumination environments. This method can effectively achieve color constancy for color cast images under multi-light source conditions and achieve a more accurate white balance effect. The specific technical solutions are as follows:

[0007] An AWB system for complex lighting environments, including a feature decoupling module, a feature reweighting module, an efficient lightweight shallow feature extraction module, and a contrast learning module based on the light-sensitive feature Loguv. Among them,

[0008] The feature decoupling module includes a depthwise convolution block and a pointwise convolution block;

[0009] The efficient lightweight shallow feature extraction module includes a 3×3 convolution layer, a 3×3 depthwise separable position encoding layer, a clustering sub-module, a depthwise separable multi-head attention mechanism model, and an upsampling sub-module;

[0010] The contrast learning module based on the light-sensitive feature Loguv includes a UNet-Decoder sub-module and a contrast learning sub-module.

[0011] Preferably, the depthwise convolution block is used to extract color context features from the input image.

[0012] Preferably, the pointwise convolution block is used to extract chromaticity information from the input image.

[0013] Preferably, the feature reweighting module combines the color context features and the chromaticity features to obtain the effective feature I'.

[0014] Preferably, the efficient lightweight shallow feature extraction module is used to obtain the first output feature X according to the effective feature I' out , and obtain a white balance image.

[0015] Preferably, the contrast learning module based on the light-sensitive feature Loguv is used to further optimize the predicted white balance image according to the first output feature X out .

[0016] An AWB method for complex lighting environments, implemented using the AWB system, includes the following steps:

[0017] S1. Use the preprocessed RAW-RGB image as the input image I;

[0018] S2. The input image is input into the feature decoupling module to extract chromaticity features and color context features;

[0019] S3. Input the chromaticity features and color context features into the feature reweighting module to obtain the effective feature I';

[0020] S4. Input the effective feature I' into the efficient lightweight shallow feature extraction module to obtain the first output feature X out ;

[0021] S5. Input the first output feature X out into the contrastive learning module based on the light-sensitive feature Loguv to obtain a white balance image.

[0022] Preferably, S2 specifically includes:

[0023] S2-1. Input the input image I into the pointwise convolution block in the feature decoupling module to extract the chromaticity feature;

[0024] S2-2. Input the input image I into the depthwise convolution block in the feature decoupling module to extract the color context feature.

[0025] Preferably, S4 specifically includes:

[0026] S4-1. Input the effective feature I′ obtained by the feature reweighting module into a 3×3 convolutional layer to obtain the input token tensor X in ;

[0027] S4-2. Use a 3×3 depthwise separable convolutional layer to perform depthwise separable convolution on the input token tensor X in and add positional encoding;

[0028] S4-3. In the clustering sub-module, cluster the input token tensor X with positional encoding added in by similarity to obtain the clustering tensor S;

[0029] S4-4. Input the clustering tensor S into the multi-head attention mechanism model, independently calculate self-attention for each channel of the clustering tensor S, and realize information interaction between channels through pointwise convolution to obtain the tensor S′;

[0030] S4-5. The upsampling sub-module performs upsampling on the tensor S′, maps it back to the original token space, and repeats S4-2 to S4-4 three times to obtain the first output feature X out 。

[0031] Preferably, S5 specifically includes:

[0032] S5-1. Use the UNet-Decoder sub-module as the light estimation layer to obtain the output image output according to the first output feature X out ;

[0033] S5-2. Input the output image output into the contrast sub-module, and use the light-sensitive feature Loguv to further improve the effect of light estimation to obtain a white balance image.

[0034] The technical solution of the present invention is specifically as follows:

[0035] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0036] 1. The AWB system for complex lighting environments of the present invention adopts a contrastive learning module based on the lighting-sensitive feature Loguv. The feature Loguv sensitive to lighting is used to reveal the essential differences between single light and multi-light when the lighting conditions change, and based on this, as a new lighting-sensitive feature, the lighting estimation effect is improved from the lighting dimension through contrastive learning.

[0037] 2. The AWB system for complex lighting environments of the present invention, through the feature decoupling module and the feature reweighting module, decouples the original features of the input image into chromaticity features and color context features, and assigns learnable weight values to the chromaticity and color context features and combines them, enabling the model to not only focus on using chromaticity features for lighting estimation but also be able to utilize color context features to cope with the influence of mixed light sources.

[0038] 3. The efficient shallow feature extraction module proposed by the AWB system for complex lighting environments of the present invention uses the SuperToken Attention technology for the first time to solve the problem of high redundancy of local representations when capturing shallow features by the previous VIT-AWB. At the same time, the depthwise separable technology is used to reduce the complexity of the model, further meeting the lightweight requirements of the multi-light color constancy network. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. By referring to the drawings, the features and advantages of the present invention can be more clearly understood. The drawings are schematic and should not be construed as imposing any limitations on the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0040] Figure 1 is the architecture diagram of the AWB system for complex lighting environments of the present invention.

[0041] Figure 2 is the workflow diagram of the feature decoupling module and the feature reweighting module of the present invention.

[0042] Figure 3 is the workflow diagram of the efficient lightweight shallow feature extraction module of the present invention.

[0043] Figure 4 is the workflow diagram of the contrastive learning module based on the lighting-sensitive feature Loguv of the present invention.

[0044] Figure 5It is the working principle diagram of the pointwise convolution block in the application scenario of Embodiment 1.

[0045] Figure 6 It is the working principle diagram of the depthwise convolution block in the application scenario of Embodiment 1.

[0046] Figure 7 It is the working principle diagram of the feature reweighting module in the application scenario of Embodiment 1.

[0047] Figure 8 It is the working principle diagram of the efficient lightweight shallow feature extraction module in the application scenario of Embodiment 1.

[0048] Figure 9 It is the working principle diagram of the contrast learning module based on the light-sensitive feature Loguv in the application scenario of Embodiment 1. Detailed implementation manners

[0049] In order to be able to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention will be further described in detail below with reference to the drawings and specific implementation manners. It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.

[0050] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.

[0051] Aiming at the problem that existing multi-light color constancy algorithms simply use angular error or other illumination-independent feature losses as the supervision of the model, and do not utilize the light-sensitive features strongly related to illumination, resulting in the input features still having coupling, which further limits the final illumination estimation effect. The present invention designs a contrast learning module based on the light-sensitive feature Loguv. By using the light-sensitive feature Loguv and taking advantage of the divergence characteristics of multi-light images on the Loguv histogram when the illumination conditions change for contrast learning, this module enables the model to pay more attention to the illumination information of the image and further improves the illumination estimation effect.

[0052] For the problem that most current deep learning-based multi-light color constancy algorithms directly use the original features of images as inputs for illumination estimation, the present invention first explores and discovers that for the multi-light color constancy task, low-level features such as chromaticity features are more conducive to estimating the color temperature of the illumination in the image compared to the original features. Moreover, previous methods did not consider the color context features. For pixels in different color contexts, the degree of influence by illumination from different positions and directions is different, resulting in different superimposed effects of different illuminations on each pixel, which were not considered by previous multi-light methods and may lead to a relatively low accuracy of current illumination estimation. The present invention designs a feature decoupling module and a feature reweighting module, which can efficiently and lightweightly extract chromaticity features and color context features from the original features of the image, and assign learnable weights to the two and combine them to obtain input features that are more effective for the multi-light source color constancy task.

[0053] Regarding the situation that current VIT-AWB does not consider the high redundancy that exists when extracting shallow features, which leads to the inability of distant tokens to be fully utilized when calculating self-attention, thereby reducing the illumination estimation effect and causing unnecessary computational costs. At the same time, due to the high model complexity of existing VITs, they do not meet the requirements of lightweight color constancy networks. The present invention designs an efficient and lightweight shallow feature extraction module, which can reduce the number of tokens and can also make good use of distant tokens, thus solving the problem of high redundancy. The depthwise separable technology is used to calculate self-attention, achieving the lightweighting of the model.

[0054] As Figure 1-2 shown, the AWB system for complex illumination environments proposed by the present invention includes a feature decoupling module, a feature reweighting module, an efficient and lightweight shallow feature extraction module, and a contrast learning module based on the illumination-sensitive feature Loguv.

[0055] The working process of the AWB system for complex illumination environments is as follows:

[0056] S1. Use the preprocessed RAW-RGB image as the input image I;

[0057] S2. The input image is input into the feature decoupling module to extract chromaticity features and color context features;

[0058] S3. Input the chromaticity features and color context features into the feature reweighting module to obtain the effective feature I′;

[0059] S4. Input the effective feature I′ into the efficient and lightweight shallow feature extraction module to obtain the first output feature X out ;

[0060] S5. Input the first output feature X out into the contrastive learning module based on the light-sensitive feature Loguv to obtain a white balance image.

[0061] Figure 2 The processing flow in the feature decoupling module is shown as follows:

[0062] S2-1. The input image I is input into the pointwise convolution block in the feature decoupling module to extract the chromaticity feature;

[0063] S2-2. The input image I is input into the depthwise convolution block in the feature decoupling module to extract the color context feature.

[0064] The processing flow in the feature reweighting module is as Figure 2 shown:

[0065] In the feature reweighting module, the chromaticity feature and the color context feature are combined through the feature reweighting mechanism to obtain an effective feature I' that is more effective for the multi-light color constancy task.

[0066] The processing flow in the efficient lightweight shallow feature extraction module is as Figure 3 shown:

[0067] S4-1. Input the effective feature I' obtained by the feature reweighting module into a 3×3 convolutional layer to obtain an input token tensor X in ;

[0068] S4-2. Use a 3×3 depthwise separable position encoding layer to perform depthwise separable convolution on the input token tensor X in to add position encoding;

[0069] S4-3. In the clustering sub-module, cluster the input token tensor X with position encoding added in by similarity to obtain a clustering tensor S;

[0070] S4-4. Input the clustering tensor S into a depthwise separable multi-head attention mechanism model, independently calculate self-attention for each channel of the clustering tensor S, and realize information interaction between channels through pointwise convolution to obtain a tensor S';

[0071] S4-5. The upsampling sub-module performs upsampling on the tensor S', maps it back to the original token space, and repeats S4-2 to S4-4 three times to obtain the first output feature X out .

[0072] The processing flow in the contrastive learning module based on the light-sensitive feature Loguv is asFigure 4 As shown in:

[0073] S5-1: Use the UNet-Decoder sub-module as the illumination estimation layer, and obtain the output image output according to the first output feature X out to obtain the output image output;

[0074] S5-2: Input the output image output into the contrast sub-module, and use the illumination-sensitive feature Loguv to further improve the effect of illumination estimation to obtain the white balance image.

[0075] Embodiment 1

[0076] Combined with the application scenario, the AWB method of the present invention for complex illumination environments is described in detail.

[0077] The preprocessed RAW-RGB image is successively passed through a feature decoupling module, a feature reweighting module, an efficient lightweight shallow feature extraction module, and a contrast learning module based on the illumination-sensitive feature Loguv to obtain a more accurate illumination estimation result and obtain a white balance image closer to the real image for further application in downstream tasks.

[0078] Step 1: Input the preprocessed RAW-RGB image I and the real white balance image GT.

[0079] The size of the input RAW-RGB image I and the real white balance image GT can be expressed as C×H×W, where C represents the number of channels of the image, H is the height of the image, which can be represented by the number of pixels in the vertical dimension of the image; W is the width of the image, which can be represented by the number of pixels in the horizontal dimension of the image.

[0080] Preferably, the size of the input RAW-RGB image I and the real white balance image GT can be 3×512×512, indicating that the image has three channels of red (Red, R), green (Green, G), and blue (Blue, B), with a height of 512 pixels (pixel, px) and a width of 512 px.

[0081] Step 2: For the RAW-RGB image I, assuming the number of pictures in a batch is 2, the scale is 2×3×512×512. As Figure 5 shown, first input the chrominance branch in the feature decoupling module. This process is implemented using pointwise convolution block operations with a stride of 2. The specific calculation formula is as follows. The output tensor represents the chrominance correlation of each pixel, and the obtained chrominance feature Y has a size of 2×6×256×256.

[0082]

[0083] Among them, X(i·s,j·s,c) represents the value of the c-th channel at position (i,j) in the input RAW image, s is the stride with a value of 2, W(1,1,c,k) represents the learned weight that connects the convolutional kernel of the c-th channel and the k-th channel, Y(i,j,k) represents the output chromaticity feature tensor of the k-th channel at position (i,j), the function σ(·) is the non-linear activation function ReLU, and C im represents the number of channels of the current input feature.

[0084] On the other side, as Figure 6 shown, the RAW-RGB image I is input into the color context branch (i.e., the depthwise convolution block) in the feature decoupling module. This process is implemented by depthwise convolution using two depth multipliers, and the specific calculation formula is as follows. Depthwise convolution is independently performed on each channel to obtain the color context feature D with a size of 2×6×256×256.

[0085]

[0086] Among them, D(·) represents the value of the m'-th output channel generated from the c-th input channel, m' represents the depth multiplier with a value of {0,1}, and each input channel c generates two output channels through different convolutional kernel weights K.

[0087] Step 3: As Figure 7 shown, after extracting the chromaticity feature and color context feature in the RAW-RGB image I through the two branches, they are combined together through the feature reweighting mechanism of the learnable adaptive weights α and β, and the sum of the weights of α and β is 1. The specific calculation formula is as follows. Finally, the effective feature I′ that is more effective for the multi-light color constancy task is obtained, with a size of 2×6×256×256.

[0088] I′(i,j,k) = α·Y(i,j,k) + β·D(i,j,k)

[0089] Step 4: As Figure 8 shown, after passing through the feature decoupling module and the feature reweighting module, the effective feature I′ that is more crucial for the multi-light color constancy task is obtained. Through the transmission of three 3×3 convolutional layers with convolutional strides of 2, 2, and 1, the input token tensor is generated where the values of C, H, and W are related to the current layer number. Subsequently, all the input token tensors X in apply 3*3 depthwise separable convolution to add positional encoding, and the size after addition is where N is the total number of original tokens. Then, the tokens with relatively high similarity among these original tokens are aggregated into a new clustering tensor The specific method is to first initialize the original tokens as supertokens by assuming that each token in X in belongs to one of the supertokens in S, and S is obtained by averaging the initialization of the tokens within a region of size h*w. Then, the similarity between each token and its surrounding 9 supertokens is calculated in a self-attention-like manner. The specific calculation formula is as follows.

[0090]

[0091] where d is equal to the number of channels C, based on the calculated similarity Q, and X represents the original tokens t represents the current iteration number, and the weighted sum of tokens is calculated as the new clustering tensor S, as shown in the following formula.

[0092] S = (Q t ) T X

[0093] By aggregating similar tokens into a new supertoken, the number of tokens is reduced, thus providing a more compact representation for shallow feature extraction in the color constancy task.

[0094] To achieve lightweight, the idea of depthwise separable is added to the multi-head attention mechanism. After aggregating tokens with high similarity into new supertokens, assuming the current supertoken has C channels, self-attention calculations are independently performed on each of the C channels, and finally the self-attention results calculated for each channel are concatenated. The concatenated self-attention results are used for channel interaction operations using 1*1 pointwise convolution, enabling information to flow between different channels, and finally obtaining the tensor S'.

[0095] Next, the input to the illumination estimation layer is obtained by mapping the new tensor S' back to the original token space, as shown in the following formula, by upsampling on the new supertokens and processing using the similarity Q.

[0096] TU(Attn(S')) = Q·Attn(S)

[0097] where Att(·) represents the self-attention calculation.

[0098] Step 5, as Figure 9As shown, after more effective shallow feature extraction is achieved through the efficient lightweight shallow feature extraction module, the extracted features are input into the UNet-Decoder sub-module for illumination estimation, and the size of the obtained white balance image is 2×3×512×512.

[0099] The RAW-RGB image, the white balance image, and the ground truth white balance image GT are subjected to overall contrast learning by applying contrast loss in the contrast sub-module. These three features are input into the MLP projection block in the efficient lightweight shallow feature extraction module and are respectively denoted as z i 、z o 、z g . The anchor sample v * is a vector randomly selected from z o . The positive sample v + is two vectors at the same position in z i and z g . The negative sample v - is a vector randomly selected from different positions in z i and z g . The decoupled infoNCE (DCE) loss is used as the contrast loss, which is based on the cross-entropy loss and calculates the similarity between features. In this way, the predicted image is further updated, and the size of the updated white balance image obtained is still 2×3×512×512.

[0100]

[0101] Among them, L contrastive represents the contrast learning loss, τ represents a fixed parameter, represents the positive sample, represents the negative sample.

[0102] Step 6: Repeat the aforementioned Steps 1-3 to train the feature decoupling module, the feature reweighting module, the efficient lightweight shallow feature extraction module, and the contrast learning module based on the light-sensitive feature Loguv. When each loss function no longer decreases and tends to be stable, stop the training.

[0103] Step 7: The single multi-light color-shifted RAW image to be processed, with a size of 3×512×512, is successively passed through the feature decoupling module, the feature reweighting module, the efficient lightweight shallow feature extraction module, and the contrast learning module based on the light-sensitive feature Loguv to obtain the final white balance image with a size of 3×512×512.

[0104] In the public dataset, the AWB method for complex illumination environments of the present invention is compared with the prior art in terms of accuracy. The results are shown in Table 1, and the present invention has achieved the best results on the LSMI test set.

[0105] Table 1

[0106] Method Average value Median Worst 25% Best 25% Gray World 11.3 8.8 20.74 4.93 White Patch 12.8 14.3 23.49 5.6 SVWB-UNet 2.31 1.91 4.24 1.01 Patch CNN 4.82 4.24 —— —— Angular GAN 4.69 3.88 —— —— TranCC 2.76 2.49 4.13 1.97 PWCC 2.83 2.7 3.8 1.86 Slot-AWB 2.71 2.34 3.31 1.77 Our method 1.57 1.53 3.03 1.16

[0107] In the present invention, unless otherwise clearly defined and limited, terms such as "installed", "connected", "joined", "fixed", etc. shall be construed in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0108] In the present invention, unless otherwise clearly defined and limited, the first feature being "on" or "under" the second feature may include the direct contact between the first and second features, or may include the situation where the first and second features are not in direct contact but in contact through additional features therebetween. Moreover, the first feature being "above", "over" and "on top of" the second feature includes that the first feature is directly above and obliquely above the second feature, or merely indicates that the horizontal height of the first feature is higher than that of the second feature. The first feature being "under", "beneath" and "underneath" the second feature includes that the first feature is directly below and obliquely below the second feature, or merely indicates that the horizontal height of the first feature is lower than that of the second feature.

[0109] In the present invention, the terms "first", "second", "third", "fourth" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. The term "plurality" means two or more, unless otherwise clearly limited.

[0110] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An AWB system for complex lighting environments, characterized in that, It includes a feature decoupling module, a feature reweighting module, an efficient lightweight shallow feature extraction module, and a contrastive learning module based on the light-sensitive feature Loguv. Among them, the feature decoupling module includes a depthwise convolution block and a pointwise convolution block; the efficient lightweight shallow feature extraction module includes a 3×3 convolutional layer, a 3×3 depthwise separable position encoding layer, a clustering sub-module, a depthwise separable multi-head attention mechanism model, and an upsampling sub-module; the contrastive learning module based on the light-sensitive feature Loguv includes a UNet-Decoder sub-module and a contrastive learning sub-module.

2. The AWB system according to claim 1, wherein The depthwise convolution block is used to extract color context features from the input image.

3. The AWB system according to claim 2, characterized in that, The pointwise convolution block is used to extract chromaticity information from the input image.

4. The AWB system according to claim 3, wherein The feature reweighting module combines the color context features and chromaticity features to obtain the effective feature I'.

5. The AWB system according to claim 1, characterized in that, The efficient lightweight shallow feature extraction module is used to obtain the first output feature X based on the effective feature I′ out , and obtain a white balance image.

6. The AWB system according to claim 4, wherein The contrastive learning module based on the light-sensitive feature Loguv is used to further optimize the predicted white balance image according to the first output feature X out ​ 7. An AWB method for complex lighting environments, characterized in that, It is implemented by using the AWB system described in any one of claims 1-6, including the following steps: S1. Use the preprocessed RAW-RGB image as the input image I; S2. The input image is input into the feature decoupling module to extract chromaticity features and color context features; S3. The chromaticity features and color context features are input into the feature reweighting module to obtain the effective feature I'; S4. The effective feature I′ is input into the efficient lightweight shallow feature extraction module to obtain the first output feature X out ; S5. Input the first output feature X out into the contrastive learning module based on the light-sensitive feature Loguv to obtain a white balance image.

8. The AWB method according to claim 7, wherein The specific content of S2 includes: S2-1. The input image I is input into the pointwise convolution block in the feature decoupling module to extract chromaticity features; S2-2. The input image I is input into the depthwise convolution block in the feature decoupling module to extract color context features.

9. The AWB method according to claim 8, wherein The specific content of S4 includes: S4-1. Input the effective feature I' obtained by the feature reweighting module into a 3×3 convolutional layer to obtain the input token tensor X in ; S4-2. Use a 3×3 depthwise separable position encoding layer to perform depthwise separable convolution on the input token tensor X in to add position encoding; S4-3. In the clustering sub-module, for the input token tensor X with positional encoding added in cluster it by similarity to obtain the clustering tensor S; S4-4. The clustering tensor S is input into the depthwise separable multi-head attention mechanism model, and self-attention is independently calculated for each channel of the clustering tensor S, and information interaction between channels is realized through pointwise convolution to obtain the tensor S'; S4-5. The upsampling sub-module performs upsampling on the tensor S′, maps it back to the original token space, and obtains the first output feature X after repeating S4-2 to S4-4 three times. out 。 10. The AWB method according to claim 9, wherein The specific content of S5 includes: S5-1. Take the UNet-Decoder sub-module as the illumination estimation layer, and obtain the output image output according to the first output feature X out ; S5-2. The output image output is input into the contrast sub-module, and the light-sensitive feature Loguv is used to further improve the effect of light estimation to obtain the white balance image.