Color constancy calculation method based on knowledge distillation

By employing a knowledge-based distillation method, utilizing teacher network training and student network distillation training, and combining data from specific image sensors for fine-tuning, the contradiction between the accuracy and model size of illumination color estimation in embedded devices is resolved. This achieves efficient and accurate color correction, applicable to scenarios such as smartphones, drones, and surveillance equipment.

CN121032876APending Publication Date: 2025-11-28QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511137713.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing methods for calculating color constancy are limited in deployment on resource-constrained embedded devices, have poor accuracy in estimating illumination color, and exhibit a trade-off between model size and performance. In particular, deep learning methods require a large amount of labeled data for training, resulting in high computational complexity.

Method used

A knowledge-based distillation approach is adopted, which involves training the teacher network and the student network through distillation, and fine-tuning with a small amount of labeled and unlabeled data from a specific image sensor to optimize the performance of the student network and achieve efficient and accurate correction of illumination color estimation.

Benefits of technology

It achieves efficient and streamlined color correction in embedded devices, reduces computing and storage resource requirements, improves the accuracy of illumination color estimation and model adaptation performance, and is suitable for miniaturized and low-power scenarios such as smartphones, drones and monitoring equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121032876A_ABST
    Figure CN121032876A_ABST
Patent Text Reader

Abstract

The invention relates to a knowledge distillation-based color constancy calculation method, which comprises the following steps of: S1, training a teacher network to realize image illumination color estimation; s2, distilling and training the student network, and migrating the illumination color estimation capability of the teacher network to the student network with a simple structure and a small number of parameters by using a distillation technology; s3, performing fine tuning to train the student network; and S4, realizing calculation color constancy, inputting test image data into the fine-tuned model to estimate the color of the light source, and further obtaining an image after color correction through an ISP process. The problem of dependence on large-scale labeled data when a special deep learning model is constructed for the image sensor is effectively solved, and the requirements for calculation and storage resources are reduced. An efficient and simple student model can be obtained, and possibility is provided for actual deployment in embedded equipment and an image signal processor chip.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of machine vision, and in particular to a computational color constancy method based on knowledge distillation. BACKGROUND

[0002] The human visual system has a remarkable visual adaptation ability, which can automatically adapt to changes in scene illumination color and perceive the true color of objects. This ability is called color constancy (CC). The goal of computational color constancy (CCC) is to enable computers to mimic this feature of the human visual system, remove color bias caused by changes in illumination in images, and provide high-quality color-bias-free images for subsequent computer vision tasks. CCC technology is widely used in image retrieval, object recognition and many other visual tasks, and is a key component of intelligent visual systems. However, under normal circumstances, the only information available comes from the sensor response of the image. Since the illumination information is closely coupled with the scene content, it is difficult to effectively separate, which makes it challenging to achieve CCC. Current CCC methods usually include a two-step process: first, estimate the illumination color in the image, and then correct the image based on the estimated results using the von Kries model. In this process, the accuracy of the illumination color estimation directly determines the quality of the correction results.

[0003] Existing illumination color estimation methods can be mainly divided into two categories: statistical-based methods and learning-based methods. Statistical-based methods rely on certain assumptions about the scene, such as the gray scale assumption, the white point assumption, etc. These methods have low computational cost, but once the image scene does not meet the assumed conditions, the performance of the method cannot be guaranteed. Learning-based methods, especially deep learning-based methods, have made significant performance improvements in recent years and have become the most advanced light source color estimation technology. However, deep learning methods generally require a large amount of labeled data for training, and the parameter quantity of the resulting model often reaches millions, with high inference calculation complexity. These characteristics have greatly limited the deployment of this type of method in resource-limited embedded devices.

[0004] The advent of knowledge distillation technology provides a possibility to solve the above problems. Knowledge distillation was first proposed by Hinton et al. in 2015, and its core idea is to migrate the knowledge contained in one or more teacher networks to a smaller student network through training, thereby significantly compressing the model parameter quantity while maintaining the model performance. With carefully designed knowledge distillation strategies, not only can the performance loss of the student network due to fewer parameters be compensated for, but in some special scenarios, the performance of the student network can even exceed that of the teacher network.

[0005] The application provides a specific image sensor-oriented scene light color estimation method based on knowledge distillation, aiming to realize the computational color constancy of the specific image sensor. With the precise estimation ability of the teacher network for light color, the student network is guided to improve the light source color estimation performance through distillation learning, and then the precise digital image color correction is realized. Finally, the student network obtained can achieve performance close to or even surpassing the teacher network with a smaller model structure and parameter amount. In addition, the method can fine-tune the student network using a small amount of labeled data and a large amount of unlabeled data of the image sensor, effectively solving the problem of dependence on large-scale labeled data when constructing a special deep learning model for the image sensor, and reducing the requirements for computing and storage resources. Distillation and fine adjustment using the method of the application can obtain an efficient and simplified student model, which provides a possibility for actual deployment in embedded devices and image signal processor (ISP) chips. SUMMARY

[0006] To solve the problems of poor light color estimation accuracy and the contradiction between model size and performance in the task of computational color constancy, the application provides an efficient and simplified color constancy deep model establishment method for deployment on resource-constrained devices. By introducing a deep learning method based on knowledge distillation, the application realizes efficient learning of the color correction ability of the teacher network by a small student network, and fine-tuning of the student network for a specific image sensor to optimize the performance of the student network. Moreover, an effective method is designed to make full use of a small amount of labeled data and a large amount of unlabeled data of the image sensor when adjusting and optimizing the student network.

[0007] A computational color constancy method based on knowledge distillation, the method comprising the following steps:

[0008] The implementation of the application comprises training of the teacher network, distillation training of the student network, fine-tuning of the student network, and realization of computational color constancy.

[0009] S1, training of the teacher network. This step trains a deep learning model as a teacher network to realize image light color estimation. The network structure uses a deep learning network with strong performance, and high-quality RAW image data and corresponding light color labels are used for training to obtain a teacher network model. The specific steps are as follows:

[0010] S11, data preprocessing. After the images in the training set composed of image data and corresponding light color labels are processed by an ISP process, data enhancement is performed, and then standardization is performed.

[0011] The ISP process refers to processing a RAW image I raw to obtain a processed sRGB image I sRGB , denoted as IsRGB = f ISP (I raw , p g ) where p g = (r g , g g , b g ) is the global illumination color of the image. Usually the global illumination takes a white light source, i.e. In this case, the ISP processing is simply denoted as I sRGB = f ISP (I raw ). Before training the teacher network and the student network, the RAW image needs to be processed by the ISP to get the image data required by the training network. The specific process includes: format conversion, linearization processing, color interpolation, white balance correction, color space conversion, gamma correction.

[0012] (1) Format conversion: use dcraw software to convert the RAW image file format to TIFF format. When converting, output linear 16-bit data, and set not to perform white balance correction.

[0013] (2) Linearization processing: linearization processing is to convert the non-linear sensor value of the image into linear luminosity value. First, read the image data from the generated TIFF file and convert it to double-precision floating-point type. Take the minimum value of the pixel in the image as the black bottom value, and the maximum value of the pixel as the saturation point value. Use the formula to normalize the pixel value to the range of 0 to 1: where I i0 , I i are the pixel values of the image before and after normalization, respectively, I max is the saturation point value, and I min is the black bottom value.

[0014] (3) Color interpolation (also known as demosaicing): is the process of reconstructing a complete RGB image from single-channel pixel data in the RAW image. The RAW image data is usually arranged in a Bayer array or similar color filter array (CFA), with four pixels in a group (e.g. BGGR arrangement), and the complete three-channel RGB color information needs to be recovered by interpolation. Specifically, the demosaic function in Matlab can be used to implement the interpolation recovery operation on the linearized RAW image data. Finally, the interpolation result is a complete three-channel RGB image.

[0015] (4) White balance correction: white balance correction is to adjust the gain ratio of each color channel of the demosaiced image using the global illumination color of the image, to eliminate color cast and ensure the color accuracy and naturalness of the final image. If the image illumination color is p g = (r g , gg ,b g A simplified von Kries model is used to correct each channel, and the formula for calculating the corrected pixel value is as follows:

[0016]

[0017] Where I raw,r I raw,g I raw,b These are the original pixel values ​​of the R, G, and B channels of the input image, respectively. corr,r I corr,g I corr,b These are the pixel values ​​of the R, G, and B channels of the corrected image, respectively.

[0018] (5) Color Space Conversion: Convert the image data to the standard sRGB color space. Specifically, first, let the conversion matrix from the standard sRGB to the CIEXYZ color space be denoted as... Next, let XYZ2Cam represent the 3×3 transformation matrix from the camera that captured the image to the camera's color space. The values ​​for XYZ2Cam can be obtained from the camera manufacturer. Then, multiply sRGB2XYZ by XYZ2Cam using matrix multiplication to obtain the transformation matrix sRGB2Cam from sRGB directly to the camera's color space. Finally, calculate the inverse matrix of sRGB2Cam to obtain the 3×3 transformation matrix Cam2sRGB from the camera's color space to sRGB. This matrix is ​​used to accurately convert the corrected image to the standard sRGB color space for correct display on various devices. The conversion formula is as follows: Where I 0,r I 0,g I 0,b These are the RGB channel pixel values ​​of the image before conversion, I r I g I b These are the RGB channel pixel values ​​of the converted image.

[0019] (6) Gamma Correction: Gamma correction makes the details in the shadows of the image more apparent while preventing the highlights from being overexposed. Specifically, the pixel value of the image before gamma correction is [I r ,I g ,I b Using a gamma value of 2.4, the value of each pixel is increased to [value missing] after correction.

[0020] The data augmentation described in step S11 refers to performing a series of transformations on the image before training the model to increase its robustness. Specifically, firstly, geometric transformations are performed, including: randomly cropping the image and resizing it to 512×512 pixels; randomly flipping the image vertically and horizontally; and randomly rotating the image. After the geometric transformations, normally distributed random noise is added to the image's color channels to simulate disturbances caused by changes in illumination. The normally distributed noise added to each color channel is:

[0021] I aug,c =I 0,c +ε,ε~N(0,σ 2 ),

[0022] Where ε represents noise; σ represents the noise standard deviation, the specific value of which is set according to the dataset; I 0,c and I aug,c Let N(0,σ) be the original and enhanced pixel color values ​​in the c channel, respectively, where c∈{r,g,b}; 2 () represents a normal distribution with a mean of 0 and a standard deviation of σ. Adding this noise enhances the diversity of light source types, which is more conducive to improving the model's generalization performance.

[0023] The normalization process described in step S11 refers to calculating the input image I0:

[0024]

[0025] Where μ c and σ c These are the mean and standard deviation of the c-channel of all images in the training set, respectively; I 0,c and I norm,c Here, c represents the original and normalized pixel color values ​​in channel c, respectively, where c ∈ {r, g, b}. Normalization ensures a consistent distribution of the input data, which typically yields better results than rescaling.

[0026] S12, Network forward propagation, minimizing loss, and updating teacher network parameters. Features are extracted using the initial convolutional layers of the teacher network, then passed through convolutional layers with relatively large kernels for dimensionality reduction, further extracting semi-dense feature maps. The semi-dense feature map contains four channels: the first three represent local light source estimates, and the fourth channel represents its confidence level. The feature map is passed to a weighted pooling layer for aggregation from local to global to generate the final illumination color estimate. The global illumination color is calculated using confidence-weighted pooling. Minimize loss function Update teacher network parameters.

[0027] The teacher network can be structured using a convolutional neural network (CNN), such as AlexNet, ResNet, or SqueezeNet. The network structure consists of multiple convolutional layers, pooling layers, and fully connected layers, ultimately outputting a 3D vector representing the global illumination color of the image. In this invention, the teacher network is based on the first five convolutional layers (conv1 to conv5) of AlexNet's feature extraction module, and uses additional convolutional layers and confidence-weighted pooling layers to estimate the global illumination color of the image, finally outputting a vector. Specifically, the teacher network can adopt the following structure:

[0028] (1) Image Input and Feature Extraction Module. Features are extracted using conv1 to conv5 of the pre-trained AlexNet convolutional layers, outputting a feature map. These layers include convolution operations, ReLU activation, and max pooling. The input is an RGB image I with dimensions w×h×3, and the feature extraction module outputs a feature map with dimensions w×h×3. 256 represents the number of channels.

[0029] (2) Semi-dense feature extraction module. The conv6 layer from the AlexNet pre-trained convolutional layer is used for further feature extraction from the input feature map. The kernel size is 6×6×64, and the output feature map size is... Furthermore, channel compression is performed using a 1×1 convolutional kernel from the conv7 layer of the AlexNet pre-trained convolutional layer, resulting in a 4-channel semi-dense feature map with an output size of [size missing]. The first three channels represent the lighting color estimation for each local area. The fourth channel represents the confidence level c for each region. i , i = 1, 2, ..., N correspond to the positions of pixels in the feature map, and N is the number of pixels in the feature map.

[0030] (3) Confidence-weighted pooling module. The confidence-weighted pooling layer combines all local illumination color estimates. and the corresponding confidence level c i Aggregation into global illumination color estimation Confidence-weighted local estimates are Where p i This is the weighted local light source estimation. Global aggregation uses the following expression: Where ∑ i p i The sum of weighted light source estimates for all local regions is used; the normalize operation is used to L2 normalize the light source estimate vector to ensure scale consistency. Through confidence-weighted pooling, the network outputs a global illumination color vector. Used to represent the lighting color of the entire image.

[0031] The goal of the teacher network is to estimate the global illumination color of an image. Make it as close as possible to the light color label loss function Defined as and p gt Angular error between: in, and ||p gt ||2 represents the L2 norm, for and p gt The dot product.

[0032] During training, this invention selects the Adam optimization algorithm; the initial learning rate is set to 10. -4 Learning rate adjustment strategies, such as learning rate decay, are employed to improve training stability. During training, the backpropagation algorithm is used to calculate the gradient of the loss function and update the network parameters θ. T The update rules are as follows: Where, η T It's the learning rate. It is the loss function relative to the network parameters θ T The gradient. The specific convergence and stopping conditions of the model are as follows: During the training process, monitor the changes in the loss function. If the loss value tends to stabilize or decreases very little after multiple iterations, the model is considered to have converged, training can be stopped, and the teacher model and parameters can be saved.

[0033] S2, Distillation Training of the Student Network. This step utilizes knowledge distillation techniques to transfer the illumination and color estimation capabilities of the teacher network to the student network, which has a simpler structure and fewer parameters. The process is as follows:

[0034] S21, Data Preparation and Teacher Guidance. A training set consisting of RAW image data and corresponding standard illumination color labels is constructed for training the student network. The images in the training set are processed using the preprocessing described in S11; the processed images are then input into the teacher network obtained in step S1 to generate guidance information, namely, global light source estimates. Used for student network model distillation.

[0035] S22, the student network performs forward propagation, minimizing the total loss consisting of distillation loss and student-owned loss, and updates the student network parameters. The preprocessed image described in S21 is input into the student network, and the source estimation values ​​of the student network are output. Calculate the total loss based on the student network's estimated values, compute the gradient and backpropagate, and update the parameters of the student network model.

[0036] The student network consists of three main components: a feature extractor (FE), a channel attention block illumination estimation (CAB-IE), and a color feature bag-based illumination refinement (CFB-IR).

[0037] The FE module aims to progressively enhance the representational power of images while preserving spatial information in the early stages of the network. The module takes an input image of size (3, 512, 512) and first passes it through a 3×3 convolutional layer with 64 filters, a stride of 1, and padding of 1, resulting in an output size of (64, 512, 512). A ReLU activation function is then applied, introducing non-linearity. Next, the feature map is fed into the first residual block. This residual block contains three convolutional operations: first, a 1×1 convolution to compress the number of channels to 16; then a ReLU activation; followed by a 3×3 convolution, maintaining the number of channels at 16; another ReLU activation; and finally, a 1×1 convolution to expand the number of channels back to 64. The residual block also includes an identity skip connection that adds the module input to the output, maintaining the output size at (64, 512, 512). The feature map then enters a 5th convolutional layer, a 3×3 convolution with 64 filters, maintaining the same output size. This is followed by the ReLU activation function. Then, the feature map is input into the second residual block, whose structure is exactly the same as the first residual block. It includes the 6th, 7th, and 8th convolutional layers in sequence to complete channel compression and restoration while keeping the spatial dimension unchanged. Next, the feature map is max-pooled using an 8×8 pooling kernel. The last layer performs adaptive average pooling to reduce the spatial dimension of the feature map, retain the main features, and unify the output size (64, 8, 8).

[0038] The CAB-IE module consists of two parts: IE (Light Source Estimation) and CAB (Channel Attention). The IE has five layers. The first layer is a 1×1 convolutional layer for feature extraction and channel adjustment. The second layer is an 8×8 max pooling layer for compressing spatial information. The third layer is another 1×1 convolutional layer that compresses the global features into three channels, corresponding to the RGB light source components. The fourth layer uses adaptive average pooling and L2 normalization. Adaptive average pooling merges the local light source color estimates of each channel into a global light source color estimate. L2 normalization standardizes the light source color estimate: This makes it independent of scale changes, ultimately producing a 3-valued vector representing the image's lighting and color. CAB performs a three-step process. First, adaptive average pooling is applied to the input feature map, resulting in an output size of (64, 1, 1). Second, two 1×1 convolutional layers are used, with ReLU and Sigmoid activation functions respectively, to generate channel attention weights. Third, these weights are applied to the input feature map through channel-wise multiplication.

[0039] The CFB-IR module consists of two parts: IR (Reference Light Source Color Estimation Fine-tuning) and CFB (Color Feature Packet). CFB integrates the prediction results of six traditional color constancy methods (WP, GW, SoG, GGW, GE1, GE2), each outputting a 3-channel RGB vector, for a total of 18 channels. The prediction results of these traditional methods are then concatenated with the 3-channel light source estimation results output by CAB-IE to form a 21-dimensional feature vector. This vector is input to the IR, which is a lightweight multilayer perceptron (MLP). The MLP contains two fully connected layers (dimensions change from 21 to 16 to 3) to further optimize the final light source prediction and perform L2 normalization. It is worth noting that this module is not enabled during distillation but is enabled during fine-tuning to improve fine-tuning accuracy and stability.

[0040] distillation loss The mean squared error (MSE) is used to calculate the difference between the student network output and the teacher network output, defined as:

[0041]

[0042] Where, r T g T b T r s g s b s According to S21 and S22 respectively The color components in T are obtained; r T g T b The distillation temperature is applied to channels R, G, and B respectively to adjust the error scale of each channel.

[0043] The students' own losses The error between the student network's estimated light source color value and the actual label is calculated and defined as:

[0044]

[0045] in, The estimated light source value for the student network, p gt The label represents the real light source, and ∥·∥2 represents solving for the L2 norm.

[0046] The total loss used in distillation training student networks is defined as:

[0047]

[0048] Where λ1 and λ2 are weighting coefficients, adjusted The relative importance of these two items.

[0049] During iterative training, the backpropagation algorithm is used based on the total loss function. Calculate the gradient, and then use this gradient to update the student network parameters. The model parameter update formula is: Where θ s η represents the parameters of the student network. S The learning rate is used. Through multiple iterations, the model performance is gradually optimized to ensure that the student network has similar or even better illumination and color estimation performance than the teacher network on a smaller scale, and the trained student network structure and parameters are saved.

[0050] S3, Fine-tuning training of the student network. To further improve the performance of the student network on a specific camera sensor, fine-tuning was performed using a small amount of labeled and unlabeled data, as detailed below:

[0051] S31, Data Preparation and Teacher Guidance. A training set consisting of RAW image data from a specific camera and corresponding standard illumination color labels is constructed. This set is used for fine-tuning of the student network. The preprocessing described in S11 is applied to the images in the training set. The processed images are then input into the teacher network obtained in step S1 to generate guidance information, namely, global light source estimates. Used for student network model distillation.

[0052] The specific camera mentioned refers to a camera of a certain model with the same type of sensor. RAW images in the training set, if they have corresponding standard lighting colors, are used as labeled data; otherwise, they are used as unlabeled data.

[0053] S32, student network forward propagation, minimizes the loss function and optimizes the student network parameters. The image prepared in S31 is input into the student network, and the output is the light source estimate from the student network. The total loss is calculated based on the estimated values ​​of the student network, and then the gradient is calculated and backpropagated to fine-tune the parameters of the student network model.

[0054] The structure of the student network is consistent with that of the student network in S2. The student network trained and saved in S2 is used as the initial student network in this step. It is worth noting that in the fine-tuning process, the parameters in the feature extraction module (FE) and the light source estimation module (IE) are frozen and the color feature package (CFB) is enabled to further increase the stability of the model.

[0055] For labeled image I w Two loss functions are defined: light source estimation error and image correction error. These two are weighted and summed to construct the total loss function. Light source estimation error The angular error used to evaluate the student network estimation and the real light source label is defined as: in, The estimated light source value for the student network, p gt Labels are for real light sources. Image correction error. Refers to the corrected image I s With real image I gt The pixel-level error between them is defined as: Where N represents the number of pixels in the image, and Representing the corrected image I s and real image I gt The RGB color value at the i-th pixel position. Corrected image I s This refers to the use of light source estimates in the training images. The image obtained through the ISP process described in step S11, i.e. Real Image I gt The training images use real light sources and are labeled p. gt Image I obtained through the ISP process described in step S11 gt =f ISP (I w ,p gt The total loss is denoted as Will By weight and Weighting

[0056] For unlabeled image I wo The light source estimation results of the teacher network As the initial pseudo-label, the image is then white-balanced using the Von Kries model described in step S11 to generate a corrected image. This corrected image is then input back into the teacher model to obtain the second light source prediction result. If the first prediction result... If the image is accurate, then the corrected image should be similar to the image taken under standard illumination. Therefore, the second predicted light source should be close to the standard light source. By calculating the angular error θ between the second predicted light source and the standard light source, the confidence level ω of the initial false label can be quantified, defined as... The smaller θ is, the higher the confidence level ω is. This invention allows... Simultaneously, light source estimation will be utilized using the teacher network. The obtained corrected image I T As a real image, let I gt =I T Based on this, the loss function is calculated for labeled images, and confidence levels are introduced to obtain a weighted loss. Furthermore, to ensure that the adjustment range of network parameters during training on unlabeled data is slightly smaller compared to the case of training on labeled data, let... The introduced correction coefficient μ can be set to 0.3–0.8. The corrected image I T This refers to the use of light source estimates in the training images. The image obtained through the ISP process described in step S11, i.e.

[0057] During iterative training, gradient updates are performed using the backpropagation algorithm, based on the total loss. Calculate the gradient, update the student network parameters, and fine-tune the student network parameters to ensure further performance improvement. The model parameter update formula is: Where θ s η represents the parameters of the student network. opt The learning rate is used. Finally, save the fine-tuned student model structure and parameters.

[0058] S4 calculates the implementation of color constancy. The test image data is input into the model obtained after fine-tuning in S3 to estimate the color of the light source, and then the color-corrected image is obtained through the ISP process.

[0059] The preprocessing described in S11 is applied to the RAW test image I. test The image is processed without data augmentation to obtain an sRGB image; this sRGB image is then input into a model finely tuned by S3 to estimate the image's illumination colors. Then, the estimated light color is used in the ISP processing flow described in S11. Perform color correction to obtain the color-corrected image, i.e. I corr That is, the color-corrected image.

[0060] The beneficial effects of this invention are:

[0061] (1) Effectively improves computational efficiency and resource utilization. This invention uses knowledge distillation technology to transfer the illumination and color estimation capabilities of the teacher network to the student network. This allows the student network to maintain high illumination and color estimation accuracy while significantly reducing the number of network model parameters and computational complexity. The student network achieves performance similar to that of the complex teacher network with a simple structure, and can run efficiently in embedded image signal processor (ISP) chips, significantly reducing hardware resource and power consumption requirements, and providing strong support for resource-constrained devices.

[0062] (2) Optimizing model performance and adapting to specific cameras enables model miniaturization and low-cost deployment. This invention first trains a high-performance but large-scale teacher model using a general dataset. Then, knowledge distillation reduces the model size. Furthermore, a small amount of labeled and unlabeled data from specific image sensors is used to fine-tune the student network, improving the deep model's adaptability to specific camera sensors and effectively enhancing image quality. Compared to traditional deep learning models, this invention significantly reduces network size and computational requirements through knowledge distillation, making it possible for the student network to adapt to embedded hardware environments. This feature not only reduces the deployment cost of the algorithm on devices but also improves the camera's real-time ability to process illumination and color estimation. This method is suitable for miniaturized, low-power scenarios such as smartphones, drones, and monitoring equipment, facilitating the widespread application of high-performance vision algorithms in consumer and industrial products.

[0063] (3) Improve the color constancy performance of imaging devices. By employing targeted optimization training and fine-tuning of the student model, this invention provides an effective solution to the contradiction between the scale and performance of deep learning models, offering an efficient and accurate color correction scheme for imaging devices. Even in scenes with complex and frequently changing lighting, this method can still maintain high color correction capability. Attached Figure Description

[0064] Figure 1 This is a flowchart of the knowledge distillation-based method for calculating color constancy according to the present invention.

[0065] Figure 2 The structure diagram of the student network of this invention mainly consists of three components: a feature extraction module (FE), a channel attention block Illumination Estimation module (CAB-IE), and a color feature bag-based Illumination Refinement module (CFB-IR), where H and W represent the height and width of the image, respectively.

[0066] Figure 3 This is a schematic diagram of the teacher network training method of the present invention. The dataset used for training the teacher model. As training samples, For the input image, For lighting color labels.

[0067] Figure 4 This is a schematic diagram of the student network distillation training method of the present invention. To distill the dataset used to train the student network, As training samples, For the input image, For the illumination color label, CFB-IR is not used in the distillation stage (i.e., CFB-IR is not used in the distillation stage, and the output of the CAB-IE module in the student network is directly used as the output of the student network; in the fine-tuning stage, the student network enables the CFB-IR module, and the output of the CAB-IE module is passed as input to the CFB-IR module, and the output of the CFB-IR module is used as the output of the student network).

[0068] Figure 5 The fine-tuning training of the student network of this invention consists of two parts: a labeled case and an unlabeled case. The figure shows the labeled case. For fine-tuning the model, a labeled dataset is needed. As training samples, For the input image, These are labels for the colors of the light. In the case where there are no labels in the image, For fine-tuning the model, unlabeled datasets, For the input image, the snowflake symbol indicates that the IE portion of the FE module and CAB-IE module was frozen during the fine-tuning process.

[0069] Figure 6 The results of the ablation experiments in this embodiment of the invention are shown. The × and √ under CAB indicate freezing and using CAB during training, respectively; the × and √ under CFB-IR indicate freezing and using CFB-IR during training, respectively. The performance differences of the four configurations are compared, and the optimal value for each evaluation metric is indicated in bold in the table.

[0070] Figure 7This section presents the experimental results of the teacher model and its corresponding distilled student model in this embodiment of the invention, before and after fine-tuning on the Cube+ dataset. FC4 indicates that the teacher network structure used in this embodiment is FC4, and the teacher model is trained based on this structure. Raw Student represents the student model obtained by knowledge distillation training based on the teacher network FC4. LD / UD refers to the fine-tuning training of the distilled model using labeled and unlabeled datasets.

[0071] Figure 8 The experimental results of the embodiments of the present invention have been compared with a number of advanced methods in recent years. However, it is worth noting that the performance of the method of the present invention is closely related to the teacher model used. The better the performance of the teacher model, the better the performance of the student model obtained by fine-tuning on the corresponding device. However, the structure of the FC4 model used as the teacher model in the embodiments of the present invention is not the best performing deep learning model structure at present. Therefore, the performance of the method of the present invention still has great potential for improvement. Detailed Implementation

[0072] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0073] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0074] Figures 1-8 This is a specific embodiment of the present invention, which is a method for calculating color constancy based on knowledge distillation, the process of which is as follows: Figure 1 As shown, the method includes:

[0075] S1. Training of the teacher network. The specific process is as follows: Figure 3 As shown, a deep learning model capable of estimating image illumination and color is trained as the teacher network. The network structure uses a high-performance deep learning network and is trained using high-quality image data and its corresponding illumination and color labels to obtain the teacher network model.

[0076] S11. Data Preprocessing: The images in the training set, which consist of image data and corresponding illumination color labels, are processed through the ISP process, then data augmentation is performed, and then standardization is carried out.

[0077] This embodiment requires obtaining an image dataset and its corresponding light source color reference value dataset. To test and verify the method of this invention, some commonly used and publicly available color constancy image datasets can be used, such as the Cube+ dataset, Gehler-Shi dataset, NUS dataset, ETH3D Dataset, and DepthAWB Dataset. This embodiment uses three datasets: Gehler-Shi, Cube+, and NUS. All three datasets contain image data files representing real-world scene photographs, and each image in the dataset is accompanied by a light source reference color value for the scene. The Gehler-Shi dataset is a collection of 568 RAW images taken with Canon 1D and Canon 5D cameras. Cube+ contains 1365 images and an additional 342 images, all taken with a Canon 550D camera. The NUS dataset contains images taken with multiple different devices, covering various scenes and light source conditions; this dataset provides color correction parameters for various devices.

[0078] The ISP process refers to the processing of RAW images. raw Preprocessing yields the processed sRGB image I sRGB , denoted as I sRGB =f ISP (I raw ,p g ). Where p g =(r g ,g g ,b g () is the global illumination color of the image, usually set to . That is, a white light source. If a white light source is used, it is abbreviated as I. sRGB =f ISP (I raw Before training the teacher network and student network, RAW images need to be processed using the ISP process to obtain the image data required by the network. The specific process includes: format conversion, linearization, color interpolation, white balance correction, color space conversion, and gamma correction.

[0079] (1) Use dcraw software for format conversion, and use the option dcraw-4-TW to convert RAW image files to TIFF format. Output linear 16-bit data during conversion, and set not to perform white balance correction.

[0080] (2) During linearization, image data is first read from the generated TIFF file and converted to double-precision floating-point type. The minimum pixel value is used as the black background value, and the maximum pixel value is used as the saturation point value. The pixel values ​​are then normalized to the range of 0 to 1 using a formula: Where I i0 I i Let I be the pixel values ​​of the image before and after normalization, respectively. max I is the saturation point value. min The values ​​are for the black background and are obtained by consulting the documentation for the corresponding datasets.

[0081] (3) Color interpolation (also known as demosaicing) is the process of reconstructing a complete RGB image from single-channel pixels in RAW image data. RAW image data is typically arranged using a Bayer array or a similar color filter array (CFA), grouped into sets of four pixels (e.g., BGGR arrangement). Interpolation is needed to recover the complete three-channel RGB color information. Specifically, the `demosaic` function in Matlab can be used to perform the interpolation recovery operation on the linearized RAW image data. Ultimately, the interpolated result is a complete three-channel RGB image.

[0082] (4) White balance correction refers to adjusting the gain ratio of each color channel in the de-mosaiced image using the global illumination color of the image to eliminate color cast and ensure accurate and natural colors in the final image. If the image illumination color is p g =(r g ,g g ,b g A simplified von Kries model is used to correct each channel, and the formula for calculating the corrected pixel value is as follows:

[0083]

[0084] Where I raw,r I raw,g I raw,b These are the original pixel values ​​of the R, G, and B channels of the input image, respectively. corr,r I corr,g I corr,b These are the pixel values ​​of the R, G, and B channels of the corrected image, respectively.

[0085] (5) Color space conversion is the process of converting image data to the standard sRGB color space. Specifically, firstly, let the conversion matrix from the standard sRGB to the CIE XYZ color space be denoted as... Next, let XYZ2Cam represent the 3×3 transformation matrix from the camera that captured the image to the camera's color space. The values ​​for XYZ2Cam can be obtained from the camera manufacturer. Then, multiply sRGB2XYZ by XYZ2Cam using matrix multiplication to obtain the transformation matrix sRGB2Cam from sRGB directly to the camera's color space. Finally, calculate the inverse matrix of sRGB2Cam to obtain the 3×3 transformation matrix Cam2sRGB from the camera's color space to sRGB. This matrix is ​​used to accurately convert the corrected image to the standard sRGB color space for correct display on various devices. The conversion formula is as follows: Where I 0,r I 0,g I 0,b These are the RGB channel pixel values ​​of the image before conversion, I r I g I b These are the RGB channel pixel values ​​of the converted image.

[0086] (6) Gamma Correction: Gamma correction makes the details in the shadows of the image more apparent while preventing the highlights from being overexposed. Specifically, the pixel value of the image before gamma correction is [I r ,I g ,I b Using a gamma value of 2.4, the value of each pixel is increased to [value missing] after correction.

[0087] The data augmentation refers to performing a series of transformations on the image before training the model to increase its robustness. Specifically, firstly, geometric transformations are performed, including: randomly cropping the image and resizing it to 512×512 pixels, randomly flipping the image vertically and horizontally, and randomly rotating it. After the geometric transformations, normally distributed random noise is added to the image's color channels to simulate disturbances caused by changes in illumination. The normally distributed noise added to each color channel is: I aug,c =I 0,c +ε,ε~N(0,σ 2 ), where ε is noise; σ is the noise standard deviation, the specific value of which is set according to the dataset; I 0,c and I aug,c Let N(0,σ) be the original and enhanced pixel color values ​​in the c channel, respectively, where c∈{r,g,b}; 2 () represents a normal distribution with a mean of 0 and a standard deviation of σ. Adding this noise enhances the diversity of light source types, which is more conducive to improving the model's generalization performance.

[0088] The standardization process refers to the calculation performed on each input image: Where μ c and σ cThese are the mean and standard deviation of the c-channel of all images in the training set, respectively; I 0,c and I norm,c Here, c represents the original and normalized pixel color values ​​for the c channel, respectively, where c ∈ {r, g, b}. Normalization ensures a consistent distribution of the input data, and this has proven to yield better results than rescaling.

[0089] S12: Network forward propagation, minimizing loss, and updating teacher network parameters. The forward propagation process of the network refers to the input data being processed through each layer of the model, progressively extracting features and generating the final output. The first few convolutional layers of the teacher network extract features from the input image. The convolutional layers apply filters (i.e., convolutional kernels) to the input image to extract image features. These features are then passed through convolutional layers with relatively large kernels for dimensionality reduction, further extracting semi-dense feature maps. The semi-dense feature map contains four channels: the first three channels represent local light source estimates, and the fourth channel represents its confidence level. These feature maps are passed to weighted pooling layers for aggregation from local to global to generate the final illumination and color estimate. The teacher network structure can use a convolutional neural network (CNN), such as AlexNet, ResNet, SqueezeNet, etc. The network structure consists of multiple convolutional layers, pooling layers, fully connected layers, etc., and finally outputs a 3D vector representing the global illumination and color of the image. In this invention, the network is based on the first five convolutional layers (conv1 to conv5) of AlexNet for feature extraction, and uses additional convolutional layers and confidence-weighted pooling layers to estimate the global illumination and color of the image, ultimately outputting a vector. Specifically, the teacher network can adopt the following structure: (1) Image input and feature extraction module. Features are extracted using conv1 to conv5 of the AlexNet pre-trained convolutional layers, and the output feature map is generated. These layers include convolution operations, ReLU activation, and max pooling. An RGB image I with size w×h×3 is input, and the feature extraction module outputs a feature map with size w×h×3. 256 is the number of channels. (2) Semi-dense feature extraction module. The conv6 layer in the AlexNet pre-trained convolutional layer is used to further extract features from the input feature map. The kernel size is 6×6×64, and the output feature map size is... Furthermore, channel compression is performed using a 1×1 convolutional kernel from the conv7 layer of the AlexNet pre-trained convolutional layer, resulting in a 4-channel semi-dense feature map with an output size of [size missing]. The first three channels represent the lighting color estimation for each local area. The fourth channel represents the confidence level c for each region. i, i = 1, 2, ..., N correspond to the positions of pixels in the feature map, and N is the number of pixels in the feature map. (3) Confidence-weighted pooling module. The confidence-weighted pooling layer combines all local illumination color estimates. and the corresponding confidence level c i Aggregation into global illumination color estimation Confidence-weighted local estimates are Where p i This is the weighted local light source estimation. Global aggregation uses the following expression: Where ∑ i p i The sum of weighted light source estimates for all local regions is used; the normalize operation is used to L2 normalize the light source estimate vector to ensure scale consistency. Through confidence-weighted pooling, the network outputs a global illumination color vector. Used to represent the lighting color of the entire image.

[0090] The goal of the teacher network is to estimate the global illumination color of an image. Make it as close as possible to the light color label This invention defines a loss function. To measure the difference between the two. Loss function. Defined as and p gt Angular error between: in, and ||p gt ||2 represents the L2 norm, which reflects the magnitude of each vector. for and p gt The dot product represents the similarity between the two. The arccos function calculates the angle between them; the smaller the angle error, the smaller the error between the network's estimated lighting color and the actual lighting color. This is achieved by minimizing the loss function... The network can progressively optimize its parameters, resulting in an output light source color estimate. More realistic lighting colors p gt .

[0091] During training, this invention selects the Adam optimization algorithm; to improve training stability, the initial learning rate is set to 10. -4 A learning rate adjustment strategy, such as learning rate decay, is employed. This strategy helps to gradually reduce the learning rate during training, ensuring that the network can finely adjust its parameters in the later stages of training and avoiding over-updating. During training, the backpropagation algorithm is used to calculate the gradient of the loss function and update the network parameters θ. T The update rules are as follows: Where, η T It's the learning rate. It is the loss function relative to the network parameters θ T The gradient of the loss function needs to be monitored during training to determine whether the model has converged. The changes in the loss function are observed. Generally, when the loss function stabilizes after multiple iterations, or the decrease in the loss value becomes very small, it indicates that the network is close to the optimal solution, and the model can be considered to have converged. At this point, the training process can be stopped to avoid meaningless training continuation. After stopping training, save the teacher model and its parameters.

[0092] S2: Distillation training of the student network. The specific process is as follows: Figure 4 As shown. This step utilizes knowledge distillation technology to transfer the illumination color estimation capabilities of the teacher network to the student network, which has a simpler structure and fewer parameters. The specific process is as follows:

[0093] S21: Data Preparation and Teacher Guidance. A training set consisting of RAW image data and corresponding standard illumination color labels is constructed for training the student network. The images in the training set are processed using the preprocessing described in S11; the processed images are then input into the teacher network obtained in step S1 to generate guidance information, namely, global light source estimates. Used for student network model distillation.

[0094] S22, the student network performs forward propagation, minimizing the total loss consisting of distillation loss and student-owned loss, and updates the student network parameters. The preprocessed image described in S21 is input into the student network, and the source estimation values ​​of the student network are output. Calculate the total loss based on the student network's estimated values, compute the gradient and backpropagate, and update the parameters of the student network model.

[0095] The student network consists of three main components: a feature extractor (FE), a channel attention block illumination estimation (CAB-IE), and a color feature bag-based illumination refinement (CFB-IR).

[0096] The FE module aims to progressively enhance the representational power of images while preserving spatial information in the early stages of the network. The module takes an input image of size (3, 512, 512) and first passes it through a 3×3 convolutional layer with 64 filters, a stride of 1, and padding of 1, resulting in an output size of (64, 512, 512). A ReLU activation function is then applied, introducing non-linearity. Next, the feature map is fed into the first residual block. This residual block contains three convolutional operations: first, a 1×1 convolution to compress the number of channels to 16; then a ReLU activation; followed by a 3×3 convolution, maintaining the number of channels at 16; another ReLU activation; and finally, a 1×1 convolution to expand the number of channels back to 64. The residual block also includes an identity skip connection that adds the module input to the output, maintaining the output size at (64, 512, 512). The feature map then enters a 5th convolutional layer, a 3×3 convolution with 64 filters, maintaining the same output size. This is followed by the ReLU activation function. Then, the feature map is input into the second residual block, whose structure is exactly the same as the first residual block. It includes the 6th, 7th, and 8th convolutional layers in sequence to complete channel compression and restoration while keeping the spatial dimension unchanged. Next, the feature map is max-pooled using an 8×8 pooling kernel. The last layer performs adaptive average pooling to reduce the spatial dimension of the feature map, retain the main features, and unify the output size (64, 8, 8).

[0097] The CAB-IE module consists of two parts: IE (Light Source Estimation) and CAB (Channel Attention). The IE has five layers. The first layer is a 1×1 convolutional layer for feature extraction and channel adjustment. The second layer is an 8×8 max pooling layer for compressing spatial information. The third layer is another 1×1 convolutional layer that compresses the global features into three channels, corresponding to the RGB light source components. The fourth layer uses adaptive average pooling and L2 normalization. Adaptive average pooling merges the local light source color estimates of each channel into a global light source color estimate. L2 normalization standardizes the light source color estimate: This makes it independent of scale changes, ultimately producing a 3-valued vector representing the image's lighting and color. CAB performs a three-step process. First, adaptive average pooling is applied to the input feature map, resulting in an output size of (64, 1, 1). Second, two 1×1 convolutional layers are used, with ReLU and Sigmoid activation functions respectively, to generate channel attention weights. Third, these weights are applied to the input feature map through channel-wise multiplication.

[0098] The CFB-IR module consists of two parts: IR (refined adjustment of light source color estimation) and CFB (color feature package). CFB integrates the prediction results of six traditional color constancy methods (WP, GW, SoG, GGW, GE1, GE2), each outputting a 3-channel RGB vector, for a total of 18 channels. The prediction results of these traditional methods are then concatenated with the 3-channel light source estimation results output by CAB-IE to form a 21-dimensional feature vector. This vector is input to the IR, which is a lightweight multilayer perceptron (MLP). The MLP contains two fully connected layers (dimensions 21→16→3) to further optimize the final light source prediction and perform L2 normalization. It is worth noting that this module is not enabled during distillation but is enabled during fine-tuning to improve fine-tuning accuracy and stability. Distillation loss. The mean squared error (MSE) is used to calculate the difference between the student network output and the teacher network output, defined as:

[0099]

[0100] Where, r T ,g T ,b T ,r s ,g s ,b s According to S21 and S22 respectively The color components in T are obtained; r ,T g ,T b The distillation temperature is applied to channels R, G, and B respectively to adjust the error scale of each channel.

[0101] The students' own losses The error between the student network's estimated light source color value and the actual label is calculated and defined as:

[0102]

[0103] in, The estimated light source value for the student network, p gt The label represents the real light source, and ∥·∥2 represents solving for the L2 norm.

[0104] The total loss used in training the student network for distillation training is defined as:

[0105]

[0106] Where λ1 and λ2 are weighting coefficients, adjusted The relative importance of these two items.

[0107] During iterative training, the backpropagation algorithm is used based on the total loss function. Calculate the gradient, and then use this gradient to update the student network parameters. The model parameter update formula is: Where θ s η represents the parameters of the student network. S The learning rate is used. Through multiple iterations, the model performance is gradually optimized to ensure that the student network has similar or even better illumination and color estimation performance than the teacher network on a smaller scale, and the trained student network structure and parameters are saved.

[0108] S3: The fine-tuning process of the student network is as follows Figure 5 As shown. To further improve the performance of the student network on a specific camera sensor, the parameters of the student network were fine-tuned using a small amount of labeled and unlabeled data, as follows:

[0109] S31, Data Preparation and Teacher Guidance. A training set consisting of RAW image data from a specific camera and corresponding standard illumination color labels is constructed. This set is used for fine-tuning of the student network. The preprocessing described in S11 is applied to the images in the training set. The processed images are then input into the teacher network obtained in step S1 to generate guidance information, namely, global light source estimates. Used for student network model distillation.

[0110] The term "specific camera" refers to a particular model of camera with the same type of sensor. These cameras share consistent imaging characteristics (such as light response and color representation), thus the invention can be trained and optimized based on the shooting data of this camera. RAW images in the training set are used as labeled data if they have a corresponding standard lighting color; otherwise, they are used as unlabeled data.

[0111] S32, student network forward propagation, minimizes the loss function and optimizes the student network parameters. The image prepared in S31 is input into the student network, and the output is the light source estimate from the student network. The total loss is calculated based on the estimated values ​​of the student network, and then the gradient is calculated and backpropagated to fine-tune the parameters of the student network model.

[0112] The structure of the student network is consistent with that of the student network in step S2. The student network trained and saved in step S2 is used as the initial student network in this step. It is worth noting that the present invention freezes the parameters in the feature extraction module (FE) and the light source estimation module (IE) and enables the color feature package (CFB) in the fine-tuning process to further increase the stability of the model.

[0113] For labeled image I wTwo loss functions are defined: light source estimation error and image correction error. These two are weighted and summed to construct the total loss function. Light source estimation error The angular error used to evaluate the student network estimation and the real light source label is defined as: in, The estimated light source value for the student network, p gt Labels are for real light sources. Image correction error. Refers to the corrected image I s With real image I gt The pixel-level error between them is defined as: Where N represents the number of pixels in the image, and Representing the corrected image I s and real image I gt The RGB color value at the i-th pixel position. Corrected image I s This refers to the use of light source estimates in the training images. The image obtained through the ISP process described in step S11, i.e. Real Image I gt The training images use real light sources and are labeled p. gt Image I obtained through the ISP process described in step S11 gt =f ISP (I w ,p gt The total loss is denoted as Will By weight and Weighting

[0114] For unlabeled image I wo The light source estimation results of the teacher network As the initial pseudo-label, the image is then white-balanced using the Von Kries model described in step S11 to generate a corrected image. This corrected image is then input back into the teacher model to obtain the second light source prediction result. If the first prediction result... If the image is accurate, then the corrected image should be similar to the image taken under standard illumination. Therefore, the second predicted light source should be close to the standard light source. By calculating the angular error θ between the second predicted light source and the standard light source, the confidence level ω of the initial false label can be quantified, defined as... The smaller θ is, the higher the confidence level ω is. This invention allows... Simultaneously, light source estimation will be utilized using the teacher network. The obtained corrected image I T As a real image, let I gt=I T Based on this, the loss function is calculated for labeled images, and confidence levels are introduced to obtain a weighted loss. Furthermore, to ensure that the adjustment range of network parameters during training on unlabeled data is slightly smaller compared to the case of training on labeled data, let... The introduced correction coefficient μ can be set to 0.3–0.8. The corrected image I T This refers to the use of light source estimates in the training images. The image obtained through the ISP process described in step S11, i.e. During iterative training, gradient updates are performed using the backpropagation algorithm, based on the total loss. Calculate the gradient, update the student network parameters, and fine-tune the student network parameters to ensure further performance improvement. The model parameter update formula is: Where θ s η represents the parameters of the student network. opt The learning rate is used. Finally, save the fine-tuned student model structure and parameters.

[0115] S4 calculates the implementation of color constancy. The test image data is input into the model finely tuned in S3 to estimate the color of the light source, and then the color-corrected image is obtained through the ISP process.

[0116] The preprocessing described in S11 is applied to the RAW test image I. test The image is processed without data augmentation to obtain an sRGB image; this sRGB image is then input into a model finely tuned by S3 to estimate the image's illumination colors. Then, the estimated light color is used in the ISP processing flow described in S11. Perform color correction to obtain the color-corrected image, i.e. I corr That is, the color-corrected image.

[0117] For performance comparison metrics, this invention uses the commonly used angular error, which is defined as follows: The angle error statistical indicators used in this invention include: mean, median, triplen, best25%, and worst25%. Mean is the average of all angle errors, reflecting the overall error level; median is the median value after sorting all angle errors from smallest to largest, more robustly reflecting typical performance; triplen is a concept in probability distribution, defined as... Q1, Q2, and Q3 are the quartiles of the angular error values ​​arranged from smallest to largest, at the 25th, 50th, and 75th percentiles, respectively. They are weighted averages of the median and the two quartiles, making them more robust than the mean and more detailed than the median. The best25% is the average error of the 25% of images with the best performance, reflecting the model's optimal performance under ideal conditions. The worst25% is the average error of the 25% of samples with the worst performance, reflecting the model's robustness or worst performance on difficult samples.

[0118] In this embodiment, during the teacher network training phase and the student network distillation phase, we introduce a series of data augmentation strategies, including random cropping (scaling range of 0.1 to 1.0), random rotation within ±30°, horizontal flipping with a 50% probability, and channel-level color perturbation (perturbation intensity range of ±20%, corresponding to scaling factors between 0.6 and 1.2).

[0119] In this embodiment, the teacher model employs a deep convolutional neural network architecture with powerful feature representation capabilities, and is trained in a supervised manner on multiple publicly available color constancy datasets described in this invention. (See...) Figure 7 The FC4 model in Table 2 (Hu Y., Wang B., Lin S. FC4: Fully Convolutional Color Constancy with Confidence-weighted Pooling. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017: 4085–4094.) was trained on the GS and NUS-8 datasets described in this invention. A uniform optimization setting was used in both the knowledge distillation and fine-tuning phases of the student model to ensure the stability of the training process. The main difference between these two phases lies in the data input strategy: in the distillation phase, training data from across devices was used to enhance the model's generalization ability; while in the fine-tuning phase, labeled and unlabeled images acquired by a specific camera were used.

[0120] All experiments in this embodiment were conducted on a computer with the following configuration: Xeon(R) Platinum 8360Y CPU, 512GB RAM, NVIDIA GeForce RTX 3090 GPU, and Ubuntu 20.04 operating system.

[0121] This embodiment validates the effectiveness of the proposed CAB-IE and CFB-IR in the light source estimation task through systematic ablation experiments on the Cube+ dataset. By progressively removing each module and observing performance changes, the actual contribution of each module to the overall performance is clarified. The CAB module enhances the ability to allocate feature attention, while the CFB-IR module improves the ability to extract color features. The combination of the two effectively optimizes the model's performance in the color constancy task. The × and √ under CAB indicate freezing and using CAB during training, respectively; the × and √ under CFB-IR indicate freezing and using CFB-IR during training, respectively. Figure 6 As shown in Table 1, to highlight the performance advantages in the comparative experimental results, the optimal values ​​for each evaluation index are indicated in bold font.

[0122] This embodiment also compares the experimental results before and after fine-tuning of the teacher model and its corresponding distillation student model. Figure 7 Table 2 shows the results of testing the method described in this embodiment on Cube+. FC4 refers to the teacher network structure used in this embodiment, which is FC4. The teacher model was trained based on this network structure in this embodiment. FC4 is a color constancy method based on a fully convolutional neural network proposed by Hu et al. (Y.Hu, B.Wang, and S.Lin, “Fc4: Fully convolutional color constancy with confidence-weighted pooling,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp.4085–4094, 2017.). This method introduces a confidence-weighted pooling mechanism, which improves the robustness to images containing uncertain regions by jointly learning the light source estimate of the region and the corresponding confidence. This network structure allows for end-to-end training and can directly process images of arbitrary sizes, exhibiting high accuracy and inference efficiency. Raw Student refers to the student model obtained through distillation training using the teacher network FC4 structure in this embodiment. Fine-tuned Student refers to the model obtained by further optimizing the distilled model in this embodiment. LD and UD refer to fine-tuning the distilled model in this embodiment using labeled and unlabeled datasets, respectively. LD / UD refers to fine-tuning the distilled model in this embodiment using labeled and unlabeled datasets.

[0123] Experimental results show that both the original student model (Raw Student) and the fine-tuned student model (except for those using only unlabeled data) consistently outperform their corresponding teacher models on core evaluation metrics. This verifies that knowledge distillation not only achieves model lightweighting but also improves performance while retaining the core capabilities of the teacher model. Furthermore, under the same network structure, the fine-tuned student model outperforms the original student model in all cases except those using only unlabeled data, highlighting the advantage of fine-tuning in adapting to the specific imaging characteristics of the device. The stronger the teacher model's modeling ability under complex lighting conditions, the more accurate the mapping relationships learned by the student model. Notably, during the fine-tuning phase, the combination of "LD+UD" (labeled data plus unlabeled data) significantly outperforms using LD or UD alone, demonstrating the complementary advantages between supervisory signals and data diversity. To highlight the performance advantages in the comparative experimental results, the optimal values ​​for each evaluation metric are indicated in bold in the table.

[0124] Furthermore, the experimental results of this embodiment were compared with those of various advanced methods in recent years, such as... Figure 8Table 3 shows the comparison of five methods, including: White-Patch (from D. H. Rainard and B. Awandell, “Analysis of the retinex theory of color vision,” Journal of the Optical Society of America A, vol. 3, no. 10, pp. 1651–1661, 1986), Grey-world (from G. Buchsbaum, “A spatial processor model for object colour perception,” Journal of the Franklin Institute, vol. 310, no. 1, pp. 1–26, 1980. [doi:10.1016 / 0016-0032(80)90058-7]), and C5 (from M. Afifi, J.T. Barron, C. Le Gendre, Y.-T. Tsai, and F. Bleibel, “Cross-camera convolutional color constancy,” in Proceedings of theIEEE / CVF International Conference on Computer Vision, pp. 1981–1990, 2021.》), FFCC (from "JTBarron and Y.-T.Tsai, "Fast fourier color constancy," in Proceedings of the IEEE Conference on Computer Vision and PatternRecognition, pp.886–894, 2017.》), FC4 (from "Y.Hu, B.Wang, and S.Lin, "Fc4: Fullyconvolutional color constancy with confidence-weighted pooling," in Proceedings of the IEEE Conference on Computer Vision and PatternRecognition, pp.4085–4094, 2017.》), One-net (from "I. D. M. and S. "One-net: Convolutional color constancy simplified," Pattern Recognition Letters, vol.159, pp.31–37, 2022."), MDLCC (from "J.Xiao, S.Gu, and L.Zhang, "Multi-domain learning for accurate and few-shot colorconstancy," in Proceedings of the IEEE / CVF Conference on Computer Vision andPattern Recognition, pp.3258–3267, 2020.[doi:10.1109 / cvpr42600.2020.00332].》), Quasi-UCC (from "S.Bianco and C.Cusano, "Quasi-unsupervised color constancy," in Proceedings of the IEEE / CVF Conference on Computer Vision and PatternRecognition, pp.12212–12221, 2019.》). It can be seen that the method of the present invention has achieved good results. However, it is worth noting that the performance of the method of the present invention is closely related to the teacher model used—the better the performance of the teacher model, the better the performance of the student model obtained by fine-tuning on the corresponding device. In this embodiment, the FC4 model used as the teacher model is not the best performing deep learning model structure at present. Therefore, the performance of the method of the present invention still has great potential for improvement.

[0125] In the method proposed in this invention, the teacher model is used to achieve accurate image illumination and color estimation. Its network structure is based on a high-performance, relatively complex deep learning network, trained using high-quality image data and their corresponding illumination and color labels. The resulting teacher model is large in scale and possesses powerful illumination and color estimation capabilities. During the distillation training of the student model, knowledge distillation technology is employed to transfer the illumination and color estimation capabilities of the teacher network to the simplified student network with fewer parameters. Through multiple iterative optimizations, the performance of the student model is gradually improved, ensuring that it can achieve or even surpass the illumination and color estimation performance of the teacher network at a smaller scale.

[0126] In this invention, to further improve the performance of the student model on a specific camera sensor, fine-tuning is performed using a small amount of labeled and unlabeled data to optimize the student model's parameters and ensure further performance improvement. Fine-tuning training is essentially a transfer learning process, where parameters are fine-tuned based on the student model obtained in step S2 of this invention. To ensure improved performance of the fine-tuned model, the learning rate setting can be adjusted during training, and some layer parameters can be frozen. Additionally, labeled and unlabeled data are mixed and shuffled during training. For unlabeled data, the gradient is calculated using the obtained loss function, and the model parameters are updated in reverse. By adding a correction coefficient to the loss value, the adjustment range of the network parameters is made slightly smaller compared to the case with labeled data.

[0127] The method provided in this embodiment is particularly suitable for training on datasets constructed for specific cameras. The obtained model can be easily embedded into the corresponding camera to efficiently perform color constancy calculation. When applied to a class of cameras with the same image sensor, this method can provide high accuracy in estimating the light source of the image scene. Because the method of this invention uses knowledge distillation technology to distill a student model with fewer parameters from a high-performance teacher model, it has the advantage of fast computation speed and is suitable for writing programs to embed into the camera image signal processor (ISP) to efficiently perform color constancy calculation.

[0128] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Any other modifications or equivalent substitutions made by those skilled in the art to the technical solutions of the present invention, as long as they do not depart from the spirit and scope of the technical solutions of the present invention, should be covered within the scope of the claims of the present invention.

Claims

1. A method for calculating color constancy based on knowledge distillation, characterized in that, Includes the following steps: S1, Training the teacher network: Training a deep learning model as the teacher network to achieve image illumination and color estimation; S2, Distillation Training of the Student Network: Constructing the student network, which consists of three components: Feature Extraction Module FE; Light Source Color Estimation Module CAB-IE with Channel Attention; and Light Source Color Estimation Fine Adjustment Module CFB-IR with Color Feature Packet. The illumination color estimation capability of the teacher network is transferred to the student network, which has a simple structure and few parameters, using distillation technology. S3, fine-tuning training of the student network, involves fine-tuning the student network using labeled and unlabeled data, with the amount of labeled data being less than the amount of unlabeled data. S4 calculates the implementation of color constancy by inputting the test image data into the fine-tuned model to estimate the light source color, and then obtaining the color-corrected image through the ISP process.

2. The method for calculating color constancy based on knowledge distillation according to claim 1, characterized in that, Specifically, S1 is: S11, Data preprocessing: The images in the training set, which consist of image data and corresponding illumination color labels, are processed through the ISP process, then data augmentation is performed, and then standardization is carried out. S12, network forward propagation, minimizing loss, updating teacher network parameters; using the initial few convolutional layers of the teacher network to extract features, then passing them through convolutional layers with relatively large kernels for dimensionality reduction, further extracting semi-dense feature maps; the semi-dense feature map contains four channels, the first three channels representing local light source estimates, and the fourth channel representing its confidence level. The feature map is passed to a weighted pooling layer to aggregate from local to global to generate the final illumination color estimate; the global illumination color is calculated through confidence-weighted pooling. Minimize loss function Update teacher network parameters.

3. The method for calculating color constancy based on knowledge distillation according to claim 2, characterized in that: The ISP process in S11 refers to the processing of a certain RAW image I raw Preprocessing yields the processed sRGB image I sRGB , denoted as I sRGB =f ISP (I raw ),p g ), where p g =(r g ,g g ,b g () is the global illumination color of the image. Global illumination typically uses a white light source, i.e. At this point, the ISP process is abbreviated as I. sRGB =f ISP (I raw Before training the teacher network and student network, RAW images need to be processed using the ISP process to obtain the image data required for training the network. The specific process includes: format conversion, linearization, color interpolation, white balance correction, color space conversion, and gamma correction. The data augmentation in S11 refers to performing a series of transformations on the image before training the model to increase its robustness. Specifically, firstly, geometric transformations are performed, including: the image is randomly cropped and resized to 512×512, and randomly flipped vertically and horizontally, and randomly rotated. After the geometric transformations, normally distributed random noise is added to the image's color channels to simulate the disturbances caused by changes in illumination. The normally distributed noise added to each color channel is: I aug,c =I 0,c +ε,ε~N(0,σ 2 ), Where ε represents noise; σ represents the noise standard deviation, the specific value of which is set according to the dataset; I 0,c and I aug,c Let N(0,σ) be the original and enhanced pixel color values ​​in the c channel, respectively, where c∈{r,g,b}; 2 ) represents a normal distribution with a mean of 0 and a standard deviation of σ; by adding this noise, the diversity of light source types is enhanced, which is more conducive to improving the generalization performance of the model; The normalization process in S11 refers to the calculation of the input image I0: Where μ c and σ c These are the mean and standard deviation of the c-channel of all images in the training set, respectively; I 0,c and I norm,c These are the original and standardized pixel color values ​​for the c channel, respectively, where c ∈ {r, g, b}. Standardization ensures a consistent distribution of the input data, which usually yields better results than rescaling.

4. The method for calculating color constancy based on knowledge distillation according to claim 2, characterized in that: The teacher network can be structured using convolutional neural networks, such as AlexNet, ResNet, and SqueezeNet. The network structure consists of multiple convolutional layers, pooling layers, and fully connected layers, ultimately outputting a 3D vector representing the global illumination color of the image. The teacher network is based on the first five convolutional feature extraction layers of AlexNet, and uses additional convolutional layers and confidence-weighted pooling layers to estimate the global illumination color of the image, finally outputting a vector. The teacher network employs the following structure: an image input and feature extraction module, a semi-dense feature extraction module, and a confidence-weighted pooling module; the goal of the teacher network is to estimate the global illumination color of the image. Make it as close as possible to the light color label loss function Defined as and p gt Angular error between: in, and ||p gt ||2 represents the L2 norm, for and p gt The dot product; During training, the Adam optimization algorithm was selected; the initial learning rate was set to 10. -4 Learning rate adjustment strategies, such as learning rate decay, are employed to improve training stability. During training, the backpropagation algorithm is used to calculate the gradient of the loss function and update the network parameters θ. T The update rules are as follows: Where, η T It's the learning rate. It is the loss function relative to the network parameters θ T The gradient; the specific convergence and stopping conditions of the model are as follows: during the training process, monitor the changes in the loss function. If the loss value tends to stabilize or decreases very little after multiple iterations, the model is considered to have converged, training can be stopped, and the teacher model and parameters can be saved.

5. The method for calculating color constancy based on knowledge distillation according to claim 1, characterized in that, Specifically, S2 is: S21, Data Preparation and Teacher Guidance: Construct a training set consisting of RAW image data and corresponding standard illumination color labels for training the student network; process the images in the training set using the preprocessing described in S11; input the processed images into the teacher network obtained in S1 to generate guidance information, i.e., global light source estimates. Used for distillation of student network models; S22, the student network performs forward propagation, minimizing the total loss consisting of distillation loss and student-owned loss, and updates the student network parameters; the preprocessed image described in S21 is input into the student network, and the light source estimate of the student network is output. Calculate the total loss based on the student network's estimated values, compute the gradient and backpropagate, and update the parameters of the student network model.

6. The method for calculating color constancy based on knowledge distillation according to claim 5, characterized in that: distillation loss The mean squared error is used to calculate the difference between the student network output and the teacher network output, defined as: Where, r T g T b T r s g s b s According to S21 and S22 respectively The color components in T are obtained; r T g T b The distillation temperature is applied to channels R, G, and B respectively to adjust the error scale of each channel. The students' own losses The error between the student network's estimated light source color value and the actual label is calculated and defined as: in, The estimated light source value for the student network, p gt For the label of a real light source, ||·||2 represents solving for the L2 norm; The total loss used in distillation training student networks is defined as: Where λ1 and λ2 are weighting coefficients, adjusted The relative importance of these two items; During iterative training, the backpropagation algorithm is used based on the total loss function. Calculate the gradient, and then use this gradient to update the student network parameters; the model parameter update formula is: Where θ s η represents the parameters of the student network. S The learning rate is used to gradually optimize the model performance through multiple iterations, ensuring that the student network has similar or even better illumination and color estimation performance than the teacher network on a smaller scale, and the trained student network structure and parameters are saved.

7. The method for calculating color constancy based on knowledge distillation according to claim 1, characterized in that, Specifically, S3 is: S31, Data Preparation and Teacher Guidance: Construct a training set consisting of RAW image data from a specific camera and corresponding standard illumination color labels, which is used for fine-tuning of the student network. The preprocessing described in S11 is then applied to the images in the training set. The processed images are input into the teacher network obtained in S1 to generate guidance information, namely, global light source estimates. Used for student network model distillation; the specific camera refers to a camera of a certain model with the same type of sensor; the RAW images in the training set, if they have corresponding standard lighting colors, are used as labeled data; if they do not have corresponding standard lighting colors, they are used as unlabeled data. S32, forward propagation of the student network, minimizing the loss function and optimizing the student network parameters; inputting the image prepared in S31 into the student network, and outputting the light source estimates of the student network. The total loss is calculated based on the estimated values ​​of the student network, and then the gradient is calculated and backpropagated to fine-tune the parameters of the student network model. The structure of the student network is consistent with that of the student network in S2. The student network trained and saved in S2 is used as the initial student network in this step. It is worth noting that the parameters in the feature extraction module and the light source estimation module (IE) are frozen and the color feature package is enabled in the fine-tuning process to further increase the stability of the model.