A method of intelligent automatic white balance

CN122845947APending Publication Date: 2026-09-29SHENZHEN UNITED OPTICAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611341393.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-09-01
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0005]针对现有技术的不足,本发明提供了一种智能自动白平衡的方法,解决了现有自动白平衡技术在复杂光源或大面积纯色场景下会偏色,且缺少异常数据拦截机制导致连续视频流出现增益跳变与画面闪烁的问题

Benefits of technology

[0025]1、本发明通过引入马氏距离计算图像特征偏离常规场景分布的程度,并结合网络预测置信度与信息分类熵动态调节人工智能白平衡增益与传统白平衡增益的混合比例;当系统输入超出神经网络模型离线训练分布边界的异常场景数据时,能够自动切断神经网络的异常输出链路并平滑过渡至传统白平衡算法,从而避免极端复杂场景下错误预测增益导致的画面整体偏色,提升系统应对未知场景的稳定性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845947A_ABST
    Figure CN122845947A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of image processing, and discloses a kind of intelligent automatic white balance method, method includes obtaining original bayer array data and outputting preprocessed image data;Extract the channel mean value of each partition of image, calculate chroma characteristic ratio and Mahalanobis distance, generate RGB format thumbnail;Output the partition semantic weight map using convolutional neural network, predict artificial intelligence white balance gain and network prediction confidence;Judge whether Mahalanobis distance exceeds preset intercept threshold, combine confidence and information classification entropy to calculate dynamic fusion weight;Based on the weight, the artificial intelligence white balance gain and the traditional white balance gain are weighted to obtain the fusion gain result;The fusion gain result is processed by interframe adaptive infinite impulse response low-pass filtering, and the final three-channel white balance gain is output.The application can automatically intercept abnormal scene data, suppress interframe gain jump, and improve the color restoration stability of the system in complex scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, specifically to an intelligent automatic white balance method. Background Technology

[0002] With the development of smartphone camera functions, users' requirements for the image signal processing capabilities of devices are constantly increasing. In traditional image signal processing systems, modules such as automatic exposure, automatic focus, and automatic white balance are used to work together to optimize image quality. Traditional automatic white balance algorithms are mainly based on calibrating the gains of the red, green, and blue channels of the original image at the color temperature under laboratory conditions, and fitting the Planck curve based on the blackbody radiation theory to determine the color temperature of the current image.

[0003] Because this traditional algorithm essentially relies on the statistical fitting of color data from the entire image or a specific region, when a large area of ​​a single color appears in the actual shooting scene, or when the ambient lighting is too complex, the distribution of the image's basic data will deviate significantly from the conventional model. This causes the system to make a mistake in determining the color temperature, resulting in inaccurate white balance reproduction of the image.

[0004] To overcome the shortcomings of pure software statistical algorithms in complex scenarios, some existing solutions choose to add an additional color temperature sensing sensor at the physical level. By introducing external hardware, the true color temperature of the current environment is directly obtained to assist the image signal processor in white balance and color reproduction. Although this solution, which relies on additional physical components, can improve the accuracy of color reproduction, it inevitably increases the hardware manufacturing cost of smart terminals and occupies the limited structural space inside the device. How to improve the accuracy and stability of automatic white balance algorithms in complex scenarios without adding additional physical sensors, combined with the computing power resources of the device's built-in neural network units, is a technical problem that urgently needs to be solved. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides an intelligent automatic white balance method that solves the problems of color cast in complex lighting or large-area solid color scenes, and the lack of an abnormal data interception mechanism that causes gain jumps and screen flickering in continuous video streams.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] This invention provides a method for intelligent automatic white balance, comprising the following steps:

[0008] The raw Bayer array data output by the sensor is acquired, and black level correction, lens shading correction, Bayer domain noise reduction and downsampling are performed to obtain preprocessed image data.

[0009] Extract the mean values ​​of each partition channel, calculate the chromaticity feature ratio and Mahalanobis distance, and generate RGB format thumbnails;

[0010] Feature vectors are extracted using convolutional neural networks, and a semantic weight map of partitions is output based on thumbnails. Classification probability distribution and information classification entropy are extracted based on feature vectors to predict artificial intelligence white balance gain and network prediction confidence.

[0011] If the Mahalanobis distance exceeds the preset interception threshold, the dynamic fusion weight is assigned a value of zero; if it does not exceed the threshold, the dynamic fusion weight is calculated based on the positive mapping of confidence and the negative suppression result of information classification entropy after truncation limit.

[0012] Using dynamic fusion weights as adjustment factors, linear interpolation weighting is performed on the artificial intelligence and traditional white balance gains to obtain the fusion gain result.

[0013] The results are subjected to inter-frame adaptive infinite impulse response low-pass filtering, and the three-channel white balance gain is output to the image signal processor.

[0014] This invention assesses the degree to which input data deviates from the empirical distribution by calculating the Mahalanobis distance of the data, adjusts the mixing ratio of artificial intelligence white balance gain and traditional white balance gain by combining semantic recognition weights and network prediction confidence, and implements low-pass filtering with adjustable smoothing coefficients on the time series, thereby suppressing abnormal gain abrupt changes and smoothing inter-frame output.

[0015] Furthermore, the steps of performing black level correction, lens shading correction, Bayer domain noise reduction, and downsampling to obtain preprocessed image data include: performing subtraction compensation operations on each pixel value of the original Bayer array data using a set black level reference value and truncating the result to zero when it is negative; performing multiplication gain compensation operations on each spatial position pixel of the data after black level correction using a pre-calibrated two-dimensional grid gain table; and, while keeping the underlying Bayer filter array structure unchanged, performing spatial domain filtering suppression on high-frequency noise points in the data after lens shading correction, and using an adjacent same-color pixel merging mechanism to perform downsampling operations according to a specified downsampling ratio to generate dimensionally reduced preprocessed data as preprocessed image data.

[0016] Furthermore, the steps of extracting the mean values ​​of each partition channel and calculating the chromaticity feature ratio and Mahalanobis distance include: dividing the preprocessed image data into multiple independent spatial blocks according to the set grid division rules, independently calculating the pixel mean value of each color channel, and removing overexposed pixels with values ​​exceeding the set saturation threshold from the mean value calculation sample; performing candidate gray point screening operation according to the preset gray point limit range, and calculating the chromaticity feature ratio of each spatial block using the red channel pixel mean value, green channel pixel mean value, and blue channel pixel mean value corresponding to the screened candidate gray points; constructing a multidimensional chromaticity feature vector by combining the chromaticity feature ratios corresponding to all spatial blocks, and calculating the Mahalanobis distance of the multidimensional chromaticity feature vector from the normal scene distribution by combining the pre-stored training set distribution parameters containing the feature mean vector and feature covariance matrix.

[0017] Furthermore, the step of extracting feature vectors using a convolutional neural network and outputting a partition semantic weight map based on the thumbnail includes: identifying the face, sky, vegetation, and indoor / outdoor semantic regions in the RGB format thumbnail; assigning higher weights to key skin regions of the face and lower weights to large solid-color areas of the sky; and outputting a partition semantic weight map. The extracted partition features are arranged into a spatial grid, and feature mapping and pooling dimensionality reduction are performed using a convolutional neural network to extract feature vectors with spatial distribution information.

[0018] Furthermore, the steps of extracting classification probability distribution and information classification entropy based on feature vectors to predict artificial intelligence white balance gain and network prediction confidence include: performing channel concatenation operation on the data dimension to form a joint input vector matrix by combining the standard light source classification probability distribution matrix and the flattened feature vectors after feature pooling dimensionality reduction; performing numerical prediction calculation on the joint input vector matrix using a fully connected regression network to output the preliminary white balance gain of each spatial block and the network prediction confidence representing the accuracy of the network prediction state; mapping and aligning the spatial resolution of the partition semantic weight map with the divided spatial block grid, and performing spatial weighted summation on the preliminary white balance gain of each spatial block according to the partition semantic weight map to calculate and generate artificial intelligence white balance gain.

[0019] In a preferred embodiment of the present invention, before the step of extracting the classification probability distribution and information classification entropy based on feature vectors, predicting the artificial intelligence white balance gain and network prediction confidence, a step of performing complete offline training of the neural network model is further included. The offline training step includes: constructing a training dataset for the neural network model; offline acquisition of underlying Bayer images and synchronous generation of corresponding RGB images; creating multi-dimensional supervised learning labels for the collected image sample data; the supervised learning labels specifically include pixel-level semantic category labels for training the semantic segmentation network, standard light source real classification labels for training the light source classification head sub-module, and real white balance gain label values ​​for training the fully connected regression network; and initializing and constructing a neural network model in a deep learning framework that includes a convolutional neural network, a semantic segmentation network, a light source classification head structure, and a fully connected regression network.

[0020] Furthermore, the step of calculating the dynamic fusion weights based on the results of forward mapping of confidence and reverse suppression of information classification entropy, after truncation and limiting, if the threshold is not exceeded, includes: extracting the quantile calibration of the offline training set data distribution to generate a preset interception threshold, and performing a numerical comparison operation between the Mahalanobis distance and the preset interception threshold; when the Mahalanobis distance is greater than the preset interception threshold, it is determined that the data features at the bottom layer of the input system seriously deviate from the empirical distribution boundary of the conventional model, blocking the prediction output link of the gain calculation submodule, and forcibly assigning the dynamic fusion weights to 0; when the Mahalanobis distance is less than or equal to the preset interception threshold, performing a forward gain mapping operation on the initial dynamic weight values ​​based on the network prediction confidence and performing a reverse suppression operation on the initial dynamic weight values ​​based on the information classification entropy parameters, and using a closed interval truncation function to process the upper and lower limit bits to generate dynamic fusion weights distributed in the real number interval from 0 to 1.

[0021] Furthermore, the step of performing linear interpolation weighting on the artificial intelligence and traditional white balance gains using dynamic fusion weights as adjustment factors to obtain the fusion gain result includes: statistically calculating the mean values ​​of the red, green, and blue channels within the global distribution range of the preprocessed image data using the gray-world algorithm; simultaneously searching for candidate reference white points that conform to the emission trajectory of a standard light source using a white point detection algorithm; and calculating and generating traditional red, green, and blue three-channel reference gain values ​​as the traditional white balance gain by combining the color channel ratio results of the candidate reference white points; and performing linear interpolation weighting fusion operation on the artificial intelligence white balance gain and the traditional red, green, and blue three-channel reference gain values ​​using dynamic fusion weights as adjustment factors to calculate and generate a fusion gain result covering three independent color channels.

[0022] Furthermore, the step of performing inter-frame adaptive infinite impulse response low-pass filtering on the results and outputting the three-channel white balance gain to the image signal processor includes: performing temporal smoothing processing on the fusion gain result corresponding to the current video frame using an infinite impulse response low-pass filter, and synchronously calculating the dynamic smoothing coefficient of the infinite impulse response low-pass filter based on the dynamic fusion weight; extracting the final white balance output gain corresponding to the previous video frame as a historical state benchmark, using the dynamic smoothing coefficient as a historical data retention weight, performing iterative calculation by combining the final white balance output gain corresponding to the previous video frame with the fusion gain result corresponding to the current video frame, and calculating the final white balance output gain corresponding to the current video frame as the final three-channel white balance gain; loading the final white balance output gain corresponding to the current video frame into the hardware Bayer domain multiplier to perform physical multiplication gain operation, and sequentially sending the original pixel data after physical multiplication gain operation processing to the de-mosaic processing unit, the color matrix correction unit, and the gamma mapping unit for pipeline closed-loop operation.

[0023] Furthermore, the step of synchronously calculating the dynamic smoothing coefficient of the infinite impulse response low-pass filter based on the dynamic fusion weight includes: calculating the dynamic smoothing coefficient by subtracting the product of the preset smoothing coefficient adjustment amount and the dynamic fusion weight from the preset base maximum smoothing coefficient; decreasing the dynamic smoothing coefficient when the value of the dynamic fusion weight approaches one, and increasing the dynamic smoothing coefficient when the value of the dynamic fusion weight approaches zero.

[0024] This invention provides a method for intelligent automatic white balance. It has the following beneficial effects:

[0025] 1. This invention calculates the degree to which image features deviate from the distribution of a normal scene by introducing Mahalanobis distance, and dynamically adjusts the mixing ratio of artificial intelligence white balance gain and traditional white balance gain by combining network prediction confidence and information classification entropy. When the system inputs abnormal scene data that exceeds the offline training distribution boundary of the neural network model, it can automatically cut off the abnormal output link of the neural network and smoothly transition to the traditional white balance algorithm, thereby avoiding the overall color cast of the image caused by incorrect prediction gain in extremely complex scenes and improving the stability of the system in dealing with unknown scenes.

[0026] 2. This invention utilizes a convolutional neural network to perform semantic region recognition on images and generate a partitioned semantic weight map. Differentiated spatial weights are assigned to key regions such as faces and large areas of solid color such as the sky. Then, spatial weighted summation is performed on the preliminary white balance gain of each partition. This allows the white balance calculation process to eliminate the interference of large areas of single solid color on global chromaticity statistics, while ensuring the accuracy of color reproduction of the core objects of interest. This solves the color shift problem of traditional global statistical algorithms against solid color backgrounds.

[0027] 3. This invention employs an infinite impulse response low-pass filter in the final output stage to perform inter-frame temporal smoothing on the fusion gain result, and synchronously adjusts the dynamic smoothing coefficient of the filter according to the dynamic fusion weight; the adaptive inter-frame filtering mechanism smooths out the inter-frame gain jump caused by local object movement or slight changes in light source in the continuous video stream, eliminates color flickering in the video stream output process, and ensures the smoothness of color transition in dynamic images. Attached Figure Description

[0028] Figure 1 This is the interaction timing diagram of the present invention;

[0029] Figure 2 This is a schematic diagram of the method flow of the present invention;

[0030] Figure 3 This is the surface curve of the dynamic fusion weights of the present invention as a function of network prediction confidence and information classification entropy;

[0031] Figure 4 This is a flowchart of the anomaly interception and weight calculation process of the present invention;

[0032] Figure 5 This is a comparison chart of the low-pass filter response curves of the inter-frame adaptive infinite impulse response of the present invention.

[0033] Figure 6 This is a comparison chart of the chromaticity angle errors of different white balance algorithms of the present invention;

[0034] Figure 7 This is a comparison chart of the temporal stability of video sequences using different white balance algorithms of this invention. Detailed Implementation

[0035] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0036] See attached document Figure 1 This invention provides an intelligent automatic white balance system, which may include:

[0037] The RAW input module is used to receive the raw Bayer array data output by the sensor and serve as the basic data input for the entire intelligent automatic white balance system.

[0038] The preprocessing module is used to sequentially perform black level correction, lens shadow correction, Bayer domain noise reduction, and downsampling on the original Bayer array data to obtain dimensionality-reduced preprocessed data. The process is used to remove sensor dark current bias, compensate for brightness attenuation around the lens, suppress high-frequency noise, and reduce the amount of data to control the computational load of subsequent stages.

[0039] The statistics extraction module is used to perform parallel dual-path processing on the dimensionality reduction preprocessed data. The statistics extraction module further includes a left-path extraction submodule and a thumbnail generation submodule. The left-path extraction submodule is used to extract the mean values ​​of the red, green and blue channels of the image partition, remove overexposed and saturated pixels, filter candidate gray points, and calculate the corresponding blue-green ratio, red-green ratio and other chromaticity ratios, as well as the Mahalanobis distance between the current frame and the training set distribution. The thumbnail generation submodule is used to perform de-mosaic processing on the dimensionality reduction preprocessed data, convert it to RGB color space and generate RGB format thumbnails of a set resolution size, while preserving the spatial topological information of the image content.

[0040] The AI ​​inference module is used to extract and learn features based on the features output by the statistical extraction module, and output the predicted AI white balance gain and confidence. The AI ​​inference module further includes a feature encoder submodule, a semantic analysis submodule, a light source classification head submodule, a gain calculation submodule, a fusion layer submodule, and an anomaly detection submodule.

[0041] The feature encoder submodule arranges the left-side statistics into a spatial grid and uses a convolutional neural network to extract feature vectors with spatial distribution information. The semantic analysis submodule takes a thumbnail as input, identifies faces, sky, vegetation, and indoor / outdoor semantic regions in the scene, and outputs a partition semantic weight map. The light source classification head submodule takes the feature vector as input, uses the Softmax function to output the probability distribution of the current scene under multiple standard light sources, and calculates the information classification entropy. The gain calculation submodule concatenates the standard light source probability distribution and the feature vector as joint input, predicts the red-green-blue-white balance gain and network prediction confidence. The fusion layer submodule weights the gain contribution of each partition according to the partition semantic weight map. To clarify the specific calculation rules of the weights, the gain calculation submodule calculates the dynamic fusion weights based on the positive mapping of the network prediction confidence and the negative suppression result of the information classification entropy through closed interval truncation limits.

[0042] The anomaly detection submodule runs in parallel with the inference main chain. It is used to determine that the current input exceeds the training distribution range when the Mahalanobis distance exceeds a preset threshold, and to force the dynamic fusion weights to be reset to zero.

[0043] The traditional white balance module is used to calculate the traditional white balance gain using the gray-world algorithm and white point detection method, providing a basic fallback gain output.

[0044] The dynamic weighted fusion module is used to perform linear interpolation weighted fusion operation on the artificial intelligence white balance gain output by the AI ​​inference module and the traditional white balance gain output by the traditional white balance module, with the dynamic fusion weight as the adjustment ratio factor.

[0045] The timing smoothing module is used to perform inter-frame adaptive infinite impulse response low-pass filtering on the final gain obtained after weighted fusion. The smoothing coefficient of the filter changes adaptively in real time according to the dynamic fusion weights to control the inter-frame convergence speed and suppress flicker.

[0046] The gain output module is used to output the final white balance gain values ​​of the red, green, and blue channels after filtering.

[0047] The ISP application module is used to apply the final white balance gain value to the Bayer domain multiplier and perform subsequent demosaicing, color matrix correction and gamma mapping processing to output the final image.

[0048] See attached document Figure 2 This invention provides a method for intelligent automatic white balance, comprising the following steps:

[0049] S1: Obtain the raw Bayer array data output by the sensor, and sequentially perform black level correction, lens shadow correction, Bayer domain noise reduction and downsampling processing to output preprocessed image data;

[0050] S2 performs parallel dual-path feature extraction on the preprocessed image data; the first path extracts the channel mean of each region of the image, removes saturated pixels and filters candidate gray points, and calculates the corresponding chromaticity feature ratio and Mahalanobis distance used to measure the degree of data deviation; the second path performs de-mosaic processing on the preprocessed image data and generates the corresponding RGB format thumbnail.

[0051] S3 arranges the partition features extracted from the first path into a spatial grid, uses a convolutional neural network for feature mapping and pooling dimensionality reduction to extract feature vectors, and performs semantic analysis on the RGB format thumbnails generated by the second path to output a partition semantic weight map; uses the Softmax function to calculate the classification probability distribution and information classification entropy of various standard light sources based on the feature vectors, and inputs the classification probability distribution and feature vectors together to predict the corresponding artificial intelligence white balance gain and network prediction confidence.

[0052] S4. When the Mahalanobis distance exceeds the preset interception threshold, the dynamic fusion weight is forcibly assigned to zero; when the Mahalanobis distance does not exceed the preset interception threshold, the dynamic fusion weight is calculated based on the confidence positive mapping and the information classification entropy reverse suppression result, after truncation limit.

[0053] S5 uses the grayscale world algorithm and white point detection algorithm to calculate and obtain the traditional white balance gain; with the dynamic fusion weight as the adjustment ratio factor, linear interpolation weighting is performed on the artificial intelligence and traditional white balance gain to obtain the fusion gain result;

[0054] S6 performs inter-frame adaptive infinite impulse response low-pass filtering on the fusion gain result. The smoothing coefficient of the filter changes in real time with the dynamic fusion weight, and outputs the three-channel white balance gain to the image signal processor.

[0055] In specific implementations, step S1 provided by this invention may include the following steps:

[0056] S11, the RAW input module receives the raw Bayer array data output by the sensor. The raw Bayer array data is a low-level image electrical signal that has not undergone color space conversion processing within the image signal processor. The pixel arrangement of the raw Bayer array data follows the standard Bayer color filter array format, consisting of alternating red, green, and blue photosensitive pixels. The RAW input module uses the raw Bayer array data as the initial input data for the intelligent automatic white balance system.

[0057] S12, the preprocessing module performs black level correction on the original Bayer array data. Since the image sensor output electrical signal has a numerical base bias, the preprocessing module uses a set black level reference value to perform subtraction compensation on the pixel values ​​of the original Bayer array data. When the calculation result is negative, it is truncated to zero. The set black level reference value is determined by reading the average output value of the non-photosensitive, light-blocking pixel area at the edge of the image sensor. In this embodiment, the range of the black level reference value is set to 64 to 256, and the specific value is dynamically matched according to the physical bit depth of the original Bayer array data. The specific calculation formula for black level correction is as follows:

[0058] ;

[0059] In the formula, coordinates The pixel value at the location after black level correction; The index of the pixel's horizontal coordinate position in the two-dimensional space of the image; The index of the pixel's position on the vertical axis in the two-dimensional space of the image; A mathematical operation function to find the maximum value; To prevent pixel values ​​from becoming negative, a zero-truncation reference lower limit constant is used; coordinates The original Bayer array data pixel values ​​at the location; The black level reference value is determined by reading the average value of the output values ​​of the non-photosensitive, light-blocking pixel area at the edge of the image sensor.

[0060] S13, the preprocessing module performs lens shadow correction processing on the data after black level correction. The light intensity in the center area of ​​the optical lens is higher than that in the edge area. The preprocessing module uses a two-dimensional grid gain table that is pre-calibrated and stored in non-volatile memory. In this embodiment, the number of grid nodes in the two-dimensional grid gain table is configured to be between 13×13 and 17×17. The preprocessing module performs multiplication gain compensation operation on the pixels at each spatial position of the data after black level correction. For the pixel gain values ​​between the grid nodes of the two-dimensional grid gain table, the preprocessing module performs numerical calculations based on the existing bilinear interpolation algorithm.

[0061] S14, the preprocessing module performs Bayer domain noise reduction on the data after lens shading correction. While keeping the underlying Bayer filter array structure unchanged, the preprocessing module performs spatial domain filtering suppression on the high-frequency noise points in the data after lens shading correction. The specific spatial domain filtering and smoothing method is based on existing image signal processing spatial domain noise reduction methods.

[0062] S15, the preprocessing module performs downsampling on the data after Bayer domain noise reduction. The preprocessing module uses an adjacent same-color pixel merging mechanism to perform downsampling on the data after Bayer domain noise reduction according to a specified downsampling ratio. The specified downsampling ratio is set to a range of one-half to one-eighth. After the preprocessing module performs downsampling, it generates dimensionality-reduced preprocessed data. The preprocessing module outputs the dimensionality-reduced preprocessed data to the statistics extraction module to perform the corresponding feature extraction operation.

[0063] In specific implementations, step S2 provided by this invention may include the following steps:

[0064] S21, the left extraction submodule divides the dimensionality reduction preprocessing data into multiple independent spatial blocks according to the set grid division rules. The left extraction submodule independently calculates the average pixel value of the red channel, the average pixel value of the green channel, and the average pixel value of the blue channel for each spatial block. Before calculating the average pixel value of the channels, the left extraction submodule performs a pixel value judgment operation. The left extraction submodule removes overexposed pixels with values ​​exceeding the set saturation threshold from the average value calculation sample. The saturation threshold is determined in combination with the bit depth structure of the dimensionality reduction preprocessing data. In this embodiment, for 10-bit bit depth dimensionality reduction preprocessing data, the value range of the saturation threshold is set to 960 to 1023.

[0065] The left-side extraction submodule performs candidate gray point filtering on each spatial block after saturation removal based on a preset gray point limit range. The preset gray point limit range includes the central region where the difference between the mean values ​​of the red and green channels and the blue and green channels in the color space are close to zero. In this embodiment, the limit ranges for the difference between the mean values ​​of the red and green channels and the blue and green channels are both set to a normalized range of -0.05 to 0.05.

[0066] S22, the left-side extraction submodule uses the average pixel values ​​of the red channel, green channel, and blue channel corresponding to the selected candidate gray points to calculate the chromaticity feature ratio of each spatial block. The chromaticity feature ratio includes the ratio of the average pixel value of the blue channel to the average pixel value of the green channel, and the ratio of the average pixel value of the red channel to the average pixel value of the green channel. The left-side extraction submodule combines the chromaticity feature ratios corresponding to all spatial blocks in spatial grid order to construct a two-dimensional chromaticity feature map. The two-dimensional chromaticity feature map records the spatial color distribution data of different blocks in the image. The number of dimensions of the two-dimensional chromaticity feature map is directly proportional to the total number of spatial blocks. The two-dimensional chromaticity feature map serves as the underlying basis for characterizing the physical properties of scene colors.

[0067] S23, the left-path extraction submodule calculates the degree of deviation from the normal scene distribution using the chroma feature ratio and pre-stored training set distribution parameters. The training set distribution parameters are obtained through statistical abstraction of a large number of normal light source scene samples. The training set distribution parameters include the feature mean vector and feature covariance matrix obtained offline. The left-path extraction submodule extracts the Mahalanobis distance between the chroma feature ratio and the feature mean vector of each spatial block based on the feature covariance matrix and performs statistical summation as the final Mahalanobis distance of the current frame. The Mahalanobis distance can eliminate the interference of correlation between feature dimensions and unify the units of measurement. The Mahalanobis distance calculation formula is as follows:

[0068] ;

[0069] In the formula, The Mahalanobis distance between the multidimensional chromaticity feature vectors and the normal scene distribution; This is a multidimensional chromaticity feature vector constructed by combining the chromaticity feature ratios corresponding to all spatial blocks; The feature mean vector in the distribution parameters of the training set obtained through offline statistics; This is the symbol for the matrix transpose operation; This is the inverse of the feature covariance matrix in the training set distribution parameters obtained through offline statistics.

[0070] S24, the thumbnail generation submodule and the left-side extraction submodule synchronously receive the dimensionality reduction preprocessed data. The thumbnail generation submodule performs de-mosaic interpolation calculation on the dimensionality reduction preprocessed data. The thumbnail generation submodule converts the dimensionality reduction preprocessed data in Bayer color filter array format into red, green, and blue three-channel color space image data. The thumbnail generation submodule scales the red, green, and blue three-channel color space image data to a set thumbnail data format, such as within the range of 128×128 to 256×256 pixels. The thumbnail data format preserves the physical spatial topology of the dimensionality reduction preprocessed data. The thumbnail generation submodule outputs the thumbnail data to the AI ​​inference module to perform scene semantic analysis.

[0071] See attached document Figure 3 In specific implementations, step S3 provided by the present invention may include the following steps:

[0072] In S31, the AI ​​inference module completes the feature encoding processing of the three-layer convolutional neural network, the partition semantic weight map output processing of the semantic analysis submodule, and the information classification entropy parameter extraction processing of the light source classification head submodule. The relevant initial logic and basic network parameters are configured according to the standard architecture. When outputting the partition semantic weight map, the system pre-sets to assign higher weights to key skin areas of the face and lower weights to large areas of solid color in the sky. The system directly performs channel concatenation operation on the data dimension to form a joint input vector matrix by combining the standard light source classification probability distribution matrix with the flattened bottom-level one-dimensional feature vector after feature pooling dimensionality reduction.

[0073] In the specific implementation of calculating the classification probability distribution, the light source classification head submodule is equipped with a classification fully connected layer and a Softmax function activation layer. After the extracted feature vector is input into the light source classification head submodule, the Softmax function is used to normalize and map the original logical value output by the neural network to the classification probability distribution matrix of each category of standard light source, and the information classification entropy is further calculated based on the probability distribution matrix.

[0074] In S32, the gain calculation submodule uses a fully connected regression network to perform numerical prediction calculations on the joint input vector matrix, outputting the preliminary white balance gain of each spatial block and the network prediction confidence, which characterizes the accuracy of the network prediction state. Subsequently, the fusion layer submodule maps and aligns the spatial resolution of the partition semantic weight map with the spatial block grid divided by the left extraction submodule, and performs spatial weighted summation on the preliminary white balance gain of each spatial block based on the partition semantic weight map output by the semantic analysis submodule, thereby calculating and outputting the final predicted red-green-blue three-channel AI white balance gain values. In the offline training phase, the AI ​​inference module adopts a dual-objective joint regression forward derivation training scheme of network prediction confidence and AI white balance gain values. The specific calculation formula for the offline confidence label data is as follows:

[0075] ;

[0076] In the formula, For offline confidence label data; molecule A reference constant for normalizing the range of label values; denominator To avoid the anomaly of zero denominator in the calculation results; The red-green-blue three-channel artificial intelligence white balance gain values ​​predicted by a fully connected regression network; To pre-determine the collected true white balance gain label values; These are mathematical operation symbols used to perform Euclidean distance square conversion on the difference between two sets of vectors.

[0077] S33, in some embodiments, the AI ​​inference module needs to perform complete offline training of the neural network model before accessing the system for real-time prediction. The specific model creation and training calibration process is as follows: First, the model training dataset is created and constructed. A large number of low-level Bayer images containing various complex mixed light sources and conventional single light source scenes are collected offline, and corresponding RGB images are generated simultaneously. For the collected image sample data, multi-dimensional supervised learning labels are created by manual calibration or measurement with professional color calibration instruments. The labels specifically include pixel-level semantic category labels for training the semantic segmentation network, standard light source real classification labels for training the light source classification head sub-module, and real white balance gain label values ​​for training the fully connected regression network. Subsequently, a neural network model containing a convolutional neural network, a semantic segmentation network, a light source classification head structure, and a fully connected regression network is initialized and constructed in the deep learning framework.

[0078] See attached document Figure 4 In specific implementations, step S4 provided by the present invention may include the following steps:

[0079] S41, the anomaly detection submodule obtains the Mahalanobis distance output by the left-path extraction submodule. The anomaly detection submodule compares the Mahalanobis distance with a preset interception threshold. The preset interception threshold is generated by extracting the 99th percentile of the offline training set data distribution. In this embodiment, the specific value range of the preset interception threshold is set to 3.0 to 5.0. When the Mahalanobis distance is greater than the preset interception threshold, the anomaly detection submodule determines that the data features at the bottom layer of the input system deviate significantly from the empirical distribution boundary of the conventional model. The anomaly detection submodule directly blocks the prediction output link of the gain calculation submodule and forces the dynamic fusion weight output by the system to be assigned a value of 0.

[0080] S42, when the Mahalanobis distance is less than or equal to the preset interception threshold, the fusion layer submodule synchronously obtains the network prediction confidence output by the gain calculation submodule and the information classification entropy parameter output by the light source classification head submodule. The network prediction confidence objectively represents the reliability of the prediction result, and the information classification entropy parameter objectively represents the disorder of the scene classification features. The fusion layer submodule uses the network prediction confidence and the information classification entropy parameter to calculate the initial dynamic weight value. The fusion layer submodule performs a positive gain mapping operation on the initial dynamic weight value based on the network prediction confidence and performs a reverse suppression operation on the initial dynamic weight value based on the information classification entropy parameter.

[0081] S43, To clarify the specific implementation rules for dynamic fusion weight calculation, the fusion layer submodule uses a closed interval truncation function to process the upper and lower limits of the initial dynamic weight values. The closed interval truncation function forcibly restricts the calculated feature data within the set absolute boundaries. The fusion layer submodule generates dynamic fusion weights that are strictly distributed within the real number interval from 0 to 1. The specific calculation formula for the dynamic fusion weights is as follows:

[0082] ;

[0083] In the formula, To calculate the generated dynamic fusion weights; A mathematical operation function to find the maximum value; This is a bit constant representing the lower limit of the dynamic fusion weights; A mathematical operation function to find the minimum value; This is a constant representing the upper limit of the dynamic fusion weights; The confidence level adjustment coefficient preset for the system; The network prediction confidence score output by the fully connected regression network; The entropy suppression coefficient preset for the system; The information classification entropy parameters are calculated using the mathematical information entropy model combined with the classification probability distribution matrix; the confidence adjustment coefficient is set to a range of 1.0 to 2.0, and the entropy suppression coefficient is set to a range of 0.1 to 0.5.

[0084] In specific implementations, step S5 provided by the present invention may include the following steps:

[0085] S51, the traditional white balance module receives the dimensionality reduction preprocessed data. The traditional white balance module uses the gray-world algorithm to statistically calculate the mean values ​​of the red channel, green channel, and blue channel within the global distribution range of the dimensionality reduction preprocessed data. Simultaneously, the traditional white balance module uses the white point detection algorithm to search for candidate reference white points that conform to the emission trajectory of the standard light source within the set physical color space. The traditional white balance module uses the mean value of the green channel as a benchmark, and divides the mean value of the green channel by the mean values ​​of the red channel and blue channel respectively. Combining the color channel ratio results of the candidate reference white points selected by the white point detection algorithm, the traditional white balance module calculates and generates the traditional red, green, and blue three-channel reference gain values. The specific statistical strategy of the gray-world algorithm and the specific white point selection and judgment rules of the white point detection algorithm are based on existing automatic white balance methods for image signal processing.

[0086] S52, to clarify the specific implementation method of weighting artificial intelligence and traditional white balance gain based on dynamic fusion weight, the dynamic weighted fusion module simultaneously extracts the artificial intelligence white balance gain values ​​of the red, green, and blue channels and the traditional red, green, and blue channel reference gain values. The dynamic weighted fusion module obtains the dynamic fusion weight, and uses the dynamic fusion weight as an adjustment factor to perform linear interpolation weighted fusion operation on the artificial intelligence white balance gain values ​​and the traditional red, green, and blue channel reference gain values, calculating and generating a fusion gain result covering three independent color channels. The linear interpolation weighted fusion calculation formula is as follows:

[0087] ;

[0088] In the formula, To calculate the fusion gain result covering three independent color channels; Dynamic fusion weights generated for pre-calculation; The predicted AI white balance gain values ​​for the red, green, and blue three channels; These are the fundamental constants for performing weighted complementary operations; The traditional red, green, and blue three-channel reference gain values ​​are generated for the underlying calculations.

[0089] S53, when the value of the dynamic fusion weight is equal to 1, the dynamic weighted fusion module outputs the red, green and blue three-channel artificial intelligence white balance gain value separately. When the Mahalanobis distance exceeds the preset interception threshold and the value of the dynamic fusion weight is forcibly assigned to 0, the dynamic weighted fusion module outputs the traditional red, green and blue three-channel reference gain value separately. The dynamic weighted fusion module maintains the continuity of the fusion gain result data output state through the pre-linear weighted fusion operation.

[0090] See attached document Figure 5 In specific implementations, step S6 provided by the present invention may include the following steps:

[0091] S61, the timing smoothing module receives the fusion gain result and dynamic fusion weights corresponding to the current video frame output by the dynamic weighted fusion module. The timing smoothing module performs timing smoothing processing on the fusion gain result corresponding to the current video frame using an infinite impulse response low-pass filter. The timing smoothing module synchronously calculates the dynamic smoothing coefficients of the infinite impulse response low-pass filter based on the dynamic fusion weights. The specific calculation formula for the dynamic smoothing coefficients is as follows:

[0092] ;

[0093] In the formula, The dynamic smoothing coefficients of the calculated infinite impulse response low-pass filter; The preset maximum smoothing coefficient; This is the preset smoothness coefficient adjustment amount; The dynamic fusion weights generated by the pre-calculation are set; the preset maximum smoothing coefficient selection range is set to 0.8 to 0.95, and the preset smoothing coefficient adjustment range is set to 0.3 to 0.5.

[0094] S62, the temporal smoothing module calculates the final white balance output gain of the current video frame by combining the final white balance output gain of the previous video frame with the fusion gain result of the current video frame. The temporal smoothing module extracts the final white balance output gain of the previous video frame as a historical state benchmark. The temporal smoothing module uses the dynamic smoothing coefficient as a weight to retain historical data and performs iterative calculations. The specific calculation formula for the final white balance output gain is as follows:

[0095] ;

[0096] In the formula, For the current number The final white balance output gain corresponding to each video frame; This is a discrete index variable representing the time series of the current video frame; The dynamic smoothing coefficients of the calculated infinite impulse response low-pass filter; For the previous one The final white balance output gain corresponding to each video frame; This is a discrete index variable representing the time series of the previous video frame; The fundamental constants for performing the smoothing coefficient complementation operation; For the current number The fusion gain result calculated from each video frame;

[0097] The temporal smoothing module reduces the dynamic smoothing coefficient when the value of the dynamic fusion weight approaches 1, and increases the dynamic smoothing coefficient and outputs the final white balance output gain corresponding to the current video frame when the value of the dynamic fusion weight approaches 0.

[0098] S63, the gain output module extracts the final white balance output gain corresponding to the current video frame output by the timing smoothing module. The gain output module loads the final white balance output gain corresponding to the current video frame into the hardware Bayer domain multiplier. The hardware Bayer domain multiplier performs physical multiplication gain operation on the original pixel data input to the intelligent automatic white balance system according to the independent red, green and blue color channels. The gain output module sends the original pixel data after physical multiplication gain operation to the subsequent demosaicing unit, color matrix correction unit and gamma mapping unit for pipeline closed-loop operation. The specific pixel interpolation strategy of the demosaicing unit, the specific color space transformation matrix parameters of the color matrix correction unit and the specific nonlinear brightness mapping curve configuration of the gamma mapping unit are based on existing conventional image signal processing rendering methods.

[0099] See attached document Figure 6 and attached Figure 7 To aid in understanding the technical solution of this invention, the following provides application examples of shooting scenarios with complex mixed light sources on smartphone terminals.

[0100] The image sensor of the smartphone terminal captures the underlying image electrical signals in an environment where indoor incandescent light and outdoor natural light meet. The RAW input module receives the underlying image electrical signals and inputs them into the preprocessing module. The preprocessing module sequentially performs black level correction, lens shading correction, and Bayer domain noise reduction. The preprocessing module performs downsampling operations according to the adjacent same-color pixel merging mechanism to generate dimensionality-reduced preprocessed data. The left-path extraction submodule divides the dimensionality-reduced preprocessed data into multiple independent spatial blocks and independently calculates the pixel mean of each color channel. The left-path extraction submodule combines the chromaticity feature ratios of each spatial block to construct a multi-dimensional chromaticity feature vector and calculates and extracts the Mahalanobis distance. The thumbnail generation submodule simultaneously generates red, green, and blue three-channel color space image data that retains the physical spatial topology and scales it to the specified RGB format thumbnail data format.

[0101] The AI ​​inference module receives multi-dimensional chromaticity feature vectors and thumbnail data formats. The semantic analysis submodule identifies faces, sky, vegetation, and indoor / outdoor semantic regions in the RGB format thumbnails, assigning higher weights to key skin areas on faces and lower weights to large areas of solid color in the sky, and outputs corresponding partition semantic weight maps. The gain calculation submodule combines the standard light source classification probability distribution matrix and multi-dimensional chromaticity feature vectors to predict the AI ​​white balance gain values ​​for the red, green, and blue channels, as well as the network prediction confidence. The anomaly detection submodule compares and determines that the Mahalanobis distance does not exceed the preset interception threshold. The fusion layer submodule calculates the dynamic fusion weights based on the confidence forward mapping and the information classification entropy reverse suppression results, after truncation and limiting. The traditional white balance module uses the gray-world algorithm and white point detection algorithm to calculate the traditional red, green, and blue channel reference gain values. The dynamic weighted fusion module uses the dynamic fusion weights as an adjustment factor to perform linear interpolation weighted fusion operations on the AI ​​white balance gain values ​​and the traditional reference gain values.

[0102] The timing smoothing module receives the fusion gain result and dynamic fusion weight covering three independent color channels. Based on the dynamic fusion weight, the timing smoothing module dynamically calculates the dynamic smoothing coefficient of the infinite impulse response low-pass filter. Combining the historical state of the final white balance output gain of the previous video frame, the timing smoothing module uses the dynamic smoothing coefficient to perform iterative calculation and output the final white balance output gain corresponding to the current video frame. The gain output module loads the final white balance output gain into the hardware Bayer domain multiplier to perform physical multiplication gain operation. The gain output module sequentially sends the processed raw pixel data to the demosaic processing unit, the color matrix correction unit, and the gamma mapping unit for pipeline closed-loop operation and outputs the final image.

[0103] The testing system is built on a standard multi-source image dataset and a dynamic video sequence dataset to create an underlying evaluation environment. The test comparison objects include the pure traditional gray-scale automatic white balance algorithm, the pure deep learning automatic white balance algorithm, and the intelligent automatic white balance system proposed in this invention. The evaluation indicators are the chromaticity angle error, which characterizes the accuracy of color reproduction, and the temporal fluctuation variance, which characterizes the smoothness of the inter-frame gain of the video stream. The smaller the physical value of the chromaticity angle error, the more accurate the objective color reproduction. The smaller the physical value of the temporal fluctuation variance, the more stable the overall brightness and color transition of the video stream.

[0104] The test system ran three comparison objects on a standard multi-source image dataset and extracted and recorded the corresponding average chromaticity angle error. The test system also ran three comparison objects on a dynamic video sequence dataset and extracted and recorded the corresponding temporal fluctuation variance values. The pure traditional gray-scale automatic white balance algorithm showed a physical state with high chromaticity angle error in complex semantic scenes, while the pure deep learning automatic white balance algorithm showed a physical state with drastic changes in temporal fluctuation variance in abnormally distributed samples.

[0105] Under the combined constraints of the Mahalanobis distance anomaly interception mechanism and the dynamic smoothing coefficient adaptive adjustment mechanism, the chromaticity angle error of the intelligent automatic white balance system proposed in this invention is maintained within the lowest physical range in all test scenarios. The temporal fluctuation variance of the intelligent automatic white balance system proposed in this invention exhibits the smoothest linear evolution trajectory. The experimental test data objectively proves that the proposed solution has definite technical advantages in both color reproduction accuracy in complex scenes and dynamic stability of video streams.

[0106] While specific embodiments of the present invention have been described above, those skilled in the art should understand that these specific embodiments are merely illustrative. Those skilled in the art can omit, substitute, and modify the details of the above methods and systems in various ways without departing from the principles and essence of the present invention. For example, combining the above method steps to perform substantially the same function and achieve substantially the same result using substantially the same method falls within the scope of the present invention.

Claims

1. A method for intelligent automatic white balance, characterized in that, Includes the following steps: The raw Bayer array data output by the sensor is acquired, and black level correction, lens shading correction, Bayer domain noise reduction and downsampling are performed to obtain preprocessed image data. Extract the mean values ​​of each partition channel, calculate the chromaticity feature ratio and Mahalanobis distance, and generate RGB format thumbnails; Feature vectors are extracted using a convolutional neural network, and a semantic weight map of partitions is output based on the thumbnail. Based on feature vector extraction of classification probability distribution and information classification entropy, predict artificial intelligence white balance gain and network prediction confidence; If the Mahalanobis distance exceeds the preset interception threshold, the dynamic fusion weight is assigned a value of zero; if it does not exceed the threshold, the dynamic fusion weight is calculated based on the positive mapping of confidence and the negative suppression result of information classification entropy after truncation limit. Using dynamic fusion weights as adjustment factors, linear interpolation weighting is performed on the artificial intelligence and traditional white balance gains to obtain the fusion gain result. The results are subjected to inter-frame adaptive infinite impulse response low-pass filtering, and the three-channel white balance gain is output to the image signal processor.

2. The intelligent automatic white balance method according to claim 1, characterized in that, The steps of performing black level correction, lens shading correction, Bayer domain noise reduction, and downsampling to obtain preprocessed image data include: The original Bayer array data is subjected to a subtraction compensation operation on each pixel value using a set black level reference value, and the result is truncated to zero when it is negative. Multiplicative gain compensation operations are performed on the pixels at each spatial location of the data after black level correction using a pre-calibrated two-dimensional grid gain table. While keeping the underlying Bayer color filter array structure unchanged, spatial domain filtering is performed to suppress high-frequency noise points in the data after lens shading correction. Adjacent same-color pixel merging mechanism is used to perform downsampling operation according to the specified downsampling ratio to generate dimension-reduced preprocessed data as the preprocessed image data.

3. The intelligent automatic white balance method according to claim 1, characterized in that, The steps of extracting the mean value of each partition channel and calculating the chromaticity feature ratio and Mahalanobis distance include: The preprocessed image data is divided into multiple independent spatial blocks according to the set grid division rules. The pixel mean of each color channel is calculated independently, and overexposed pixels with values ​​exceeding the set saturation threshold are removed from the mean calculation sample. Candidate gray points are filtered based on a preset gray point limit range. The chromaticity feature ratio of each spatial block is calculated using the average pixel values ​​of the red channel, green channel, and blue channel corresponding to the filtered candidate gray points. A multidimensional chromaticity feature vector is constructed by combining the chromaticity feature ratios corresponding to all spatial blocks. The Mahalanobis distance between the multidimensional chromaticity feature vector and the normal scene distribution is calculated by combining the pre-stored training set distribution parameters containing the feature mean vector and feature covariance matrix.

4. The intelligent automatic white balance method according to claim 1, characterized in that, The steps of extracting feature vectors using a convolutional neural network and outputting a partition semantic weight map based on the thumbnail include: Identify the face, sky, vegetation, and indoor / outdoor semantic regions in the RGB format thumbnail, assign higher weights to key skin areas of the face, and assign lower weights to large areas of solid color in the sky, and output the partition semantic weight map. The extracted partition features are arranged into a spatial grid, and a convolutional neural network is used for feature mapping and pooling dimensionality reduction to extract feature vectors with spatial distribution information.

5. The intelligent automatic white balance method according to claim 1, characterized in that, The steps for extracting the classification probability distribution and information classification entropy based on feature vectors, and predicting the artificial intelligence white balance gain and network prediction confidence include: The standard light source classification probability distribution matrix and the flattened feature vectors after feature pooling and dimensionality reduction are combined in the data dimension to form a joint input vector matrix; Numerical prediction calculations are performed on the joint input vector matrix using a fully connected regression network, and the preliminary white balance gain of each spatial block and the network prediction confidence, which characterizes the accuracy of the network prediction state, are output. The spatial resolution of the partition semantic weight map is mapped and aligned with the divided spatial block grid. Based on the partition semantic weight map, the preliminary white balance gain of each spatial block is spatially weighted and summed to calculate and generate the artificial intelligence white balance gain.

6. The intelligent automatic white balance method according to claim 1, characterized in that, Before the step of extracting the classification probability distribution and information classification entropy based on feature vectors, and predicting the artificial intelligence white balance gain and network prediction confidence, a step of performing complete offline training of the neural network model is also included. The offline training step includes: The training dataset for the neural network model is constructed by offline acquisition of the underlying Bayer image and synchronous generation of the corresponding RGB image. For the collected image sample data, multi-dimensional supervised learning labels are created. The supervised learning labels specifically include pixel-level semantic category labels for training the semantic segmentation network, standard light source real classification labels for training the light source classification head sub-module, and real white balance gain label values ​​for training the fully connected regression network. Initialize and build a neural network model within a deep learning framework, including a convolutional neural network, a semantic segmentation network, a light source classification head structure, and a fully connected regression network.

7. The intelligent automatic white balance method according to claim 1, characterized in that, If the threshold is not exceeded, the steps for calculating the dynamic fusion weights based on the confidence forward mapping and the information classification entropy reverse suppression result, after truncation and limiting, include: Extract the quantiles of the offline training set data distribution to generate the preset interception threshold, and perform a numerical comparison operation between the Mahalanobis distance and the preset interception threshold; When the value of the Mahalanobis distance is greater than the preset interception threshold, it is determined that the data features at the bottom layer of the input system deviate significantly from the empirical distribution boundary of the conventional model, the prediction output link of the gain calculation submodule is blocked, and the dynamic fusion weight is forcibly assigned to 0. When the Mahalanobis distance is less than or equal to the preset interception threshold, a positive gain mapping operation is performed on the initial dynamic weight value based on the network prediction confidence, and a reverse suppression operation is performed on the initial dynamic weight value based on the information classification entropy parameter. The upper and lower limit bits are processed using the closed interval truncation function to generate the dynamic fusion weight distributed in the real number interval from 0 to 1.

8. The intelligent automatic white balance method according to claim 1, characterized in that, The step of performing linear interpolation weighting on the artificial intelligence and traditional white balance gains using dynamic fusion weights as adjustment scaling factors to obtain the fusion gain result includes: The mean values ​​of the red, green, and blue channels in the global distribution range of the preprocessed image data are statistically calculated using the grayscale world algorithm. Simultaneously, a white point detection algorithm is used to search for candidate reference white points that conform to the emission trajectory of a standard light source. Based on the color channel ratio results of the candidate reference white points, the traditional red, green, and blue three-channel reference gain values ​​are calculated and generated as the traditional white balance gain. Using the dynamic fusion weight as an adjustment factor, a linear interpolation weighted fusion operation is performed on the artificial intelligence white balance gain and the traditional red, green and blue three-channel reference gain values ​​to calculate and generate the fusion gain result covering the three independent color channels.

9. The intelligent automatic white balance method according to claim 1, characterized in that, The steps of performing inter-frame adaptive infinite impulse response low-pass filtering on the results and outputting the three-channel white balance gain to the image signal processor include: The fusion gain result corresponding to the current video frame is subjected to temporal smoothing using an infinite impulse response low-pass filter, and the dynamic smoothing coefficient of the infinite impulse response low-pass filter is calculated synchronously based on the dynamic fusion weight. Extract the final white balance output gain corresponding to the previous video frame as the historical state benchmark, use the dynamic smoothing coefficient as the historical data retention weight, combine the final white balance output gain corresponding to the previous video frame with the fusion gain result corresponding to the current video frame to perform iterative calculation, and calculate the final white balance output gain corresponding to the current video frame as the three-channel white balance gain. The final white balance output gain corresponding to the current video frame is loaded into the hardware Bayer domain multiplier to perform physical multiplication gain operation. The original pixel data after physical multiplication gain operation is then sequentially sent to the demosaic processing unit, color matrix correction unit, and gamma mapping unit for pipeline closed-loop operation.

10. The intelligent automatic white balance method according to claim 9, characterized in that, The steps for synchronously calculating the dynamic smoothing coefficients of the infinite impulse response low-pass filter based on the dynamically fused weights include: The dynamic smoothing coefficient is calculated by subtracting the product of the preset smoothing coefficient adjustment amount and the dynamic fusion weight from the preset base maximum smoothing coefficient. The dynamic smoothing coefficient is decreased when the value of the dynamic fusion weight approaches one, and the dynamic smoothing coefficient is increased when the value of the dynamic fusion weight approaches zero.