End-to-end color imaging system chromaticity characterization method based on hyperspectral calibration
Through hyperspectral calibration and end-to-end deep learning methods, the visual distortion and performance bottleneck problems of color imaging systems in high-resolution imaging are solved, and device-independent high-precision colorimetry and real-time conversion are achieved, which is suitable for quality inspection in multiple fields.
Patent Information
- Application Number
- CN202510966043.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-10-17
AI Technical Summary
When performing high-resolution imaging, existing color imaging systems rely on traditional methods to calibrate standard color cards, which have limited sampling coverage. This results in large extrapolation errors in highly saturated dark colors and severe visual distortion. Furthermore, pixel-by-pixel calculation becomes a performance bottleneck, making it difficult to achieve device-independent high-precision colorimetry.
Hyperspectral calibration is combined with system spectral sensitivity and color matching function to construct a calibration module. The RGB to XYZ color space conversion is realized through an end-to-end deep learning model. The entire image is mapped using a fully convolutional network or a generative adversarial network, and training is performed by combining numerical error with visual consistency constraints.
It achieves high-resolution color space conversion in milliseconds, breaks away from the constraints of standard color cards, takes into account color difference accuracy, visual naturalness and system real-time performance, and is suitable for cross-media color management between different devices and quality inspection in the fields of environment, industry, agriculture, food, etc.
Smart Images

Figure CN120807660A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of color imaging, display and colorimetry, and particularly relates to an end-to-end color imaging system colorimetric characterization method based on hyperspectral calibration. BACKGROUND
[0002] In order to keep consistent colors among imaging, display, industrial and agricultural detection, environmental food monitoring and remote medical treatment and other scenes, realize color management and scene colorimetry, the color management and colorimetry must accurately map the device-related RGB data to the device-independent XYZ standard space; however, the industry still adopts the standard color card calibration to collect data, and realizes color space conversion, color management and colorimetry by using the scheme of polynomial or matrix regression. This scheme has the disadvantages of high dependence on calibration conditions (controlled light, dark room and calibration colorimetry), limited sampling coverage and the like for single-point color measurement, resulting in large extrapolation error of high saturation dark color; for large-area high-resolution imaging systems, the above method only makes empirical fitting, ignores the physical constraints between the camera spectral sensitivity and the CIE color matching function, and thus causes visual distortion when the data linear relationship is strong, although the color difference is low; and the traditional matrix regression pixel-by-pixel calculation becomes a performance bottleneck for high-resolution high-speed conversion. The industry urgently needs an integrated colorimetric characterization scheme which simultaneously introduces "end-to-end calibration" and "generated visual effect fidelity".
[0003] Scholars have carried out a large amount of colorimetric characterization research, and the main methods include polynomial model [1,2] , three-dimensional look-up table [3] and neural network [4,5,6] and the like. The above methods all use standard color cards for system colorimetric calibration, and generally face the problems of insufficient sampling quantity, limited sample coverage range and model optimization of only pixel precision and loss of visual naturalness, which leads to the phenomenon of large prediction deviation and unnatural visual effect of the model when facing wide color gamut or rich texture scenes [7,8] , and in the actual landing process, the physical color card is high in price and easy to age, and the cost is high after one round of lighting, shooting and measuring; and in high resolution, pixel-by-pixel matrix operation becomes a bottleneck of real-time processing, and it is difficult to meet the delay index of industrial pipeline or mobile terminal.
[0004] Core concept: RGB is the three-channel electrical signal corresponding to the digital value output directly by the camera chip, XYZ and Lab are color value definitions proposed by the International Commission on Illumination, which are decoupled from the device; hyperspectral calibration refers to replacing the limited color card with dozens of waveband information to generate true values directly through fixed spectral integration formula; end-to-end deep mapping means that the network processes the entire image at one time rather than point-by-point cycle; two-stage training is to first do numerical "rough adjustment" and then do visual "fine adjustment", which can effectively avoid the oscillation caused by mixing multiple losses at one time. The invention takes hyperspectral calibration as the main body and realizes fast, accurate and device-independent colorimetric characterization with whole-image end-to-end network.
[0005] References
[0006] [1] Su M, Li S Q, Yang L M, et al. Color Reproduction of Qianlong Color Spectrum across Devices Based on Colorimetric Characterization[J]. Journal of Textile Research, 2023, 44(10): 104-112. DOI:10.13475 / j.fzxb.20220905601.
[0007] [2] Cao Q. Multispectral Dimensionality Reduction Algorithm Based on Second-order Polynomial Regression and Weighted Principal Component Analysis[J]. Optical Technique, 2023, 49(02): 250-256. DOI:10.13741 / j.cnki.11-1879 / o4.2023.02.015.
[0008] [3] Li R J, Deng Q. Research on RGB to XYZ Color Space Conversion Based on Three-dimensional Look-up Table[J]. Packaging Engineering, 2012, 33(13): 116-119. DOI:10.19554 / j.cnki.1001-3563.2012.13.030.
[0009] [4] Yan P G. Color Constancy Algorithm Combining Colorimetric Characterization and Lightweight Neural Network[J]. Computer Knowledge and Technology, 2023, 19(20): 47-50. DOI:10.14004 / j.cnki.ckt.2023.0998.
[0010] [5] Li Y S, Liao N F, Li Y M, et al. Colorimetric Conversion Simulation of Multi-primary Display System Based on Brightness Factor Graded BP Neural Network[J]. Optical Technique, 2023, 49(03): 257-263. DOI:10.13741 / j.cnki.11-1879 / o4.2023.03.001.
[0011] [6]Tong P W,Jen J C,Wang C T.Colorimetric Characterization of Color Image Sensors Based on Convolutional Neural Network Modeling[J].Sensors and Materials,2019,31(5):1513-1513.
[0012] [7]Kirchner, Eric, et al. "Exploring the limits of color accuracy in technical photography." Heritage Science 9.1 (2021): 57.
[0013] [8]Finlayson, Graham D., and Yuteng Zhu. "Designing color filters that make cameras more colorimetric." IEEE Transactions on Image Processing 30 (2020): 853-867. SUMMARY
[0014] In order to solve the above-mentioned problems, the present application proposes an end-to-end color imaging system colorimetric characterization method based on hyperspectral calibration, that is, the color chart data is collected by hyperspectral imaging, the system spectral sensitivity and color matching function (CMF) are combined to calibrate the color space data, and a calibration module is constructed; and an end-to-end deep learning color space conversion model is proposed, which realizes the conversion from device-dependent RGB value to device-independent color space. The technology can be applied to cross-media color management between different devices; it can also be applied to color measurement of samples to be measured, and can be used for quality detection and evaluation in the fields of environment, industry, agriculture, food, etc., to realize CIE XYZ, CIE LAB or Lxy color imaging measurement, and realize high-resolution imaging colorimetric output. This method can not only get rid of the shackles of standard color cards and artificial calibration, but also can complete high-resolution inference in milliseconds, truly taking into account color difference accuracy, visual naturalness and system real-time performance, providing visual support for film and television production and wearable display; at the same time, it provides reliable device-independent color (CIE XYZ and CIELAB) output for online detection in the fields of industry, agriculture, environment and food.
[0015] In order to achieve the above-mentioned purposes, the present application adopts the following technical solutions:
[0016] S1, collecting spectral radiance of a target scene or obtaining a multi-channel spectral cube data containing spatial resolution and band information from an existing data set;
[0017] S2, determining a three-channel spectral sensitivity curve and a scaling coefficient of an imaging device; as a weight basis for subsequent calculation;
[0018] S3, inputting the spectral cube data and the spectral sensitivity curve into a mapping module, the mapping module can be a single-layer linear MLP with fixed weights, one-time matrix multiplication, a lookup table, sparse convolution or other equivalent integral implementation, performing a weighted integral operation to obtain pixel-level RGB data;
[0019] S4, performing the same integral operation on the spectral cube data and the international lighting commission color matching function to obtain a pixel-level XYZ three-stimulus value image;
[0020] S5, performing amplitude normalization processing, uniform size on the RGB data and the XYZ data, dividing the training set, the validation set and the test set according to a preset proportion, and establishing a data loader based on the division result;
[0021] S6, constructing an end-to-end color mapping model and completing training, the model can use a full convolutional network, a generative adversarial network or other deep frameworks with whole image mapping capability; in the preferred embodiment, an adversarial structure containing an encoder-decoder generator and a local discriminator is adopted, and numerical error constraints and perceptual consistency constraints are considered at the same time during the training process until the model parameters converge;
[0022] S7, loading the converged model weight to an inference environment, inputting real-time RGB data, and outputting corresponding XYZ data in a single inference, and simultaneously determining single-frame delay and throughput to evaluate running performance;
[0023] S8, calculating CIE DE2000 color difference, mean absolute error and mean square error of the output three-stimulus value image, and evaluating visual naturalness by using LPIPS, SSIM and other recognized perceptual indicators, and comprehensively verifying the technical effects of the method in numerical accuracy and subjective consistency;
[0024] The above method can further comprise the following steps:
[0025] S11, according to the application requirement of the target imaging system, setting the band sampling mode: 31 visible light center bands with equal interval of 10 nm can be selected, and the near-infrared segment can also be extended; at the same time, the spatial resolution of single-frame imaging is determined. The band and resolution configuration is suitable for self-sampling hyperspectrum and screening public data sets;
[0026] S12, obtaining the spectral cube by field self-collection or using a public data set;
[0027] S13, interpolate the spectral cube data obtained in step S12 in the wavelength dimension to the same waveband center as the subsequent integral weight (spectral sensitivity curve or CMF); then perform geometric correction on distortion and parallax, crop to retain the effective field of view, and finally output a spectral cube file meeting the specifications of S11, and the file format can adopt.npy,.mat or ENVI standard header file;
[0028] The method described above, optionally, S2 specifically comprises:
[0029] S21, for the R, G and B three channels of the imaging device to be calibrated, obtain the original wavelength-response curve through document extraction or actual measurement calibration;
[0030] S22, uniformly process the three-channel curve obtained in step S21, if the original sampling interval is inconsistent with the waveband center of the hyperspectral cube, then use one-dimensional linear or cubic spline interpolation to resample the curve to an array consistent with the waveband set in step S1;
[0031] S23, calibrate the channel gain (scaling factor) for analog camera white balance and analog-to-digital conversion gain;
[0032] S24, construct a sensitivity matrix, multiply the three-channel spectral sensitivity column vector of step S22 by the gain coefficient obtained in step S23, and splice it into a sensitivity matrix Sgain, which will be used as a fixed weight in spectral integration in step S3 and will not be updated; at the same time, store the gain value and resampled wavelength in the metadata file for subsequent archiving and tracing;
[0033] The method described above, optionally, S3 specifically comprises:
[0034] S31, solidify the three-channel sensitivity-gain matrix Sgain generated in S24 as the weight of the mapping unit, and the mapping unit can be optionally one-time matrix multiplication, single-layer MLP with fixed weight, or other equivalent implementation;
[0035] S32, perform weighted integration on the spectral cube data, and after all pixels are calculated, reshape the three-channel tensor to obtain pixel-level RGB data;
[0036] S33, if the spectral radiance Lλ is in absolute power units, then the RGB data naturally has physical dimensions, at which time it can be directly entered into subsequent network training and inference, without the need for 0-255 quantization or proportional compression. If only relative radiance is obtained in the simulation environment and needs to be connected with the image pipeline, the maximum brightness can be multiplied by a uniform coefficient for quantization compression in the subsequent S5 step, by the whole image or batch;
[0037] The method described above, optionally, S4 specifically comprises:
[0038] S41, read the international lighting commission CIE 1931 color matching function, form a waveband-channel matrix, and solidify it as a mapping unit weight. Similar to S3, this unit can use one-time matrix multiplication, fixed weight single-layer MLP, or other equivalent implementation methods, and the weight remains unchanged during operation;
[0039] S42, after the calculation of the spectral cube data, reshape it into a three-channel tensor to obtain pixel-level XYZ data;
[0040] S43, since the CIE matching function implicitly weights the spectrum, step S4 does not introduce additional gain coefficients; its output is in the same physical dimension as the RGB data of step S3. If the radiance is only a relative value and the experimental process needs to be consistent with the 0-255 range of the RGB data quantization link, the XYZ data can be uniformly scaled in the optional scheme in the subsequent S5 step;
[0041] The above method, optionally, S6 specifically includes:
[0042] S61, in the case of mapping the entire RGB to XYZ, the present application allows the use of any deep network framework that can process the entire image at once, including but not limited to:
[0043] 1) pure fully convolutional network (FCN), encoder-decoder network, U-Net and its variants;
[0044] 2) generative adversarial network (GAN) series: Pix2Pix, CycleGAN, StyleGAN-based structure;
[0045] 3) diffusion model, hybrid attention-convolutional network, etc.
[0046] The model input is three-channel RGB data, and the output is three-channel XYZ data. The network weight is updated uniformly through backpropagation iteration;
[0047] S62, the model can use a single-stage joint loss or a phased strategy, specifically including but not limited to:
[0048] 1) pixel-level L1 / L2 loss, used to ensure the absolute error of XYZ data;
[0049] 2) perceptual loss, feature matching loss or adversarial loss, used to improve visual naturalness;
[0050] It can be optimized together at one time, or it can be pre-trained for several rounds using numerical errors first, and then introduce perceptual or adversarial loss for refinement until the model converges on the validation set;
[0051] The optimizer can choose Adam, RMSProp, or other adaptive methods. The learning rate, batch size, and loss weight can be flexibly set according to the data scale and hardware conditions. After training is completed, all network parameters are frozen and the weight file is saved for subsequent inference deployment.
[0052] In the above method, optionally, S7 specifically includes:
[0053] S71. Import the end-to-end mapping model weights into the inference environment. During the loading phase, the network structure description file, normalization parameters, and spectral metadata are read simultaneously to ensure that the input tensors are completely consistent with the scale and arrangement during training.
[0054] S72. Convert the three-channel RGB image output by the camera in real time into a network input tensor using the same normalization logic as S5, and directly obtain the entire XYZ tristimulus image in a single forward propagation. After inference is completed, a high-resolution timer is used to record the end-to-end latency. If the hardware supports pipeline parallelism, the average frame rate or number of images per second under continuous stable output conditions can also be calculated.
[0055] In the above method, optionally, S8 specifically includes:
[0056] S81. The inferred tristimulus value image is first compared point by point with the true value at the pixel level.
[0057] S82. To verify whether the network meets the numerical standards while avoiding local false colors or texture loss, the present invention simultaneously calculates three perception indicators: LPIPS, SSIM, and FID. LPIPS measures the difference in deep representations using the distance of pre-trained convolutional features;
[0058] The beneficial effects of the present invention are:
[0059] This paper proposes an end-to-end color space conversion method based on hyperspectral calibration and deep learning. This method uses hyperspectral imaging to acquire color reference data, combines system spectral sensitivity and color matching functions (CMF) to calibrate the color space data, and constructs a calibration module. It also proposes an end-to-end deep learning color space conversion model to achieve device-dependent RGB to device-independent color space conversion. This technology can be applied to cross-media color management between different devices and to the color measurement of test samples. This technology is targeted at quality inspection and evaluation in environmental, industrial, agricultural, and food applications, enabling CIEXYZ, CIE LAB, or Lxy color imaging measurements and achieving large-area, high-resolution imaging colorimetric output.
[0060] This method can not only break free from the constraints of standard color cards and manual calibration, but also complete high-resolution inference in milliseconds, truly taking into account color difference accuracy, visual naturalness and system real-time performance, providing visual support for film and television production and wearable displays; at the same time, it can provide reliable device-independent color (CIE XYZ and CIELAB) output for online detection in industry, agriculture, environment and food.
[0061] This invention addresses the current industry practice of calibrating data using standard color charts and implementing color space conversion, color management, and colorimetric measurement using polynomial or matrix regression. This approach relies heavily on controlled lighting, darkroom conditions, and manual photography, resulting in limited sampling coverage and large extrapolation errors for high-saturation and dark areas. Furthermore, the algorithm relies solely on empirical fitting, completely ignoring the physical constraints between the camera's spectral sensitivity and the CIE color matching function. This often results in visual distortion despite low color difference when the data has a strong linear relationship. Furthermore, the pixel-by-pixel measurement required by traditional matrix regression poses a performance bottleneck for high-speed measurement.
[0062] The advantage of the present invention is that it is a color space conversion method based on hyperspectral physical calibration and end-to-end deep learning. This method uses hyperspectral acquisition combined with camera spectral sensitivity curves and color matching functions to synchronously convert scene color information into RGB and XYZ true value image pairs required for training, and uses this data for training, thereby achieving high-precision and real-time color space conversion without the need for standard color card calibration.
[0063] By embedding hyperspectral physical priors into an end-to-end deep learning network, we achieve end-to-end mapping from RGB to XYZ. The main technical benefits are: (1) eliminating expensive equipment such as standard color cards, darkrooms, and spectrometers, simplifying the on-site calibration process; being robust to extremely saturated and low-light scenes, and avoiding the streaking artifacts and distortion common in traditional methods. (2) When applying a generative adversarial network, through the L1 combined with the adversarial loss stage training scheme, parameters such as CIE DE2000 that reflect numerical accuracy are reduced by more than 60% on average compared to traditional polynomial regression, achieving PSNR > 48dB, and simultaneously improving subjective visual indicators such as LPIPS, SSIM, and FID. (3) The end-to-end mapping inference speed is fast, with the single-frame inference latency at 512×512 resolution reaching milliseconds. (4) The generative adversarial network has only 32MB of parameters, making it lightweight to deploy. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 -Simulation calibration module diagram;
[0065] Figure 2 -Two-stage adversarial mapping module diagram;
[0066] Figure 3 -End-to-end deep learning colorimetry feature extraction method. DETAILED DESCRIPTION
[0067] In order to make the technical means and purposes of the present application easy to understand, the present application is further described below in combination with specific embodiments. A colorimetric characterization method of an end-to-end color imaging system based on hyperspectral calibration includes the following steps:
[0068] S1, collecting spectral radiance of a target scene or obtaining multi-channel spectral cube data containing spatial resolution and band information from an existing data set;
[0069] S2, determining the three-channel spectral sensitivity curve and the calibration coefficient of the imaging device; as the weight basis for subsequent calculation;
[0070] S3, inputting the spectral cube data and the spectral sensitivity curve into a mapping module. The mapping unit can be a single-layer linear MLP with fixed weights, one-time matrix multiplication, a lookup table, sparse convolution, or other equivalent integral implementation, to perform a weighted integral operation to obtain a pixel-level RGB image;
[0071] S4, performing the same integral operation on the spectral cube data and the international lighting commission color matching function to obtain a pixel-level XYZ three-stimulus value image;
[0072] S5, performing amplitude normalization processing, uniform size, and dividing the training set, validation set, and test set according to a preset ratio on the RGB image and the XYZ image, and constructing a data loader;
[0073] S6, constructing an end-to-end color mapping model and completing training. The model can use a full convolutional network, a generative adversarial network, or other deep frameworks with whole-image mapping capability. In the preferred embodiment, an adversarial structure containing an encoder-decoder generator and a local discriminator is used. Numerical error constraints and perceptual consistency constraints are considered simultaneously during the training process until the model parameters converge;
[0074] S7, loading the converged model weight to an inference environment, inputting a real-time RGB image, and outputting a corresponding three-stimulus value image in a single inference. The single-frame delay and throughput rate are simultaneously determined to evaluate the running performance;
[0075] S8, calculating the CIE DE2000 color difference, mean absolute error, and mean square error of the output three-stimulus value image, and simultaneously evaluating the visual naturalness using recognized perceptual indicators such as LPIPS, SSIM, and FID to comprehensively verify the technical effects of the method in terms of numerical accuracy and subjective consistency;
[0076] Further, S1 specifically includes:
[0077] S11, according to the application requirement of the target imaging system, preset the waveband sampling mode: 31 visible light center wavebands with equal interval of 10 nm can be selected, and it can also be extended to the near infrared segment; At the same time, the spatial resolution of single frame imaging is determined. The waveband and resolution configuration is suitable for both self-hyperspectral imaging and screening public data sets;
[0078] S12, the spectral cube can be obtained in two ways:
[0079] 1) Use public data sets: retrieve existing spectral cube data that meets the configuration of step S11 from ICVL, CAVE, Harvard and other databases, and record the sampling interval, waveband order and radiance unit;
[0080] 2) On-site self-collection: use a hyperspectral camera to scan the target scene; In order to eliminate environmental stray light, the distance is kept perpendicular to the lens entrance pupil and the object plane, and the working distance is not less than three times the focal length;
[0081] S13, the spectral cube data obtained in step S12 is interpolated in the wavelength dimension to the same waveband center as the subsequent integral weight (spectral sensitivity curve or CMF); Then, geometric correction is performed on distortion and parallax, and the effective field of view is retained, and finally the spectral cube file meeting the S11 specification is output, and the file format can adopt.npy,.mat or ENVI standard header file;
[0082] Further, S2 specifically includes:
[0083] S21, for the R, G, B three channels of the imaging device to be calibrated, the wavelength-response curve can be obtained in two ways:
[0084] 1) Document extraction: directly call the channel spectral sensitivity curve given in the manufacturer's data manual or third-party evaluation report;
[0085] 2) Actual measurement and calibration: under darkroom conditions, use a monochromator with an integrating sphere to expose the camera wave by wave, scan 400nm to 700nm at 5nm or 10nm intervals, record the digital count of each channel and deduct the dark current to obtain the original response curve;
[0086] S22, the three-channel curves obtained in step S21 are uniformly processed, if the original sampling interval is inconsistent with the waveband center of the hyperspectral cube, one-dimensional linear or cubic spline interpolation is used to resample the curve to an array consistent with the waveband set in step S1;
[0087] S23, for simulating camera white balance and analog-to-digital conversion gain, the channel gain (scaling coefficient) is calibrated, which can be done in the following way:
[0088] 1) Let the camera take a picture of an 18% gray card or an integrating sphere light source, record the average digital counts of RGB channels;
[0089] 2) Set the target gray card brightness value (such as Y = 0.18), and obtain the gain coefficient that makes the three-channel output of the gray card equal in brightness by linear least squares or direct proportion;
[0090] 3) If the manufacturer's white balance coefficient is used, the RGB gain corresponding to the daylight-D65 item in the AWB table can be directly read;
[0091] S24, construct the sensitivity matrix, multiply the three-channel spectral sensitivity column vector of step S22 by the gain coefficient obtained in step S23, and splice into the sensitivity matrix Sgain, as shown in formula (1)
[0092]
[0093] Where, SR, SG, SB are the channel sensitivities sampled at 31 wavelengths λ b , k R , k G , k B are the calibration coefficients in step S23, the number of rows is 31 corresponding to the number of wavebands, and the number of columns is 3 corresponding to the RGB three channels;
[0094] This matrix will be involved in spectral integration as a fixed weight in step S3 and will not be updated; At the same time, the gain value and the resampling wavelength are stored in the metadata file for subsequent archiving and tracing;
[0095] Further, S3 specifically includes:
[0096] S31, solidify the three-channel sensitivity-gain matrix Sgain generated in S24 as the weight of the mapping unit, and the mapping unit can be optional one-time matrix multiplication, fixed weight single-layer MLP or other equivalent implementation;
[0097] S32, perform weighted integration on the spectral cube data, and after all pixels are calculated, reshape it into a three-channel tensor to obtain pixel-level RGB data, as shown in formula (2). For example, the response of pixel (x, y) can be expressed as discrete weighted integration:
[0098]
[0099] S33, if the spectral radiance L λSince the data is in absolute power units, the corresponding RGB data naturally has physical dimensions and can be directly used for subsequent network training and inference without the need for 0-255 quantization or proportional compression. If only relative radiance is obtained in the simulation environment and it needs to be connected to the image pipeline, quantization compression can be performed in the subsequent S5 step by multiplying the maximum brightness of the entire image or batch by a uniform fixed coefficient; this compression is optional and is not a necessary limitation of this method.
[0100] Furthermore, S4 specifically includes:
[0101] S41 reads the CIE 1931 color matching functions, constructs a band-channel matrix, and solidifies it into mapping unit weights. Similar to S3, this unit can use a one-time matrix multiplication, a fixed-weight single-layer MLP, or other equivalent implementation methods. The weights remain unchanged during operation.
[0102] S42, performing weighted integration on the spectral cube data, and reshaping it into a three-channel tensor after all pixels are calculated to obtain pixel-level XYZ data;
[0103] S43. Because the CIE matching function already implicitly includes spectral weighting, step S4 does not introduce an additional gain coefficient; its output is in the same physical dimension as the RGB data from step S3. If the radiance is only a relative value and the experimental process needs to be consistent with the 0–255 range RGB data quantization chain, the XYZ data can be uniformly scaled in the optional solution in the subsequent step S5; this scaling is an engineering measure to achieve compatibility and is not a necessary limitation of the present invention.
[0104] Furthermore, S5 specifically includes:
[0105] S51. (Optional) If the spectral radiance used is a relative value or needs to be connected to an image pipeline, the RGB data of step S3 and the XYZ data of step S4 can be amplitude compressed using the same scaling factor, with the maximum brightness value in the data set corresponding to 255 (G channel) and 100 (Y channel), and other pixels are scaled using the same factor.
[0106] S52. Divide all RGB-XYZ paired samples according to a preset ratio (e.g., 70% / 20% / 10%) by random sampling, scene grouping, or k-fold cross validation;
[0107] S53. Based on the partitioning results, a data loader is established to support batch reading, random shuffling, GPU asynchronous transmission, and necessary data enhancement. At the same time, information such as the band center, channel gain coefficient, and whether S51 compression is performed is recorded in the metadata for the end-to-end model training call in step S6.
[0108] Furthermore, S6 specifically includes:
[0109] S61、In the scenario of mapping the entire RGB data to XYZ data, the application allows the use of any deep network framework capable of processing the entire image at once, including but not limited to:
[0110] 1) Pure fully convolutional network (FCN), encoder-decoder network, U-Net and its variants;
[0111] 2) Generative adversarial network (GAN) series: Pix2Pix, CycleGAN, StyleGAN-based structure;
[0112] 3) Diffusion model, hybrid attention-convolutional network, etc.
[0113] The model input is three-channel RGB data, and the output is three-channel XYZ data. The network weights are updated through uniform iteration by back propagation.
[0114] S62、The model can use a single-stage joint loss or a multi-stage strategy, including but not limited to:
[0115] 1) Pixel-level L1 / L2 loss, used to ensure the absolute error of XYZ data,
[0116] 2) Perceptual loss, feature matching loss or adversarial loss, used to improve visual naturalness;
[0117] It can be optimized together at once, or it can be pre-trained for several rounds using numerical error, and then introduce perceptual or adversarial loss for refinement until the model converges on the validation set;
[0118] The optimizer can be Adam, RMSProp or other adaptive method. The learning rate, batch size and loss weight can be flexibly set according to the data size and hardware conditions. After training, all network parameters are frozen and the weight file is saved for subsequent inference and deployment;
[0119] A possible example is as follows:
[0120] As shown in Figure 2 , a conditional generative adversarial network (CGAN) is used as a representative implementation. The generator part uses an encoder-decoder U-Net as the backbone: input 3-channel RGB data, extract multi-scale features through three times of downsampling, and concatenate 6 residual blocks at the bottleneck position to enhance the expression ability; then, the features are gradually upsampled and the spatial details are restored through skip connection, and finally the normalized XYZ data is output through 3x3 convolution and tanh activation. This structure not only guarantees the global context, but also preserves the edge and texture information, suitable for simultaneous optimization of numerical accuracy and visual naturalness.
[0121] The discriminator adopts the PatchGAN local structure, and the receptive field is about 70*70 pixels. All convolution layers apply spectral normalization to suppress gradient explosion, and periodically add R1 gradient penalty to the real branch to improve the stability of adversarial training. The loss function consists of three parts: the pixel-level L1 loss ensures the numerical accuracy of the tristimulus value, as shown in equation (3), where N = H*W is the number of image pixels, r represents the input RGB, x corresponds to the true value XYZ, and G represents the generator;
[0122]
[0123] The adversarial loss adopts the Hinge-GAN form to enhance the visual authenticity; the adversarial term at the generator end takes the negative mean of the discriminator output, as shown in equation (4);
[0124] L adv = -E r [D(r,G(r))](4)
[0125] To reduce the distance between the generated result and the true value in the feature space of the discriminator, a multi-layer feature matching loss is introduced, as shown in equation (5). The feature matching loss further constrains the local consistency of the generated result by comparing the intermediate feature distribution of the discriminator, where D(l) is the lth layer feature map of the discriminator, n l is the number of its elements;
[0126]
[0127] where the Hinge-GAN form of the discriminator target is used in the adversarial stage, as shown in equation (6);
[0128] L D = E (R,X) [max(0,1-D(r,x))]+E R [max(0,1-D(r,G(r)))](6)
[0129] The training adopts a phased strategy. In the first phase, only the generator is enabled, and the pixel-level L1 loss is used to regress for 20 rounds, so that the output XYZ quickly fits the numerical scale with the true value. In the second phase, the discriminator is enabled, and the Hinge-GAN adversarial loss and feature matching loss are introduced; the weights of the two new losses are linearly raised from zero to 0.1 in the next 30 rounds, and then remain constant to continue iteration until 400 rounds. The total objective function of the generator in the adversarial stage is shown in equation (7);
[0130] L G = L L1 +ω adv (e)L adv +ω fm (e)L fm (7)
[0131] wherein the epoch-dependent weight is set by a linear ramping strategy as shown in equation (8);
[0132]
[0133] The optimizer is selected as Adam, and the initial learning rates of the generator and the discriminator are set to 1x10-4 and 2x10-4, respectively, and are decayed by 0.3 times every 100 rounds. During the training process, an image of “input RGB | true value XYZ | predicted XYZ” is saved every iteration for manual inspection, and the latest weights are written at the end of each round;
[0134] CIE DE2000, LPIPS, and SSIM are monitored in real time on the validation set. When the average CIE DE2000 is less than or equal to 2, LPIPS is less than or equal to 0.1, and SSIM is greater than or equal to 0.95, the model is considered to have converged. After meeting the conditions, all network parameters are frozen and the weight file is exported for online inference in step S7. If the indicators do not meet the standards, the training of the next stage can be appropriately extended or the weights of the adversarial and feature matching can be adjusted until the standards are met;
[0135] Further, S7 specifically includes:
[0136] S71, after completing the convergence determination of S6, the finally frozen end-to-end mapping model weight is imported into the inference environment, and the present application does not limit the specific running environment. The network structure description file, normalization parameters, and spectral metadata are read simultaneously during the loading stage to ensure that the input tensor is consistent with the scale and arrangement during training;
[0137] S72, the three-channel RGB image output by the camera is converted into a network input tensor according to the same normalization logic of S5, and the entire XYZ tristimulus value image is obtained directly within a single forward propagation. After inference, the end-to-end delay is recorded by a high-resolution timer; if the hardware supports pipeline parallelism, the average frame rate or the number of images per second in a stable output state also needs to be counted;
[0138] Taking the above example as an example, under the inference conditions of a resolution input of 512*512, an NVIDIA RTX 3080ti, and a CPU Intel-i5-14600kf, the single-frame delay can be controlled at about 16ms, and the throughput rate is about 60fps;
[0139] Further, S8 specifically includes:
[0140] S81, the tristimulus value image after inference is first compared with the true value point by point at the pixel level. The present application uses CIEDE2000 and CIE Lab as the main color error indicators, and also provides the mean absolute error (MAE) and the root mean square error (RMSE) as auxiliary parameters;
[0141] ΔE2000≤2 as the line of no perceptible difference for the human eye; if a more stringent standard is required, the threshold can be lowered to 1.5 or the 95% pixel quantile is used instead of the average value for judgment;
[0142] S82, in order to test whether the network is in numerical compliance while avoiding local false colors or texture loss, the application synchronously calculates three perception indicators: LPIPS, SSIM and FID. LPIPS measures the deep representation difference by means of pre-trained convolutional feature distance; SSIM evaluates local consistency in terms of three components of brightness, contrast and structure; FID uses Inception depth statistics to measure the distance between the whole distribution and the true value distribution. The system default threshold is LPIPS≤0.10, SSIM≥0.95 and FID≤15. When all three indicators meet the threshold and the numerical error passes the S81 detection, it is determined that the output has both chroma accuracy and visual naturalness;
[0143] The application is based on the physical mechanism of hyperspectral radiance-three stimulus value mapping in color imaging system. Firstly, it is determined that the emission spectrum of a scene can be expressed in the subspace spanned by the spectral sensitivity of the three channels of the camera. Then an end-to-end chroma characterization model based on hyperspectral calibration is proposed. Step one, use the spectral cube with uniform band sampling and the measured three-channel spectral sensitivity curve to perform fixed weight linear integration to obtain pixel-level RGB representation; Step two, integrate the same spectral cube with the color matching function to generate pixel-level XYZ true value; Step three, input the RGB-XYZ paired samples into the deep network, first perform numerical regression correction, then introduce the perceptual consistency constraint (this example uses a conditional generative adversarial network containing an encoder-decoder generator and a local discriminator, which can also be equivalent to a convolution-attention hybrid network, a visual transformer or a diffusion model), finally output the whole XYZ image, realizing end-to-end chroma characterization of the color imaging system. To verify the feasibility of the method, select multiple scene hyperspectral data and camera spectral sensitivity curves to construct a dataset, divide the samples into training set and test set, fit the network weight through the training set, and then evaluate the numerical color difference (ΔE (2000) ) and perception indicators (LPIPS, SSIM, FID) with the test set; the experimental results show that the application is significantly better than the traditional pixel-by-pixel regression scheme in terms of numerical accuracy and visual naturalness, confirming the effectiveness and high precision of the method.
[0144] The principle of the method of the application is as follows:
[0145] An end-to-end colorimetric characterization method for color imaging systems, the core idea of which is "imaging link physical model calibration + end-to-end mapping". Based on a hyperspectral dataset, combined with the pre-measured camera three-channel spectral sensitivity curve and CIE 1931 color matching function, and then through a fixed weight mapping module to generate a matching pixel resolution of RGB and XYZ true value image pair, so as to realize the construction of simulation calibration dataset covering wide color gamut and complex texture.
[0146] Based on the above dataset, an end-to-end colorimetric characterization deep learning network is designed. The method can be realized by using generative adversarial network and U-shaped network, and an example of a method for realizing end-to-end colorimetric characterization using generative adversarial network is as follows:
[0147] First, build a system calibration module, as shown in Figure 1 The multi-channel spectral data obtained by the hyperspectral camera is input into the fixed weight mapping module through the camera three-channel spectral sensitivity curve (taking Canon 600D as an example) and CIE 1931 color matching function (CMF). The module completes spectral integration operation with a preset matrix inside, thereby synchronously outputting RGB image and corresponding XYZ true value image at the same resolution, and directly obtaining the corresponding RGB and XYZ data pair.
[0148] The network structure is as follows: the generator uses a U-Net encoding-decoding framework stacked with six residual blocks to perform full convolution mapping on the input RGB image and directly output pixel-level XYZ; the discriminator uses a 70x70 Patch structure to jointly judge the authenticity of the (RGB, XYZ) pair in the local receptive field, and applies R1 gradient penalty to the real branch to stabilize the training.
[0149] As shown in Figure 2As shown, the system enters the "two-stage adversarial mapping module". In the first link of this module, the RGB image is input into the Generator using the U-Net encoding-decoding skeleton and superimposed with six residual blocks to obtain the preliminary XYZ prediction; at the same time, the prediction result is cropped to generate PatchMap for subsequent local authenticity determination. In the second link, the tensor obtained by splicing the real XYZ and RGB, and the tensor obtained by splicing the RGB and the predicted XYZ, are input into the Discriminator of the PatchGAN structure, and the discriminator outputs the authenticity probability of each local patch and imposes R1 gradient penalty on the real branch to stabilize the training. The training adopts a two-stage strategy: in the first stage, only the numerical alignment of the Generator is performed with the L1 loss; from a certain epoch, the second stage starts, and the adversarial loss and feature matching loss based on Hinge-GAN are introduced, and the weights of the two items are linearly increased to the set upper limit in the next 30 epochs, so as to further approximate the real color gamut distribution and texture details, and only the converged Generator weight is retained after the training is completed.
[0150] As shown in Figure 3 Due to the high parallelism of the convolutional network, a single image can obtain the entire XYZ through multiple layers of convolution at one time, realizing end-to-end overall image prediction, saving the cycle overhead of pixel-level point-by-point regression, and significantly improving the inference speed compared with traditional models, which embodies the significant advantage of the present application in real-time or near real-time application scenarios. The real-time RGB frame output by the on-site camera can be pixel-level mapped in milliseconds through the network, realizing high-precision, real-time and cross-device consistent color output.
[0151] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can make equivalent replacement or change according to the technical solution and concept of the present application within the technical range disclosed by the present application, which should be covered within the protection scope of the present application.
Claims
1. A method for color characterization of an end-to-end color imaging system based on hyperspectral calibration, characterized in that: The steps include: S1. Collect the spectral radiance of the target scene, or obtain a spectral cube containing spatial resolution and band information from a public dataset; S2, measuring the R, G, and B three-channel spectral sensitivity curves of the imaging device to be calibrated and obtaining the calibration coefficients; S3, inputting the spectral cube and the sensitivity curve into a mapping module, performing a fixed weight integral operation, and obtaining pixel-level RGB data; S4, performing the same type of integration operation on the spectral cube and the CIE 1931 color matching function to obtain pixel-level XYZ tristimulus value data; S5. Normalize the amplitude and size of the RGB data and the XYZ data, divide them into training set, validation set and test set according to the preset ratio, and then establish a data loader; S6. Build an end-to-end color mapping model and complete training. During the training process, both numerical error constraints and perceptual consistency constraints are introduced until the model converges. S7. Load the converged model weights into the inference environment, input real-time RGB data, output the corresponding XYZ data for a single inference, and measure the single-frame latency and throughput. S8. Calculate the CIE DE2000, mean absolute error, and mean square error between the inference results and the true value, and use the LPIPS and SSIM indicators to evaluate the visual naturalness.
2. The method for color characterization of an end-to-end color imaging system based on hyperspectral calibration according to claim 1, wherein: The step S1 comprises: S11. Set the band sampling method and spatial resolution according to application requirements; S12. Obtain spectral cubes through field acquisition or public datasets; S13. Interpolate the spectral cube in the wavelength dimension to the center of the band consistent with the subsequent integration weight, and complete the geometric correction and effective field of view clipping.
3. The method for color characterization of an end-to-end color imaging system based on hyperspectral calibration according to claim 1, wherein: The step S2 comprises: S21, obtaining three-channel original wavelength-response curves; S22, resampling the curve to the center of the band set in step S1 using linear or cubic spline interpolation; S23, calibrating channel gain coefficients to simulate camera white balance and analog-to-digital conversion gain; S24. Multiply the resampling curve by the corresponding gain coefficient and concatenate the result by columns to obtain a sensitivity-gain matrix Sgain.
4. The method for color characterization of an end-to-end color imaging system based on hyperspectral calibration according to claim 1, wherein: The step S3 comprises: S31. Solidify the Sgain into a mapping unit weight, where the mapping unit is a one-time matrix multiplication, a fixed-weight single-layer linear MLP, a lookup table, or other equivalent methods; S32, performing weighted integration on the spectral cube and reshaping it into a three-channel tensor to obtain pixel-level RGB data; S33. When the spectral radiance is a relative value and needs to be connected to the image pipeline, the RGB data is compressed in a uniform ratio in step S5.
5. The method for color characterization of an end-to-end color imaging system based on hyperspectral calibration according to claim 1, wherein: The step S4 comprises: S41, read and resample the CIE 1931 color matching function and construct a color matching matrix; S42, performing weighted integration on the spectral cube and reshaping it into a three-channel tensor to obtain pixel-level XYZ data; S43. The XYZ data and the RGB data obtained in step S3 are in the same physical dimension. If quantization compression is required, the two use the same proportional coefficient.
6. The method for color characterization of an end-to-end color imaging system based on hyperspectral calibration according to claim 1, wherein: The step S6 comprises: S61. End-to-end color mapping models are deep network frameworks that can process the entire image at once, including fully convolutional networks, generative adversarial networks, or diffusion models. S62. Model training adopts a phased strategy: first pre-train with pixel-level L1 or L2 loss, then introduce adversarial loss or feature matching loss for refinement until the validation set indicators converge.
7. The method for color characterization of an end-to-end color imaging system based on hyperspectral calibration according to claim 1, wherein: The step S7 comprises: S71. Load the model weights and normalization parameters in the inference environment to ensure that the input tensors are consistent with those during training. S72. Convert the real-time RGB image into a network input tensor and complete a forward inference, recording the end-to-end latency and average throughput.
8. The method for color characterization of an end-to-end color imaging system based on hyperspectral calibration according to claim 1, wherein: The step S8 comprises: S81. Calculate the pixel-level difference between the inference output XYZ and the true value XYZ to obtain CIEDE2000, MAE, and RMSE. S82. Calculate the LPIPS and SSIM perception indicators, and use ΔE2000≤2, LPIPS≤0.10, and SSIM≥0.95 as the thresholds for achieving the combined standard of color accuracy and visual naturalness.
9. The method for color characterization of an end-to-end color imaging system based on hyperspectral calibration according to claim 1, wherein: The end-to-end color mapping model preferably includes a conditional generative adversarial network of an encoder-decoder generator and a local discriminator, wherein the generator adopts a U-Net backbone and superimposes a residual block at the bottleneck, and the discriminator adopts a Patch structure and adds an R1 gradient penalty to the real branch to improve training stability.
Citation Information
Cited By
Real-time color measurement system and method for spectral level calibration
CN121409411A
A real-time color measurement system and method with spectral level calibration
CN121409411B