Video image color correction method and system based on artificial intelligence

By introducing artificial intelligence technology into the video image color correction method, using Retinex-Net and VGG models for lighting and style correction, combining the generation adversarial network and optical flow vector field for color change prediction, the problem of color inconsistency in dynamic videos is solved, and the efficient color consistency and visual effect improvement of videos is achieved.

CN120182155AInactive Publication Date: 2025-06-20WUHAN HENGJI INTELLIGENT CLOUD NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510292766.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-20
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When processing dynamic videos, existing color correction methods lack effective considerations for video timing consistency, resulting in color inconsistency between consecutive frames, resulting in unsmooth visual effects, jumps and flickering.

Method used

Using the video image color correction method based on artificial intelligence, reflectivity and lighting data are extracted through the Retinex-Net model, and the lighting adjustment coefficient is calculated for pixel correction and color temperature correction; combining the target style image frame data, the style features are extracted using the VGG deep learning model, and the style mapping is used for the generation adversarial network GAN, the optical flow vector field is calculated for color change prediction, and the timing smoothing factor is calculated to ensure color consistency between frames.

Benefits of technology

It significantly improves the color consistency and visual effect of the video, avoids color jumps and flickering between frames, and ensures the consistency of the video style and the color accuracy of local areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182155A_ABST
    Figure CN120182155A_ABST
Patent Text Reader

Abstract

The invention discloses a video image color correction method and system based on artificial intelligence, and relates to the technical field of color correction, and the method comprises the steps: collecting video data, carrying out the image size normalization and denoising processing, and obtaining the preprocessed video frame sequence data; and extracting reflectivity data and illumination data by using a Retinex-Net deep learning model, calculating an illumination adjustment coefficient, carrying out pixel correction, and carrying out color temperature correction on the pixel data of the video frame subjected to pixel correction according to the illumination adjustment coefficient. According to the method, the preliminary color mapping matrix is generated through the color histogram matching method, the color distribution of the image can be optimized, the image is enabled to better conform to the color features of the target reference frame, the color temperature adjustment is performed on the image according to the illumination adjustment coefficient and the illumination threshold, the cold and warm tones of the video are enabled to be accurately adjusted, and the image quality is improved. Through successively carrying out pixel correction, color correction and local area correction, a stable color adjustment frame is provided, so that the video styles are consistent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of color correction, and particularly to a method and system for video image color correction based on artificial intelligence. Background Art

[0002] Video image color correction technology has experienced significant development. Early image color correction technologies mostly relied on traditional image processing methods, such as histogram equalization, white balance correction, and color mapping. These methods achieved good results in static image processing, but when dealing with dynamic videos, especially under different lighting conditions and complex scenes, problems such as inconsistent colors and local distortion often occurred. With the rise of deep learning technology, especially the application of convolutional neural network CNN in image processing, the accuracy and adaptability of color correction have been significantly improved;

[0003] Although deep learning methods have made significant progress in the field of video image color correction, existing color correction methods usually focus on the processing of single-frame images and lack effective consideration of the temporal consistency of videos. Due to the complex lighting and color changes between video frames, traditional methods often fail to fully ensure color consistency between consecutive frames, easily leading to uneven visual effects of videos, or even jump and flicker phenomena. In addition, most existing deep learning models rely on the style mapping and correction of the overall image. When dealing with dynamic videos, the color performance of the target area often differs from that of the background and other parts. Therefore, global adjustment may cause color distortion in local areas. Summary of the Invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides a method for video image color correction based on artificial intelligence to solve the problems that existing color correction methods usually focus on the processing of single-frame images and lack effective consideration of the temporal consistency of videos. Due to the complex lighting and color changes between video frames, traditional methods often fail to fully ensure color consistency between consecutive frames, easily leading to uneven visual effects of videos, or even jump and flicker phenomena. In addition, most existing deep learning models rely on the style mapping and correction of the overall image. When dealing with dynamic videos, the color performance of the target area often differs from that of the background and other parts. Therefore, global adjustment may cause color distortion in local areas.

[0006] To solve the above technical problems, the present invention provides the following technical solutions:

[0007] In a first aspect, the present invention provides a method for video image color correction based on artificial intelligence, which includes:

[0008] Collect video data, perform image size normalization and denoising processing to obtain preprocessed video frame sequence data;

[0009] Use the Retinex-Net deep learning model to extract reflectance data and illumination data, calculate the illumination adjustment coefficient, perform pixel correction, and perform color temperature correction on the pixel data of the pixel-corrected video frames according to the illumination adjustment coefficient;

[0010] Extract the pixel data of the target area from the video frames to obtain the target area histogram, perform target area marking, calculate the color adjustment error with the minimum error to obtain the local color transformation matrix, and perform local optimization of the target area;

[0011] Combine the target style image frame data, use the VGG deep learning model to extract style features respectively, perform style mapping using the generative adversarial network GAN, calculate the optical flow vector field of the consecutive frames after mapping, perform color change prediction, and calculate the temporal smoothing factor to obtain the final style adjustment frame data;

[0012] Perform image sharpening processing and contrast enhancement processing, and perform color consistency adjustment, perform video data compression and integrity verification, and store the data through the cloud.

[0013] As a preferred solution of the video image color correction method based on artificial intelligence according to the present invention, wherein: the use of the Retinex-Net deep learning model to extract reflectance data and illumination data, calculate the illumination adjustment coefficient, perform pixel correction, and perform color temperature correction on the pixel data of the pixel-corrected video frames according to the illumination adjustment coefficient includes:

[0014] Based on the preprocessed video frame sequence and the corresponding sequence of pixel coordinates, use the Retinex-Net deep learning model to extract reflectance data and illumination data from the pixel data of the video frame sequence, calculate the mean and standard deviation of the illumination data of the entire frame, compare with the standard reference illumination, and calculate the illumination adjustment coefficient;

[0015] Perform color space conversion according to the pixel data of the video frame sequence to obtain a color histogram, and use the color histogram matching method to generate a preliminary color mapping matrix;

[0016] Set the optimization goal to minimize the color deviation, calculate the Euclidean distance between the color histogram of the current frame and the reference color histogram, and calculate the optimal color transformation matrix according to the color adjacency and Euclidean distance of all pixel points of the video frame;

[0017] Use the gradient descent method to update the color transformation matrix, stop the iteration when the iteration loss calculated during the iteration no longer decreases significantly, obtain the optimal color transformation matrix, and apply it to all pixel points of the video frame for pixel correction;

[0018] Take the ratio of the standard reference light mean value to the illumination mean value as the illumination threshold, perform color temperature correction on the pixel data of the video frame corrected by the illumination adjustment coefficient, and update all pixel points of the video frame after color temperature correction.

[0019] As a preferred solution of the artificial intelligence-based video image color correction method described in the present invention, wherein: extracting the pixel data of the target area from the video frame, obtaining the target area histogram, performing target area marking, calculating the color adjustment error with the minimum error to obtain the local color transformation matrix, and performing local optimization of the target area, including:

[0020] Use the target detection model YOLOv5 to detect the target object in the video frame and extract the pixel data of the target area;

[0021] Perform color space conversion according to the pixel data of the target area, obtain the target area histogram, and calculate the Euclidean distance of the pixel deviation between the target area and the standard color;

[0022] Take the sum of the historical mean value and the standard deviation of the Euclidean distance value as the distance threshold. If the calculated Euclidean distance is greater than or equal to the distance threshold, the target area is marked as having a large deviation;

[0023] For the marked target area, traverse the pixel points of the target area to calculate the color adjustment error with the minimum error;

[0024] Solve the optimal color transformation matrix by the least square method, calculate the optimal local color transformation matrix for correcting the marked target area, and obtain the pixel data of the video frame after local target area optimization.

[0025] As a preferred solution of the artificial intelligence-based video image color correction method described in the present invention, wherein: combining the target style image frame data, using the VGG deep learning model to extract style features respectively, performing style mapping by using the generative adversarial network GAN, calculating the optical flow vector field of the consecutive frames after mapping, performing color change prediction, and calculating the temporal smoothing factor to obtain the final style adjustment frame data, including:

[0026] Based on the pixel data of the video frame after local target area optimization and the target style image frame data, output the style features of the video frame after local optimization and the style features of the target style image frame through the pre-trained VGG deep learning model, and calculate the style matching error according to the Euclidean distance of the two feature means;

[0027] Use the generator of the generative adversarial network GAN to generate stylized features according to the two features, and calculate the Euclidean distance between the stylized features and the features of the optimized video frame as the content loss value;

[0028] The discriminator of the generative adversarial network (GAN) determines the adversarial loss between the stylized features and the features of the target style image frame. The generator is optimized based on minimizing the weighted sum of the style loss, content loss, and adversarial loss, while the discriminator is optimized to maximize the adversarial loss. When the calculated loss of the optimization result no longer changes significantly in consecutive optimizations, the optimization is stopped, and the optimized generative adversarial network (GAN) is completed. The pixel data of the video frame optimized for the local target area is subjected to style mapping to obtain the stylized pixel data based on the target style image frame;

[0029] According to the sequence of the stylized pixel data, the pixels of two consecutive frames are converted into grayscale images, and the Farneback method is used to calculate the optical flow vector field. A pixel grid is constructed to calculate the new positions of the pixels in the previous frame in the current frame based on the optical flow vector field, and the pixel colors of the previous frame are resampled using bilinear interpolation to obtain the color information of the previous frame after optical flow mapping as the optical flow mapping frame data;

[0030] Using the pre-trained spatio-temporal Transformer (STT) model, color change prediction is performed based on the sequence of the stylized pixel data and the optical flow mapping frame data of every two frames, and the temporal smoothing factor is calculated;

[0031] Based on the temporal smoothing factor, per-pixel scalar calculation is performed to obtain the final style adjustment frame data.

[0032] As a preferred solution of the video image color correction method based on artificial intelligence according to the present invention, wherein: the image sharpening process and contrast enhancement process are performed, and color consistency adjustment is performed, including:

[0033] For the final style adjustment frame sequence data, the Laplace filter is used for image sharpening, and the adaptive histogram equalization method (CLAHE) is used for contrast enhancement;

[0034] The average color histogram of the processed frame sequence data is calculated, and the ratio to the color temperature reference value of the current frame is calculated as the image frame data after color consistency adjustment.

[0035] As a preferred solution of the video image color correction method based on artificial intelligence according to the present invention, wherein: the video data is collected, and image size normalization and denoising processing are performed to obtain the preprocessed video frame sequence data, including:

[0036] Frame data is extracted from consecutive frames of the video data using the equal-interval sampling method;

[0037] The average resolution of all frames is calculated, and bicubic interpolation is used for resolution scaling to perform image size normalization;

[0038] Frame data denoising is performed using a Gaussian filter and Gamma correction is carried out to obtain a preprocessed video frame sequence.

[0039] As a preferred solution of the video image color correction method based on artificial intelligence according to the present invention, wherein: performing video data compression and integrity verification and storing data in the cloud means that for the processed image frame data, the H.265 compression algorithm is used for video compression, the key frame interval GOP is used for inter-frame compression, the RAID 1 / 5 storage structure is adopted for storing image frame data, and the hash check SHA-256 is used for integrity verification, and it is transmitted to the cloud through wireless transmission technology for data backup storage.

[0040] In a second aspect, the present invention provides a system for a video image color correction method based on artificial intelligence, including

[0041] A data acquisition and processing module that acquires video data from an original video source and performs preprocessing, including image size normalization, denoising processing, and frame sampling;

[0042] A lighting adjustment module that uses the Retinex-Net deep learning model to extract reflectance and lighting data and corrects the color temperature of video frames according to the lighting adjustment coefficient;

[0043] A local optimization module that uses an object detection algorithm to extract the target area in the video, calculates the color histogram of the target area, and performs local color optimization;

[0044] A style mapping optimization module that combines the target style image frame data and performs style mapping using a deep learning model;

[0045] A visual processing module that performs image sharpening processing and contrast enhancement processing and performs color consistency adjustment according to the calculated color temperature reference value;

[0046] A compression and storage module that compresses video data, performs data integrity verification and secure storage.

[0047] In a third aspect, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and wherein: when the computer program is executed by the processor, any step of the video image color correction method based on artificial intelligence as described in the first aspect of the present invention is implemented.

[0048] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and wherein: when the computer program is executed by the processor, any step of the video image color correction method based on artificial intelligence as described in the first aspect of the present invention is implemented.

[0049] The beneficial effects of the present invention are as follows: By generating a preliminary color mapping matrix through the color histogram matching method, the color distribution of the image can be optimized to make it more conform to the color characteristics of the target reference frame. The color temperature of the image is adjusted according to the illumination adjustment coefficient and the illumination threshold, so that the warm and cold tones of the video can be accurately adjusted. By performing pixel correction, color correction, and local area correction successively, a stable color adjustment framework is provided, making the video style consistent. The local optimization ensures that the key target areas will not be negatively affected by the global adjustment, so that the entire video has been significantly improved in terms of illumination change, color consistency, and visual detail retention. By extracting style features through the VGG deep learning model and using the generative adversarial network GAN for style mapping, it is ensured that both the content information of the video can be retained and the target style can be matched during the style transfer process. By combining the optical flow vector calculation and the temporal consistency optimization, it can effectively ensure the smooth style change between adjacent frames and reduce the color jump between frames, making the stylized video have a more natural transition effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0051] Figure 1 It is a schematic flow chart of the method for color correction of video images based on artificial intelligence in Embodiment 1.

[0052] Figure 2 It is a schematic structural diagram of the system for color correction of video images based on artificial intelligence in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings of the specification.

[0054] Many specific details are set forth in the following description in order to provide a thorough understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0055] Secondly, the "one embodiment" or "embodiment" referred to herein means a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it an individual or selectively mutually exclusive embodiment with other embodiments.

[0056] Embodiment 1, referring to Figure 1 and Figure 2 , is the first embodiment of the present invention. This embodiment provides an artificial intelligence-based video image color correction method, including the following steps:

[0057] S1. Collect video data, perform image size normalization and denoising processing to obtain preprocessed video frame sequence data;

[0058] Preferably, collecting video data, performing image size normalization and denoising processing to obtain preprocessed video frame sequence data includes:

[0059] Extract frame data from consecutive frames of video data using the equal-interval sampling method;

[0060] Calculate the average resolution of all frames, perform resolution scaling using bicubic interpolation, and perform image size normalization;

[0061] Use a Gaussian filter to denoise the frame data and perform Gamma correction to obtain a preprocessed video frame sequence.

[0062] Extracting consecutive frames of video data by the equal-interval sampling method can effectively reduce redundant data and improve calculation efficiency. Calculating the average resolution of all frames and using bicubic interpolation for resolution scaling can unify video frames of different resolutions into a standard resolution, ensuring consistency in subsequent processing. By using a Gaussian filter for denoising processing, high-frequency noise in the video frames can be effectively removed, maintaining the smoothness of the image and improving the image quality.

[0063] S2. Use the Retinex-Net deep learning model to extract reflectance data and illumination data, calculate the illumination adjustment coefficient, perform pixel correction, and perform color temperature correction on the pixel data of the pixel-corrected video frames according to the illumination adjustment coefficient;

[0064] Preferably, using the Retinex-Net deep learning model to extract reflectance data and illumination data, calculate the illumination adjustment coefficient, perform pixel correction, and perform color temperature correction on the pixel data of the pixel-corrected video frames according to the illumination adjustment coefficient includes:

[0065] Based on the pre - processed video frame sequence and the corresponding pixel coordinates of the sequence, use the Retinex - Net deep learning model to extract the reflectance data and illumination data from the pixel data of the video frame sequence, calculate the mean and standard deviation of the illumination data for the entire frame, compare with the standard reference illumination, and calculate the illumination adjustment coefficient, expressed as:

[0066] ;

[0067] where represents the illumination adjustment coefficient, represents the illumination mean, represents the standard reference light mean, represents the illumination standard deviation, represents the standard reference light standard deviation;

[0068] Perform color space conversion based on the pixel data of the video frame sequence to obtain a color histogram, and use the color histogram matching method to generate a preliminary color mapping matrix;

[0069] Set the optimization goal to minimize the color deviation, calculate the Euclidean distance between the color histogram of the current frame and the reference color histogram, and calculate the optimal color transformation matrix based on the color adjacency and Euclidean distance of all pixel points of the video frame, expressed as:

[0070] ;

[0071] where represents the optimal color transformation matrix, T represents the current color transformation matrix (used to adjust the color of the current frame to be close to the color of the reference frame), represents the summation over all pixel points (x, y), represents the color vector of the current video frame at the pixel coordinate (x, y), which can be expressed as , represents the color vector of the standard reference frame at the pixel coordinate (x, y), represents the Euclidean distance between the color histograms of the current frame and the standard reference frame, and respectively represent the color histograms of the current frame and the reference frame, represents the color correction intensity factor, which can be expressed as ;

[0072] Use the gradient descent method to update the color transformation matrix. Stop the iteration when the iteration loss calculated during the iteration no longer decreases significantly, obtain the optimal color transformation matrix, and apply it to all pixel points of the video frame for pixel correction;

[0073] Taking the ratio of the standard reference light mean value to the illumination mean value as the illumination threshold, perform color temperature correction on the pixel data of the video frame corrected by pixels according to the illumination adjustment coefficient, and update all pixel points of the video frame after color temperature correction, expressed as:

[0074] ;

[0075] where represents the color temperature correction matrix, represents the illumination threshold. If the illumination adjustment coefficient is greater than or equal to the illumination threshold, then the green channel G is multiplied by S to reduce the yellow tendency, and the blue channel B is multiplied by to enhance the cool tone, and the red channel is multiplied by to reduce the red channel. If the illumination adjustment coefficient is less than the illumination threshold, then the green channel G is multiplied by S to enhance the yellow tendency, and the blue channel B is multiplied by to reduce the cool tone, and the red channel is multiplied by to enhance the red channel.

[0076] By extracting the reflectance and illumination data, the illumination and reflectance components in the image can be effectively separated, which helps to retain the original details of the video frame under various illumination conditions, while eliminating the impact of overexposed or underexposed areas on the image quality, and can ensure that subsequent processing only focuses on the reflectance data, thereby improving the clarity and detail retention of the video frame;

[0077] By calculating the illumination adjustment coefficient, the change of the image illumination can be quantified and compared with the standard reference illumination, so as to calculate the most appropriate illumination adjustment parameters. The key to this operation is to dynamically adjust according to the illumination conditions of the actual scene, so that the image presents a more natural effect under different illumination environments, avoiding the influence of too strong or too weak illumination, and ensuring the authenticity and naturalness of the image;

[0078] By generating a preliminary color mapping matrix through the color histogram matching method, the color distribution of the image can be optimized to make it more in line with the color characteristics of the target reference frame. In the calculation process of the optimal color transformation matrix, by using Euclidean distance optimization and gradient descent method, the color of the current frame can be accurately adjusted to ensure that the color of the video frame is as close as possible to the reference frame. In terms of color temperature correction, the color temperature of the image is adjusted according to the illumination adjustment coefficient and the illumination threshold, so that the cool and warm tones of the video are accurately adjusted. Specifically, when the illumination adjustment coefficient is greater than or equal to the threshold, by appropriately adjusting the green, blue and red channels, the yellow tendency is reduced and the cool tone is enhanced, so that the picture presents a more calm and clear visual effect. When the illumination adjustment coefficient is less than the threshold, the color temperature adjustment enhances the yellow and red tendencies.

[0079] S3. Extract the pixel data of the target area from the video frame, obtain the target area histogram, perform target area marking, calculate the color adjustment error with the minimum error to obtain the local color transformation matrix, and perform local optimization of the target area;

[0080] Preferably, extracting the pixel data of the target area from the video frame, obtaining the target area histogram, performing target area marking, calculating the color adjustment error with the minimum error to obtain the local color transformation matrix, and performing local optimization of the target area includes:

[0081] Use the target detection model YOLOv5 to detect the target object in the video frame and extract the pixel data of the target area;

[0082] Perform color space conversion according to the pixel data of the target area to obtain the target area histogram, and calculate the Euclidean distance of the pixel deviation between the target area and the standard color, expressed as:

[0083] ;

[0084] where represents the Euclidean distance between the target area histogram and the standard color histogram, represents the target area In the color channel, the histogram frequency of the k-level color value, represents the frequency of the color value k in the standard color histogram;

[0085] Based on the sum of the historical mean and standard deviation of the Euclidean distance value as the distance threshold, if the calculated Euclidean distance is greater than or equal to the distance threshold, the target area is marked as having a large deviation;

[0086] For the marked target area, traverse the pixel points of the target area to calculate the color adjustment error with the minimum error, expressed as:

[0087] ;

[0088] where represents the local color transformation matrix, represents the candidate color transformation matrix (a matrix used to adjust pixel colors, which will be continuously adjusted during the solution process to minimize the error), represents the color value of the image frame after color temperature correction at the pixel point , represents the color value of the target area under ideal color conditions;

[0089] Solve the optimal color transformation matrix by the least squares method, calculate the optimal local color transformation matrix for correcting the marked target area, and obtain the pixel data of the video frame after local target area optimization.

[0090] The target object detection of video frames is performed through the YOLOv5 target detection model, and the pixel data of the target area is extracted, which can ensure the accuracy of color adjustment, making the color optimization not only limited to global adjustment, but also performing fine-grained correction for key target areas. During the color space conversion and the calculation of the target area histogram, the pixel distribution of the target area is calculated and compared with the standard color histogram. Using the Euclidean distance as the measurement standard, this process can effectively identify the color deviation of the target area, enabling the system to accurately find the areas with large color deviations;

[0091] For the target areas marked with large deviations, traverse their pixel points and calculate the color adjustment error with the minimum error to ensure the accuracy of color adjustment, which can ensure that the colors of the target areas will not change abruptly, and at the same time reduce artifacts or color unevenness problems caused by color matching errors. After the local target area is optimized, the pixel data of the optimized video frame can effectively maintain the global color consistency while performing fine color adjustment for the target area to ensure the color accuracy of the target area and prevent color distortion problems due to the overall color mapping;

[0092] The combination of local target area optimization and global color adjustment ensures the color consistency of the video frame, and at the same time avoids the problem of excessive color deviation in key target areas caused by global color transformation. By performing pixel correction, color correction, and local area correction in sequence, a stable color adjustment framework is provided, making the video style consistent, while local optimization ensures that key target areas will not be negatively affected by global adjustment, resulting in significant improvements in the video in terms of light change, color consistency, and visual detail retention, thereby improving the intelligence level of color correction and enhancing the realism and viewing experience of the video.

[0093] S4. Combining the target style image frame data, use the VGG deep learning model to extract style features respectively, adopt the generative adversarial network GAN for style mapping, calculate the optical flow vector field of the consecutive frames after mapping, perform color change prediction, and calculate the temporal smoothing factor to obtain the final style adjustment frame data;

[0094] Preferably, combining the target style image frame data, use the VGG deep learning model to extract style features respectively, adopt the generative adversarial network GAN for style mapping, calculate the optical flow vector field of the consecutive frames after mapping, perform color change prediction, and calculate the temporal smoothing factor to obtain the final style adjustment frame data, including:

[0095] Based on the video frame pixel data optimized for the local target area and the target style image frame data, the style features of the locally optimized video frame and the style features of the target style image frame are output through a pre-trained VGG deep learning model. According to the Euclidean distance of the means of the two features, the style matching error is calculated, expressed as:

[0096] ;

[0097] where represents the style loss value, o represents the layer index of the deep learning model, represents the mean, represents the style feature of the locally optimized video frame at the o-th layer of the VGG model, represents the style feature of the target style image frame at the o-th layer of the VGG model;

[0098] The generator of the generative adversarial network GAN generates stylized features based on the two features, and calculates the Euclidean distance between the stylized features and the features of the optimized video frame as the content loss value;

[0099] The discriminator of the generative adversarial network GAN determines the adversarial loss between the stylized features and the features of the target style image frame. Based on the weighted sum minimization of the style loss, content loss, and adversarial loss, the generator is optimized, and at the same time, the discriminator is optimized to maximize the adversarial loss. When the calculated loss of the optimization result no longer changes significantly in continuous optimization, the optimization is stopped and the optimized generative adversarial network GAN is completed. The video frame pixel data optimized for the local target area is subjected to style mapping to obtain the stylized pixel data based on the target style image frame;

[0100] According to the sequence of the stylized pixel data, the pixels of two consecutive frames are converted into grayscale images, and the Farneback method is used to calculate the optical flow vector field. A pixel grid is constructed to calculate the new position of the pixels in the previous frame in the current frame based on the optical flow vector field, and the pixel color of the previous frame is resampled using bilinear interpolation to obtain the color information of the previous frame after optical flow mapping as the optical flow mapping frame data;

[0101] Using a pre-trained spatio-temporal Transformer STT model, color change prediction is performed based on the sequence of the stylized pixel data and the optical flow mapping frame data of every two frames, and the temporal smoothing factor is calculated, expressed as:

[0102] ;

[0103] where represents the temporal smoothing factor, and respectively represent the stylized pixel data at times t and t - 1, Color change prediction frame data representing time t;

[0104] Perform per-pixel scalar calculation based on the temporal smoothing factor to obtain the final style adjustment frame data , expressed as:

[0105] .

[0106] Extract the style features of the locally optimized video frame and the target style image frame through the VGG deep learning model, and calculate the style matching error, which can ensure the accuracy of style mapping, make the style conversion more in line with the target style features, and avoid style distortion caused by overfitting. Through the optimization of the generative adversarial network GAN, the generated stylized features can not only conform to the target style but also prevent the locally optimized video frame from losing its original content structure. The weighted minimization of style loss, content loss, and adversarial loss enables the generator to perform style transfer more precisely, while the maximization of the discriminator's adversarial loss ensures the authenticity of the stylized result, thereby improving the overall visual quality of the stylized pixel data;

[0107] By calculating the optical flow vector field, resampling the pixels of the previous frame and generating optical flow mapping frame data, the temporal consistency of the stylized video frame is ensured. Calculate the motion trajectory of the pixels based on the optical flow information and perform resampling through bilinear interpolation, making the color change between adjacent frames smooth and natural, reducing the phenomenon of sudden color change between frames caused by style mapping, and improving the viewing experience of the stylized video. Use the spatio-temporal Transformer STT model for color change prediction and calculate the temporal smoothing factor, making the color change of the entire stylized video smoother. The calculation of the temporal smoothing factor ensures the stability of frame-to-frame color conversion, making the style transfer not only reflected in the single-frame effect but also consistent in the time dimension, optimizing the final output video in terms of style consistency, visual fluency, and detail retention;

[0108] Combined with the three steps of global color correction, local optimization of the target area, and style mapping and temporal smoothing, the entire color correction and style transfer process is highly optimized in terms of global consistency, local accuracy, and temporal stability, making the color performance of the video more natural, coordinated, and meeting the requirements of the target style. Global color correction ensures the balance of video frames in terms of brightness, contrast, and overall color distribution, making the color matching of the video frame sequence more unified. Introducing local optimization of the target area can effectively and precisely adjust local areas with large color deviations, ensuring that the target area will not be distorted due to global adjustment. After completing global and local optimization, the introduction of style mapping can convert the entire video frame sequence into the visual effect of the target style, making all video frames highly unified in color style. By extracting style features through the VGG deep learning model and using a generative adversarial network (GAN) for style mapping, it is ensured that both the content information of the video can be retained and the target style can be matched during the style transfer process. By combining optical flow vector calculation and temporal consistency optimization, it can effectively ensure the smooth style change between adjacent frames, reduce color jumps between frames, and make the stylized video have a more natural transition effect.

[0109] S5. Perform image sharpening processing, contrast enhancement processing, and color consistency adjustment, and perform video data compression and integrity verification, and store the data through the cloud.

[0110] Preferably, performing image sharpening processing, contrast enhancement processing, and color consistency adjustment includes:

[0111] For the final style-adjusted frame sequence data, use a Laplacian filter for image sharpening processing and use the Contrast Limited Adaptive Histogram Equalization (CLAHE) method for contrast enhancement processing.

[0112] Calculate the average color histogram of the processed frame sequence data and calculate the ratio with the color temperature reference value of the current frame as the image frame data after color consistency adjustment. The color temperature reference value can be obtained based on the average brightness of the image target area and mapped according to experience.

[0113] By using a Laplacian filter for image sharpening processing, the edge details in the image can be enhanced, the clarity of the image can be improved, and the structure and texture of the image can be made more obvious. By using the Contrast Limited Adaptive Histogram Equalization (CLAHE) method for contrast enhancement, the low-contrast areas in the image can be effectively improved, making the visual effect of the image clearer and more layered. Calculating the average color histogram of the processed frame sequence data and comparing it with the ratio of the color temperature reference value can ensure the consistency of the color distribution of the entire frame sequence.

[0114] Furthermore, video data compression and integrity verification are performed, and data is stored in the cloud. Specifically, for the processed image frame data, the H.265 compression algorithm is used for video compression, GOP (Group of Pictures) with key frame intervals is used for inter-frame compression, the RAID 1 / 5 storage structure is adopted for storing image frame data, and the SHA-256 hash check is used for integrity verification. Then, the data is transmitted to the cloud through wireless transmission technology for backup storage.

[0115] By using the H.265 compression algorithm to compress the processed image frame data, the storage space of the data can be significantly reduced, and the efficiency of video storage and transmission can be improved. By adjusting the GOP settings, the compression efficiency can be optimized without significantly sacrificing image quality, thereby further reducing storage requirements and improving data transmission efficiency. Using the RAID 1 / 5 storage structure to store image frame data can provide high reliability and redundant backup. Through the SHA-256 hash check for integrity verification, it can be ensured that the data has not been tampered with or damaged during storage and transmission. Transmitting the data to the cloud through wireless transmission technology for backup storage can achieve remote backup and improve the accessibility and flexibility of the data.

[0116] This embodiment also provides a system for an artificial intelligence-based video image color correction method, including

[0117] A data acquisition and processing module that acquires video data from the original video source and performs preprocessing, including image size normalization, denoising, and frame sampling;

[0118] A lighting adjustment module that uses the Retinex-Net deep learning model to extract reflectance and lighting data and corrects the color temperature of the video frames according to the lighting adjustment coefficient;

[0119] A local optimization module that uses an object detection algorithm to extract the target area in the video, calculates the color histogram of the target area, and performs local color optimization;

[0120] A style mapping and optimization module that combines the target style image frame data and uses a deep learning model for style mapping;

[0121] A visual processing module that performs image sharpening and contrast enhancement processing and adjusts color consistency according to the calculated color temperature reference value;

[0122] A compression and storage module that compresses the video data, performs data integrity verification, and ensures secure storage.

[0123] This embodiment also provides a computer device, which is applicable to the case of the video image color correction method based on artificial intelligence, and includes: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the video image color correction method based on artificial intelligence as proposed in the above embodiment.

[0124] The computer device can be a terminal. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad set on the outer shell of the computer device, or an external keyboard, a touchpad, or a mouse, etc.

[0125] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the video image color correction method based on artificial intelligence as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM for short), Electrically Erasable Programmable Read-Only Memory (EEPROM for short), Erasable Programmable Read-Only Memory (EPROM for short), Programmable Read-Only Memory (PROM for short), Read-Only Memory (ROM for short), magnetic memory, flash memory, a magnetic disk, or an optical disc.

[0126] In summary, the present invention generates a preliminary color mapping matrix through the color histogram matching method, which can optimize the color distribution of the image to make it more conform to the color characteristics of the target reference frame. The color temperature of the image is adjusted according to the illumination adjustment coefficient and the illumination threshold, so that the warm and cold tones of the video are accurately adjusted. By performing pixel correction, color correction, and local area correction in sequence, a stable color adjustment framework is provided, making the video style consistent. The local optimization ensures that the key target areas will not be negatively affected by the global adjustment, so that the entire video has been significantly improved in terms of illumination change, color consistency, and visual detail retention. The style features are extracted through the VGG deep learning model, and the generative adversarial network GAN is used for style mapping to ensure that both the content information of the video and the target style can be retained during the style transfer process. By combining the optical flow vector calculation and the temporal consistency optimization, it can effectively ensure the smooth style change between adjacent frames and reduce the color jump between frames, making the stylized video have a more natural transition effect.

[0127] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A video image color correction method based on artificial intelligence, characterized in that: include: Collect video data, perform image size normalization and denoising, and obtain preprocessed video frame sequence data; Use the Retinex-Net deep learning model to extract reflectance data and illumination data, calculate the illumination adjustment coefficient, perform pixel correction, and perform color temperature correction on the pixel-corrected video frame pixel data based on the illumination adjustment coefficient; Extract pixel data of the target area from the video frame, obtain the target area histogram, mark the target area, calculate the color adjustment error with the minimum error to obtain the local color transformation matrix, and perform local optimization of the target area; Combined with the target style image frame data, the VGG deep learning model is used to extract style features respectively, and the generative adversarial network GAN is used for style mapping. The optical flow vector field of the mapped continuous frames is calculated, the color change prediction is performed, and the temporal smoothing factor is calculated to obtain the final style adjustment frame data; Perform image sharpening and contrast enhancement, adjust color consistency, compress video data and verify integrity, and store data in the cloud.

2. The video image color correction method based on artificial intelligence as claimed in claim 1, characterized in that: The method uses the Retinex-Net deep learning model to extract reflectance data and illumination data, calculates the illumination adjustment coefficient, performs pixel correction, and performs color temperature correction on the pixel-corrected video frame pixel data according to the illumination adjustment coefficient, including: Based on the preprocessed video frame sequence and the pixel coordinates of the corresponding sequence, the Retinex-Net deep learning model is used to extract reflectance data and illumination data from the pixel data of the video frame sequence, and the mean and standard deviation of the illumination data of the entire frame are calculated, compared with the standard reference illumination, and the illumination adjustment coefficient is calculated; Perform color space conversion on the pixel data of the video frame sequence to obtain a color histogram, and use the color histogram matching method to generate a preliminary color mapping matrix; Set the optimization goal to minimize color deviation, calculate the Euclidean distance between the color histogram of the current frame and the reference color histogram, and calculate the optimal color transformation matrix based on the color proximity and Euclidean distance of all pixels in the video frame; Use the gradient descent method to update the color transformation matrix. If the iterative loss calculated during the iteration process no longer decreases significantly, the iteration is stopped to obtain the optimal color transformation matrix, which is then applied to all pixels of the video frame for pixel correction. The ratio of the standard reference light mean to the illumination mean is used as the illumination threshold, the color temperature of the pixel data of the pixel-corrected video frame is corrected according to the illumination adjustment coefficient, and all the pixels of the video frame after the color temperature correction are updated.

3. The video image color correction method based on artificial intelligence as claimed in claim 2, characterized in that: The method extracts pixel data of the target area from the video frame, obtains a histogram of the target area, marks the target area, calculates a color adjustment error with a minimum error to obtain a local color transformation matrix, and performs local optimization of the target area, including: Use the target detection model YOLOv5 to detect target objects in video frames and extract pixel data of the target area; Perform color space conversion according to the pixel data of the target area to obtain the histogram of the target area, and calculate the Euclidean distance of the pixel deviation between the target area and the standard color; The sum of the historical mean and standard deviation of the Euclidean distance value is used as the distance threshold. If the calculated Euclidean distance is greater than or equal to the distance threshold, it is marked as a target area with a large deviation; For the marked target area, traverse the pixel points in the target area to calculate the color adjustment error with the minimum error; The optimal color transformation matrix is ​​solved by the least square method, and the optimal local color transformation matrix for correcting the marked target area is calculated to obtain the video frame pixel data optimized for the local target area.

4. The video image color correction method based on artificial intelligence as claimed in claim 3, characterized in that: The target style image frame data is combined, the style features are extracted respectively using the VGG deep learning model, the style mapping is performed using the generative adversarial network GAN, and the optical flow vector field of the mapped continuous frames is calculated, the color change prediction is performed, and the temporal smoothing factor is calculated to obtain the final style adjustment frame data, including: Based on the video frame pixel data and the target style image frame data that have been optimized in the local target area, the style features of the locally optimized video frame and the style features of the target style image frame are output through the pre-trained VGG deep learning model, and the style matching error is calculated based on the Euclidean distance between the means of the two features. The generator of the Generative Adversarial Network (GAN) is used to generate stylized features based on the two features, and the Euclidean distance between the stylized features and the features of the optimized video frame is calculated as the content loss value; The discriminator of the generative adversarial network (GAN) determines the feature adversarial loss between the stylized features and the target style image frame, and optimizes the generator based on the weighted sum minimization of the style loss, content loss and adversarial loss. At the same time, the discriminator is optimized to maximize the adversarial loss. When the calculation loss of the optimization result no longer changes significantly in the continuous optimization, the optimization is stopped and the optimized generative adversarial network (GAN) is completed. The style mapping is performed on the pixel data of the video frame optimized in the local target area to obtain the stylized pixel data based on the target style image frame. According to the sequence of stylized pixel data, the pixels of two consecutive frames are converted into grayscale images, and the Farneback method is used to calculate the optical flow vector field. The pixel grid is constructed to calculate the new position of the pixels in the previous frame in the current frame based on the optical flow vector field. The pixel color of the previous frame is resampled using bilinear interpolation to obtain the color information of the previous frame after optical flow mapping as the optical flow mapping frame data; Use the pre-trained spatiotemporal TransformerSTT model to predict color changes based on the sequence of stylized pixel data and the optical flow mapping frame data of every two frames, and calculate the temporal smoothing factor; A pixel-by-pixel scalar calculation is performed based on the temporal smoothing factor to obtain the final style-adjusted frame data.

5. The video image color correction method based on artificial intelligence as claimed in claim 4, characterized in that: The image sharpening and contrast enhancement processing and color consistency adjustment include: For the final style adjustment of the frame sequence data, the Laplacian filter is used for image sharpening, and the adaptive histogram equalization method CLAHE is used for contrast enhancement. The average color histogram of the processed frame sequence data is calculated, and the ratio to the color temperature reference value of the current frame is calculated as the image frame data after color consistency adjustment.

6. The video image color correction method based on artificial intelligence as claimed in claim 5, characterized in that: The collecting of video data, performing image size normalization and denoising processing to obtain preprocessed video frame sequence data includes: For the continuous frames of the video data, the frame data is extracted by using the equal interval sampling method; Calculate the average resolution of all frames, use bicubic interpolation to scale the resolution, and normalize the image size; A Gaussian filter is used to denoise the frame data and perform Gamma correction to obtain a preprocessed video frame sequence.

7. The video image color correction method based on artificial intelligence as claimed in claim 6, characterized in that: The video data compression and integrity verification and data storage through the cloud refer to the processed image frame data, using the H.265 compression algorithm to compress the video, using the key frame interval GOP to perform inter-frame compression, using the RAID 1 / 5 storage structure to store the image frame data, and using the hash check SHA-256 for integrity verification, and transmitting the data to the cloud through wireless transmission technology for data backup storage.

8. A system for a video image color correction method based on artificial intelligence, based on the video image color correction method based on artificial intelligence according to any one of claims 1 to 7, characterized in that: include, The data acquisition and processing module acquires video data from the original video source and performs preprocessing, including image size normalization, denoising, and frame sampling; The lighting adjustment module uses the Retinex-Net deep learning model to extract reflectance and lighting data, and performs color temperature correction on video frames based on the lighting adjustment coefficient; The local optimization module uses the target detection algorithm to extract the target area in the video, calculates the color histogram of the target area, and performs local color optimization; The style mapping optimization module combines the target style image frame data and uses a deep learning model for style mapping; The visual processing module performs image sharpening and contrast enhancement, and adjusts color consistency based on the calculated color temperature reference value; The compression and storage module compresses the video data and performs data integrity verification and secure storage.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the video image color correction method based on artificial intelligence described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the video image color correction method based on artificial intelligence described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Production process visualized agricultural product quality tracing generation method

    CN120746408A

  • Method and system for detecting and correcting brightness and color difference of LED liquid crystal display screen

    CN120833763A

  • Image sequence flicker elimination method and system based on clustering and multi-scale histogram matching

    CN121147028A

  • An image sequence flicker elimination method and system based on clustering and multi-scale histogram matching

    CN121147028B

  • Artificial intelligence-based automatic calibration method and system for creative and cultural patterns

    CN121280220A