Video color enhancement method, device and equipment based on deep learning guidance multi-lookup table fusion and storage medium

Through the multi-finding table fusion method based on deep learning, the problem of color jump between frames after video color enhancement is solved, and a more stable and efficient video color enhancement effect is achieved.

CN119996638APending Publication Date: 2025-05-13HAIWEI ZHIZAO TECH (WUHAN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510017207.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing video color enhancement method based on artificial intelligence drives has poor video stability after color enhancement, and color jumps are prone to occur between frames.

Method used

Using a multi-finding table fusion method based on deep learning guidance, multiple target three-dimensional lookup tables are obtained through the image classification network, pixel adaptive weights are obtained through the perception network, initial pixel values ​​of the original video data are obtained, and pending pixel values ​​are determined based on multiple target three-dimensional lookup tables. Finally, pending pixel values ​​are processed according to the pixel adaptive weights to obtain the target pixel values.

Benefits of technology

It effectively reduces the color jump between video frames after color enhancement, and improves video stability and color enhancement effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996638A_ABST
    Figure CN119996638A_ABST
Patent Text Reader

Abstract

The invention discloses a video color enhancement method and device based on deep learning guidance multi-lookup table fusion, equipment and a storage medium, and relates to the technical field of video enhancement, and the video color enhancement method based on deep learning guidance multi-lookup table fusion comprises the steps: obtaining a plurality of target three-dimensional lookup tables through an image classification network; obtaining a pixel adaptive weight through a sensing network; obtaining an initial pixel value of original video data, and determining a plurality of undetermined pixel values corresponding to the initial pixel value according to the plurality of target three-dimensional lookup tables; and processing the plurality of undetermined pixel values according to the pixel adaptive weights to obtain a target pixel value. According to the invention, the color hopping phenomenon between video frames after color enhancement can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video enhancement technology, and in particular to a video color enhancement method, device, equipment and storage medium based on deep learning-guided multi-lookup table fusion. Background Art

[0002] Video color enhancement technology is not only used in film post-production, video surveillance, medical imaging and other fields, but also widely used in online video playback and digital image processing. At present, the integration of artificial intelligence technology makes video color enhancement more efficient. Although video color enhancement technology has made significant progress, the AI-driven method still faces technical defects. When the color is enhanced, the video stability is poor and color jumps are prone to occur between frames. Therefore, how to reduce the color jump phenomenon between video frames after color enhancement is a problem that still needs to be solved.

[0003] The above contents are only used to assist in understanding the technical solution of the present application and do not constitute an admission that the above contents are prior art. Summary of the invention

[0004] The main purpose of this application is to provide a video color enhancement method, device, equipment and storage medium based on deep learning-guided multi-lookup table fusion, aiming to solve the technical problem of how to reduce the color jump phenomenon between video frames after color enhancement.

[0005] To achieve the above objectives, the present application proposes a video color enhancement method based on deep learning-guided multi-lookup table fusion, the method comprising: Obtain multiple target three-dimensional lookup tables through an image classification network; Get pixel adaptive weights through the perception network; Acquire an initial pixel value of the original video data, and determine a plurality of pending pixel values ​​corresponding to the initial pixel value according to the plurality of target three-dimensional lookup tables; The multiple pending pixel values ​​are processed according to the pixel adaptive weights to obtain a target pixel value.

[0006] In one embodiment, the step of obtaining a plurality of target three-dimensional lookup tables through an image classification network comprises: Construct an image classification network through downsampling layers and convolution layers; Determining input requirements of a three-dimensional lookup table through a downsampling layer of the image classification network; Extracting training set features through the convolutional layer of the image classification network and predicting the weights of the three-dimensional lookup table; A plurality of initial three-dimensional lookup tables are obtained, and a plurality of target three-dimensional lookup tables are obtained according to the input requirements, the three-dimensional lookup table weights and the initial three-dimensional lookup tables.

[0007] In one embodiment, the step of obtaining a plurality of target three-dimensional lookup tables according to the input requirement, the three-dimensional lookup table weights and the initial three-dimensional lookup table comprises: Obtaining a target loss function, wherein the target loss function includes a mean square error loss function, a smoothing regularization loss function, a monotonicity regularization loss function, and a multi-scale structural similarity loss function; The initial three-dimensional lookup table is processed according to the target loss function, the input requirement, and the three-dimensional lookup table weight to obtain a plurality of target three-dimensional lookup tables.

[0008] In one embodiment, the step of obtaining pixel adaptive weights through a perceptual network includes: Construct a perceptual network through convolutional layers and adaptive average pooling layers; A training set image is obtained, and the perception network is trained using the training set image to predict pixel adaptive weights.

[0009] In one embodiment, the step of training the perception network using training set images to predict pixel adaptive weights includes: Dividing the training set image into regions to obtain multiple image regions; Acquire the color temperature and hue of the image area; The perception network is trained by the color temperature and hue to predict pixel adaptive weights.

[0010] In one embodiment, the step of determining a plurality of pending pixel values ​​corresponding to the initial pixel value according to the plurality of target three-dimensional lookup tables comprises: Determine vertex data of a cube in the plurality of target three-dimensional lookup tables according to the initial pixel value; A plurality of undetermined pixel values ​​corresponding to the initial pixel value are determined by interpolation according to the vertex data.

[0011] In one embodiment, the step of processing the multiple pending pixel values ​​according to the pixel adaptive weight to obtain the target pixel value includes: Multiplying the multiple undetermined pixel values ​​by the pixel adaptive weights to obtain multiple target pixel value components; Add the multiple target pixel value components to obtain the target pixel value.

[0012] In addition, to achieve the above purpose, the present application also proposes a video color enhancement device based on deep learning to guide the fusion of multiple lookup tables, and the video color enhancement device based on deep learning to guide the fusion of multiple lookup tables includes: A classification module, used to obtain a plurality of target three-dimensional lookup tables through an image classification network; The perception module is used to obtain pixel adaptive weights through the perception network; A determination module, configured to obtain an initial pixel value of the original video data, and determine a plurality of pending pixel values ​​corresponding to the initial pixel value according to the plurality of target three-dimensional lookup tables; A processing module is used to process the multiple pending pixel values ​​according to the pixel adaptive weight to obtain a target pixel value.

[0013] In addition, to achieve the above-mentioned objectives, the present application also proposes a video color enhancement device based on deep learning guided multi-lookup table fusion, the device comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program being configured to implement the steps of the video color enhancement method based on deep learning guided multi-lookup table fusion as described above.

[0014] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the video color enhancement method based on deep learning guided multi-lookup table fusion as described above are implemented.

[0015] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the video color enhancement method based on deep learning guided multi-lookup table fusion as described above.

[0016] The present application provides a video color enhancement method based on deep learning-guided multi-lookup table fusion. The present application obtains multiple target three-dimensional lookup tables through an image classification network; obtains pixel adaptive weights through a perception network; obtains the initial pixel value of the original video data, and determines multiple pending pixel values ​​corresponding to the initial pixel value according to the multiple target three-dimensional lookup tables; processes the multiple pending pixel values ​​according to the pixel adaptive weights to obtain the target pixel value.

[0017] In summary, the present application roughly calculates the mapping space of color enhancement of different images by using multiple three-dimensional lookup table weight networks, and then uses a pixel weight network to accurately adjust the lookup table results, and finally obtains a color-enhanced image, thereby reducing the color jump phenomenon between video frames after color enhancement. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0019] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0020] Figure 1 A flowchart diagram of a first embodiment of a method for video color enhancement based on deep learning-guided multi-lookup table fusion in this application; Figure 2 A schematic diagram of the perception network structure of a video color enhancement method based on deep learning-guided multi-lookup table fusion provided in Example 1 of the present application; Figure 3 A brief schematic diagram of a video color enhancement method based on deep learning-guided multi-lookup table fusion provided in Example 1 of the present application; Figure 4 A flowchart diagram of Embodiment 2 of the video color enhancement method based on deep learning-guided multi-lookup table fusion provided in this application; Figure 5 A schematic diagram of an image classification network structure of a video color enhancement method based on deep learning-guided multi-lookup table fusion provided in Example 2 of the present application; Figure 6 This is a schematic diagram of the module structure of a video color enhancement device based on deep learning-guided multi-lookup table fusion according to an embodiment of the present application; Figure 7 Schematic diagram of the device structure of the hardware operating environment involved in the video color enhancement method based on deep learning guided multi-lookup table fusion in the embodiment of the present application.

[0021] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0022] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0023] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0024] The main solution of the present application is to obtain multiple target three-dimensional lookup tables through an image classification network; obtain pixel adaptive weights through a perception network; obtain initial pixel values ​​of original video data, and determine multiple pending pixel values ​​corresponding to the initial pixel values ​​according to the multiple target three-dimensional lookup tables; process the multiple pending pixel values ​​according to the pixel adaptive weights to obtain target pixel values.

[0025] At present, the integration of artificial intelligence technology has made video color enhancement more efficient. Although video color enhancement technology has made significant progress, the AI-driven method still faces technical defects. When the video is color enhanced, the stability is poor and color jumps are likely to occur between frames. Therefore, how to reduce the color jump phenomenon between video frames after color enhancement is a problem that still needs to be solved.

[0026] This application uses multiple three-dimensional lookup table weight networks to roughly calculate the mapping space of color enhancement of different images, and then uses a pixel weight network to accurately adjust the lookup table results, and finally obtains a color-enhanced image, thereby reducing the color jump phenomenon between video frames after color enhancement.

[0027] Based on this, the embodiment of the present application provides a video color enhancement method based on deep learning-guided multi-lookup table fusion, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the video color enhancement method based on deep learning guided multi-lookup table fusion in this application.

[0028] In this embodiment, the video color enhancement method based on deep learning-guided multi-lookup table fusion includes steps S10 to S40: Step S10: obtaining a plurality of target three-dimensional lookup tables through an image classification network; It should be noted that the execution subject of this embodiment may be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of realizing the above functions, a video color enhancement device based on deep learning to guide the fusion of multiple lookup tables, etc. The following takes the video color enhancement device based on deep learning to guide the fusion of multiple lookup tables as an example to illustrate this embodiment and the following embodiments.

[0029] It should be noted that the 3D Lookup Table (3D-LUT) is a tool used for color correction and management in video and image processing. It achieves efficient and accurate adjustment of image color by pre-calculating and storing color mapping data. Therefore, LUT is obtained through deep learning methods. However, since not all scenes are suitable for the same LUT, multiple LUT tables and adaptive weights are used to dynamically combine multiple lookup tables to generate LUTs for different image contents. LUT weights can be obtained using classification networks.

[0030] Step S20: Obtain pixel adaptive weights through the perception network; It should be noted that the same pixel value on the same image corresponds to the same mapping value in the LUT, that is, it is enhanced to the same pixel value. However, for the image as a whole, the same pixel value needs to be enhanced to different values ​​for color temperature and hue depending on the person area and background. Therefore, pixel adaptive weights are added to dynamically adjust the pixel values ​​after LUT enhancement.

[0031] In a feasible manner, the step of obtaining pixel adaptive weights through a perception network includes: Construct a perceptual network through convolutional layers and adaptive average pooling layers; A training set image is obtained, and the perception network is trained using the training set image to predict pixel adaptive weights.

[0032] It should be noted that in the video color enhancement task, it is usually necessary to make different adjustments to the color temperature and hue for different regions. The step of training the perception network through the training set images and predicting the pixel adaptive weights specifically includes: dividing the training set images into regions to obtain multiple image regions; obtaining the color temperature and hue of the image regions; training the perception network through the color temperature and hue to predict the pixel adaptive weights. Therefore, the contextual information of the image pixels should be carefully considered during the enhancement process. Therefore, a local contextual information perception network is proposed to predict the pixel adaptive weights and accurately adjust the lookup table results. The specific structure of the network can be referred to Figure 2 , consisting of 7 convolutional layers and 1 adaptive average pooling layer, where Input represents input, Conv represents convolution kernel, ReLU and Sigmoid represent activation functions, AdaptiveAvePool represents adaptive average pooling layer, and AdaptiveAvePool, Conv3 and Conv4 constitute the channel self-attention mechanism. Considering that the size of the pixel weight map is the same as the input image, the network operation amount and reasoning speed are determined by the size of the input map. If the method is to achieve real-time, the network size should be strictly controlled. Therefore, except for Conv1, which is a 3*3 convolution kernel, all others are 1*1 convolution kernels. In addition, in order to prevent the last few layers of the network from forgetting the feature information of the input image, a channel self-attention structure is used to eliminate this effect. The structure consists of 1 adaptive average pooling layer and 2 convolution layers. The structure of the perception network can be referred to Table 1. Table 1 shows the specific content of the number of output channels, convolution kernel size and step size of each convolution layer, and N is the number of lookup tables.

[0033] Table 1

[0034] Step S30: acquiring an initial pixel value of the original video data, and determining a plurality of pending pixel values ​​corresponding to the initial pixel value according to the plurality of target three-dimensional lookup tables; It is understandable that after the initial pixel values ​​of the original video data are obtained, a plurality of pending pixel values ​​corresponding to any one of the initial pixel values ​​may be obtained through the target three-dimensional lookup table.

[0035] In a feasible manner, the step of determining a plurality of pending pixel values ​​corresponding to the initial pixel value according to the plurality of target three-dimensional lookup tables comprises: Determine vertex data of a cube in the plurality of target three-dimensional lookup tables according to the initial pixel value; A plurality of undetermined pixel values ​​corresponding to the initial pixel value are determined by interpolation according to the vertex data.

[0036] Step S40: Processing the multiple pending pixel values ​​according to the pixel adaptive weights to obtain target pixel values.

[0037] It is understandable that after obtaining a plurality of pending pixel values, the plurality of pending pixel values ​​need to be quantified by adaptive weights to obtain a final target pixel value.

[0038] In a feasible manner, the step of processing the multiple pending pixel values ​​according to the pixel adaptive weight to obtain the target pixel value includes: Multiplying the multiple undetermined pixel values ​​by the pixel adaptive weights to obtain multiple target pixel value components; Add the multiple target pixel value components to obtain the target pixel value.

[0039] It is understandable that each undetermined pixel value corresponds to a certain adaptive weight, and the final target pixel value can be obtained by multiplying the undetermined pixel value and the adaptive weight and then adding them together. At this time, color enhancement is completed. For a specific schematic diagram, please refer to Figure 3 , Figure 3 In the method, the pixel values ​​of the initial image are processed by three three-dimensional lookup tables and the lookup table weights and pixel weights to obtain the target image after color enhancement.

[0040] This embodiment obtains multiple target three-dimensional lookup tables through an image classification network; obtains pixel adaptive weights through a perception network; obtains initial pixel values ​​of original video data, and determines multiple pending pixel values ​​corresponding to the initial pixel values ​​according to the multiple target three-dimensional lookup tables; processes the multiple pending pixel values ​​according to the pixel adaptive weights to obtain target pixel values.

[0041] To summarize, this embodiment roughly calculates the mapping space of color enhancement of different images by using multiple three-dimensional lookup table weight networks, and then uses a pixel weight network to accurately adjust the lookup table results, and finally obtains a color-enhanced image, thereby reducing the color jump phenomenon between video frames after color enhancement.

[0042] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction, and will not be repeated in the following. Figure 4 , step S10 also includes steps S101 to S104: Step S101: construct an image classification network through a downsampling layer and a convolution layer; It should be noted that downsampling is an operation that reduces the sampling rate of a signal or image by reducing the number of data points. In image processing, downsampling is usually used to reduce the resolution or size of an image, that is, to reduce the number of pixels in the image. In deep learning, downsampling helps reduce the dimensionality of feature maps and the amount of computation while retaining important feature information.

[0043] Step S102: determining the input requirements of the three-dimensional lookup table through the downsampling layer of the image classification network; It is understandable that by downsampling the input image size to 256x256, the number of parameters for subsequent convolutions is effectively reduced and the overall network operation speed is improved. 3D grid composed of elements, where M represents the sampling interval in each color channel. Given an input image I, the table lookup method is used to transform the enhanced output O. The transformation operation is:

[0044] in, It means that the pixel-to-pixel mapping is defined by a lookup table. The common setting is M=33, which contains 108K parameters and can achieve a good balance between inference speed and enhancement quality.

[0045] Step S103: extracting training set features through the convolution layer of the image classification network, and predicting the weights of the three-dimensional lookup table; It is understandable that the training set features are extracted through the convolutional layer of the image classification network, and the overall clues in the image are captured and the adaptive weights of the lookup table are predicted by deep learning. Adjust the table lookup results, and the transformation operation of the LUT weight is described as follows:

[0046] in, represents the image enhancement process, is the mapping operation of the nth lookup table, where n=3. The network flow of the lookup table adaptive weight can be referred to Figure 5 , Figure 5In the example, Input is the input, Downsample is the downsampling layer, the input image size is limited to 256x256, Conv is the convolution kernel, LeaklyReLU is the activation function, and InstanceNorm is the instance normalization. This network is a simple image classification network consisting of 6 convolution layers and one downsampling layer. First, the input image size is limited to 256x256 by downsampling, which greatly reduces the number of parameters for subsequent convolutions and improves the overall network operation speed. Then, 6 convolution layers are used to extract overall image clues and features, predict N LUT weights, and dynamically adjust the enhancement quality. For the specific content of the number of output channels, convolution kernel size, and step size of each convolution layer, please refer to Table 2.

[0047] Table 2

[0048] Step S104: obtaining a plurality of initial three-dimensional lookup tables, and obtaining a plurality of target three-dimensional lookup tables according to the input requirements, the three-dimensional lookup table weights and the initial three-dimensional lookup tables.

[0049] It is understandable that when a lookup table is used to enhance video color information, the MSE loss is used , smooth regularization loss , monotonicity regularization loss and MSSIM loss As a basic loss term, it can be expressed as:

[0050] in, Network prediction value, is the true value.

[0051]

[0052] in, and Represent the mean and variance of the predicted and real images respectively, and represents a constant, N is a multi-scale component, represents the covariance between the predicted and true images. and Represents the trade-off coefficients of different multi-scale components.

[0053]

[0054] in, Represents the 3D-LUT output value, Represents the lookup table weights, which is the output of the lookup table weight network.

[0055]

[0056] in, It is defined as the standard ReLU operation, i.e. g(a) = max(0, a).

[0057]

[0058] in, and is the trade-off coefficient, set to = , =10. and Ensure the content and color consistency between the enhanced results and the target photo. Smooth Regularization and monotonicity regularization It is used to ensure that the output values ​​of the LUT are smooth and monotonous, and to maintain the relative brightness and saturation of the input RGB values, ensuring a natural enhancement result.

[0059] This embodiment constructs an image classification network through a downsampling layer and a convolution layer; determines the input requirements of a three-dimensional lookup table through the downsampling layer of the image classification network; extracts training set features through the convolution layer of the image classification network, and predicts the weights of the three-dimensional lookup table; obtains multiple initial three-dimensional lookup tables, and obtains multiple target three-dimensional lookup tables based on the input requirements, the three-dimensional lookup table weights and the initial three-dimensional lookup table.

[0060] In summary, this embodiment predicts the adaptive weights of the three-dimensional lookup table through the image classification network to obtain adaptive weights for subsequent color enhancement, thereby reducing the color jump phenomenon between video frames after color enhancement.

[0061] This application also provides a video color enhancement device based on deep learning-guided multi-lookup table fusion, please refer to Figure 6 , the video color enhancement device based on deep learning-guided multi-lookup table fusion includes: A classification module 10, used to obtain a plurality of target three-dimensional lookup tables through an image classification network; A perception module 20, used to obtain pixel adaptive weights through a perception network; A determination module 30, configured to obtain an initial pixel value of the original video data, and determine a plurality of pending pixel values ​​corresponding to the initial pixel value according to the plurality of target three-dimensional lookup tables; The processing module 40 is used to process the multiple pending pixel values ​​according to the pixel adaptive weight to obtain a target pixel value.

[0062] This embodiment obtains multiple target three-dimensional lookup tables through an image classification network; obtains pixel adaptive weights through a perception network; obtains initial pixel values ​​of original video data, and determines multiple pending pixel values ​​corresponding to the initial pixel values ​​according to the multiple target three-dimensional lookup tables; processes the multiple pending pixel values ​​according to the pixel adaptive weights to obtain target pixel values.

[0063] This embodiment roughly calculates the mapping space of color enhancement of different images by using multiple three-dimensional lookup table weight networks, and then uses a pixel weight network to accurately adjust the lookup table results, and finally obtains a color-enhanced image, thereby reducing the color jump phenomenon between video frames after color enhancement.

[0064] In one embodiment, the classification module 10 is also used to construct an image classification network through a downsampling layer and a convolution layer; determine the input requirements of the three-dimensional lookup table through the downsampling layer of the image classification network; extract training set features through the convolution layer of the image classification network, and predict the weights of the three-dimensional lookup table; obtain multiple initial three-dimensional lookup tables, and obtain multiple target three-dimensional lookup tables based on the input requirements, the three-dimensional lookup table weights and the initial three-dimensional lookup table.

[0065] In one embodiment, the classification module 10 is also used to obtain a target loss function, which includes a mean square error loss function, a smoothing regularization loss function, a monotonicity regularization loss function, and a multi-scale structural similarity loss function; the initial three-dimensional lookup table is processed according to the target loss function, the input requirements, and the three-dimensional lookup table weights to obtain multiple target three-dimensional lookup tables.

[0066] In one embodiment, the perception module 20 is further used to construct a perception network through a convolution layer and an adaptive average pooling layer; obtain a training set image, and train the perception network through the training set image to predict pixel adaptive weights.

[0067] In one embodiment, the perception module 20 is further used to divide the training set image into regions to obtain multiple image regions; obtain the color temperature and hue of the image region; train the perception network through the color temperature and hue to predict pixel adaptive weights.

[0068] In one embodiment, the determination module 30 is further used to determine the vertex data of the cube in the multiple target three-dimensional lookup tables according to the initial pixel value; and determine multiple pending pixel values ​​corresponding to the initial pixel value by interpolation method according to the vertex data.

[0069] In one embodiment, the processing module 40 is further configured to multiply the multiple pending pixel values ​​by the pixel adaptive weights to obtain multiple target pixel value components; and add the multiple target pixel value components to obtain the target pixel value.

[0070] The video color enhancement device based on deep learning-guided multi-lookup table fusion provided in the present application adopts the video color enhancement method based on deep learning-guided multi-lookup table fusion in the above-mentioned embodiment, which can solve the technical problem of how to reduce the color jump phenomenon between video frames after color enhancement. Compared with the prior art, the beneficial effects of the video color enhancement device based on deep learning-guided multi-lookup table fusion provided in the present application are the same as the beneficial effects of the video color enhancement method based on deep learning-guided multi-lookup table fusion provided in the above-mentioned embodiment, and the other technical features of the video color enhancement device based on deep learning-guided multi-lookup table fusion are the same as the features disclosed in the above-mentioned embodiment method, which will not be repeated here.

[0071] The present application provides a video color enhancement device based on deep learning-guided multi-lookup table fusion, and the video color enhancement device based on deep learning-guided multi-lookup table fusion includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the video color enhancement method based on deep learning-guided multi-lookup table fusion in the above-mentioned embodiment one.

[0072] Reference below Figure 7 , which shows a schematic diagram of the structure of a video color enhancement device based on deep learning-guided multi-lookup table fusion suitable for implementing the embodiment of the present application. The video color enhancement device based on deep learning-guided multi-lookup table fusion in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions: tablet computers), PMPs (Portable Media Players: portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The video color enhancement device based on deep learning-guided multi-lookup table fusion shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0073] like Figure 7As shown, the video color enhancement device based on deep learning guiding multi-lookup table fusion may include a processing device 1001 (such as a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM: ReadOnly Memory) 1002 or the program loaded from the storage device 1003 to the random access memory (RAM: RandomAccess Memory) 1004. Various programs and data required for the operation of the video color enhancement device based on deep learning guiding multi-lookup table fusion are also stored in RAM1004. The processing device 1001, ROM1002 and RAM1004 are connected to each other through a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 1003 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 1009. The communication device 1009 can allow the video color enhancement device based on deep learning-guided multi-lookup table fusion to communicate wirelessly or wired with other devices to exchange data. Although the figure shows a video color enhancement device based on deep learning-guided multi-lookup table fusion with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or have instead.

[0074] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0075] The video color enhancement device based on deep learning-guided multi-lookup table fusion provided by the present application adopts the video color enhancement method based on deep learning-guided multi-lookup table fusion in the above-mentioned embodiment, which can solve the technical problem of how to reduce the color jump phenomenon between video frames after color enhancement. Compared with the prior art, the beneficial effects of the video color enhancement device based on deep learning-guided multi-lookup table fusion provided by the present application are the same as the beneficial effects of the video color enhancement method based on deep learning-guided multi-lookup table fusion provided by the above-mentioned embodiment, and the other technical features of the video color enhancement device based on deep learning-guided multi-lookup table fusion are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.

[0076] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0077] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0078] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the video color enhancement method based on deep learning-guided multi-lookup table fusion in the above-mentioned embodiment.

[0079] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM: Random Access Memory), a read-only memory (ROM: Read Only Memory), an erasable programmable read-only memory (EPROM: Erasable Programmable Read Only Memory or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM: CD-Read Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency: Radio Frequency), etc., or any suitable combination of the above.

[0080] The above-mentioned computer-readable storage medium can be included in the video color enhancement device based on deep learning guided multi-lookup table fusion; or it can exist independently without being assembled into the video color enhancement device based on deep learning guided multi-lookup table fusion.

[0081] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by a video color enhancement device based on deep learning guided multi-lookup table fusion, the video color enhancement device based on deep learning guided multi-lookup table fusion: obtains multiple target three-dimensional lookup tables through an image classification network; obtains pixel adaptive weights through a perception network; obtains initial pixel values ​​of original video data, and determines multiple pending pixel values ​​corresponding to the initial pixel values ​​according to the multiple target three-dimensional lookup tables; processes the multiple pending pixel values ​​according to the pixel adaptive weights to obtain target pixel values.

[0082] Computer program code for performing the operations of the present application may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0083] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0084] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.

[0085] The readable storage medium provided in the present application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned video color enhancement method based on deep learning-guided multi-lookup table fusion, and can solve the technical problem of how to reduce the color jump phenomenon between video frames after color enhancement. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in the present application are the same as the beneficial effects of the video color enhancement method based on deep learning-guided multi-lookup table fusion provided in the above-mentioned embodiment, and will not be repeated here.

[0086] The present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the above-mentioned video color enhancement method based on deep learning-guided multi-lookup table fusion.

[0087] The computer program product provided by the present application can solve the technical problem of how to reduce the color jump phenomenon between video frames after color enhancement. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as the beneficial effects of the video color enhancement method based on deep learning-guided multi-lookup table fusion provided by the above embodiment, which will not be repeated here.

[0088] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A video color enhancement method based on deep learning-guided multi-lookup table fusion, characterized in that: The method includes: Obtain multiple target three-dimensional lookup tables through an image classification network; Get pixel adaptive weights through the perception network; Acquire an initial pixel value of the original video data, and determine a plurality of pending pixel values ​​corresponding to the initial pixel value according to the plurality of target three-dimensional lookup tables; The multiple pending pixel values ​​are processed according to the pixel adaptive weights to obtain a target pixel value.

2. The method according to claim 1, characterized in that The step of obtaining a plurality of target three-dimensional lookup tables through an image classification network comprises: Construct an image classification network through downsampling layers and convolution layers; Determining input requirements of a three-dimensional lookup table through a downsampling layer of the image classification network; Extracting training set features through the convolutional layer of the image classification network and predicting the weights of the three-dimensional lookup table; A plurality of initial three-dimensional lookup tables are obtained, and a plurality of target three-dimensional lookup tables are obtained according to the input requirements, the three-dimensional lookup table weights and the initial three-dimensional lookup tables.

3. The method according to claim 2, characterized in that The step of obtaining a plurality of target three-dimensional lookup tables by combining the three-dimensional lookup table weights with the initial three-dimensional lookup table according to the input requirement comprises: Obtaining a target loss function, wherein the target loss function includes a mean square error loss function, a smoothing regularization loss function, a monotonicity regularization loss function, and a multi-scale structural similarity loss function; The initial three-dimensional lookup table is processed according to the target loss function, the input requirement, and the three-dimensional lookup table weight to obtain a plurality of target three-dimensional lookup tables.

4. The method according to claim 1, characterized in that The step of obtaining pixel adaptive weights through the perception network includes: Construct a perceptual network through convolutional layers and adaptive average pooling layers; A training set image is obtained, and the perception network is trained using the training set image to predict pixel adaptive weights.

5. The method according to claim 4, characterized in that The step of training the perception network through the training set images and predicting pixel adaptive weights comprises: Dividing the training set image into regions to obtain multiple image regions; Acquire the color temperature and hue of the image area; The perception network is trained by the color temperature and hue to predict pixel adaptive weights.

6. The method according to claim 1, characterized in that The step of determining a plurality of pending pixel values ​​corresponding to the initial pixel value according to the plurality of target three-dimensional lookup tables comprises: Determine vertex data of a cube in the plurality of target three-dimensional lookup tables according to the initial pixel value; A plurality of undetermined pixel values ​​corresponding to the initial pixel value are determined by interpolation according to the vertex data.

7. The method according to claim 1, characterized in that The step of processing the multiple undetermined pixel values ​​according to the pixel adaptive weight to obtain the target pixel value comprises: Multiplying the multiple undetermined pixel values ​​by the pixel adaptive weights to obtain multiple target pixel value components; Add the multiple target pixel value components to obtain the target pixel value.

8. A video color enhancement device based on deep learning-guided multi-lookup table fusion, characterized in that: The device comprises: A classification module, used to obtain a plurality of target three-dimensional lookup tables through an image classification network; The perception module is used to obtain pixel adaptive weights through the perception network; A determination module, configured to obtain an initial pixel value of the original video data, and determine a plurality of pending pixel values ​​corresponding to the initial pixel value according to the plurality of target three-dimensional lookup tables; A processing module is used to process the multiple pending pixel values ​​according to the pixel adaptive weight to obtain a target pixel value.

9. A video color enhancement device based on deep learning-guided multi-lookup table fusion, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the video color enhancement method based on deep learning-guided multi-lookup table fusion as described in any one of claims 1 to 7.

10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the video color enhancement method based on deep learning guided multi-lookup table fusion as described in any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Image color enhancement method and device, equipment and storage medium

    CN120976082A