A method and device for intelligent scene color adjustment based on feature recognition

By constructing a multi-dimensional feature vector library and LUT mapping table, accurate and fast scene category recognition and LUT matching are achieved, solving the problems of low color grading efficiency and high manual dependence in existing technologies, and improving the automation and flexibility of film and television post-production.

CN120375107BActive Publication Date: 2025-09-23CHENGDU SOBEY DIGITAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510884414.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-09-23
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

The existing LUT color grading technology has low efficiency and high dependence on manual labor in the video color grading process. The color grading process is complex and has a low degree of automation.

Method used

By building a multi-dimensional feature vector library of color, semantics, and scene categories, and establishing a mapping table between scene categories and LUTs, accurate and fast scene category recognition and LUT matching can be achieved. It also supports dynamic addition of scene categories and LUT resources, improving the flexibility and speed of the film and television post-production process.

Benefits of technology

It achieves accurate and fast scene category recognition and LUT matching, improves color grading efficiency, reduces manual intervention, and enhances the automation level of the color grading process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375107B_ABST
    Figure CN120375107B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computer vision and computer graphics, and more specifically, to a method and device for intelligent scene color adjustment based on feature recognition. The method comprises the following steps: S1: pre-define a scene category library and construct a corresponding LUT resource pool; S2: construct a mapping table between a multi-dimensional feature vector of a scene category and a LUT; S3: divide the input video into multiple video clips according to scene changes; S4: extract the multi-dimensional feature vectors of key frames in the video clips; S5: match the scene category that meets the multi-dimensional feature vector in the mapping table, realize intelligent scene classification and dynamically load the corresponding LUT; it can realize accurate and rapid scene category recognition and LUT matching, and can improve the flexibility and speed of the film and television post-production process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical fields of computer vision and computer graphics, and in particular to a method and device for intelligent scene color adjustment based on feature recognition. Background Art

[0002] LUT (Look Up Table) color grading technology is widely used in film and television post-production. It achieves efficient color stylization through predefined color mapping relationships. With the evolution of digital media technology, LUT color grading technology has found widespread application in filmmaking, live broadcasting, and short video.

[0003] Although the method of using LUT color grading technology for video color grading is relatively mature, there are still problems in the video color grading process, such as low color grading efficiency, high dependence on manual labor, complex color grading process and low degree of automation. Summary of the Invention

[0004] The present invention aims to provide an intelligent scene color grading method and device based on feature recognition. By constructing a multi-dimensional feature vector library of color, semantics, and scene categories and establishing a mapping table between scene categories and LUTs, the method can achieve accurate and rapid scene category recognition and LUT matching. It also supports the dynamic addition of scene categories and LUT resources, and can arbitrarily adjust scene categories, LUTs, and feature categories according to actual usage requirements, thereby improving the flexibility and speed of the film and television post-production process.

[0005] The purpose of this application is achieved through the following technical solutions:

[0006] In the first aspect, the present application proposes an intelligent scene color adjustment method based on feature recognition, comprising the following steps:

[0007] S1: Predefine scene category library and build corresponding LUT resource pool;

[0008] S2: Construct a mapping table between the multi-dimensional feature vector of the scene category and the LUT;

[0009] S3: Split the input video into multiple video segments according to scene changes;

[0010] S4: Extract multi-dimensional feature vectors of key frames in video clips;

[0011] S5: Match the scene category that meets the multi-dimensional feature vector in the mapping table, realize intelligent scene classification and dynamically load the corresponding LUT.

[0012] Preferably, the scene category library includes at least one of program opening, program ending, indoor, outdoor, and studio.

[0013] Preferably, the LUT resource pool is constructed by historical color adjustment data.

[0014] Preferably, the multidimensional feature vector includes color features of main hue, saturation, and brightness based on HSV, as well as semantic features of sky, grass, and human face extracted by a pre-trained segmentation model, and scene category features of indoor, outdoor, and studio obtained by scene detection.

[0015] Preferably, step S2 includes:

[0016] Constructing a multidimensional feature vector: First, each extracted feature value is added to a list in a preset fixed order. This list is the multidimensional feature vector.

[0017] Construct a mapping table between multi-dimensional feature vectors and LUTs: For a scene category, set a standard range for each feature value in the multi-dimensional feature vector and set the corresponding LUT for the scene category.

[0018] Preferably, the characteristic value includes a specific numerical value, a Boolean type, and a category number.

[0019] Preferably, step S3 adopts an AutoShot-based scene segmentation algorithm, which predicts the probability that each input frame is a scene segmentation point, and determines the segmentation point by setting a threshold, thereby achieving scene-based segmentation of the video into several video segments.

[0020] Preferably, step S4 includes the following steps:

[0021] S41, extracting one or more frames of the video clip as key frames;

[0022] S42, uses the Seaformer semantic segmentation model to extract semantic features of key frames;

[0023] S43, calculating the color features of the key frame in HSV;

[0024] S44, scene category features are detected using deep features extracted from pre-trained ViT and ResNet models;

[0025] S45, fuses color, semantic and scene category features into a unified multi-dimensional feature vector.

[0026] Preferably, step S5 includes:

[0027] The matching degree is calculated in sequence between the multi-dimensional feature vector extracted from the video clip and the preset standard range of each scene category in the mapping table. The formula is:

[0028] ;

[0029] ;

[0030] in represents a multi-dimensional feature vector, represents the eigenvalue number, and Represent the lower and upper bounds of the standard range, respectively. Represents the total number of features in the multi-dimensional feature vector, Represents the final matching degree, The corresponding The result of feature matching;

[0031] After calculating the matching degree of all scene categories and multi-dimensional feature vectors, the scene with the maximum matching degree is taken. When the matching degree of the scene is greater than the empirical value of 0.8, the match is successful and the corresponding LUT is loaded into the video clip. Otherwise, the scene is judged as not belonging to the known scene category.

[0032] In the second aspect, the present application proposes an intelligent scene color adjustment device based on feature recognition, wherein the intelligent scene color adjustment device based on feature recognition includes: a processor and a memory, wherein the memory stores a computer program, and the computer program is loaded and executed by the processor to implement an intelligent scene color adjustment method based on feature recognition as described in the first aspect.

[0033] The above-mentioned main solution of this application and its various further options can be freely combined to form multiple solutions, all of which are solutions that can be adopted and protected by this application. Moreover, in this application, (non-conflicting options) can also be freely combined with each other and with other options. After understanding the solution of this application, those skilled in the art will understand that there are many combinations based on existing technology and common knowledge, all of which are technical solutions to be protected by this application, and this is not an exhaustive list.

[0034] The beneficial effects of this application are:

[0035] The present invention provides an intelligent scene color adjustment method and device based on feature recognition, which can combine color, semantics and scene features to achieve accurate and rapid identification of scene categories, and support flexible expansion of scene categories and LUTs to adapt to the needs of new scene categories, thereby solving the problems of poor scene adaptability, frequent manual intervention and low color adjustment efficiency in the existing technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.

[0037] Figure 1 This is a schematic diagram of the process of an intelligent scene color adjustment method based on feature recognition in an embodiment of the present application.

[0038] Figure 2 This is a schematic diagram of a method for intelligent scene color adjustment based on feature recognition according to an embodiment of the present application.

[0039] Figure 3 Matching instance graph for a single video clip. DETAILED DESCRIPTION

[0040] The following describes the embodiments of the present application through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other unless they conflict.

[0041] Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative work shall fall within the scope of protection of this application.

[0042] An embodiment of the present invention provides an intelligent scene color adjustment method based on feature recognition.

[0043] refer to Figure 1-3 , an intelligent scene color adjustment method based on feature recognition, comprising the following steps:

[0044] S1: Predefine scene category library and build corresponding LUT resource pool.

[0045] The scene category library includes but is not limited to program opening, program ending, indoor, outdoor, studio, etc. The LUT resource pool is constructed based on historical color grading data and expert experience.

[0046] S2: Construct a mapping table between the multi-dimensional feature vector of the scene category and the LUT.

[0047] Each scene category is associated with at least one LUT. The multi-dimensional feature vector includes, but is not limited to, color features such as primary hue, saturation, and brightness based on the HSV (HueSaturation Value) color space model; semantic features such as sky, grass, and faces extracted through a pre-trained segmentation model; and scene category features such as indoor, outdoor, and studio obtained through scene detection.

[0048] During the construction of a multidimensional feature vector, each extracted eigenvalue is first added to a list in a predetermined, fixed order. This list becomes the multidimensional feature vector. Eigenvalues ​​can be specific numeric values, such as the mean brightness, Boolean values, such as whether the broadcast is a studio, or category numbers, such as the dominant color hue. This construction method is scalable, allowing for the addition and removal of desired feature types.

[0049] The mapping table between the multi-dimensional feature vectors of each scene category and the LUT is constructed by setting a standard range for each eigenvalue in the multi-dimensional feature vector and configuring the corresponding LUT for that scene category. This mapping table features a dynamic expansion mechanism, allowing the desired number and categories of scenes to be detected to be controlled simply by modifying the mapping table, without modifying or affecting other steps in the process.

[0050] Figure 2 This is a schematic diagram of the process of an intelligent scene color adjustment method based on feature recognition in an embodiment of the present application, which realizes intelligent color adjustment through an automated progressive processing structure: starting from the input video, "segment segmentation" is first performed to generate continuous video segments, and each segment is converted into a feature vector through feature extraction; these feature vectors are matched with the feature variables and scene categories to output the corresponding LUT, and the matching LUT identifier (LUT-1 to LUT-N) is dynamically output by comparing with the preset scene LUT mapping table. Finally, the corresponding LUT is loaded into the original video segment to complete the full process of automated color processing.

[0051] S3: Divide the input video into multiple video segments according to scene changes.

[0052] The input video is divided into multiple video segments according to scene changes using the AutoShot-based scene segmentation algorithm. The AutoShot-based scene segmentation algorithm predicts the probability that each input frame is a scene segmentation point, and determines the segmentation point by setting a threshold, thereby achieving scene-based segmentation of the video into several video segments.

[0053] S4: Extract multi-dimensional feature vectors of key frames in video clips.

[0054] Specifically, step S4 includes the following steps S41-S45:

[0055] S41: extract one or more frames of the video clip as key frames.

[0056] S42, uses the Seaformer semantic segmentation model to extract semantic features of key frames.

[0057] S43, calculating the color features of the key frame in HSV.

[0058] S44, detects scene category features using deep features extracted from pre-trained models such as ViT and ResNet.

[0059] S45, fuses color, semantic and scene category features into a unified multi-dimensional feature vector.

[0060] In actual application, a multi-process parallel method can be used to handle feature extraction tasks. A separate process is assigned to each feature detection for processing, which can improve processing efficiency and facilitate the scalability of the feature detection method. Adding feature detection methods will not overly affect the processing speed.

[0061] It should be understood that feature detection is not limited to the methods proposed and used in this patent, and appropriate methods need to be adopted according to the category of the required features. For example, for color features, pixel statistics methods or histogram analysis can be used; semantic features can select specific algorithm models according to the desired semantic categories; scene features can be deep feature matched based on templates, and for specific scenes, specific deep network models can also be added for detection.

[0062] S5: Match the scene category that meets the multi-dimensional feature vector in the mapping table, realize intelligent scene classification and dynamically load the corresponding LUT.

[0063] In this step, the specific method is to calculate the matching degree between the multi-dimensional feature vector extracted from the video clip and the standard range preset for each scene category in the mapping table in turn. The formula is:

[0064] ;

[0065] ;

[0066] in represents a multi-dimensional feature vector, represents the eigenvalue number, and Represent the lower and upper bounds of the standard range, respectively. Represents the total number of features in the multi-dimensional feature vector, Represents the final matching degree, The corresponding The result of feature matching.

[0067] After calculating the matching degree of all scene categories and multi-dimensional feature vectors, the scene with the maximum matching degree is taken. When the matching degree of the scene is greater than the empirical value of 0.8, the match is successful and the corresponding LUT is loaded into the video clip. Otherwise, the scene is judged as not belonging to the known scene category.

[0068] Figure 3 The single video clip matching instance graph starts with the video clip input and is analyzed in multiple dimensions by the feature vector extraction module: keyframe extraction locates representative frames, and color detection, semantic detection, and scene detection are performed in sequence. The above detection results are structured and stored as a feature vector table, which clearly lists the feature names and their corresponding feature values, forming a machine-parseable digital segment portrait. The feature vector table is input to the matching module with the scene category LUT mapping table, and the output color adjustment strategy is determined by dynamic rules:

[0069] Mapping table structure: stores preset scene categories (category 1 to category N). Each category defines conditional rules and corresponding LUT ID.

[0070] Matching logic: The system compares the eigenvalues ​​in the feature vector table with the mapping table rules (such as R1 (hue), R2 (face), R3 (sky), R4 (outdoors)) one by one. When all the rules are met, the LUTID corresponding to the scene category is immediately output to drive subsequent color loading operations.

[0071] It should be understood that the matching calculation is not limited to the above method and can be modified according to the usage situation. For example, certain features can be set as required and the original method can be used to calculate other necessary features. Another example is that several parameters can be combined to calculate the distance as a matching reference. This shows that the scene matching process also has strong scalability and flexibility.

[0072] Based on the same inventive concept, an embodiment of the present invention further provides an intelligent scene color adjustment device based on feature recognition.

[0073] A feature-recognition-based intelligent scene color adjustment device includes: a processor, a memory, and a communication bus. The communication bus is used to achieve a communication connection between the processor and the memory. The processor is used to execute a computer program stored in the memory to implement the above-mentioned feature-recognition-based intelligent scene color adjustment method.

[0074] The feature recognition-based intelligent scene color grading device can be implemented in various forms, including mobile phones, tablet computers, PDAs, laptop computers, and desktop computers.

[0075] The memory may be used to store instructions, programs, codes, code sets, or instruction sets. The memory may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for at least one function, and instructions for implementing the feature recognition-based intelligent scene color adjustment method provided in the above embodiment. The data storage area may store data involved in the feature recognition-based intelligent scene color adjustment method provided in the above embodiment.

[0076] The processor may include one or more processing cores. The processor calls the data stored in the memory by running or executing the instructions, programs, code sets or instruction sets stored in the memory, performs the various functions of the present invention and processes data. The processor can be at least one of an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller and a microprocessor. It is understandable that for different devices, the electronic device used to implement the above-mentioned processor function can also be other, and the embodiment of the present invention is not specifically limited.

[0077] In summary, a method and device for intelligent scene color grading based on feature recognition is provided. The method comprises the following steps: S1: pre-defining a scene category library and constructing a corresponding LUT resource pool; S2: constructing a mapping table between the multi-dimensional feature vectors of the scene categories and the LUTs; S3: dividing the input video into multiple video segments according to scene changes; S4: extracting the multi-dimensional feature vectors of the key frames in the video segments; S5: matching the scene categories that conform to the multi-dimensional feature vectors in the mapping table, realizing intelligent scene classification and dynamically loading the corresponding LUTs; the method can achieve accurate and rapid scene category recognition and LUT matching, and can improve the flexibility and speed of the film and television post-production process.

[0078] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. An intelligent scene color adjustment method based on feature recognition, characterized in that: The following steps are involved: S1: Predefine scene category library and build corresponding LUT resource pool; S2: Construct a mapping table between the multi-dimensional feature vector of the scene category and the LUT; S3: Split the input video into multiple video segments according to scene changes; S4: Extract multi-dimensional feature vectors of key frames in video clips; The step S4 comprises the following steps: S41, extracting one or more frames of the video clip as key frames; S42, uses the Seaformer semantic segmentation model to extract semantic features of key frames; S43, calculating the color features of the key frame in HSV; S44, scene category features are detected using deep features extracted from pre-trained ViT and ResNet models; S45, fuses color, semantic and scene category features into a unified multi-dimensional feature vector; S5: Match the scene category that meets the multi-dimensional feature vector in the mapping table, realize intelligent scene classification and dynamically load the corresponding LUT; The step S5 comprises: The matching degree is calculated in sequence between the multi-dimensional feature vector extracted from the video clip and the preset standard range of each scene category in the mapping table. The formula is: ; ; in represents a multi-dimensional feature vector, represents the eigenvalue number, and Represent the lower and upper bounds of the standard range, respectively. Represents the total number of features in the multi-dimensional feature vector, Represents the final matching degree, The corresponding The result of feature matching; After calculating the matching degree of all scene categories and multi-dimensional feature vectors, the scene with the maximum matching degree is taken. When the matching degree of the scene is greater than the empirical value of 0.8, the match is successful and the corresponding LUT is loaded into the video clip. Otherwise, the scene is judged as not belonging to the known scene category.

2. The intelligent scene color adjustment method based on feature recognition according to claim 1, characterized in that: The scene category library includes at least one of program opening, program ending, indoor, outdoor, and studio.

3. The intelligent scene color adjustment method based on feature recognition according to claim 2, characterized in that: The LUT resource pool is constructed using historical color adjustment data.

4. The intelligent scene color adjustment method based on feature recognition according to claim 1, characterized in that: The multidimensional feature vector includes color features of main hue, saturation, and brightness based on HSV, semantic features of sky, grass, and human face extracted by a pre-trained segmentation model, and scene category features of indoor, outdoor, and studio obtained by scene detection.

5. The intelligent scene color adjustment method based on feature recognition according to claim 4, characterized in that: The step S2 comprises: Constructing a multidimensional feature vector: First, each extracted feature value is added to a list in a preset fixed order. This list is the multidimensional feature vector. Construct a mapping table between multi-dimensional feature vectors and LUTs: For a scene category, set a standard range for each feature value in the multi-dimensional feature vector and set the corresponding LUT for the scene category.

6. The intelligent scene color adjustment method based on feature recognition according to claim 5, characterized in that: The characteristic value includes a specific numerical value, a Boolean type, and a category number.

7. The intelligent scene color adjustment method based on feature recognition according to claim 1, characterized in that: The step S3 adopts the AutoShot-based scene segmentation algorithm, which predicts the probability that each input frame is a scene segmentation point, and determines the segmentation point by setting a threshold, thereby segmenting the video into several video segments based on the scene.

8. An intelligent scene color adjustment device based on feature recognition, characterized in that: The intelligent scene color adjustment device based on feature recognition includes: a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the intelligent scene color adjustment method based on feature recognition according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • GPU computing power management method, medium, device and system

    CN114661482A

  • Short video annotation method based on video questions and answers

    CN118968383A