Automatic color matching system and method for record based on deep learning

Through a deep learning-based method, the visual and semantic features of documentary images are extracted and tone constraints are determined, which solves the problem of inconsistent tone styles in documentary color tuning, and the visual and narrative logic coordination in the automatic color tuning process is achieved, which improves the color tuning effect and audience experience.

CN120529028APending Publication Date: 2025-08-22CHONGQING JIUGUANG CULTURE MEDIA CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510628775.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

The existing documentary color tuning technology is difficult to achieve consistency and visual coherence of tone styles in long-term, multi-scene, and strong narrative content, and lacks the recognition and utilization of semantic relationships between segments, resulting in a sudden visual change in color tuning results, affecting the audience's understanding and immersion.

Method used

Using a deep learning-based method, the shallow visual features and scene semantic features of each frame of the documentary image are extracted, and the tone constraint conditions are determined through the style similarity and brightness difference between adjacent frames. Combining the cross-correlation and fuzzy association of scene semantic features, the cost of chroma reconstruction is constructed and the automatic color tuning processing is realized.

Benefits of technology

It improves the adaptability of the color tuning strategy to the original material content, reduces sudden jumps between frames, enhances the naturalness and style unity of visual transition between shots, ensures the consistency and coherence of the color tuning results in visual presentation and narrative logic, and improves the color expression quality and visual coherence of the documentary.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120529028A_ABST
    Figure CN120529028A_ABST
Patent Text Reader

Abstract

The invention provides a deep learning-based automatic toning system and method for a record. The method comprises the following steps of: determining a toning constraint condition of the record in visual constraints according to style similarity of shallow visual features between adjacent recorded image frames in the record and brightness difference between the adjacent recorded image frames; performing fuzzy correlation on the scene semantic elements of the narrative segments in the record through the cross correlation among the scene semantic features to obtain the element membership of the scene semantics in the narrative segments, and further determining the scene attention of different narrative segments in the record according to all the element membership; and according to the scene attention and the hue constraint conditions of the different narrative segments, determining the chroma reconstruction cost of the record in the toning process, and carrying out automatic toning processing on the record based on the chroma reconstruction cost. According to the scheme, the visual features and the semantic elements can be fused in the automatic toning process of the record, and toning coordination with narrative logic as the guide is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image extraction technology, and more specifically, to a documentary automatic color grading system and method based on deep learning. Background Art

[0002] In the documentary color grading process, image extraction plays a fundamental and guiding role. Its core significance lies in providing accurate and structured image information support for subsequent color grading strategies. As an image form that combines artistic expression and authenticity, the visual quality of a documentary directly affects the clarity of information conveyed and the audience's immersive experience. In traditional documentary production, color grading work is highly dependent on manual operation. Not only is the work intensity high and the cost high, it also relies heavily on the professional experience of the colorist. It is difficult to achieve consistency in tonal style in long-term, multi-scene materials. Especially in the context of the increasing trend of self-service content production, realizing intelligent processing of documentary color style is of great practical significance for improving production efficiency, reducing labor costs, and ensuring the uniformity of image style.

[0003] Existing color grading technologies have significant flaws when dealing with documentaries, which have long durations, multiple scenes, and strong narrative content. They mostly use a method based on overall image style transfer for color grading, ignoring the continuity and differences in visual effects between adjacent image frames, resulting in abrupt visual jumps in the color grading results, destroying the tonal consistency and visual coherence between shots. In addition, existing technologies generally lack the recognition and utilization of semantic relationships between clips, and are unable to differentiate the tonal style according to the status of different scenes in the narrative logic and the relevance of content, resulting in a disconnect between color processing and content expression, affecting the audience's understanding and immersion. Therefore, how to integrate visual features and semantic elements in the automatic color grading process of documentaries to achieve tonal coordination guided by narrative logic has become a difficult problem facing the industry. Summary of the Invention

[0004] This application provides a documentary automatic color grading system and method based on deep learning, which can integrate visual features and semantic elements in the automatic color grading process of documentaries to achieve tonal coordination guided by narrative logic.

[0005] In a first aspect, the present application provides a method for automatic color grading of documentaries based on deep learning, comprising the following steps: Receive the target documentary to be color graded, and then extract the shallow visual features of each frame of the target documentary and the scene semantic features of different narrative segments based on the pre-trained deep learning model; Determining a correlation relationship of tonal consistency between different recorded image frames in a target documentary based on the style similarity of shallow visual features between adjacent recorded image frames, and then determining a tonal constraint condition of the target documentary in a visual constraint based on the correlation relationship and the brightness difference between adjacent recorded image frames; The scene semantic elements of each narrative segment in the target documentary are fuzzily associated through the mutual correlation between the semantic features of each scene, and the element membership of the scene semantics in each narrative segment is obtained. Then, the scene attention of different narrative segments in the target documentary is determined based on all the element memberships. The chroma reconstruction cost of the target documentary during the color grading process is determined according to the scene attention of different narrative segments and the color hue constraint condition, and then the target documentary is automatically color-graded based on the chroma reconstruction cost.

[0006] Preferably, the extraction of shallow visual features of each frame of recorded image and scene semantic features of different narrative segments in the target documentary based on the pre-trained deep learning model specifically includes: Set shallow feature modules and deep feature modules in the pre-trained deep learning model; The shallow visual features of the recorded image frames in the target documentary are extracted frame by frame through the shallow feature module; Divide the target documentary into narrative segments to obtain multiple narrative segments; The scene semantics of each narrative segment are labelled by the deep feature module to obtain the scene semantic labels of each narrative segment; The scene semantic features of different narrative segments are determined based on the global semantic vector of the target documentary and the scene semantic labels of each narrative segment.

[0007] Preferably, determining the correlation relationship of the tonal consistency between different recorded image frames in the target documentary based on the style similarity of the shallow visual features between adjacent recorded image frames specifically includes: Determining the style similarity of shallow visual features between adjacent recorded image frames; Determine the tonal constraint matrix of the target documentary based on all the style similarities; The tone constraint matrix is ​​used to perform correlation analysis on the tone consistency between different recorded image frames in the target documentary, so as to obtain the correlation relationship of the tone consistency between different recorded image frames in the target documentary.

[0008] Preferably, determining the hue constraint condition of the target documentary in the visual constraint based on the association relationship and the brightness difference between adjacent recorded image frames specifically includes: Performing a differential operation on the energy of the brightness channels between adjacent recorded image frames to obtain the brightness difference between the adjacent recorded image frames, and then determining a brightness difference map between the adjacent recorded image frames; Determine a correlation map of tones between adjacent recorded image frames according to the correlation relationship; The color tone stability between adjacent recorded image frames is constrained by analyzing the difference map and the correlation map, and the color tone constraint conditions of the target documentary in the visual constraint are obtained.

[0009] Preferably, the scene semantic elements of each narrative segment in the target documentary are fuzzily associated through the mutual correlation between the semantic features of each scene, and the element membership of the scene semantics in each narrative segment is obtained, which specifically includes: Initialize the initial membership of each narrative segment corresponding to the scene semantic elements; According to the mutual correlation between the semantic features of each scene, the mutual correlation matrix of scene semantics between narrative segments is constructed; Mapping the values ​​in the mutual correlation matrix into semantic fuzzy sets based on Gaussian membership function; Constructing a fuzzy association matrix of scene semantics through all initial memberships and the semantic fuzzy set; The expression of each narrative segment in the dimension of scene semantic elements is fuzzily evaluated according to the fuzzy association matrix to obtain the element membership of the scene semantics in each narrative segment.

[0010] Preferably, determining the scene attention of different narrative segments in the target documentary based on the membership of all elements specifically includes: Obtain the scene semantic elements of each narrative segment in the target documentary, and then determine the information entropy of the scene semantic elements of each narrative segment; The scene attention of different narrative segments in the target documentary is determined by the element membership of the scene semantics in each narrative segment and the corresponding information entropy.

[0011] Preferably, the target documentary to be color-graded is received through a file input interface.

[0012] In a second aspect, the present application provides a documentary automatic color grading system based on deep learning, comprising: An extraction module, which receives the target documentary to be color-graded and extracts shallow visual features of each frame of the target documentary and scene semantic features of different narrative segments based on a pre-trained deep learning model; a processing module for determining, based on the style similarity of shallow visual features between adjacent recorded image frames, a correlation relationship of tonal consistency between different recorded image frames in a target documentary, and further determining, based on the correlation relationship and the brightness difference between adjacent recorded image frames, a tonal constraint condition of the target documentary in a visual constraint; The processing module is further configured to perform fuzzy association on the scene semantic elements of each narrative segment in the target documentary based on the mutual correlation between the semantic features of each scene, thereby obtaining the element membership of the scene semantics in each narrative segment, and then determining the scene attention of different narrative segments in the target documentary based on all the element memberships; An execution module is used to determine the chromaticity reconstruction cost of the target documentary during the color grading process based on the scene attention of different narrative segments and the color tone constraint condition, and then automatically perform color grading on the target documentary based on the chromaticity reconstruction cost.

[0013] In a third aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory stores a code, and the processor is configured to obtain the code and execute the above-mentioned deep learning-based documentary automatic color grading method.

[0014] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned deep learning-based documentary automatic color grading method.

[0015] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects: In an embodiment of the present application, a target documentary to be color-adjusted is first received, and then shallow visual features of each frame of recorded image in the target documentary and scene semantic features of different narrative segments are extracted based on a pre-trained deep learning model; the correlation relationship of hue consistency between different recorded image frames in the target documentary is determined according to the style similarity of the shallow visual features between adjacent recorded image frames, and then the hue constraint condition of the target documentary in the visual constraint is determined according to the correlation relationship and the brightness difference between adjacent recorded image frames; the scene semantic elements of each narrative segment in the target documentary are fuzzy associated through the mutual correlation between each scene semantic feature, and the element affiliation of the scene semantics in each narrative segment is obtained, and then the scene attention of different narrative segments in the target documentary is determined based on all the element affiliations; the chromaticity reconstruction cost of the target documentary in the color adjustment process is determined according to the scene attention of different narrative segments and the hue constraint condition, and then the target documentary is automatically color-adjusted based on the chromaticity reconstruction cost.

[0016] It can be seen that the present application determines the chromaticity reconstruction cost of the target documentary in the chromaticity adjustment process through the scene attention and chromaticity constraints of different narrative segments, and automatically performs color adjustment on the target documentary based on the chromaticity reconstruction cost; first, based on the pre-trained deep learning model, the shallow visual features of each frame of the recorded image in the target documentary and the scene semantic features of different narrative segments are extracted, and the structured modeling of the image frames in the two dimensions of the perception layer and the semantic layer is realized, which provides an accurate data basis for the subsequent construction of chromaticity constraints and semantic associations, and significantly improves the adaptability of the color adjustment strategy to the original material content; secondly, the chromaticity constraint conditions of the target documentary in the visual constraint are determined by the correlation relationship between the chromaticity consistency between different recorded image frames in the target documentary and the brightness difference between adjacent recorded image frames, which can reduce the impact of abrupt jumps caused by ignoring the continuity between frames in traditional color adjustment, thereby enhancing the naturalness of the visual transition between shots. and style unity; then, based on the mutual correlation between the scene semantic features in each narrative segment, a fuzzy membership and scene attention mechanism is constructed, so that the color adjustment strategy can flexibly adjust the local color distribution according to the narrative relationship between the segments, ensuring that the color style has both overall coordination and local differentiated expression, avoiding the separation between content expression and visual style; finally, the chromaticity reconstruction cost of the target documentary in the chromaticity adjustment process is determined through the scene attention and chromaticity constraints of different narrative segments, and the target documentary is automatically color-adjusted based on the chromaticity reconstruction cost, which not only realizes the dynamic adaptive control of the color adjustment process, but also ensures the consistency and coherence of the color adjustment results in visual presentation and narrative logic, thereby comprehensively improving the color performance quality and viewing coherence of the documentary; in summary, the present application scheme can integrate visual features and semantic elements in the automatic color adjustment process of documentaries to achieve narrative logic-oriented color coordination. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is an exemplary flow chart of a method for automatic color grading of documentaries based on deep learning according to some embodiments of the present application; Figure 2 is a schematic diagram of a process for determining hue constraints according to some embodiments of the present application; Figure 3 is a schematic diagram of a process for determining element membership according to some embodiments of the present application; Figure 4 1 is a schematic diagram of a deep learning-based automatic color grading system for documentaries according to some embodiments of the present application; Figure 5 This is a structural diagram of a computer device for implementing a deep learning-based automatic color grading method for documentaries according to some embodiments of the present application. DETAILED DESCRIPTION

[0018] In order to better understand the technical solution of the present application, the technical solution of the present application will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0019] refer to Figure 1 , which is an exemplary flow chart of a documentary automatic color grading method based on deep learning according to some embodiments of the present application. The documentary automatic color grading method 100 based on deep learning mainly includes the following steps: In step 101, a target documentary to be color-graded is received, and then shallow visual features of each frame of recorded image and scene semantic features of different narrative segments in the target documentary are extracted based on a pre-trained deep learning model.

[0020] It should be noted that the target documentary in this application refers to the original documentary material to be automatically color-processed, specifically including continuous video images, sound tracks and the narrative content they express. The video part, as the processing object of this solution, is the basic unit that constitutes the visual expression of the documentary. It has a certain temporal continuity and narrative logic, and is the basic data source for color adjustment and semantic analysis. In specific implementation, the target documentary to be color-processed can be received through the file input interface, supporting common video formats (such as MP4, MOV), and using a video decoding module (such as FFmpeg) to decode it into an image sequence.

[0021] In some embodiments, extracting shallow visual features of each recorded image frame and scene semantic features of different narrative segments in a target documentary based on a pre-trained deep learning model can be achieved by using the following steps: Set shallow feature modules and deep feature modules in the pre-trained deep learning model; The shallow visual features of the recorded image frames in the target documentary are extracted frame by frame through the shallow feature module; Divide the target documentary into narrative segments to obtain multiple narrative segments; The scene semantics of each narrative segment are labelled by the deep feature module to obtain the scene semantic labels of each narrative segment; The scene semantic features of different narrative segments are determined based on the global semantic vector of the target documentary and the scene semantic labels of each narrative segment.

[0022] It should be noted that the shallow feature module in this application is a module used to extract basic visual information from images; the deep feature module in this application is a module used to extract scene semantic information from narrative clips; the shallow visual features in this application are feature indicators that reflect the basic visual information in recorded images; the scene semantic features in this application are feature indicators that reflect the scene semantic information in narrative clips; the scene semantic labels in this application refer to the category labels of the scene semantics corresponding to the narrative clips.

[0023] In specific implementation, setting shallow feature modules and deep feature modules in the pre-trained deep learning model can be achieved in the following way, namely: the network structure of the deep learning model is functionally divided according to the hierarchical feature abstraction capability. In this embodiment, the first three convolutional layers (such as conv1 to conv3 in ResNet) can be used as shallow feature modules to extract the basic visual elements of the recorded image, including edge texture and local color distribution, and the subsequent convolutional layers (such as conv4, conv5) and their fully connected layers can be used as deep feature modules to extract scene-level features with global semantic generalization capabilities; the shallow visual features of the recorded image frames in the target documentary are extracted frame by frame through the shallow feature module, which can be achieved in the following way, namely: by The video is extracted frame by frame (usually read at 25 frames / second) and input into the shallow feature module, and the output feature map is the shallow visual feature; the target documentary is divided into narrative segments, and the acquisition of multiple narrative segments can be achieved in the following way, namely: a content-based dynamic segmentation method (such as an algorithm for detecting shot switching) can be used to segment the target documentary to obtain multiple narrative segments with temporal continuity; the scene semantics of each narrative segment is labeled by the deep feature module to obtain the scene semantic label of each narrative segment. This can be achieved in the following way, namely: for each narrative segment, the image frame set within the time period is input into the deep feature module, and its high-order semantic features are extracted by the deep feature module. The deep features are processed by Global After Average Pooling processing, the data is input into the Softmax classification head, which is connected to a pre-defined set of scene categories (such as street, office, nature, and war) to generate a probability distribution, and one or more scene semantic labels are assigned to each narrative segment based on high-order semantic features. The scene semantic features of different narrative segments can be determined based on the global semantic vector of the target documentary and the scene semantic labels of each narrative segment. This can be achieved by: using the deep feature module to extract the deep semantic features of all image frames in the entire documentary, and performing global average pooling on them to obtain a global semantic vector representing the overall semantic features of the documentary. Secondly, for each narrative segment, based on its scene semantic label, the typical semantic vector of the corresponding category is screened from the deep feature space, and combined with the global semantic vector for weighted fusion to form a context-enhanced semantic representation of the narrative segment. The fused semantic vector is used as the scene semantic feature of the narrative segment, thereby obtaining the scene semantic feature of each narrative segment.It should be further explained that weighted fusion in this application refers to linearly combining the typical semantic vectors of the semantic labels corresponding to the narrative segments with the global semantic vector of the target documentary at a certain weight ratio. The fusion method can adopt a weighted summation strategy, setting an adjustable parameter a∈[0,1] to control the proportion of local semantics to global semantics, that is, the fusion vector = a×local semantic vector+(1−a)×global semantic vector, thereby enhancing the contextual relevance of the segment semantics.

[0024] In step 102, the association relationship of the tonal consistency between different recorded image frames in the target documentary is determined based on the style similarity of the shallow visual features between adjacent recorded image frames, and then the tonal constraint condition of the target documentary in the visual constraint is determined based on the association relationship and the brightness difference between adjacent recorded image frames.

[0025] In some embodiments, determining the association relationship of the tonal consistency between different recorded image frames in the target documentary based on the style similarity of the shallow visual features between adjacent recorded image frames can be achieved by using the following steps: Determining the style similarity of shallow visual features between adjacent recorded image frames; Determine the tonal constraint matrix of the target documentary based on all the style similarities; The tone constraint matrix is ​​used to perform correlation analysis on the tone consistency between different recorded image frames in the target documentary, so as to obtain the correlation relationship of the tone consistency between different recorded image frames in the target documentary.

[0026] It should be noted that the style similarity in this application is an indicator for measuring the degree of similarity in visual style between different recorded image frames; the tone constraint matrix in this application is a symmetric matrix that represents the strength of the pairwise tone consistency relationship between each image frame in the target documentary based on style similarity, wherein each matrix element represents the degree of similarity in tone between the corresponding two frames of image; the association relationship in this application is an indicator for measuring the degree of consistency and coherence between different image frames in the documentary in the tone dimension.

[0027] In specific implementation, determining the style similarity of shallow visual features between adjacent recorded image frames can be achieved in the following way, namely: the style similarity between each pair of adjacent recorded image frames can be calculated through the Gram matrix (i.e., the inner product between channels of shallow visual features). The closer the Gram matrix is, the more consistent the style between adjacent recorded image frames is; determining the tone constraint matrix of the target documentary based on all style similarities can be achieved in the following way, namely: the style similarities between all adjacent frames are combined into a two-dimensional matrix according to the time sequence of the documentary, and the two-dimensional matrix is ​​used as the tone constraint matrix of the target documentary, wherein the matrix elements represent the strength of the tone consistency between different recorded image frames; the tone consistency between different recorded image frames in the target documentary is correlated and analyzed through the tone constraint matrix, and the correlation relationship of the tone consistency between different recorded image frames in the target documentary is obtained. This can be achieved in the following way, namely: first, each frame image in the documentary is used as a graph node, and the style similarity between frames is used as the relationship between graph nodes through the constructed tone constraint matrix. , and then all graph nodes are connected by all edge weights to obtain a weighted undirected graph; secondly, a graph clustering algorithm (such as spectral clustering) is applied to segment the weighted undirected graph, and image frames with high hue similarity are classified into the same hue-related cluster. In this process, spectral clustering performs low-dimensional embedding and clustering on the hue space structure by calculating the Laplacian matrix eigenvector of the hue constraint matrix, which can effectively extract the hue consistency pattern between image frames; then, the image frames in the same hue-related cluster are regarded as a set of associated frames with strong hue consistency, and a strong correlation value is assigned to the hue consistency correlation relationship between all recorded image frames in the same associated frame set. In the embodiment of the present application, the strong correlation value can be 0.9. A weak correlation value is assigned to the hue consistency correlation relationship between image frames recorded in different associated frame sets. In the embodiment of the present application, the weak correlation value can be 0.6. In other embodiments, in order to distinguish the hue consistency correlation relationship between different recorded image frames, the difference between the strong correlation value and the weak correlation value can be appropriately increased, which is not specifically limited here.

[0028] In some embodiments, reference Figure 2 As shown in FIG. 1 , this figure is a flow chart of determining the hue constraint condition in some embodiments of the present application. In this embodiment, the hue constraint condition of the target documentary in the visual constraint is determined by the association relationship and the brightness difference between adjacent recorded image frames, which can be implemented by the following steps: In step 1021, a differential operation is performed on the energy of the brightness channel between adjacent recorded image frames to obtain the brightness difference between the adjacent recorded image frames, and then a brightness difference map between the adjacent recorded image frames is determined; In step 1022, a correlation map of tones between adjacent recorded image frames is determined based on the correlation relationship; In step 1023, a constraint analysis is performed on the tone stability between adjacent recorded image frames using the difference map and the correlation map to obtain tone constraint conditions of the target documentary in the visual constraint.

[0029] It should be noted that the brightness difference in this application is an indicator for measuring the degree of brightness difference between adjacent image frames; the brightness difference map in this application is a temporal reference for measuring the degree of brightness change between adjacent image frames, and is used to identify areas of sudden brightness changes; the hue association map in this application is structured information used to represent the strength of the relationship between the consistency of hue styles between adjacent image frames, and is used to identify continuous areas of narrative style; the hue constraint condition in this application is a control parameter used to limit the amplitude of hue changes between image frames during automatic color adjustment, and is used to ensure the smoothness of color transitions and style consistency.

[0030] In specific implementation, a differential operation is performed on the energy of the brightness channel between adjacent recorded image frames to obtain the brightness difference between adjacent recorded image frames, and then the brightness difference map between adjacent recorded image frames can be determined in the following manner, namely: the documentary video can be extracted frame by frame and converted into Lab color space, and the energy value in the brightness channel is extracted, and then the energy values ​​in the brightness channel between adjacent image frames are compared, and the difference obtained by comparison is used as the brightness difference between adjacent recorded image frames. The brightness difference between all adjacent recorded image frames can be obtained in the above manner, and the brightness difference between all adjacent recorded image frames can be further arranged in the order of frame numbers, and the one-dimensional time series map obtained by the arrangement is used as the brightness difference map between adjacent recorded image frames; the association map of the color tone between adjacent recorded image frames according to the association relationship can be determined in the following manner, namely: the association relationship of the color tone consistency between adjacent recorded image frames is obtained, and then all the association relationships are arranged in the order of frame numbers, and the one-dimensional relationship map obtained by the arrangement is used as the brightness difference map between adjacent recorded image frames. The spectrum is used as a correlation map of the tones between adjacent recorded image frames; the tones between adjacent recorded image frames are constrained and analyzed by the difference map and the correlation map, and the tones constraint conditions of the target documentary in the visual constraints can be obtained in the following way, namely: traverse the frame pair sequence, and for the frame pairs whose difference in the brightness difference map is greater than the set threshold, further search in the tones correlation map to see whether their style similarity is higher than the tones similarity threshold. If these two conditions are met, it indicates that the frame pair has a sudden change in brightness but the visual style remains consistent, and has narrative continuity, and should be regarded as a tones transition sensitive area. On this basis, the tones adjustment range of the frame pair can be restricted, such as setting the hue adjustment not to exceed a set change amount, or applying a gradient smoothing function constraint to the region during the automatic color adjustment process to suppress drastic tones changes while maintaining overall style consistency. Finally, the frame pairs that meet the conditions of large brightness change and strong tones correlation and their corresponding tones adjustment restriction parameters are used as the tones constraint conditions of the target documentary in the visual constraints.

[0031] In step 103, the scene semantic elements of each narrative segment in the target documentary are fuzzy associated through the mutual correlation between the semantic features of each scene, and the element affiliation of the scene semantics in each narrative segment is obtained, and then the scene attention of different narrative segments in the target documentary is determined based on all the element affiliations.

[0032] In some embodiments, reference Figure 3 As shown in FIG. 1 , this figure is a schematic diagram of a process for determining element membership in some embodiments of the present application. In this embodiment, the scene semantic elements of each narrative segment in the target documentary are fuzzily associated through the mutual correlation between the semantic features of each scene. The element membership of the scene semantics in each narrative segment can be obtained by the following steps: In step 1031, the initial membership of each narrative segment corresponding to the scene semantic element is initialized; In step 1032, a scene semantic inter-correlation matrix between narrative segments is constructed based on the inter-correlation between the semantic features of each scene; In step 1033, the values ​​in the mutual correlation matrix are mapped into semantic fuzzy sets based on a Gaussian membership function; In step 1034, a fuzzy association matrix of scene semantics is constructed using all initial membership degrees and the semantic fuzzy set; In step 1035, a fuzzy evaluation is performed on the expression of each narrative segment in the dimension of scene semantic elements according to the fuzzy association matrix to obtain the element membership of the scene semantics in each narrative segment.

[0033] It should be noted that the scene semantic elements in this application refer to the basic elements that constitute the scene semantic content of the narrative fragment; the initial membership in this application is the initial quantitative value used to characterize the degree of direct semantic matching between the narrative fragment and the scene semantic elements, which is used to provide the basic expression of fuzzy semantic modeling; the element membership in this application is the final value used to characterize the expression intensity of the semantic elements of the narrative fragment after integrating the global semantic correlation.

[0034] In specific implementation, initializing the initial membership of the scene semantic elements corresponding to each narrative segment can be achieved in the following way, namely: for each narrative segment, a pre-trained deep learning model can be used to extract semantic features of the text and image of the narrative segment to obtain the graphic semantic vector of the narrative segment, obtain the scene semantic label of the narrative segment, and convert the scene semantic label into the embedding vector of the scene semantic element corresponding to the narrative segment, and then calculate the cosine similarity between the graphic semantic vector and the embedding vector of the scene semantic element, and then use the cosine similarity as the initial membership of the scene semantic element corresponding to the narrative segment; constructing the scene semantic mutual correlation matrix between narrative segments according to the mutual correlation between each scene semantic feature can be achieved in the following way, namely: for every two scene semantic features, the Pearson correlation coefficient between the two scene semantic features is used as a descriptive indicator of the mutual correlation between the two scene semantic features, and then the Pearson correlation coefficients of the scene semantic features between all narrative segments are arranged in a matrix form to obtain a symmetric matrix, and the matrix is ​​used as the scene semantic mutual correlation matrix between narrative segments, wherein each element of the matrix represents the degree of correlation between the scene semantics of the two segments; based on the Gaussian membership function, the elements in the mutual correlation matrix are combined into a matrix. The numerical mapping of the value into a semantic fuzzy set can be achieved in the following way, namely: selecting a Gaussian membership function as a fuzzy mapping tool, the Gaussian membership function can output the corresponding membership according to the degree of proximity between the value and the center value. In this embodiment, each value in the mutual correlation matrix can be regarded as a similarity measure between the semantic features of a certain scene, and the Gaussian membership function is input in turn to obtain the fuzzy membership corresponding to each value. All the memberships are further recombined according to the original matrix structure to form a fuzzy set that describes the strength of the relationship between the semantic features of each scene; the fuzzy association matrix of the scene semantics is constructed by all the initial memberships and the semantic fuzzy set. This can be achieved in the following way: the initial membership of the scene semantic elements corresponding to each narrative segment can be used as the basic expression to represent the initial association strength between the segment and the scene semantic elements; the fuzzy relationship between the semantic features of each scene is then extracted from the semantic fuzzy set to measure the similarity between different scene semantic elements; and a weighted superposition strategy is adopted to fuse the initial membership with the fuzzy membership in the fuzzy set to obtain the fuzzy expression of the narrative segment in each semantic dimension; finally, the fuzzy expression of each narrative segment is organized into a matrix structure according to the order of the narrative segments, and the matrix structure is used as the fuzzy association matrix of the scene semantics.

[0035] In addition, in a specific implementation, a fuzzy evaluation is performed on the expression of each narrative segment in the dimension of scene semantic elements according to the fuzzy association matrix, and the element membership of the scene semantics in each narrative segment can be obtained in the following manner, namely: first, a fuzzy association matrix is ​​obtained, wherein each row of the matrix corresponds to a narrative segment, and each column corresponds to a scene semantic element, and the numerical value in the matrix represents the fuzzy expression intensity of the narrative segment in the corresponding semantic element; second, a fuzzy comprehensive evaluation is performed on each row vector, namely, by introducing the importance weight of the semantic element, a weighted fusion is performed on the fuzzy expression of the narrative segment in all semantic element dimensions, and the weight can be set according to the frequency of occurrence of the scene semantic element in the entire film; then, a fuzzy normalization method is used to map the fusion result to the interval [0, 1] to ensure that the semantic expression between different narrative segments has a unified scale, wherein the maximum and minimum normalization method can be used for processing; finally, the normalized result is recorded, and the normalized result is used as the element membership of the scene semantics in each narrative segment.

[0036] In some embodiments, determining the scene attention of different narrative segments in the target documentary based on all element memberships can be achieved by using the following steps: Obtain the scene semantic elements of each narrative segment in the target documentary, and then determine the information entropy of the scene semantic elements of each narrative segment; The scene attention of different narrative segments in the target documentary is determined by the element membership of the scene semantics in each narrative segment and the corresponding information entropy.

[0037] It should be noted that the scene attention in this application refers to the weight value used to measure the relative importance of each narrative segment in the overall semantic expression of the documentary. Its role is to guide the processing priority of different segments in the subsequent color adjustment process.

[0038] In specific implementation, the scene semantic elements of each narrative segment in the target documentary are obtained, and then the information entropy of the scene semantic elements of each narrative segment is determined. This can be achieved in the following way: for each narrative segment, all scene semantic elements of the narrative segment are obtained, and the Shannon entropy of all scene semantic elements is calculated, and the calculated Shannon entropy is used as the information entropy of the scene semantics of the narrative segment, and then the information entropy of the scene semantics of each narrative segment is obtained. The information entropy can be used to measure the uncertainty of its distribution in the documentary as a whole; the scene attention of different narrative segments in the target documentary is determined by the element membership of the scene semantics in each narrative segment and the corresponding information entropy. This can be achieved in the following way: the element membership of each narrative segment and the information entropy of the corresponding semantic element are weighted, and the weighted result is normalized, and the result obtained by the normalization is used as the scene attention value of the narrative segment, so as to distinguish the semantic importance of each narrative segment.

[0039] In step 104, the chromaticity reconstruction cost of the target documentary during the color grading process is determined based on the scene attention of different narrative segments and the hue constraint conditions, and then the target documentary is automatically color-graded based on the chromaticity reconstruction cost. The target documentary is automatically color-graded based on all chromaticity reconstruction costs.

[0040] In some embodiments, determining the chroma reconstruction cost of the target documentary during the color grading process based on the scene attention of different narrative segments and the hue constraint condition can be achieved by using the following steps: determining the tone costs of adjacent recorded image frames in the target documentary according to the tone constraint condition; Assign different scene attention weights to the hue costs of adjacent recorded image frames based on scene attention of different narrative segments; The chroma reconstruction cost of the target documentary during the color grading process is determined by all hue costs and the corresponding scene attention weights.

[0041] It should be noted that the hue cost in this application refers to a measurement value used to measure the degree of difference between adjacent image frames at the hue level; the chroma reconstruction cost in this application refers to the total cost of evaluating the overall chroma adjustment effect based on the comprehensive hue difference and the degree of semantic attention, which can represent the optimization cost of the color adjustment scheme for visual coherence.

[0042] In specific implementation, the hue cost of adjacent recorded image frames in the target documentary is determined according to the hue constraint condition, which can be achieved in the following manner, namely: initialize a hue cost model, the cost model is constructed based on the mean square error, and is used to measure the hue difference between image frames, extract the limit parameter of the hue adjustment between adjacent recorded image frames from the hue constraint condition, and use the limit parameter as the weight adjustment factor in the cost model to regulate the evaluation sensitivity of the hue difference, and then uniformly convert the adjacent image frames into a color space with higher perceptual consistency (such as Lab color space) to enhance the discernibility of hue difference, further expand each frame of the image into a vector at the pixel level, extract the components related to hue (such as the a channel and b channel in the Lab space), calculate and square the difference between the corresponding pixels of the two image frames in the hue component, and finally perform a hue calculation on all pixel points. The square differences of the two images are averaged to obtain the mean square error value of the pair of image frames, and the mean square error value is used as the hue cost of the corresponding adjacent recorded image frames; it should be noted that the weight adjustment factor in this application is a parameter introduced in the cost model based on the mean square error, which is used to adjust the weight sensitivity of the hue difference in the cost evaluation. Its essential function is to dynamically control the tolerance of hue changes between image frames in different scenes. In technical implementation, the weight adjustment factor is extracted based on the hue constraint condition and used as the adjustment coefficient in the mean square error calculation process to weight the square value of the hue difference of pixels between image frames, so that the hue change can be moderately amplified or suppressed while meeting the visual consistency, thereby achieving a balance between flexible evaluation of hue cost and style maintenance. The introduction of the weight adjustment factor ensures the adaptability of chroma adjustment to content semantics and scene requirements.

[0043] In addition, in a specific implementation, different scene attention weights are assigned to the hue costs of adjacent recorded image frames based on the scene attention of different narrative segments. This can be achieved in the following way, namely: for each pair of adjacent recorded image frames, the scene attention of their respective corresponding narrative segments is obtained, and the average of these two scene attentions is calculated, and the average value of the two scene attentions is used as the scene attention weight of the hue cost of the pair of adjacent recorded image frames. Specifically, when two adjacent recorded image frames are in the same narrative segment, the scene attention of the narrative segment is used as the scene attention weight of the hue cost of the pair of adjacent recorded image frames; when two adjacent recorded image frames are in different narrative segments, the scene attention of the narrative segment is used as the scene attention weight of the hue cost of the pair of adjacent recorded image frames. The average value of the corresponding scene attention of the two narrative clips is used as the scene attention weight of the hue cost of the pair of adjacent recorded image frames; the chromaticity reconstruction cost of the target documentary in the color grading process is determined by all the hue costs and the corresponding scene attention weights. This can be achieved in the following way, namely: multiplying the hue cost of each pair of image frames by the corresponding scene attention weight to obtain a weighted hue cost, which is used to reflect the influence intensity of the image frame pair in the overall chromaticity reconstruction process; then summing up the weighted hue costs of all image frame pairs, and optionally performing normalization to eliminate the influence of the frame number difference; finally, the obtained sum is used as the chromaticity reconstruction cost of the target documentary under the current color grading conditions.

[0044] It should be noted that the automatic color adjustment of the target documentary based on the chromaticity reconstruction cost in this application refers to automatically adjusting the hue parameters of each image frame of the documentary to improve the overall color consistency and visual coherence under the optimization goal of minimizing the chromaticity reconstruction cost.

[0045] On the other hand, in some embodiments, the present application provides a documentary automatic color grading system based on deep learning, referring to Figure 4 , which is a schematic diagram of the structure of a documentary automatic color grading system based on deep learning according to some embodiments of the present application. The documentary automatic color grading system based on deep learning 400 includes: an extraction module 401, a processing module 402, and an execution module 403, which are described as follows: Extraction module 401, in this application, is mainly used to receive the target documentary to be color-graded, and then extract the shallow visual features of each frame of the target documentary and the scene semantic features of different narrative segments based on the pre-trained deep learning model; Processing module 402, in this application, is used to determine a correlation relationship of tonal consistency between different recorded image frames in a target documentary based on the style similarity of shallow visual features between adjacent recorded image frames, and further determine a tonal constraint condition of the target documentary in a visual constraint based on the correlation relationship and the brightness difference between adjacent recorded image frames; The processing module 402 in the present application is further configured to perform fuzzy association on the scene semantic elements of each narrative segment in the target documentary through the mutual correlation between the semantic features of each scene, thereby obtaining the element membership of the scene semantics in each narrative segment, and then determining the scene attention of different narrative segments in the target documentary based on all the element memberships; Execution module 403, in this application, execution module 403 is mainly used to determine the chromaticity reconstruction cost of the target documentary during the color grading process based on the scene attention of different narrative segments and the hue constraint conditions, and then automatically perform color grading on the target documentary based on the chromaticity reconstruction cost.

[0046] In addition, the present application also provides a computer device, which includes a memory and a processor, the memory storing a code, and the processor being configured to obtain the code and execute the above-mentioned deep learning-based documentary automatic color grading method.

[0047] In some embodiments, reference Figure 5 , which is a schematic diagram of the structure of a computer device for implementing a method for automatically adjusting the color of a documentary based on deep learning according to some embodiments of the present application. The method for automatically adjusting the color of a documentary based on deep learning in the above embodiment can be achieved by Figure 5 The computer device 500 shown in FIG. 5 is implemented as shown in FIG. 5 . The computer device 500 includes at least one processor 501 , a communication bus 502 , a memory 503 , and at least one communication interface 504 .

[0048] The processor 501 may be a general-purpose central processing unit (CPU) or an application-specific integrated circuit (ASIC).

[0049] The communication bus 502 may be used to transmit information between the aforementioned components.

[0050] The memory 503 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, a random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, an optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. The memory 503 may be independent and connected to the processor 501 via the communication bus 502. The memory 503 may also be integrated with the processor 501.

[0051] Memory 503 is used to store program code for executing the solution of the present application, and is controlled by processor 501 for execution. Processor 501 is used to execute the program code stored in memory 503. The program code may include one or more software modules. The deep learning-based documentary automatic color grading method in the above embodiment can be implemented by processor 501 and one or more software modules in the program code in memory 503.

[0052] The communication interface 504 uses any device such as a transceiver to communicate with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area network (WLAN), etc.

[0053] In a specific implementation, as an example, a computer device may include multiple processors, each of which may be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. A processor herein may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0054] The aforementioned computer device can be a general-purpose computer device or a dedicated computer device. In a specific implementation, the computer device can be a desktop computer, a portable computer, a network server, a personal digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, a communication device, or an embedded device. The embodiments of this application do not limit the type of computer device.

[0055] In addition, the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned deep learning-based documentary automatic color grading method.

[0056] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0057] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A documentary automatic color grading method based on deep learning, characterized in that: The steps include: Receive the target documentary to be color graded, and then extract the shallow visual features of each frame of the target documentary and the scene semantic features of different narrative segments based on the pre-trained deep learning model; Determining a correlation relationship of tonal consistency between different recorded image frames in a target documentary based on the style similarity of shallow visual features between adjacent recorded image frames, and then determining a tonal constraint condition of the target documentary in a visual constraint based on the correlation relationship and the brightness difference between adjacent recorded image frames; The scene semantic elements of each narrative segment in the target documentary are fuzzily associated through the mutual correlation between the semantic features of each scene, and the element membership of the scene semantics in each narrative segment is obtained. Then, the scene attention of different narrative segments in the target documentary is determined based on all the element memberships. The chroma reconstruction cost of the target documentary during the color grading process is determined according to the scene attention of different narrative segments and the color hue constraint condition, and then the target documentary is automatically color-graded based on the chroma reconstruction cost.

2. The method according to claim 1, wherein The pre-trained deep learning model is used to extract the shallow visual features of each frame of the target documentary and the scene semantic features of different narrative segments, including: Set shallow feature modules and deep feature modules in the pre-trained deep learning model; The shallow visual features of the recorded image frames in the target documentary are extracted frame by frame through the shallow feature module; Divide the target documentary into narrative segments to obtain multiple narrative segments; The scene semantics of each narrative segment are labelled by the deep feature module to obtain the scene semantic labels of each narrative segment; The scene semantic features of different narrative segments are determined based on the global semantic vector of the target documentary and the scene semantic labels of each narrative segment.

3. The method according to claim 1, wherein The correlation relationship of the tonal consistency between different recorded image frames in the target documentary is determined based on the style similarity of the shallow visual features between adjacent recorded image frames. Specifically, the following are included: Determining the style similarity of shallow visual features between adjacent recorded image frames; Determine the tonal constraint matrix of the target documentary based on all the style similarities; The tone constraint matrix is ​​used to perform correlation analysis on the tone consistency between different recorded image frames in the target documentary, so as to obtain the correlation relationship of the tone consistency between different recorded image frames in the target documentary.

4. The method according to claim 1, wherein The hue constraint conditions of the target documentary in the visual constraint are determined based on the association relationship and the brightness difference between adjacent recorded image frames, specifically including: Performing a differential operation on the energy of the brightness channels between adjacent recorded image frames to obtain the brightness difference between the adjacent recorded image frames, and then determining a brightness difference map between the adjacent recorded image frames; Determine a correlation map of tones between adjacent recorded image frames according to the correlation relationship; The color tone stability between adjacent recorded image frames is constrained by analyzing the difference map and the correlation map, and the color tone constraint conditions of the target documentary in the visual constraint are obtained.

5. The method according to claim 1, wherein The scene semantic elements of each narrative segment in the target documentary are fuzzily associated through the mutual correlation between the semantic features of each scene, and the element membership of the scene semantics in each narrative segment is obtained, which specifically includes: Initialize the initial membership of each narrative segment corresponding to the scene semantic elements; According to the mutual correlation between the semantic features of each scene, the mutual correlation matrix of scene semantics between narrative segments is constructed; Mapping the values ​​in the mutual correlation matrix into semantic fuzzy sets based on Gaussian membership function; Constructing a fuzzy association matrix of scene semantics through all initial memberships and the semantic fuzzy set; The expression of each narrative segment in the dimension of scene semantic elements is fuzzily evaluated according to the fuzzy association matrix to obtain the element membership of the scene semantics in each narrative segment.

6. The method according to claim 1, wherein The scene attention of different narrative segments in the target documentary is determined based on the affiliation of all elements, specifically including: Obtain the scene semantic elements of each narrative segment in the target documentary, and then determine the information entropy of the scene semantic elements of each narrative segment; The scene attention of different narrative segments in the target documentary is determined by the element membership of the scene semantics in each narrative segment and the corresponding information entropy.

7. The method according to claim 1, wherein The target documentary to be color graded is received through the file input interface.

8. A documentary automatic color grading system based on deep learning, characterized by: include: An extraction module, which receives the target documentary to be color-graded and extracts shallow visual features of each frame of the target documentary and scene semantic features of different narrative segments based on a pre-trained deep learning model; a processing module for determining, based on the style similarity of shallow visual features between adjacent recorded image frames, a correlation relationship of tonal consistency between different recorded image frames in a target documentary, and further determining, based on the correlation relationship and the brightness difference between adjacent recorded image frames, a tonal constraint condition of the target documentary in a visual constraint; The processing module is further configured to perform fuzzy association on the scene semantic elements of each narrative segment in the target documentary based on the mutual correlation between the semantic features of each scene, thereby obtaining the element membership of the scene semantics in each narrative segment, and then determining the scene attention of different narrative segments in the target documentary based on all the element memberships; An execution module is used to determine the chromaticity reconstruction cost of the target documentary during the color grading process based on the scene attention of different narrative segments and the color tone constraint condition, and then automatically perform color grading on the target documentary based on the chromaticity reconstruction cost.

9. A computer device comprising a memory and a processor, wherein the memory stores a code, wherein: The processor is configured to obtain the code and execute the deep learning-based documentary automatic color grading method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the deep learning-based automatic color grading method for documentaries is implemented.

Citation Information

Cited By

  • Multi-agent inquiry method and system fusing fuzzy semantics and focus image features

    CN121885166A