Intelligent transcoding and color consistency calibration method supporting multi-format film and television materials

By using synchronous dual-path analysis of dynamic color space perception and mapping core, and integrating metadata and visual features to generate quantitative parameters, intelligent transcoding and color consistency calibration of multi-format film and television materials are achieved. This solves the problems of long processing cycles and difficulty in ensuring consistency in existing technologies, and improves the automation processing capability.

CN121865007APending Publication Date: 2026-04-14CHANGCHUN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-12
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

In existing film and television production, the transcoding and color calibration processes for multi-format video materials are separated and rely on manual operation, resulting in lengthy processing cycles and difficulty in ensuring consistent results, especially when metadata is missing or the scene is complex and the effect is unstable.

Method used

By constructing a dynamic color space perception and mapping core, synchronous dual-path analysis is performed, metadata and visual features are integrated to generate quantitative parameters, a color mapping model is dynamically synthesized, and color consistency calibration is achieved in a single encoding process.

Benefits of technology

It significantly reduces processing time, increases automation and throughput, reduces reliance on human experience, and ensures the stability and consistency of color calibration, meeting the demands of modern high-efficiency, high-quality content production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121865007A_ABST
    Figure CN121865007A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of digital video signal processing, and particularly discloses an intelligent transcoding and color consistency calibration method supporting multi-format film and television materials. The method comprises the following steps: decoding a source video, inputting the decoded source video into a dynamic color space perception and mapping core for synchronous double-path analysis, analyzing color metadata in one path, and extracting color statistical features in the other path through unsupervised visual feature analysis; fusing the two through an adaptive decision module to determine a real color attribute; dynamically synthesizing a color mapping model based on a target color space specified by a user and a calibration intention, and applying the color mapping model to pixel data conversion; meanwhile, the influence of conversion on image characteristics is analyzed, and a coding guidance strategy is generated; and finally, completing compression of the converted data by using the strategy in a single coding process, and outputting videos with uniform formats and consistent colors. According to the invention, deep integration of transcoding and color calibration is realized, the processing efficiency and the automation level are improved, and the color coherence of batch processing results is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital video signal processing technology, and more specifically, to a method for intelligent transcoding and color consistency calibration that supports multi-format film and television materials. Background Technology

[0002] In film and television production, streaming services, and digital content creation, processing multi-format video footage from various camera models, mobile devices, and computers has become commonplace. These materials typically employ different encoding formats and color science systems, necessitating standardized transcoding and visual color consistency calibration before post-production or cross-platform distribution. This requirement is fundamental for achieving efficient workflows and ensuring the final visual quality of the finished product, and thus has significant practical implications.

[0003] Current technical solutions typically treat transcoding and color calibration as two independent and sequential steps. First, transcoding tools convert various compressed formats to intermediate or target editing formats. This process primarily focuses on bitrate, resolution, and encoding efficiency, lacking in-depth analysis and processing of embedded or implicit color information. Subsequently, post-production software relies on manual operation, using experience to perform color space conversion, basic correction, and matching for each segment of footage, or with the assistance of limited automatic matching plugins. The main drawbacks of this approach are: firstly, the fragmented workflow leads to lengthy processing cycles, failing to meet the demands of rapid production; secondly, color calibration is highly dependent on operator skills and subjective judgment, making it difficult to guarantee consistent results during batch processing; and thirdly, existing automated tools typically rely on limited metadata or simple image statistics for matching, resulting in unstable calibration effects when metadata is missing, erroneous, or the scene is complex, often requiring extensive manual correction. The entire process lacks intelligent perception of the true color attributes of the footage and task-oriented automated mapping capabilities.

[0004] Therefore, this paper proposes an intelligent transcoding and color consistency calibration method that supports multi-format film and television materials to address the above problems. The technical problem to be solved is: how to design an integrated method that can intelligently understand and unify the color performance of materials from different sources during the video transcoding process, overcome the strong dependence on human experience, improve processing efficiency and batch result consistency, and adapt to the needs of modern high-efficiency and high-quality content production. Summary of the Invention

[0005] In order to overcome the above-mentioned deficiencies of the prior art, embodiments of the present invention provide an intelligent transcoding and color consistency calibration method for supporting multi-format film and television materials, so as to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a method for intelligent transcoding and color consistency calibration of multi-format film and television materials, comprising the following steps: S1 decodes the input source video with different encoding formats and color space attributes to obtain the decoded raw pixel data; S2, the raw pixel data is input into a dynamic color space perception and mapping core, which performs synchronous dual-path analysis: the first path extracts and parses structured color metadata from the video bitstream, the color metadata including color primary coordinates, white point and photoelectric conversion function identifier; the second path performs unsupervised visual feature parsing on the raw pixel data to extract scene-specific color statistical features, the parsing including calculating the multidimensional histogram distribution of the image sequence in a specific color space; S3, through an adaptive decision module, compares and fuses the color metadata with the color statistical features, and generates a set of quantified parameters to describe the true color attributes of the source video; S4. Based on the user-specified target color space and high-order calibration intent, and combined with the real color attributes described by the quantization parameters, a task-oriented color mapping model is dynamically synthesized within the core. The model is embodied as a programmable shader function or three-dimensional lookup table data. S5, apply the color mapping model to convert the original pixel data to the target color space to generate intermediate pixel data. At the same time, analyze the changes in image characteristics caused by the model to generate quantization parameter adjustment strategies and bitrate allocation suggestions to guide encoding. S6. During a single encoding process, the encoder parameters are configured according to the quantization parameter adjustment strategy and bitrate allocation suggestion, and the intermediate pixel data is compressed and encoded to output a target video with a uniform format that meets the high-order calibration intent.

[0007] Preferably, in step S3, the step of the adaptive decision module comparing and fusing color metadata with color statistical features specifically includes: Calculate the deviation between the actual image characteristics reflected by the color statistical features and the standard characteristics claimed by the color metadata; When the deviation is below the first threshold, the true color attribute is determined based on the color metadata; when the deviation is above the second threshold, the true color attribute is inferred based on the color statistical features and through a preset reverse mapping rule library; when the deviation is between the first and second thresholds, the two are weighted and fused to determine the true color attribute.

[0008] Preferably, in step S2, the unsupervised visual feature parsing of the original pixel data further includes: Representative keyframes are extracted from image sequences by scene boundary detection based on sliding window and frame difference analysis. For each keyframe, calculate its joint histogram in the color space formed by its luminance component Y and chrominance components Cb and Cr. An unsupervised color grouping process is performed on the joint histogram to identify at least one dominant color set, and the center value, distribution variance and spatial distribution density of each dominant color set in the image are calculated. The scene-specific color statistical features consist of the central value, distribution variance, and spatial distribution density of the dominant color set.

[0009] Preferably, in step S4, when the higher-order calibration intent is stylistic uniformity and the user provides a reference image, the step of dynamically synthesizing a task-oriented color mapping model includes: Perform the synchronous dual-path analysis on the reference image to obtain its reference real color attributes and reference color statistical characteristics; Within a selected, perceptually uniform color space, the source video and the center of the dominant color set of the reference image are matched and corresponded respectively. Based on the matching correspondence, calculate the color transformation matrix required to migrate each dominant color set of the source video to the corresponding color set of the reference image; The color transformation matrix is ​​synthesized, and a smoothing constraint is applied to optimize the color transition, forming the final color mapping model.

[0010] Preferably, in step S4, when the higher-order calibration intent is dynamic range intelligent remapping, the step of dynamically synthesizing a task-oriented color mapping model includes: The absolute brightness range of the source video is determined based on the photoelectric conversion function identifier in the true color attributes. The difference between the absolute brightness range and the standard brightness range of the target color space is compared to determine the mapping compression ratio; Based on the pixel ratio of highlight and shadow areas and texture complexity in the color statistical features, the basis functions of multiple preset tone mapping curves are adaptively selected or fused. The generated adaptive tone mapping curve is concatenated with the standard color space transformation matrix to form the color mapping model.

[0011] Preferably, in step S5, the generation of the quantization parameter adjustment strategy for guiding the encoding specifically includes: The temporal complexity of the image sequence is evaluated after applying the color mapping model, and the change is measured by calculating the average energy of the inter-frame residuals; If the inter-frame residual energy increases, the generation strategy guides the encoder to reduce the quantization step size of the predicted frame or shorten the GOP length. Simultaneously, the local texture complexity changes in the image spatial domain are evaluated, and a bitrate weight map related to spatial location is generated to guide the encoder to perform non-uniform bitrate allocation.

[0012] Preferably, the method further includes an iterative optimization step S7 for the color mapping model: S7.1 For the target video output from the initial processing, after decoding, extract the visual features of its key frames as the actual output features; S7.2, compare the actual output characteristics with the expected output characteristics derived from the higher-order calibration intention, and calculate the characteristic difference; S7.3, This feature difference is used as a feedback signal to adjust the fusion weights in the adaptive decision module or the parameters of the inverse mapping rule base.

[0013] Preferably, when processing an editing project consisting of multiple source videos, the method further includes a project-level color coherence constraint step S8: S8.1 maintains a project color context and records the color statistical characteristics of the last frame of the processed video segment; S8.2, When processing the current source video, read the end feature of the most recent segment in the project color context; S8.3 When dynamically synthesizing the color mapping model of the current video, a constraint is introduced to ensure that the color features of the starting frame of the current video smoothly transition with the ending features of the most recent segment after mapping. S8.4 After processing is complete, update the features of the current video's last frame to the project's color context.

[0014] Preferably, the deviation is calculated based on the difference between the luminance and chromaticity distribution moments in the color statistical features and the ideal distribution moments defined by the color metadata; the weights in the weighted fusion algorithm are dynamically calculated based on the embedded confidence labels of the color metadata or the semantic classification results of the image content.

[0015] Preferably, the dynamic color space perception and mapping core is implemented as a microservice unit that can be processed in parallel; the single encoding process supports H.264, H.265, AV1 encoding standards and ProRes, DNxHD encoding formats; the method provides services through RESTful API or SDK, receiving a task description containing the source video address, target parameters, and calibration intent descriptor.

[0016] The technical effects and advantages of this invention are as follows: Compared to existing technologies that separate transcoding and color calibration in a serial workflow, this invention constructs a dynamic color space perception and mapping core, simultaneously completing color analysis and mapping within a single decode-encode pipeline. This core performs synchronous dual-path analysis on the decoded pixel data, fusing objective metadata from the bitstream with visual statistical features from the image content to accurately infer the true color attributes of the material. Subsequently, based on the user-defined target color space and high-order calibration intent, a targeted color mapping model is dynamically synthesized and applied, while simultaneously generating strategies to guide encoding optimization. This approach deeply integrates two traditionally independent steps, eliminating multiple read / write operations and processing of intermediate files, significantly shortening the overall time from raw material to usable finished product, and improving the automation level and processing throughput of the workflow.

[0017] Compared to existing color calibration methods that rely on manual experience or simple metadata matching, this invention addresses complex situations by introducing an adaptive decision-making module and an unsupervised visual feature parsing mechanism. This method not only parses standard color metadata but also extracts contextualized color statistical features through joint histogram analysis of image sequences and identification of dominant color sets. When metadata is unreliable, the system can inversely infer color attributes based on visual features using a pre-set rule base and generate a reliable mapping foundation accordingly. When matching reference styles or performing dynamic range remapping, the system dynamically synthesizes a nonlinear transformation model based on quantified color set relationships or brightness distribution characteristics. This process reduces reliance on operator expertise, improves robustness and automation in handling abnormal or missing metadata, and makes color calibration results more objective and stable.

[0018] Compared to existing technologies that struggle to guarantee color consistency across multiple clips or long time sequences, this invention ensures global consistency by introducing project-level color coherence constraints and iterative optimization mechanisms. When processing a series of clips, the system maintains a continuously updated project color context, recording the ending features of processed content. This context serves as a constraint influencing the mapping model of the starting frame of subsequent clips, ensuring smooth visual transitions. Furthermore, the system compares the initial processing results with the expected target and uses the differences as feedback signals to adaptively optimize core decision parameters. This closed-loop design allows the system to not only focus on the processing quality of individual clips but also manage color narrative from a project-wide perspective, effectively ensuring overall coherence and unity in color style and tonal transitions in the final product. Attached Figure Description

[0019] Figure 1 This is an overall flowchart of the method of the present invention.

[0020] Figure 2 This is a flowchart illustrating the core processing steps of the present invention.

[0021] Figure 3 This is a flowchart illustrating the optimization and extension of the present invention. Detailed Implementation

[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] Example 1 As attached Figures 1 to 3 The method shown supports intelligent transcoding and color consistency calibration of multi-format film and television materials. This method deeply integrates the traditionally separate steps of video format conversion and professional color calibration by constructing an integrated processing pipeline. At the core of this pipeline is a dynamic color space perception and mapping engine, which is responsible for intelligently inferring the color attributes of the material, color mapping for specific visual targets, and co-optimizing with the encoding process in a single processing flow.

[0024] The entire process begins with the decoding of the input source video. The input source video can cover various encoding formats such as H.264, H.265, AV1, ProRes, and DNxHD, and may contain different color gamuts such as Rec.709, Rec.2020, and DCI-P3, as well as various photoelectric conversion characteristics such as S-Log, Log-C, HLG, and PQ. The decoding step removes the compression layer, extracting the original pixel data sequence and the structured metadata embedded in the bitstream. This metadata typically conforms to SMPTE, ITU-R, or specific manufacturer specifications, and includes key parameters such as primary color coordinates, white point coordinates, and photoelectric conversion function identifiers.

[0025] The decoded raw pixel data and metadata are simultaneously fed into the dynamic color space perception and mapping engine. This engine executes two analysis paths in parallel. The first path parses and verifies the extracted structured color metadata. The second path performs unsupervised visual feature parsing on the raw pixel data itself, including but not limited to: scene boundary detection and adaptive keyframe extraction by calculating the inter-frame histogram Bach distance; and converting the pixel data to... A color space is constructed and a three-dimensional joint histogram is generated. A density-based peak search algorithm is used to identify the dominant color set in the image. The center coordinates, distribution covariance, and spatial distribution density of each dominant color set are calculated to form a set of quantified scene-specific color statistical features.

[0026] Subsequently, an adaptive decision-making module compares and fuses the outputs of the two paths. This module calculates the quantified deviation between visual statistical features and the metadata claim standard. The deviation can be calculated based on the differences between the mean, variance, skewness, and other higher-order moments of the brightness distribution and the standard model. The decision logic is graded according to preset thresholds: when the deviation is below the first threshold, the metadata is accepted; when the deviation is above the second threshold, the metadata is deemed unreliable, and color attributes are inferred from visual features through an inverse mapping rule base; when the deviation is between the two thresholds, the two are weighted and fused. The fusion weights can be dynamically adjusted based on the confidence labels embedded in the metadata or the scene semantic information obtained through lightweight image classification. The module ultimately outputs a set of quantified parameters that accurately describe the true color attributes of the source video, including its actual color space model and photoelectric conversion characteristics.

[0027] Based on the user-specified target color space and higher-order calibration intent, combined with the aforementioned determined real color attributes, the engine dynamically synthesizes a task-oriented color mapping model. This calibration intent can be defined as a specific visual objective such as "stylization uniformity," "dynamic range intelligent remapping," or "project-level color narrative continuity." The specific method of model synthesis depends on the intent. For the purpose of "stylization uniformity", the system performs the same analysis on the reference image provided by the user to obtain its color features. Then, it establishes the correspondence between the dominant color set of the source video and the reference image in a perceptually uniform color space. A global color transformation matrix is ​​calculated by solving a weighted least squares problem, and edge-aware smoothing constraints are integrated to avoid color gradation breaks.

[0028] For the "Dynamic Range Intelligent Remapping" intent, the system determines the absolute brightness range of the source video based on the real color attributes, analyzes the pixel ratio and texture complexity of the highlight and shadow areas of the image, adaptively selects or fuses a content-aware tone mapping curve from the preset basis function library, and cascades it with the standard color space transformation matrix.

[0029] After applying the color mapping model to transform the original pixel data to the target color space, the engine proactively analyzes the impact of this transformation on the temporal and spatial statistical properties of the image. Specifically, it assesses the increase in temporal complexity by calculating the change in inter-frame residual energy; and assesses the redistribution of spatial complexity by calculating the change in texture gradients of each macroblock. Based on this analysis, specific coding guidance strategies are generated, including but not limited to: adjusting the image group length, correcting the quantization parameters of the predicted frames, and generating spatially non-uniform bitrate allocation weight masks.

[0030] Finally, during a single encoding pass, the encoder receives pixel data converted to the target color space and configures its internal parameters according to the aforementioned encoding guidance strategy. During rate-distortion optimization, the encoder references a bitrate weight mask, allocating more bitrate resources to regions with increased complexity. This optimized encoding outputs a final video file that meets both the target format and bitrate requirements, while achieving high color consistency and adherence to intended intent.

[0031] The following is a more detailed implementation method for the above process: Furthermore, the calculation and decision-making process of the deviation in the adaptive decision-making module can be specifically described as follows: The system pre-sets a standard color space parameter knowledge base. For keyframes extracted from the visual analysis path, its global brightness histogram is calculated, and the average brightness is derived from it. , highlight points (e.g., the 99th percentile) and shaded areas (For example, the 1st percentile). Simultaneously, query the knowledge base for the ideal value under the corresponding standard based on the metadata. , , Deviation Quantified by the following formula: ; in , , Preset weighting coefficients. Set a low threshold. With high threshold .like If the metadata description is true color attribute, then the metadata description should be adopted; if If this is not the case, the inverse mapping rule base is activated. This rule base contains heuristic rules that infer possible photoelectric conversion functions and color matrices from specific visual statistical patterns (such as low average brightness, narrow dynamic range, and specific chromaticity distribution). Then the metadata inference results Results of visual feature inference Perform weighted fusion, fusion result The weight and Negative correlation, for example , This is an adjustment factor. This quantitative decision-making mechanism ensures that even when metadata is incomplete, erroneous, or inconsistent with the image content, the system can still reliably infer the true color science parameters, laying the foundation for subsequent accurate color conversion.

[0032] Furthermore, the extraction process of the dominant color set in the unsupervised visual feature parsing is as follows: First, the Bach distance of the histograms in the YUV color space between consecutive frames is calculated using a sliding window. ,when When the threshold is exceeded by the sequence content, it is determined to be a scene change, and the frames before and after the change point are selected as keyframes; for long shots, sampling is performed at fixed time intervals.

[0033] For each keyframe, convert its pixel values ​​from RGB to Color space, and construct a 3D histogram. Its three dimensions correspond to brightness respectively. and chromaticity , A density-based peak search algorithm is employed: traversing each cell of the histogram. Calculate its neighborhood Density within , will those It is a local maximum and exceeds the global average density. A certain proportion (e.g.) The cells are labeled as candidate peak points. Centered on these peak points, similar neighboring units in the color space are aggregated using an iterative region growing method to form several dominant color sets. For each set Calculate its color center With covariance matrix Simultaneously, the coordinate distribution of pixels within the set on the original image plane is analyzed, and their spatial clustering degree is calculated. Ultimately, this keyframe is represented as a set. This process abstracts high-dimensional pixel data into low-dimensional color semantic descriptors with clear statistical meaning, which facilitates subsequent feature matching and difference calculation.

[0034] Furthermore, when the calibration intent is "stylization uniformity," the dynamic synthesis of the color mapping model follows these steps. Let the dominant color set of the source video be... The corresponding set of reference images is First, put all and Switch to A color space is used to achieve better perceptual uniformity. Subsequently, a bipartite graph matching problem is constructed, with the goal of finding the mapping... Minimize the total sensing distance: ; in for Euclidean distance in space, weight Based on the optimal matching pair Calculate a weighted least squares problem by solving a affine transformation matrix : ; To prevent discontinuities in certain color regions caused by global linear transformation, an edge-aware smoothing filter is introduced in the post-processing stage of the model. This filter performs bilateral filtering or guided filtering on the image, applying stronger color smoothing to flat areas to eliminate potential blockiness, while maintaining filtering strength at texture edges to preserve details. Finally, the color mapping model is composed of a transformation matrix. Defined together with the smoothing filter parameters, this method achieves semantic transfer of color style from the reference image to the source video, rather than simple pixel value matching, ensuring the naturalness and harmony of the results.

[0035] Furthermore, when the calibration intent is "Dynamic Range Intelligent Remapping," the synthesis of the mapping model focuses on generating a content-aware tone mapping curve. Let the absolute luminance range of the source video be... The standard brightness range of the target color space is System analysis of highlight areas in visual features (XS represents pixels) and shadow areas (XS represents pixels) Pixel percentage , and average local texture entropy , The system has a pre-defined basis function containing multiple tone mapping curves. Libraries such as Reinhard functions, adjustable Gamma functions, and S-curves are available. The model synthesizer dynamically selects or fuses basis functions based on the scene content. One implementation involves constructing a parameterized adaptive curve. : ; Mixed weights Related to scene classification, parameters yes The function. For example, for High and In high-resolution scenarios (such as those involving complex cloud layers), the weights of the Reinhard function basis increase, and its parameters... Adjusted to compress highlights more smoothly; for High and In low-light scenes (such as flat shadows), the weights of the S-curve base are increased to effectively brighten shadows while suppressing noise. The final generated curve... With the linear transformation matrix from the source color gamut to the target color gamut Cascaded to form a complete mapping function ,in It is a pixel In application The result is obtained after adjusting for its luminance component. This process ensures both intelligent dynamic range compression and visual fidelity.

[0036] Furthermore, the generation of the encoding guidance strategy is based on predictive analysis of the content of the color-mapped image. The temporal complexity change is achieved by calculating the absolute difference between the luminance components of consecutive frames in the mapped sequence. mean To evaluate and compare with the baseline value estimated based on motion vectors before mapping. Compare.

[0037] like ,in If the sensitivity coefficient is used, the generation strategy suggests that the encoder use a shorter image group length. =Round down[ And reduce the quantization parameter of the predicted frame. , Positively correlated with the residual energy growth rate. Spatial complexity analysis calculates each macroblock. Changes in the sum of the Sobel gradient magnitudes of its internal pixels before and after mapping .based on Generate bitrate weight mask : ; in This serves as a reference value for the global average gradient change. To control the intensity coefficient, this weighted mask... This will serve as the basis for adjusting the Lagrange multiplier in encoder rate-distortion optimization, causing the bitrate allocation to tend to protect areas where detail is enhanced due to color mapping. This co-optimization ensures that coding efficiency and reconstruction quality are maintained while pursuing specific visual color effects.

[0038] Furthermore, the method may include a feedback-based iterative optimization mechanism. After completing the initial processing of a batch of materials, the system outputs the video. Sampling and decoding are performed, and unsupervised visual feature parsing is executed again to obtain the actual output features. Meanwhile, the system, based on the initial calibration intent... Derivation of expected output features Calculate feature differences The difference The feedback signal is sent back to the adaptive decision module and the color mapping model synthesizer. For the adaptive decision module, the weighting coefficients in its deviation calculation can be adjusted. Or decision threshold , For example, minimizing using stochastic gradient descent The gradients of these parameters. For the color mapping synthesizer, the weights in style matching can be adjusted. Hybrid weights in dynamic range remapping With parameter function Through this closed-loop feedback, the system can adaptively optimize its internal parameters, gradually improving the consistency between the output results and the expected goals when processing similar materials, demonstrating its ability to continuously learn and improve itself.

[0039] Furthermore, when processing editing projects containing multiple shot sequences, the system introduces project-level color coherence constraints. The system maintains a project-wide state variable called the color context. It records the color feature vector of the last frame of the most recently processed segment. When processing the current segment, the system first reads... In In the color mapping model of the synthesized current segment. In this case, a coherence penalty term is added to the original optimization objective (such as style matching error): ; in It is a feature of the starting frame of the current segment. It is a hyperparameter that controls the strength of coherence. The optimization process will simultaneously minimize the style matching error and This allows the color state at the beginning of the current segment to smoothly transition to the end of the previous segment. After processing the current segment, the features of its last frame are... Updated to This mechanism ensures the continuity and consistency of color evolution across shots over time, improving the overall narrative fluency of the automated processing output.

[0040] Furthermore, the deviation calculation and weight fusion can further incorporate high-level semantic information to enhance decision-making accuracy. The system can integrate a lightweight convolutional neural network or a scene classifier based on bag-of-words visual language to perform real-time semantic classification of the current keyframe and output scene category labels. (e.g., "outdoor sunlight", "indoor tungsten filament lamps", "night scene", "green screen"). The system has a pre-defined semantic-confidence mapping table. This table defines the prior confidence levels of various metadata fields (such as white balance and exposure index) under different scenarios. During the weighted fusion stage, the metadata inference results... weight Not only with deviation Related, and also with Modulation by multiplication: For example, for "green screen" scenes, white balance metadata... The value can be set to 0.2, indicating extremely low confidence, thus significantly reducing its weight in the fusion process. This hybrid decision-making model, which combines low-level statistical moment analysis with high-level semantic understanding, improves the system's robustness and judgment accuracy when facing complex and special shooting scenarios.

[0041] Furthermore, the entire technical solution can be deployed on a cloud-native architecture. The dynamic color space perception and mapping engine is encapsulated as an independent microservice, capable of horizontal scaling. The encoding optimization module, as another microservice, supports multiple encoding standards. The two services communicate via a high-speed RPC protocol. The system provides a unified RESTful API interface, receiving task descriptions in JSON format, which must include the source video access path, target format specification, target color space identifier, and optional calibration intent descriptor. After receiving the task, the microservice cluster coordinates the internal services to complete the processing pipeline, ultimately writing the output video to the designated storage and sending a callback notification. This architecture enables this method to be integrated into various media production workflows as a highly available and elastically scalable service.

[0042] To more clearly demonstrate the sequential collaboration of the various steps in the technical solution of this invention, the following describes the complete implementation process from task submission to result output, using a specific and simplified application example. This example assumes a common video production scenario: the creator needs to process two clips—one shot with a DSLR camera in S-Log3 mode (denoted as Source_A) and the other shot with a smartphone in Rec.709 mode (denoted as Source_B)—into H.264 MP4 format for online distribution, requiring both clips to maintain a consistent color style with a professionally color-graded reference still image (denoted as Ref_Image).

[0043] Step 1: Task Submission and Initialization 1. Users create tasks through client software (such as plugins or web interfaces). Task parameters are set as follows: Target format: H.264 High Profile, 1080p (1920x1080), constant bitrate 8Mbps.

[0044] Target color space: Rec.709 (BT.1886Gamma).

[0045] Higher-order calibration intent: style_unify (stylization unification).

[0046] Reference object: Uploaded Ref_Image.jpg.

[0047] 2. The client encapsulates the above parameters and the storage paths of the two source files into a structured JSON task description file, and sends it to the cloud processing service API gateway deployed with this embodiment of the invention via HTTPS protocol.

[0048] Step 2: Parallel Decoding and Dual-Path Analysis The cloud scheduler creates separate processing instances for Source_A and Source_B, respectively. Each instance executes: 3. Decoding: Call the corresponding decoder (such as the FFmpeg library) to output the raw pixel data (e.g., YUV4:2:010-bit sequence).

[0049] 4. Metadata extraction: Parse color metadata from file containers or SEI (Supplemental Enhancement Information).

[0050] Source_A: Parsed out, ColorPrimaries=BT.709,TransferCharacteristics=SLog3,MatrixCoefficients=BT.709.

[0051] Source_B: Parsed out, ColorPrimaries=BT.709,TransferCharacteristics=BT.1886,MatrixCoefficients=BT.709.

[0052] 5. Visual feature analysis: This involves analyzing the decoded pixel sequence.

[0053] a. Keyframe extraction: Sampling is performed at 1-second intervals, supplemented by scene detection. Assume that the keyframe of Source_A is Frame_A5 at the 5th second, and the keyframe of Source_B is Frame_B3 at the 3rd second.

[0054] b. Calculate the joint histogram and dominant set: Transform Frame_A5 and Frame_B3 to... In space, a 3D histogram is calculated, and three dominant color sets are identified through density clustering. The center of each set is recorded. .For example: Frame_A5 set 1 (sky): ; Frame_A5 set 2 (vegetation): ; Frame_B3 set 1 (face): ; c. Simultaneously perform the same analysis on Ref_Image to obtain its dominant color set center, for example, referencing the skin tone set: .

[0055] Step 3: Adaptive decision-making to determine true color attributes 6. For Source_A, calculate the deviation of its visual features (Frame_A5 is generally grayish and has low contrast) from the "SLog3" standard model. Assume the calculated deviation... System default , .because The system then enters weighted fusion mode. Based on the high confidence level of the metadata, the true color attributes of Source_A are finally confirmed as: color gamut = BT.709, photoelectric conversion function = SLog3, with additional black and white point offset values ​​finely adjusted according to visual characteristics.

[0056] 7. For Source_B, calculate the deviation of its visual features from the "Rec.709 / BT.1886" standard model. .because The system directly adopts its metadata and determines the true color attribute as standard Rec.709.

[0057] Step 4: Dynamically synthesize a color mapping model 8. Based on the intent of "stylization unification", the system uses the features of Ref_Image as the target to synthesize mapping models for Source_A and Source_B respectively.

[0058] a. Establish color correspondence: In Spatial approximation of the skin tone set of Frame_A5 and the skin tone set of Ref_Image. Matching is the primary mapping relationship. Similarly, for Source_B... Establish and The corresponding one.

[0059] b. Solve for the transformation matrix: For Source_A, based on its dominant set... Construct a weighted least squares problem using the corresponding reference set: Solving for the results Color transformation matrix This matrix maps the dark tones of SLog3 to a bright, saturated tone close to the reference.

[0060] c. Integrated smoothing processing: In Based on this, a lookup table LUT_A with edge-aware smoothing is generated as the final color mapping model for Source_A. Similarly, a transformation matrix is ​​generated for Source_B. The transformation magnitude of LUT_B is usually smaller than that of Source_A.

[0061] Step 5: Generate coding guidance strategy 9. Simulate the pixel mapping from LUT_A to Source_A. Analysis reveals: Temporal domain: The significant contrast improvement from SLog3 to Rec.709 resulted in inter-frame residual energy in some scenes. Compared to baseline It increased by approximately 40%.

[0062] Airspace: Texture details in vegetation and sky areas are highlighted by enhanced contrast, with local gradient variations. Significantly positive.

[0063] 10. Based on this, a strategy is generated: It is recommended that the encoder use a shorter GOP (e.g., 64 frames) for Source_A and reduce the QP value of the predicted frames by 2.

[0064] Generate a bitrate weight mask and label the weights in the vegetation and sky regions. The other regions are 1.0.

[0065] Step Six: Single-pass Optimization of Encoding and Output 11. Encoder microservice receives instructions: For Source_A, first apply LUT_A to convert the pixels to the target Rec.709 color space, then encode at a bitrate of 8Mbps and GOP=64, and apply a bitrate weight mask for non-uniform bitrate distribution during the encoding process. The same applies to Source_B, applying LUT_B and then encoding with standard parameters.

[0066] 12. After encoding is complete, the two output video files, Output_A.mp4 and Output_B.mp4, are uploaded to the specified storage location. The processing service sends a task completion notification to the client, along with a link to access the result files.

[0067] Step 7: Result Verification (Optional) 13. Users or the system can optionally perform rapid analysis on the output video to extract keyframe color features. Assume the analysis finds that the average skin tone value of Output_A is similar to that of Ref_Image. exist The distance in space is (If the difference is less than the perceptible difference threshold of 5), it indicates successful color matching. This feedback can be recorded and used to optimize processing parameters for similar SLog3 footage in the future.

[0068] In summary, the embodiments of the present invention construct an intelligent and automated multi-format video processing pipeline through dynamic color space perception and mapping core, adaptive decision fusion, intent-oriented model dynamic synthesis, coding collaborative optimization, and optional project-level coherence constraints and iterative feedback.

[0069] Finally, the following points should be noted: First, in the description of this application, it should be noted that, unless otherwise specified and limited, the terms "installation", "connection", and "linkage" should be interpreted broadly, and can be mechanical or electrical connections, or internal connections between two components, or direct connections. "Up", "down", "left", "right", etc. are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may change. Secondly: The accompanying drawings of the embodiments disclosed in this invention only involve the structures involved in the embodiments disclosed in this invention. Other structures can refer to the general design. In the absence of conflict, the same embodiment and different embodiments of this invention can be combined with each other. In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for intelligent transcoding and color consistency calibration of multi-format film and television materials, characterized in that, Includes the following steps: S1 decodes the input source video with different encoding formats and color space attributes to obtain the decoded raw pixel data; S2, the raw pixel data is input into a dynamic color space perception and mapping core, which performs synchronous dual-path analysis: the first path extracts and parses structured color metadata from the video bitstream, the color metadata including color primary coordinates, white point and photoelectric conversion function identifier; The second path performs unsupervised visual feature parsing on the original pixel data to extract scene-specific color statistical features. The parsing includes calculating the multidimensional histogram distribution of the image sequence in a specific color space. S3, through an adaptive decision module, compares and fuses the color metadata with the color statistical features, and generates a set of quantified parameters to describe the true color attributes of the source video; S4. Based on the user-specified target color space and high-order calibration intent, and combined with the real color attributes described by the quantization parameters, a task-oriented color mapping model is dynamically synthesized within the core. The model is embodied as a programmable shader function or three-dimensional lookup table data. S5, apply the color mapping model to convert the original pixel data to the target color space to generate intermediate pixel data. At the same time, analyze the changes in image characteristics caused by the model to generate quantization parameter adjustment strategies and bitrate allocation suggestions to guide encoding. S6. During a single encoding process, the encoder parameters are configured according to the quantization parameter adjustment strategy and bitrate allocation suggestion, and the intermediate pixel data is compressed and encoded to output a target video with a uniform format that meets the high-order calibration intent.

2. The intelligent transcoding and color consistency calibration method for supporting multi-format film and television materials according to claim 1, characterized in that, In step S3, the adaptive decision module compares and fuses color metadata with color statistical features as follows: Calculate the deviation between the actual image characteristics reflected by the color statistical features and the standard characteristics claimed by the color metadata; When the deviation is lower than the first threshold, the true color attribute is determined based on the color metadata. When the deviation is higher than the second threshold, the true color attribute is inferred based on the color statistical features by using a preset reverse mapping rule library. When the deviation is between the first and second thresholds, the two are weighted and fused to determine the true color attributes.

3. The intelligent transcoding and color consistency calibration method for supporting multi-format film and television materials according to claim 1, characterized in that, In step S2, the unsupervised visual feature parsing of the original pixel data further includes: Representative keyframes are extracted from image sequences by scene boundary detection based on sliding window and frame difference analysis. For each keyframe, calculate its joint histogram in the color space formed by its luminance component Y and chrominance components Cb and Cr. An unsupervised color grouping process is performed on the joint histogram to identify at least one dominant color set, and the center value, distribution variance and spatial distribution density of each dominant color set in the image are calculated. The scene-specific color statistical features consist of the central value, distribution variance, and spatial distribution density of the dominant color set.

4. The intelligent transcoding and color consistency calibration method for supporting multi-format film and television materials according to claim 1 or 3, characterized in that, In step S4, when the higher-order calibration intent is stylistic uniformity and the user provides a reference image, the step of dynamically synthesizing a task-oriented color mapping model includes: Perform the synchronous dual-path analysis on the reference image to obtain its reference real color attributes and reference color statistical characteristics; Within a selected color space with uniform perception, the source video is matched with the center of the dominant color set of the reference image. Based on the matching correspondence, calculate the color transformation matrix required to migrate each dominant color set of the source video to the corresponding color set of the reference image; The color transformation matrix is ​​synthesized, and a smoothing constraint is applied to optimize the color transition, forming the final color mapping model.

5. The intelligent transcoding and color consistency calibration method for supporting multi-format film and television materials according to claim 1, characterized in that, In step S4, when the higher-order calibration intent is dynamic range intelligent remapping, the step of dynamically synthesizing a task-oriented color mapping model includes: The absolute brightness range of the source video is determined based on the photoelectric conversion function identifier in the true color attributes. The difference between the absolute brightness range and the standard brightness range of the target color space is compared to determine the mapping compression ratio; Based on the pixel ratio of highlight and shadow areas and texture complexity in the color statistical features, the basis functions of multiple preset tone mapping curves are adaptively selected or fused. The generated adaptive tone mapping curve is concatenated with the standard color space transformation matrix to form the color mapping model.

6. The intelligent transcoding and color consistency calibration method for supporting multi-format film and television materials according to claim 1, characterized in that, In step S5, the generation of the quantization parameter adjustment strategy for guiding the encoding specifically includes: The temporal complexity of the image sequence is evaluated after applying the color mapping model, and the change is measured by calculating the average energy of the inter-frame residuals; If the inter-frame residual energy increases, the generation strategy guides the encoder to reduce the quantization step size of the predicted frame or shorten the GOP length. Simultaneously, the local texture complexity changes in the image spatial domain are evaluated, and a bitrate weight map related to spatial location is generated to guide the encoder to perform non-uniform bitrate allocation.

7. The intelligent transcoding and color consistency calibration method for supporting multi-format film and television materials according to claim 1, characterized in that, The method also includes an iterative optimization step S7 for the color mapping model: S7.1 For the target video output from the initial processing, after decoding, extract the visual features of its key frames as the actual output features; S7.2, compare the actual output characteristics with the expected output characteristics derived from the higher-order calibration intention, and calculate the characteristic difference; S7.3, This feature difference is used as a feedback signal to adjust the fusion weights in the adaptive decision module or the parameters of the inverse mapping rule base.

8. The intelligent transcoding and color consistency calibration method for supporting multi-format film and television materials according to claim 1, characterized in that, When processing an editing project consisting of multiple source videos, the method further includes a project-level color coherence constraint step S8: S8.1 maintains a project color context and records the color statistical characteristics of the last frame of the processed video segment; S8.2, When processing the current source video, read the end feature of the most recent segment in the project color context; S8.3 When dynamically synthesizing the color mapping model of the current video, a constraint is introduced to ensure that the color features of the starting frame of the current video smoothly transition with the ending features of the most recent segment after mapping. S8.4 After processing is complete, update the features of the current video's last frame to the project's color context.

9. The intelligent transcoding and color consistency calibration method for supporting multi-format film and television materials according to claim 2, characterized in that, The deviation is calculated based on the difference between the luminance and chromaticity distribution moments in the color statistical features and the ideal distribution moments defined by the color metadata; the weights in the weighted fusion algorithm are dynamically calculated based on the embedded confidence labels of the color metadata or the semantic classification results of the image content.

10. The intelligent transcoding and color consistency calibration method for supporting multi-format film and television materials according to claim 1, characterized in that, The dynamic color space perception and mapping core is implemented as a microservice unit that can be processed in parallel; the single encoding process supports H.264, H.265, AV1 encoding standards and ProRes, DNxHD encoding formats; the method provides services through RESTful API or SDK, and receives a task description containing the source video address, target parameters and calibration intent descriptor.