Live broadcast processing method and system, and storage medium
By constructing a dynamic metadata set and a color reference block timetable, the problem of color performance imbalance in high-definition/ultra-high-definition simultaneous broadcast scenarios was solved, realizing real-time color synchronization and efficient color adjustment of multi-format video signals, adapting to the color processing requirements of 4K/8K ultra-high-definition resolution, and improving the quality and efficiency of live broadcasting.
Patent Information
- Application Number
- CN202511806223.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-02
- Publication Date
- 2026-02-10
AI Technical Summary
In high-definition/ultra-high-definition simultaneous broadcast scenarios, existing technologies struggle to dynamically adapt to real-time changes in live content. Manual color adjustment is inefficient and cannot meet the color processing requirements of 4K/8K ultra-high-definition resolutions, leading to color imbalances and visual differences that negatively impact the viewing experience on the terminal.
By acquiring multi-format video signals from live streams, a dynamic metadata set is constructed, and a target dynamic color reference block timetable is generated. Based on this timetable, the video signals are processed to achieve color signal fusion output. A dynamic color reference block timetable and a seamless switching caching mechanism are adopted to adapt to real-time changes in live stream content.
It achieves real-time and accurate color synchronization of multi-format video signals, reduces latency, improves color grading efficiency, adapts to the color processing requirements under 4K/8K ultra-high-definition resolution, and enhances the quality and real-time performance of live streaming content.
Smart Images

Figure CN121509746A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of live streaming data processing technology, and in particular to a live streaming processing method, system, and storage medium. Background Technology
[0002] In high-definition / ultra-high-definition simultaneous broadcast scenarios, static 3D lookup table (3DLUT) templates are often used to achieve different output modes. Due to the significant differences in color reproduction between different modes, the color performance of the simultaneous broadcast content becomes unbalanced, especially in scenarios such as skin tone reproduction and highlight details, where visual differences are prominent. Therefore, it is difficult to dynamically adapt to the real-time changes of the live broadcast content through the above methods, while manual color adjustment is inefficient and cannot meet the color processing requirements under 4K / 8K ultra-high-definition resolution. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a live streaming processing method, system, and storage medium to solve the problems in the prior art that make it difficult to dynamically adapt to real-time changes in live streaming content, and that manual color adjustment is inefficient and unable to meet the color processing requirements under 4K / 8K ultra-high-definition resolution.
[0004] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:
[0005] The first aspect illustrates a live streaming processing method, the method comprising:
[0006] During the live stream, acquire multi-format video signals from each video frame in the live video;
[0007] The video signals in multiple formats are parsed to obtain a dynamic metadata set;
[0008] Based on the dynamic metadata set, the video frame images of the video signal corresponding to the signal type are processed to generate a target dynamic color reference block timetable;
[0009] Based on the target dynamic color reference block timetable, the video signals of different formats are processed to obtain the fused output color signal, which is then output.
[0010] Optionally, the video signals in multiple formats are parsed to obtain a dynamic metadata set, including:
[0011] Identify video signals of multiple formats and determine the signal type of each video signal;
[0012] For each type of video signal, the video signal is subjected to header separation processing to obtain header data;
[0013] Extract metadata of the video signal from the header data;
[0014] The metadata is converted into metadata in a preset format;
[0015] The metadata of different signal types is combined to obtain a dynamic metadata set.
[0016] Optionally, based on the dynamic metadata set, the video frame images of the video signal corresponding to the signal type are processed to generate a target dynamic color reference block timetable, including:
[0017] The color-coded YUV values are determined according to different signal types to match the standard dynamic range and color gamut standards of the corresponding signals.
[0018] According to the standard dynamic range and color gamut standards corresponding to different signal types, the video frame images of the video signals corresponding to the signal types are cropped to obtain video frame images cropped for different signal types;
[0019] For each signal type, the video frame is cropped and then divided into blocks according to the dynamic metadata set to determine the histogram of each block under different signal types.
[0020] For each block, a target dynamic color reference timetable is constructed based on the histogram corresponding to the block and the corresponding metadata in the dynamic metadata set.
[0021] Optionally, for each signal type, the video frame after cropping is divided into blocks according to the dynamic metadata set to determine the histogram of each block under different signal types, including:
[0022] For each signal type, the video frame of the video signal corresponding to the signal type is divided into blocks according to the block characteristics of the metadata in the dynamic metadata set, resulting in multiple blocks and corresponding block IDs;
[0023] For each block of each signal type, a corresponding histogram is constructed based on the bars in the block, and the number of histograms is 3;
[0024] The histogram corresponding to the block is optimized based on the block segmentation characteristics of the metadata corresponding to the block in the dynamic metadata set.
[0025] Optionally, for each block, a target dynamic color reference timetable is constructed based on the histogram corresponding to the block and the corresponding metadata, including:
[0026] For each block, the histogram is associated with the corresponding metadata within a preset time period to obtain a feature time series.
[0027] The pre-set color parameter prediction model is invoked to process the feature time series to obtain the color parameter prediction value of each block within a preset time period. The color parameter prediction model is trained based on historical feature time series.
[0028] A target dynamic color reference timetable is constructed based on the predicted values of the color parameters within a preset time period for each block.
[0029] Optionally, a target dynamic color reference timetable is constructed based on the predicted color parameter values within a preset time period for each block, including:
[0030] The predicted color parameters of each block within a preset time period are mapped to an initial dynamic color reference timetable according to a preset timestamp.
[0031] A real-time block feature table is generated based on the block features of each block obtained from real-time sampling, and associated with the initial dynamic color reference time table through the block ID;
[0032] Based on the preset feature threshold library and the real-time block feature table, determine whether the parameters of each block in the initial dynamic color reference time table need to be calibrated;
[0033] If necessary, the block can be used as the block to trigger calibration;
[0034] The blocks that trigger calibration are calibrated based on the real-time block feature table, and the initial dynamic color reference time table is adjusted to obtain the target dynamic color reference time table.
[0035] If none of these are required, the initial dynamic color reference timetable shall be used as the target dynamic color reference timetable.
[0036] Optionally, the video signals of different formats are processed based on the target dynamic color reference block timetable to obtain a fused output color signal, including:
[0037] Get the target color parameters of the target dynamic color reference timetable of the block at the current timestamp and the target dynamic color reference timetable of the previous timestamp at the same location;
[0038] Based on the real-time block feature table, calculate the parameter difference of the target color parameter of the block at the same location at the current timestamp and the previous timestamp;
[0039] For each block, transition parameters are determined based on the parameter differences;
[0040] The transition parameters corresponding to each block are matched with the blocks of the current video frame according to the block ID of each block.
[0041] The actual parameters of the previous video frame are obtained based on the block ID, and the parameters of the current video frame are determined by processing the actual parameters of the previous video frame and the transition parameters corresponding to the block ID.
[0042] A transition lookup table (LUT) is generated based on the current video frame parameters and the new lookup table (LUT) parameters, wherein the new lookup table (LUT) parameters are the difference table parameters that match the color conversion of the current video frame;
[0043] The target dynamic color reference timetable and the transition lookup table (LUT) are fused together to obtain the fused output color signal.
[0044] Optional, also includes:
[0045] Based on the LUT (Learning Undefined Table), the corresponding transition curve is generated from the encoding matrix produced by rotational encoding fusion.
[0046] The color difference of the fused output color signal is calculated by judging the transition curve, the new color reference timetable and the old color reference timetable. The old color reference timetable refers to the target dynamic color reference timetable, and the new color reference timetable refers to the color reference timetable obtained by combining the target dynamic color reference timetable with the transition lookup table (LUT).
[0047] If the color difference of the fused output color signal is greater than the preset color difference, then return to the step of processing the video frame of the video signal corresponding to the signal type based on the dynamic metadata set to generate the target dynamic color reference block timetable.
[0048] The second aspect illustrates a live streaming processing system, which includes a global static color processing module, a loading controller, and a non-dynamic encoding processing module.
[0049] The global static color processing module is used to acquire multi-format video signals of each video frame in the live video during the live broadcast; parse the video signals of multiple formats to obtain a dynamic metadata set; and process the video frame images of the video signals corresponding to the signal types based on the dynamic metadata set to generate a target dynamic color reference block timetable.
[0050] The loading controller and non-dynamic encoding processing module are used to process the video signals of different formats based on the target dynamic color reference block timetable to obtain the fused output color signal and output it.
[0051] The third aspect discloses a storage medium comprising a stored program, wherein, when the program is executed, it controls the device on which the storage medium resides to perform the live streaming processing method as described in any one of claims 1-8.
[0052] Based on the live streaming processing method, system, and storage medium provided in the above embodiments of the present invention, for multi-format video signals acquired during live streaming, the method first parses them to construct a dynamic metadata set; then, for different signal types, the video frames of different format video signals are processed sequentially to obtain a target dynamic color reference block timetable constructed from the feature parameters of different video signals; then, combined with a seamless switching buffer mechanism, the multi-format video signals are fused and encoded based on the target dynamic color reference block timetable to obtain the fused output color signal; this covers the conversion requirements of multi-format video signals such as SDR, HLG, or PQ, thereby adapting to the real-time changes of live streaming content, improving color grading efficiency, and being able to cope with the color processing requirements under 4K / 8K ultra-high-definition resolution. Attached Figure Description
[0053] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0054] Figure 1 This is a schematic diagram of the structure of a live streaming processing system according to an embodiment of the present invention;
[0055] Figure 2 This is a schematic flowchart illustrating a live streaming processing method according to an embodiment of the present invention;
[0056] Figure 3 This is a schematic diagram illustrating the specific architecture of live streaming processing according to an embodiment of the present invention. Detailed Implementation
[0057] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0058] As the background technology indicates, in high-definition / ultra-high-definition simultaneous broadcast scenarios, the commonly used static 3D lookup table (3DLUT) template cannot adapt to dynamic scene switching in live broadcasts, resulting in significant color differences between HDR / SDR outputs. This means that color consistency calibration between Standard Dynamic Range (SDR) and HDR signals presents a significant challenge. Furthermore, fully loading the 3DLUT can easily cause stuttering and high switching latency, posing a significant challenge to the real-time performance of live broadcasts. In such cases, frequent manual color adjustments are required, leading to high reliance on manual processes and low efficiency in live broadcast processing. In addition, the SDR to HDR conversion of overlaid elements such as subtitles and images can easily result in color style fragmentation, failing to preserve the original artistic intent and severely impacting the viewing experience on the terminal.
[0059] See Figure 1 This is a schematic diagram of the structure of a live streaming processing system according to an embodiment of the present invention. The system includes a global static color processing module 10, a loading controller 20, and a non-dynamic encoding processing module 30.
[0060] Among them, the global static color processing module 10, the loading controller 20 and the non-dynamic encoding processing module 30 are interconnected; the loading controller 20 has a built-in dynamic encoding and transcoding switching module.
[0061] The global static color processing module 10 is used to acquire multi-format video signals of each video frame in the live video during the live broadcast; parse the video signals of multiple formats to obtain a dynamic metadata set; process the video frame images of the video signals corresponding to the signal types based on the dynamic metadata set to generate a target dynamic color reference block timetable; and load the controller 20 and the non-dynamic encoding processing module 30 to process the video signals of different formats based on the target dynamic color reference block timetable to obtain a fused output color signal and output it.
[0062] Optionally, the global static color processing module 10, which parses video signals of multiple formats to obtain a dynamic metadata set, is specifically used for:
[0063] Multiple video signal formats are identified to determine the signal type of each video signal; for each signal type, the video signal is subjected to header separation processing to obtain header data; metadata of the video signal is extracted from the header data; the metadata is converted into metadata of a preset format; and the metadata of different signal types is combined to obtain a dynamic metadata set.
[0064] Optionally, the global static color processing module 10, which processes the video frames of the video signal corresponding to the signal type based on the dynamic metadata set to generate the target dynamic color reference block timetable, is specifically used for:
[0065] Based on different signal types, the color-coded YUV values are matched to the standard dynamic range and color gamut standards of the corresponding signals. Video frames of the corresponding video signals are cropped according to the standard dynamic range and color gamut standards for different signal types to obtain cropped video frames for different signal types. For each cropped video frame, the video frame is divided into blocks according to the dynamic metadata set to determine the histogram of each block under different signal types. For the histogram corresponding to each block, a target dynamic color reference timetable is constructed based on the histogram corresponding to the block and the corresponding metadata in the dynamic metadata set.
[0066] Specifically, for each signal type, the video frame after cropping is segmented according to the dynamic metadata set to determine the histogram of each block under different signal types, including:
[0067] For each signal type, the video frame of the video signal corresponding to the signal type is divided into blocks according to the block segmentation characteristics of the metadata in the dynamic metadata set, resulting in multiple blocks and corresponding block IDs; for each block of each signal type, a corresponding histogram is constructed based on the bars in the block, and the number of histograms is 3; the histogram corresponding to the block is optimized based on the block segmentation characteristics of the metadata corresponding to the block in the dynamic metadata set.
[0068] Specifically, for each block's corresponding histogram, a target dynamic color reference time table is constructed based on the histogram and corresponding metadata of the block, including:
[0069] For each block, the histogram is associated with the corresponding metadata within a preset time period to obtain a feature time series. A pre-set color parameter prediction model is called to process the feature time series to obtain the color parameter prediction value within a preset time period for each block. The color parameter prediction model is trained based on historical feature time series. A target dynamic color reference time table is constructed based on the color parameter prediction value within the preset time period for each block.
[0070] The process of constructing a target dynamic color reference timetable based on the predicted color parameter values within a preset time period for each block includes:
[0071] The predicted color parameters of each block within a preset time period are mapped to an initial dynamic color reference timetable according to a preset timestamp; a real-time block feature table is generated based on the block features of each block obtained from real-time sampling, and associated with the initial dynamic color reference timetable through the block ID; it is determined whether the parameters of each block in the initial dynamic color reference timetable need to be calibrated according to a preset feature threshold library and the real-time block feature table; if so, the block is used as the trigger calibration block; the trigger calibration block is calibrated based on the real-time block feature table, and the initial dynamic color reference timetable is adjusted to obtain the target dynamic color reference timetable; if neither is required, the initial dynamic color reference timetable is used as the target dynamic color reference timetable.
[0072] Optionally, the controller 20 and the non-dynamic encoding processing module 30 are loaded, specifically for:
[0073] The loading controller 20 is used to obtain the target color parameters of the target dynamic color reference timetable of the block at the current timestamp and the target dynamic color reference timetable of the previous timestamp for the same location; calculate the parameter difference of the target color parameters of the block at the current timestamp and the previous timestamp based on the real-time block feature table; for each block, determine the transition parameter based on the parameter difference; match the corresponding transition parameter of each block with the block of the current video frame according to the block ID; obtain the actual block parameter of the previous video frame according to the block ID, and process it based on the actual block parameter of the previous video frame and the transition parameter corresponding to the block ID to determine the parameters of the current video frame; generate a transition lookup table (LUT) based on the parameters of the current video frame and the loaded new lookup table (LUT) parameters.
[0074] The non-dynamic encoding processing module 30 is used to perform fusion processing based on the target dynamic color reference timetable and the transition lookup table (LUT) to obtain the fused output color signal.
[0075] Optionally, the non-dynamic encoding processing module 30 is also used for:
[0076] Based on the target dynamic color reference timetable, the encoding matrix produced by rotational encoding fusion generates a corresponding transition curve; the color difference of the fused output color signal is calculated by judging the transition curve, the new transition lookup table (LUT), and the old transition lookup table (LUT); if the color difference of the fused output color signal is greater than the preset color difference, then the video frame of the video signal corresponding to the signal type is processed based on the dynamic metadata set to generate the target dynamic color reference block timetable.
[0077] In this embodiment of the invention, a high-definition / ultra-high-definition simultaneous broadcast color management method based on dynamic color reference blocks is provided. This method aims to address issues such as significant color differences in HDR / SDR output during real-time conversion of multi-format signals (SDR / HLG / PQ) in live broadcast scenarios; high switching latency caused by full loading of 3DLUTs leading to stuttering, posing a significant challenge to live broadcast real-time performance; and heavy reliance on manual intervention. By establishing a dynamic color reference block time table parameter prediction model, a seamless 3DLUT caching mechanism, and through manual intervention, the method ultimately achieves accurate and low-latency synchronized color output for multi-format signals, supporting the industrial application of 4K / HDR ultra-high-definition live broadcasting.
[0078] See Figure 2 The above is a flowchart illustrating a video processing method according to an embodiment of the present invention. The method includes:
[0079] Step S201: During the live broadcast, acquire the multi-format video signal of each video frame in the live video;
[0080] It should be noted that multi-format video signals include high-definition SDR signals, 4K HLG signals, and 4K PQ signals.
[0081] In the specific implementation step S201, in the high-definition / ultra-high-definition simultaneous broadcast scenario, the global static color processing module synchronously receives the high-definition SDR signal, 4K HLG signal and 4K PQ signal produced by each video frame during the live broadcast.
[0082] Step S202: Parse the video signals in multiple formats to obtain a dynamic metadata set;
[0083] It should be noted that the specific implementation of step S202 includes the following steps.
[0084] Step S11: Identify multiple video signal formats and determine the signal type of each video signal.
[0085] In the specific implementation step S11, each video signal is identified by the signal channel ID to distinguish the signal type, thereby obtaining the signal type of each video signal. Specifically, firstly, according to the SMPTEST 2086 standard, the signal type is distinguished by a 1-byte binary identifier: "0x00" for high-definition SDR signals, "0x01" for 4K HLG signals, and "0x02" for 4K PQ signals. That is, for each video signal, if its video signal contains the byte 0x00, it indicates that its corresponding signal type is high-definition SDR; similarly, if its video signal contains the byte 0x01, it indicates that its corresponding signal type is 4K HLG; and if its video signal contains the byte 0x02, it indicates that its corresponding signal type is 4K PQ.
[0086] Step S12: For each type of video signal, perform head separation processing on the video signal to obtain head data.
[0087] The header data includes the GOP header data.
[0088] In the specific implementation of step S12, the GOP header decoding and parsing process of different signals is clearly distinguished by packet.belong_to(signal_channel). Specifically, for each type of video signal, the Group of Pictures (GOP) header data in the video encoding of the video signal is located by the Payload Type field of the Real-time Transport Protocol (RTP) packet.
[0089] It should be noted that if the video signal type is a high-definition SDR signal, the length of the GOP header is fixed at 128 bytes due to the low encoding bitrate (usually 8-15Mbps). In this case, the complete header data is determined and extracted by the packet.payload_length field, that is, the GOP header data of the high-definition SDR signal is located by the Payload Type field of the Real-time Transport Protocol (RTP) packet.
[0090] If the video signal type is 4K HLG or 4K PQ, because the signal contains wide color gamut and high dynamic range parameters, the length of the GOP header is extended to 256 bytes. Therefore, the complete header data is determined and extracted by the packet.payload_length field, that is, the GOP header data of the 4K HLG / PQ signal is located by the Payload Type field of the Real-Time Transport Protocol (RTP) packet.
[0091] The header data refers to the GOP header data, which specifically includes the start position, length, and end position of the GOP header.
[0092] Step S13: Extract the metadata of the video signal from the header data;
[0093] It should be noted that the signal types include HD SDR type, which is the signal type corresponding to HD SDR signal; 4K HLG type, which is the signal type corresponding to 4K HLG signal; and 4K PQ type, which is the signal type corresponding to 4K PQ signal.
[0094] In the specific implementation of step S13, if it is a high-definition SDR type, dynamic range parameter parsing is not required, and fixed fields are directly read. That is, fixed fields are directly read from the header data of the corresponding video signal and used as metadata. If it is a 4K HLG type, the intra-frame luminance statistics are associated with the construction verification mode in the header data of the corresponding video signal and used as metadata.
[0095] If it is a 4K PQ type, the brightness range parameter of the PQ curve is extracted from the header data of the corresponding video signal and used as metadata. This is for subsequent verification of the integrity of dynamic metadata.
[0096] It should be noted that the verification mode refers to the 32nd byte of the header data being the HLG flag bit, with 0x01 indicating that it is enabled.
[0097] The brightness range parameter of the PQ curve is the brightness interval field corresponding to bytes 40 to 48 in the header data;
[0098] The metadata includes dynamic range identifiers, resolution markers, peak brightness (nit), color coverage (Rec.2020 percentage), dynamic range grading data, and chunking features. Here, chunking features refer to the chunking features of the GOP header.
[0099] If it is a high-definition SDR type, dynamic range parameter parsing is not required, and fixed fields are read directly. That is, fixed fields are read directly from the header data of the corresponding video signal and used as metadata, including:
[0100] First, the high-definition SDR type is matched with the type in the 3DLUT library, and the dynamic range identifier in the 3DLUT library that matches the high-definition SDR type is used as the dynamic range identifier of the high-definition SDR signal corresponding to that high-definition SDR type.
[0101] At the same time, it is necessary to ensure that the SDR signal calls the Rec.709 3DLUT library;
[0102] Next, the pixel size and scanning method of the video signal are extracted from the header data and used as resolution markers so that they can be pushed to the dynamic encoding and transcoding switching module to achieve adaptive switching between 4K signal ultra-high-definition block processing and SDR signal conventional block processing.
[0103] It should be noted that the resolution specification of a high-definition SDR signal, i.e., the pixel size and scanning method, is generally 1920×1080i50.
[0104] Then, the peak brightness of the video signal is extracted from the header data so that it can be used as a basis for determining the block granularity.
[0105] Next, the signal color gamut of the video signal is extracted from the header data, and then the overlap ratio between the signal color gamut and the Rec.2020 standard is calculated using the CIE1931 chromaticity diagram to obtain the color coverage.
[0106] It should be noted that the Rec.2020 standard is based on pre-set parameters. Generally speaking, the color coverage of a high-definition SDR signal can be 35%.
[0107] Next, the header data is classified according to the preset dynamic range to obtain dynamic range classified data;
[0108] It should be noted that the preset brightness dynamic range refers to the preset maximum brightness / preset minimum brightness. The brightness dynamic range of the high-definition SDR signal can be the preset maximum brightness 100 / preset minimum brightness 1.
[0109] The dynamic range classification data includes three levels of data: the first level data is data that is greater than or equal to the preset maximum brightness, the second level data is data that is less than the preset maximum brightness but greater than the preset minimum brightness, and the third level data is data that is less than or equal to the preset minimum brightness.
[0110] The first, second, and third levels decrease in rank from left to right.
[0111] It can be used to provide feature weights for the prediction model (the higher the level, the greater the prediction weight of brightness change).
[0112] Finally, the block features of the video signal are extracted from the header data.
[0113] The segmentation features include motion intensity, brightness uniformity, and color complexity;
[0114] Specifically, based on the pixels displayed in the header data, the difference between two adjacent pixels is calculated and used as the motion intensity. Next, the luminance Y channel data in the YUV color space is extracted from the header data, and the standard deviation of the Y channel data is calculated to obtain the luminance uniformity, i.e., luminance uniformity is calculated based on the Y channel standard deviation. Finally, the two chroma channels in the YUV color space are extracted from the U / V channels in the header data, and the standard deviation of the chroma U / V channels is calculated to obtain the color complexity, i.e., color complexity is calculated based on the U / V channel variance.
[0115] To support adaptive image segmentation: when the motion intensity is greater than the first value, such as 60, the corresponding area is divided into 8×8 blocks to capture details; when the color complexity is greater than the second value, such as 50, the corresponding area is enabled for color detail enhancement.
[0116] Among them, the value range of motion intensity can be 0 to 100; the value range of brightness uniformity can be 0 to 50; and the value range of color complexity can be 0 to 80.
[0117] If it is a 4K HLG or 4K PQ type, the process of determining the metadata includes:
[0118] First, the 4K HLG type or 4K PQ type is matched with the type in the 3DLUT library, and the dynamic range identifier in the 3DLUT library that matches the 4K HLG type or 4K PQ type is used as the dynamic range identifier of the corresponding type of signal.
[0119] It is also necessary to ensure that the 4K HLG signal or 4K PQ signal calls the Rec.2020 3DLUT library;
[0120] Next, the pixel size and scanning method of the video signal are extracted from the header data and used as resolution markers;
[0121] It should be noted that the resolution specification for 4K HLG or 4K PQ signals, i.e., pixel size and scanning mode, is generally 3840×2160p50.
[0122] Then, the peak brightness corresponding to the corresponding signal is found in the header data, and the peak brightness of the video signal is extracted so that it can be used as the basis for determining the block granularity in the future.
[0123] It should be noted that the peak brightness of 4K HLG signals can be 1000 to 2000 nits, and the peak brightness of 4K PQ signals can be 1000 to 4000 nits.
[0124] Next, the signal color gamut of the video signal is extracted from the header data, and then the overlap ratio between the signal color gamut and the Rec.2020 standard is calculated using the CIE1931 chromaticity diagram to obtain the color coverage.
[0125] It should be noted that the Rec.2020 standard is based on pre-set parameters. Generally speaking, the color coverage of 4K HLG signals is 90%-95%; the color coverage of 4K PQ signals is 95%-100%.
[0126] Next, the header data is classified according to the preset dynamic range to obtain dynamic range classified data;
[0127] It should be noted that the preset brightness dynamic range refers to the preset maximum brightness / preset minimum brightness. The brightness dynamic range of the 4K HLG signal can be the preset maximum brightness of 2000 / preset minimum brightness of 1; the brightness dynamic range of the 4K PQ signal can be the preset maximum brightness of 8000 / preset minimum brightness of 1.
[0128] Finally, the block features of the video signal are extracted from the header data.
[0129] The segmentation features include motion intensity, brightness uniformity, and color complexity;
[0130] It should be noted that the specific implementation process of dynamic range hierarchical data and block features for 4K HLG or 4K PQ types is the same as the implementation process of the high-definition SDR type described above, and they can be referred to each other.
[0131] Step S14: Convert the metadata into metadata in a preset format;
[0132] In the specific implementation of step S14, for the metadata of each signal type, the extracted metadata is converted into a preset format that can be recognized by the global static color processing module, but the exclusive fields of the metadata of different signal types need to be retained to support subsequent color conversion.
[0133] In other words, for high-definition SDR type, the rec709_gamma field needs to be added during conversion, meaning the format corresponding to the rec709_gamma field remains unchanged; for 4K HLG type, the hlg_oetf_flag field is additionally carried, meaning the format corresponding to the hlg_oetf_flag field remains unchanged; for 4K PQ type, the 4K PQ signal additionally carries the pq_max_luminance field, meaning the format corresponding to the pq_max_luminance field remains unchanged.
[0134] The function rec709_gamma field has a fixed value of 2.2.
[0135] It should be noted that the preset format is set in advance based on multiple experiments, and can generally be set to JSON format;
[0136] Step S15: Combine the metadata of different signal types to obtain a dynamic metadata set.
[0137] Step S203: Process the video frame images of the video signal corresponding to the signal type based on the dynamic metadata set to generate a target dynamic color reference block timetable;
[0138] It should be noted that the specific implementation of step S203 includes the following steps.
[0139] Step S21: Determine the standard dynamic range and color gamut standard of the corresponding signal based on the color encoding YUV values according to different signal types;
[0140] In the specific implementation step S21, for the high-definition SDR type, the differential interval parameter is called to determine the standard dynamic range of luminance Y in the color-coded YUV value as [0, 255]; (this standard dynamic range is based on 8-bit quantization), the standard dynamic range of chrominance U in the color-coded YUV value is [-128, 127]; the standard dynamic range of chrominance V in the color-coded YUV value is [-128, 127], and there is no luminance compression, and its corresponding color gamut standard is adapted to the Rec.709 narrow color gamut;
[0141] For the 4K HLG type, the standard dynamic range of luminance Y in the color-coded YUV value is determined by calling the differential interval parameter as [16, 940]; (this standard dynamic range is based on 10-bit quantization, corresponding to 0~1000 nit), the standard dynamic range of chrominance U in the color-coded YUV value is [-512, 511]; the standard dynamic range of chrominance V in the color-coded YUV value is [-512, 511], and there is no luminance compression. Its corresponding color gamut standard is adapted to the Rec.2020 wide color gamut.
[0142] For the 4K PQ type, the standard dynamic range of luminance Y in the color-coded YUV value is determined by calling the differential interval parameter as [64, 940]; (this standard dynamic range is based on 10-bit quantization, corresponding to 0.005~4000 nit), the standard dynamic range of chrominance U in the color-coded YUV value is [-512, 511]; the standard dynamic range of chrominance V in the color-coded YUV value is [-512, 511], matching the luminance range of the PQ perceptual quantization curve and avoiding noise interference.
[0143] The limit_yuv_range function performs range clipping to ensure that all sampled pixel values fall within the valid range without overflow or distortion.
[0144] Step S22: Crop the video frame images of the video signal corresponding to the signal type according to the standard dynamic range and color gamut standards corresponding to different signal types to obtain video frame images after cropping for different signal types;
[0145] In the specific implementation of step S22, for the video frame of the video signal corresponding to each signal type, firstly, the video frame is cropped according to the standard dynamic range and color gamut standard set in step 21 by the limit_yuv_range function to ensure that all sampled pixel values fall within the valid range without overflow or distortion, thereby obtaining the video frame after cropping for different signal types.
[0146] Step S23: For each signal type of cropped video frame, the video frame is divided into blocks according to the dynamic metadata set to determine the histogram of each block under different signal types.
[0147] It should be noted that the specific implementation of step S22 includes the following steps.
[0148] Step S31: For each signal type, the video frame of the video signal corresponding to the signal type is divided into blocks according to the block characteristics of the metadata in the dynamic metadata set, to obtain multiple blocks and corresponding block IDs.
[0149] In the specific implementation step S31, for the metadata corresponding to each signal type in the dynamic metadata set, if the signal type is a high-definition SDR type, the video frame of the high-definition SDR signal corresponding to the high-definition SDR type is divided into 32×32 pixel blocks according to the block characteristics in the metadata, thereby obtaining multiple blocks; and the block IDs are separated in order, and then the high-definition SDR type signal_type is bound to each block;
[0150] If the signal type is 4K HLG or 4K PQ, the video frame of the video signal corresponding to the 4K HLG or 4K PQ type is divided into 16×16 pixel blocks according to the block characteristics in the metadata, so as to obtain multiple blocks. The block IDs are separated in order, and then the signal type signal_type is bound to each block.
[0151] Optionally, a corresponding block ID can be set for each block.
[0152] Step S32: For each block of each signal type, construct the corresponding histogram based on the bars in the block;
[0153] It should be noted that each block has three histograms, specifically three independent histograms for the Y, U, and V channels.
[0154] Specifically, first, different sampling frequencies are determined according to different signal types, and then, the bars in the block are sampled according to different sampling frequencies.
[0155] Different types of video signals are acquired at different sampling frequencies. For high-definition SDR, the blocks are acquired at a sampling rate of 1 pixel / 4×4 grid (i.e., a sampling rate of 25%). For 4K HLG or 4K PQ, the blocks are acquired at a sampling rate of 1 pixel / 2×2 grid (i.e., a sampling rate of 100%) to balance efficiency and accuracy.
[0156] If it is a high-definition SDR type, the sampling frequency corresponding to the signal type is used to collect 256 bins of the Y channel, and the number of bins of the U / V channels is matched according to the signal quantization bit number; and the sampling pixel frequency in each bin is counted. Based on the pixel frequency corresponding to each channel of luminance Y, chrominance U, and chrominance V, the corresponding histogram is constructed, that is, the independent histograms of the three channels of luminance Y, chrominance U, and chrominance V are output.
[0157] For example: If it is a high-definition SDR type, the corresponding block output of the SDR type will be 3 groups of 256-dimensional frequency arrays. Based on the sampling pixel frequency in each bin, the output will be a Y, U, V three-channel histogram.
[0158] It should be noted that the process of constructing histograms for 4K HLG or 4K PQ types is the same as that for constructing histograms for high-definition SDR types.
[0159] Specifically, for 4K HLG type, the number of bins in the Y channel acquired is 462; for 4K HLG type, the number of bins in the Y channel acquired is 462.
[0160] Step S33: Optimize the histogram corresponding to the block based on the block segmentation characteristics of the metadata corresponding to the block in the dynamic metadata set;
[0161] In the specific implementation step S33, firstly, metadata with the same signal type as the block is determined from the dynamic metadata set, and the histogram of each block under this type is optimized based on the motion intensity in the block features under the metadata. Specifically, for the histogram of the Y channel, when it is determined that there is a block with the motion intensity greater than the first threshold, it indicates that the block belongs to the high-speed motion area. At this time, the first preset weight is obtained, and the pixel frequency in the histogram, that is, the frequency weight of the Y channel histogram, is multiplied by the first preset weight to enhance the priority of brightness statistics and adapt to the rapidly changing dynamic picture in the live broadcast scene, that is, to optimize the corresponding histogram.
[0162] For the histogram of the U / V channel, when it is determined that there is a block with a color complexity greater than the second threshold, it indicates that the block belongs to a multi-color mixing region. At this time, the second preset weight is obtained, and the pixel frequency in the histogram, i.e. the frequency weight of the U / V channel histogram, is multiplied by the second preset weight to improve the accuracy of color statistics and ensure that key colors such as skin color and stage lighting are without deviation, i.e., optimize the corresponding histogram.
[0163] The first threshold and the second threshold are set in advance based on multiple experiments. The first threshold can generally be set to 60, and the second threshold can be set to 50.
[0164] The first and second preset weights are both set in advance based on multiple experiments or experience. The first preset weight can be set to 1.2 and the second preset weight can be set to 1.5.
[0165] Step S24: For each block, construct a target dynamic color reference time table based on the histogram corresponding to the block and the corresponding metadata;
[0166] The specific implementation of step S24 includes the following steps.
[0167] Step S41: For the histogram corresponding to each block, associate the histogram and the corresponding metadata of the block within a preset time period to obtain the feature time series.
[0168] In the specific implementation step S41, for the histogram corresponding to each block, the features of the histogram corresponding to the block within a preset time period (e.g., 3 GOP periods) are extracted in chronological order, namely the YUV histogram features; then, the YUV histogram features are associated with the corresponding metadata to form the feature time series of 3 GOPs, thus obtaining the feature time series of the block.
[0169] Step S42: Call the pre-set color parameter prediction model to process the feature time series and obtain the color parameter prediction value of each block within the preset time period.
[0170] Next, the segmented feature time series is used as input to the color parameter prediction model so that the color parameter prediction model can be trained on the segmented feature time series, learn the color change rules under different scenarios, and output the color parameter prediction value of each block within a preset time period.
[0171] It should be noted that the preset time period is set in advance, and can generally be set to the color parameter prediction values within the next 1 to 3 GOP periods.
[0172] The color parameter prediction model can be a model trained on historical block feature time series based on the LSTM time series prediction algorithm.
[0173] The predicted color parameters include the target brightness Y_init, chromaticity U_init / V_init, saturation S_init, and brightness fluctuation range ΔY_range_init.
[0174] Step S43: Construct a target dynamic color reference timetable based on the predicted color parameter values within a preset time period for each block.
[0175] The specific implementation of step S43 includes the following:
[0176] Step S51: Map the predicted color parameters of each block within a preset time period to an initial dynamic color reference timetable according to the preset timestamp;
[0177] It should be noted that the preset timestamp refers to one time node corresponding to each GOP cycle.
[0178] Specifically, the predicted color parameters for each block over the next 1-3 GOP periods are mapped to an initial dynamic color reference timetable using timestamps. The parameters of each block are bound to the block ID for easy subsequent calibration association.
[0179] Next, a real-time feature acquisition thread synchronized with the GOP cycle is started to acquire the motion intensity, brightness uniformity, and color complexity of each block.
[0180] Step S52: Generate a real-time block feature table based on the block features of each block obtained from real-time sampling, and associate it with the initial dynamic color reference time table through the block ID.
[0181] Block features include motion intensity, brightness uniformity, and color complexity.
[0182] The preset sampling frequency is equal to the GOP frame rate, such as 50fps, which corresponds to sampling once every 20ms.
[0183] Specifically, firstly, for each block, the Y channel pixels and U / V channel pixels are collected in real time according to the preset sampling frequency; the average value of the difference between adjacent frames of Y channel pixels in the block is calculated and used as the motion intensity; the standard deviation σ_Y of Y channel pixels in the block is calculated; and it is substituted into formula (1) for conversion to obtain the brightness uniformity K; finally, the average value of the variance of U / V channel pixels in the block is calculated to obtain the color complexity.
[0184] Formula (1):
[0185] Among them, the average pixel difference of the Y channel can be quantized from 0 to 100 (i.e., 0 represents static and 100 represents high-speed motion); the brightness uniformity K can be quantized from 0 to 50 (i.e., 0 represents extremely large brightness difference and 50 represents complete uniformity); and the color complexity can be quantized from 0 to 80 (i.e., 0 represents a monochrome area and 80 represents a multi-color mixed area).
[0186] Finally, a real-time block feature table is generated based on the real-time motion intensity, brightness uniformity, and color complexity of each block, and then associated with the initial timetable, i.e., the initial dynamic color reference timetable, through the block ID.
[0187] Step S53: Determine whether each block in the initial dynamic color reference time table needs calibration based on the preset feature threshold library and the real-time block feature table. If it needs calibration, proceed to step S54; otherwise, proceed to step S55.
[0188] In the specific implementation of step S53, for each block, a preset feature threshold library is searched according to the signal type corresponding to the block to determine the corresponding feature threshold condition; then, it is determined whether the motion intensity, brightness uniformity and color complexity of the same block ID in the real-time block feature table meet the corresponding feature threshold conditions respectively; if any threshold condition is met, the block is used as the block to trigger calibration, and calibration is triggered, and step S54 is executed; otherwise, step S55 is executed.
[0189] For example: if the signal type corresponding to the block is 4K PQ signal type, and the color complexity of the same block ID in the block feature table is greater than 60, it means that it meets the corresponding feature threshold condition and step S62 is executed.
[0190] The signal type corresponding to the block is high-definition SDR, and the feature thresholds for motion intensity, brightness uniformity, and color complexity can be set to greater than 40. The signal type corresponding to the block is 4K HLG, and the feature thresholds for motion intensity, brightness uniformity, and color complexity can be set to greater than 45. The signal type corresponding to the block is 4K PQ, and the feature thresholds for motion intensity, brightness uniformity, and color complexity can be set to greater than 30, less than 15, and greater than 50.
[0191] Furthermore, the feature threshold conditions of color complexity greater than 50 under 4K PQ type, signal brightness uniformity less than 15 under 4K HLG type, and motion intensity greater than 60 under HD SDR type are all defined as high-priority calibration and must be executed first.
[0192] Step S54: Use the block as the trigger calibration block, calibrate the trigger calibration block based on the real-time block feature table, and adjust the initial dynamic color reference time table to obtain the target dynamic color reference time table.
[0193] In the specific implementation of step S54, the feature dimension to be adjusted is determined according to the feature threshold satisfied in step S61. If the feature dimension to be adjusted is motion intensity, the brightness Y_cal is calculated by substituting the motion intensity of the block in the block feature table into formula (2).
[0194] Formula (2):
[0195]
[0196] Wherein, k_m is the calibration coefficient, with k_m being 0.005 for SDR type, 0.006 for HLG type, and 0.007 for PQ type; threshold_m is the preset motion intensity threshold for the corresponding signal; motion_intensity is the motion intensity of the block in the real-time block feature table; and Y_init can be the preset brightness.
[0197] Then, the block that triggered the calibration is calibrated according to the adjusted brightness Y_cal, and the brightness parameter of the block in the initial dynamic color reference timetable is adjusted based on the adjusted brightness Y_cal to obtain the target dynamic color reference timetable.
[0198] For example: In the block feature table, the motion intensity of a 4K PQ type block is 40, the preset motion intensity threshold of the corresponding signal is 30, and Y_init is 500nit. Then, the brightness is adjusted to Y_cal=500×(1+0.007×10)=535nit.
[0199] If the feature dimension to be adjusted is brightness uniformity, the brightness uniformity of the block in the block feature table is substituted into formula (3) for calculation to obtain the brightness fluctuation range ΔY_range. Then, the block that triggers calibration is calibrated according to the brightness fluctuation range ΔY_range, and the brightness fluctuation range of the block in the initial dynamic color reference timetable is adjusted based on the brightness fluctuation range ΔY_range to obtain the target dynamic color reference timetable.
[0200] Formula (3):
[0201]
[0202] Where threshold_l is the corresponding signal brightness uniformity threshold, luminance_uniformity is the brightness uniformity of the block in the real-time block feature table, and ΔY_range_init is the standard brightness fluctuation range, which is preset.
[0203] For example: In the block feature table, the brightness uniformity of a 4K PQ type block is 12, the corresponding signal brightness uniformity threshold is 15, and ΔY_range_init is 200nit. Then the brightness fluctuation range ΔY_range = 200 × (1 - 0.01 × 3) = 194nit.
[0204] If the feature dimension to be adjusted is color complexity calibration, the color complexity of the block in the block feature table is substituted into formula (4) for calculation to obtain the adjusted chroma saturation S_cal. Then, the block that triggers calibration is calibrated according to the adjusted chroma saturation S_cal, and the chroma saturation of the block in the initial dynamic color reference timetable is adjusted based on the adjusted chroma saturation S_cal to obtain the target dynamic color reference timetable.
[0205] Formula (4):
[0206]
[0207] Wherein, k_c is the calibration coefficient, with k_c being 0.003 for SDR type, 0.004 for HLG type, and 0.005 for PQ type; threshold_c is the corresponding signal color complexity threshold; color_complexity is the color complexity of the block in the real-time block feature table; S_init is the preset chroma saturation; and S_cal is the adjusted chroma saturation.
[0208] For example, in the block feature table, the color complexity of a 4K PQ type block is 60, that is, threshold_c is 50, S_init is 0.8, then S_cal = 0.8 × (1 + 0.005 × 10) = 0.84.
[0209] Step S55: Use the initial dynamic color reference timetable as the target dynamic color reference timetable.
[0210] Optionally, after performing step S54 to calibrate the triggered block based on the real-time block feature table, it is also necessary to determine the effect of the calibration. Therefore, it further includes:
[0211] For the block that triggers calibration, obtain the chromaticity of the calibrated block and the chromaticity of the real-time video frame; and calculate the difference between the chromaticity of the real-time video frame and the calibrated chromaticity to obtain the color difference;
[0212] Determine whether the color difference is less than or equal to a preset color difference;
[0213] If it is less than or equal to, it means that it meets the live color consistency requirements, that is, the calibration is passed, and the initial dynamic color reference timetable is executed and adjusted to obtain the target dynamic color reference timetable.
[0214] If the value is greater than the specified value, the calibration coefficient is adjusted, and the block that triggered the calibration is recalibrated based on the real-time block feature table.
[0215] Specifically, the color difference ΔE between the corresponding blocks before and after calibration is calculated using a color algorithm, i.e., the deviation between the calibrated parameters and the real-time image color is compared. It is then determined whether the color difference is less than or equal to a preset color difference. If it is less, it indicates that it meets the live broadcast color consistency requirements, meaning the calibration is successful. The process then proceeds to update the initial dynamic color reference timetable based on parameters such as Y_cal, ΔY_range, or S_cal to obtain the target dynamic color reference timetable. If the color difference is greater than the preset color difference, the calibration coefficient is adjusted, and the process returns to step S54 to re-perform the calibration.
[0216] The process of adjusting the calibration coefficients includes: increasing the calibration coefficients (i.e., k_m, k_l, or k_c) by 10% to obtain new calibration coefficients; and re-performing calibration based on the new calibration coefficients, i.e., returning to step S54. This continues until the color difference is less than or equal to the preset color difference, ensuring that the calibrated parameters accurately match the real-time image characteristics.
[0217] It should be noted that the preset color difference is set in advance based on multiple experiments or experience, and can generally be set to 3.
[0218] Optionally, the target dynamic color reference timetable, i.e. the dynamically calibrated timetable, will be pushed to the 3DLUT loading controller and the dynamic encoding / transcoding switching module to provide a real-time reference standard for accurate color synchronization of multi-format signals.
[0219] Step S204: Process the video signals of different formats based on the target dynamic color reference block timetable to obtain the fused output color signal, and output it;
[0220] It should be noted that the specific implementation of step S204 includes the following steps:
[0221] Step S61: Obtain the target color parameters of the target dynamic color reference timetable of the block at the current timestamp and the target dynamic color reference timetable of the previous timestamp for the same location.
[0222] In the specific implementation step S61, the color reference timetable is divided into timelines according to the GOP period (20ms / GOP for 50fps live streaming). Each timestamp Tn in the timeline (n is the GOP number) is bound to the target color parameters (Y) of all blocks. n U n / V n Color gamut coefficient K n );
[0223] For blocks at the same position with adjacent timestamps, the corresponding target color parameters are obtained from the current target dynamic color reference timetable and the previous target dynamic color reference timetable, i.e., the previous round of switching results, based on the block ID corresponding to that block.
[0224] Among them, obtaining the corresponding target color parameters from the previous target dynamic color reference time table refers to The time, that is, the previous timestamp block, includes , / , ; Taken from the previous round of switching results;
[0225] Obtaining the corresponding target color parameters from the current target dynamic color reference timetable refers to... The moment, i.e., the current timestamp block. , / ,and .
[0226] Step S62: Calculate the parameter difference of the target color parameters of the blocks at the same location at the current timestamp and the previous timestamp based on the real-time block feature table.
[0227] In the specific implementation of step S62, firstly, for each block of the frame-by-frame image, the corresponding difference calculation method is determined based on the motion intensity and color complexity of the collected real-time block feature table.
[0228] If the motion intensity of the block feature table is determined to be greater than the first threshold, it indicates that the block is a high-speed moving block. Integer difference calculation is then used to calculate... Target color parameters at time and The parameter difference of the target color parameters at any given time, i.e. , , as well as It also rounds down to the nearest integer to quickly respond to changes in brightness in moving scenes and avoid motion blur.
[0229] If the motion intensity of the block feature table is determined to be less than or equal to the first threshold, it indicates that the block is a static / low-speed block. A floating-point difference operation is then used to calculate... Target color parameters at time and The parameter difference of the target color parameters at any given time, i.e. , , as well as It retains one decimal place, smoothly transitions static area parameters, and avoids flickering;
[0230] If the color complexity of the block feature table is determined to be greater than the second threshold, it indicates that the block is a multi-color block. ΔU and ΔV are then calculated independently for each channel to determine the color complexity. Target color parameters at time and The parameter difference of the target color parameters at any given time, i.e. , , as well as To accurately preserve color details;
[0231] If the color complexity of the block feature table is determined to be less than or equal to the second threshold, it indicates that the block is a monochrome block. ΔU and ΔV are then combined for calculation. Target color parameters at time and The parameter difference of the target color parameters at any given time, i.e. , , as well as Next, calculate ΔUV=(ΔU+ΔV) / 2 to simplify the calculation and improve efficiency.
[0232] The first threshold and the second threshold are set in advance based on multiple experiments. The first threshold can generally be set to 60, and the second threshold can be set to 50.
[0233] Step S63: For each block, determine the transition parameters based on the parameter difference.
[0234] In the specific implementation of step S63, for each block, the total difference of a block within a GOP period is first calculated, and the average value is calculated based on the total difference of the block, which is the frame-by-frame transition step size.
[0235] The frame-by-frame transition step size is a transition parameter.
[0236] For example, the total difference in blocks (e.g., ΔY=35nit) is decomposed into "frame-by-frame transition step size" based on the number of frames in the GOP period (e.g., 12 frames). The formula is: Step size per frame = total difference / number of transition frames.
[0237] For high-speed moving blocks (blocks with motion intensity greater than 80), the number of transition frames is halved (e.g., 12 frames become 6 frames), and the step size is doubled to balance real-time performance and smoothness.
[0238] For example: the total block difference ΔY corresponding to a 4K PQ signal is 35 nits. One GOP cycle contains 12 frames, meaning the brightness step size per frame, or transition parameter, is approximately 2.92 nits / frame. The brightness of the first frame is... +2.92, the brightness of the first frame is +2.92*2, ..., the brightness of the 12th frame is To achieve a gradual transition.
[0239] It should be noted that the number of frames contained within a GOP period is the number of transition frames.
[0240] Step S64: Match the corresponding transition parameters of each block with the blocks of the current video frame according to the block ID of each block;
[0241] In the specific implementation of step S64, the block transition step size is read frame by frame, and it is matched to all blocks of the current video frame according to the block ID.
[0242] Step S65: Obtain the actual block parameters of the previous video frame according to the block ID, and process the actual block parameters of the previous video frame and the transition parameters corresponding to the block ID to determine the parameters of the current video frame.
[0243] It should be noted that the actual parameters of the block include the brightness Y, chroma U / V, saturation, brightness fluctuation range, etc. of the block described in the previous video frame.
[0244] In the specific implementation step S65, the actual parameters of the previous video frame are obtained according to the block ID, and the transition step size is superimposed on the actual parameters of the previous video frame to obtain the parameters of each block; for example, the current video frame Y = the previous video frame Y + 2.92nit), and it is checked whether the parameters of the current video frame conform to the signal YUV interval (such as Y∈[64,940] of PQ). If it exceeds the interval, it is cropped; otherwise, it is retained.
[0245] After all block switching is completed, the parameters of each block of the current video frame are combined to obtain the current frame parameters. The current frame parameters are then pushed to the timestamp synchronization module for subsequent interpolation calculation of the old and new LUTs, ensuring that there is no switching delay or screen tearing within the frame.
[0246] Step S66: Generate a Transition Lookup Table (LUT) based on the current video frame parameters and the new LUT parameters;
[0247] In the specific implementation of step S66, firstly, the lookup table parameters matching the color conversion of the current video frame are obtained from the preset 3DLUT library built into the upconversion / downconversion 3DLUT loading controller, which supports multiple format conversions such as SDR to HLG and SDR to PQ, and these parameters are used as the new lookup table LUT parameters. Based on the current video frame parameters and the new lookup table LUT parameters, LUT interpolation calculation is performed to obtain the transition weight α. Then, based on the transition weight, the parameters in the target dynamic color reference block timetable are multiplied to obtain the transition lookup table LUT.
[0248] The transition lookup table (LUT) includes the progressive transition brightness and chroma of each block in each video frame for each signal type. Specifically, it includes block association information, progressive transition color parameters, and dynamic mapping relationships, all of which serve the purpose of "smooth transition of multi-format signals".
[0249] Among them, the block association information refers to the block ID bound to each entry, ensuring that it can be accurately matched with the specific block of the current video frame;
[0250] Progressive color parameters refer to the brightness, chroma U / V, and saturation of each block, which are transitioned frame by frame. For example, the brightness of a certain block gradually transitions from "the brightness of the previous frame + 2.92 nits" to the target value in each frame to avoid abrupt color changes.
[0251] The dynamic mapping relationship includes the color mapping rules corresponding to the transition weight α. For example, when α=0.3 in a certain frame, the color value is the parameter of the current video frame, that is, the old LUT parameter × 0.7 + the new LUT parameter × 0.3, to ensure smooth color without stuttering when switching between the old and new LUTs.
[0252] It should be noted that the transition weight changes gradually from 0 to 1.
[0253] Then, the switching instant adopts dual-thread parallel processing: a three-level cache loading technology is established. The main thread outputs the transition LUT result, and the child thread completes the full loading of the new LUT as shown in Table (1), thus achieving low switching latency.
[0254]
[0255] Step S67: Perform fusion processing based on the target dynamic color reference timetable and the transition lookup table (LUT) to obtain the fused output color signal;
[0256] In the specific implementation step 67, the target dynamic color reference timetable is fused with the video signal by rotation encoding and then shifted and rotated in real time, as shown in formula (5), to obtain the fused output color signal. In the process of formula (5), the transition lookup table LUT is used to make the fusion result of the standard dynamic range signal SDR, the mixed log-gamma high dynamic range signal HLG, and the perceptual quantization high dynamic range signal PQ, etc., smoother, so as to achieve smooth fusion of multi-format video signals and meet the requirements of live real-time performance and color consistency.
[0257] Formula (5):
[0258] .
[0259] Output is the final color signal output by fusion; SDR is the standard dynamic range signal; HLG is the hybrid log-gamma high dynamic range signal; Rotate(HLG,θ) is the rotation transformation of the HLG signal by an angle of θ; PQ is the perceptual quantization high dynamic range signal; a is the weighting coefficient of the high-definition SDR type in fusion, which can generally be 0.4~0.5; b is the weighting coefficient of the 4K HLG type in fusion, which can generally be 0.5~0.8; c is the weighting coefficient of the 4K PQ type in fusion, which can generally be 0.1~0.3.
[0260] Among them, the weight coefficients a, b, and c are preset.
[0261] Optionally, the block feature parameters required for the color reference block timetable are comprehensive parameters formed by combining the block features in the GOP header with the dynamic block features of the real-time block feature table.
[0262] The RS(255,239) forward error correction algorithm is used to repair transmission errors. Simultaneously, the FPGA hardware performs real-time comparisons of the matching degree between the GOP header block features and dynamic block features, as well as the Rec.2020 color gamut coverage parameter. If any indicator exceeds a threshold, the block feature parameter is deemed abnormal, triggering a pre-marking mechanism to ensure accurate basic parameters are provided for the color reference block timetable.
[0263] Optionally, dual-path scan conversion technology and a field-programmable gate array (FPGA) are deployed in the global static color module. The backup unit preloads multiple sets of dynamic LUT mapping tables generated by the cache, i.e., dynamic color reference timetables, to achieve low-latency fault switching through caching. A consistent three-level cache architecture is adopted, and the caching strategy is optimized using the LRU-K algorithm (where the data set D = { , …, In the context of}, "data" specifically refers to dynamically generated 3DLUT mapping table data, including color conversion parameters for different scenes, LUT indexes corresponding to block features, etc.; each data K access timestamp set ={ , ,…, After sorting by time, the elimination rule is: Eliminated data = argmin dᵢ∈D This ensures that the LUT loading required during switching achieves near real-time (low latency) performance.
[0264] Optionally, for the dynamic range remapping stage of the color reference table, the primary and backup encoding channels buffer YUV4:2:2 intermediate data in real time. When jitter occurs in the output bitstream of the primary channel, the system switches to the backup channel within 100ms using SR-IOV technology. The phase-locked loop (PLL) control logic is a mechanism used to maintain the synchronization of color parameters during remapping. Its core is to receive the CIEDE2000 color difference ΔE result from the color algorithm output and adjust the mapping parameters of the HDR / SDR dual-channel output in real time (such as the transition curve alpha and LUT switching timing) to ensure that the color deviation of the image is always controlled within the threshold (ΔE≤3) before and after the primary / backup channel switch and during the remapping process, thus achieving "phase-locked" synchronization of color parameters.
[0265] Optionally, based on the live streaming processing method shown above, this application also illustrates a specific architecture diagram of the live streaming processing system, such as... Figure 3 As shown, the loading controller includes an up-conversion 3DLUT loading controller and a down-conversion 3DLUT loading controller;
[0266] The global static color processing module is used to implement the GOP header decoding function call, dynamic metadata set generation, LSTM time series prediction model, and color reference mapping table generation processes, namely, steps S201 to S203.
[0267] The up-conversion 3DLUT loading controller and down-conversion 3DLUT loading controller, i.e., the 3DLuts library, cover the signal conversion needs of multiple formats such as SDR / HLG / PQ; it has a built-in dynamic encoding and transcoding switching module that supports intelligent scheduling based on the scene. By reading color configuration parameters and performing global color tampering preprocessing on the screen, combined with the seamless switching 3DLuts caching mechanism, the risk of full loading stuttering is reduced, thus realizing the process in step S203;
[0268] The non-dynamic encoding processing module, namely the linkage fusion encoding output end, performs full-link reconstruction and deep optimization of the color management process during multi-format signal conversion, breaks through the application limitations of traditional static adaptation in live dynamic scenarios, and achieves accurate color synchronization and low-latency output, that is, the process of step S203.
[0269] In this embodiment of the invention, for multi-format video signals acquired during live streaming, they are first parsed to construct a dynamic metadata set; then, for different signal types, the video frames of different format video signals are processed sequentially to obtain a target dynamic color reference block timetable constructed from the feature parameters of different video signals; then, a seamless switching buffer mechanism is combined to fuse and encode the multi-format video signals based on the target dynamic color reference block timetable to obtain the fused output color signal; this covers the conversion requirements of multi-format video signals such as SDR, HLG, or PQ, thereby adapting to the real-time changes of live streaming content, improving color grading efficiency, and being able to cope with the color processing requirements under 4K / 8K ultra-high-definition resolution.
[0270] Optionally, based on the live streaming processing method shown in the above embodiments of the present invention, it further includes the following steps:
[0271] Step S71: Generate the corresponding transition curve based on the encoding matrix produced by the rotation encoding fusion of the LUT;
[0272] During step S71 within the specific time frame, the encoding matrix is obtained by rotating and encoding the transition lookup table (LUT), i.e., the transition color reference time table; and the transition curve alpha corresponding to the encoding matrix is drawn so as to form a closed-loop control chain through the remapping of the transition curve (alpha).
[0273] Step S72: Determine the color difference of the fused output color signal by judging the transition curve, the new color reference timetable T, and the old color reference timetable; wherein, the old color reference timetable refers to the target dynamic color reference timetable, and the new color reference timetable T refers to the color reference timetable obtained by combining the target dynamic color reference timetable with the transition lookup table LUT, that is, the target dynamic color reference timetable is dynamically smoothed and optimized by the transition lookup table LUT to form the reference table for the current round.
[0274] In the specific implementation step S72, the current timestamp t is first substituted into formula (6) to calculate the transition coefficient alpha of the transition curve; then, according to the color algorithm, the CIEDE2000 color difference of the HDR / SDR dual-channel output is generated, that is, the transition coefficient alpha, the new color reference timetable and the old color reference timetable are substituted into formula (7) to calculate the color difference Output_ΔE of the fused output color signal.
[0275] Formula (6):
[0276]
[0277] Formula (7):
[0278]
[0279] Wherein, alpha is the transition coefficient (range 0~1), used to control the smoothness of the switch between the old and new LUTs; t is the time parameter (unit: frames), which changes dynamically with the live broadcast process; Output_ΔE is the color difference output value calculated in real time during the transition process; old_LUT is the old color reference timetable; new_LUT is the new color reference timetable (new parameter); ΔE is the internationally recognized color difference quantification index.
[0280] Step S73: Determine if the color difference of the fused output color signal is greater than the preset color difference. If so, return to step S203. Otherwise, determine the color changes of the fused output color signal that are adapted to different video signal formats.
[0281] After returning to step S203, it is necessary to adjust the historical color feature sequence of the color parameter prediction model, i.e., the training parameters; the weight coefficients of the block feature extraction, i.e., the first preset weight and the second preset weight; and the window length of the time series prediction, before returning to step S203.
[0282] In this embodiment of the invention, a corresponding transition curve is generated by rotating and encoding the coding matrix produced by the target dynamic color reference timetable; the color difference of the fused output color signal is calculated by judging the transition curve, the new transition lookup table (LUT), and the old transition lookup table (LUT); if the color difference of the fused output color signal is greater than a preset color difference, the video frame of the video signal corresponding to the signal type is processed based on the dynamic metadata set to generate a target dynamic color reference block timetable until the color difference of the fused output color signal meets the expectation, thereby achieving accurate synchronization of colors for multiple format signals.
[0283] The specific principles and execution processes of each unit in the live streaming processing system disclosed in the above embodiments of the present invention are the same as the corresponding contents in the live streaming processing method provided in the above embodiments of the present invention. Please refer to the corresponding parts in the live streaming processing method disclosed in the above embodiments of the present invention, and they will not be repeated here.
[0284] This application provides an electronic device, which includes a processor and a memory. The memory is used to store live streaming processing program code and data, and the processor is used to call the program instructions in the memory to execute the steps shown in the live streaming processing method in the above embodiments.
[0285] This invention provides a storage medium, namely a computer-readable storage medium, which includes the electronic device provided in the above-described embodiments of this application. The electronic device is used to execute the live streaming processing method disclosed in the embodiments of this application.
[0286] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A live streaming processing method, characterized in that, The method includes: During the live stream, acquire multi-format video signals from each video frame in the live video; A dynamic metadata set is obtained by parsing video signals in multiple formats. Based on the dynamic metadata set, the video frame images of the video signal corresponding to the signal type are processed to generate a target dynamic color reference block timetable; Based on the target dynamic color reference block timetable, the video signals of different formats are processed to obtain the fused output color signal, which is then output.
2. The method according to claim 1, characterized in that, Based on the parsing of video signals in multiple formats, a dynamic metadata set is obtained, including: Identify video signals of multiple formats and determine the signal type of each video signal; For each type of video signal, the video signal is subjected to header separation processing to obtain header data; Extract metadata of the video signal from the header data; The metadata is converted into metadata in a preset format; The metadata of different signal types is combined to obtain a dynamic metadata set.
3. The method according to claim 1, characterized in that, Based on the dynamic metadata set, the video frames of the video signal corresponding to the signal type are processed to generate a target dynamic color reference block timetable, including: The color-coded YUV values are determined according to different signal types to match the standard dynamic range and color gamut standards of the corresponding signals. According to the standard dynamic range and color gamut standards corresponding to different signal types, the video frame images of the video signals corresponding to the signal types are cropped to obtain video frame images cropped for different signal types; For each signal type, the video frame is cropped and then divided into blocks according to the dynamic metadata set to determine the histogram of each block under different signal types. For each block, a target dynamic color reference timetable is constructed based on the histogram corresponding to the block and the corresponding metadata in the dynamic metadata set.
4. The method according to claim 3, characterized in that, For each signal type, the cropped video frame is segmented according to the dynamic metadata set to determine the histogram of each block under different signal types, including: For each signal type, the video frame of the video signal corresponding to the signal type is divided into blocks according to the block characteristics of the metadata in the dynamic metadata set, resulting in multiple blocks and corresponding block IDs; For each block of each signal type, construct the corresponding histogram based on the bars in the block; The histogram corresponding to the block is optimized based on the block segmentation characteristics of the metadata corresponding to the block in the dynamic metadata set.
5. The method according to claim 3, characterized in that, For each block, a target dynamic color reference time table is constructed based on the histogram corresponding to the block and the corresponding metadata, including: For each block, the histogram is associated with the corresponding metadata within a preset time period to obtain a feature time series. The pre-set color parameter prediction model is invoked to process the feature time series to obtain the color parameter prediction value of each block within a preset time period. The color parameter prediction model is trained based on historical feature time series. A target dynamic color reference timetable is constructed based on the predicted values of the color parameters within a preset time period for each block.
6. The method according to claim 5, characterized in that, A target dynamic color reference timetable is constructed based on the predicted color parameter values within a preset time period for each block, including: The predicted color parameters of each block within a preset time period are mapped to an initial dynamic color reference timetable according to a preset timestamp. A real-time block feature table is generated based on the block features of each block obtained from real-time sampling, and associated with the initial dynamic color reference time table through the block ID; Based on the preset feature threshold library and the real-time block feature table, determine whether the parameters of each block in the initial dynamic color reference time table need to be calibrated; If necessary, the block can be used as the block to trigger calibration; The blocks that trigger calibration are calibrated based on the real-time block feature table, and the initial dynamic color reference time table is adjusted to obtain the target dynamic color reference time table. If none of these are required, the initial dynamic color reference timetable shall be used as the target dynamic color reference timetable.
7. The method according to claim 1, characterized in that, Based on the target dynamic color reference block timetable, the video signals of different formats are processed to obtain a fused output color signal, including: Get the target color parameters of the target dynamic color reference timetable of the block at the current timestamp and the target dynamic color reference timetable of the previous timestamp at the same location; Based on the real-time block feature table, calculate the parameter difference of the target color parameter of the block at the same location at the current timestamp and the previous timestamp; For each block, transition parameters are determined based on the parameter differences; The transition parameters corresponding to each block are matched with the blocks of the current video frame according to the block ID of each block. The actual parameters of the previous video frame are obtained based on the block ID, and the parameters of the current video frame are determined by processing the actual parameters of the previous video frame and the transition parameters corresponding to the block ID. A transition lookup table (LUT) is generated based on the current video frame parameters and the new lookup table (LUT) parameters, wherein the new lookup table (LUT) parameters are the difference table parameters that match the color conversion of the current video frame; The target dynamic color reference timetable and the transition lookup table (LUT) are fused together to obtain the fused output color signal.
8. The method according to claim 7, characterized in that, Also includes: Based on the LUT (Learning Undefined Table), the corresponding transition curve is generated from the encoding matrix produced by rotational encoding fusion. The color difference of the fused output color signal is calculated by judging the transition curve, the new color reference timetable and the old color reference timetable. The old color reference timetable refers to the target dynamic color reference timetable, and the new color reference timetable refers to the color reference timetable obtained by combining the target dynamic color reference timetable with the transition lookup table (LUT). If the color difference of the fused output color signal is greater than the preset color difference, then return to the step of processing the video frame of the video signal corresponding to the signal type based on the dynamic metadata set to generate the target dynamic color reference block timetable.
9. A live streaming processing system, characterized in that, The live streaming processing system includes a global static color processing module, a loading controller, and a non-dynamic encoding processing module. The global static color processing module is used to acquire multi-format video signals of each video frame in the live video during the live broadcast; and to parse the video signals of multiple formats to obtain a dynamic metadata set. Based on the dynamic metadata set, the video frame images of the video signal corresponding to the signal type are processed to generate a target dynamic color reference block timetable; The loading controller and non-dynamic encoding processing module are used to process the video signals of different formats based on the target dynamic color reference block timetable to obtain the fused output color signal and output it.
10. A storage medium, characterized in that, The storage medium includes a stored program, wherein, when the program is executed, it controls the device where the storage medium is located to perform the live streaming processing method as described in any one of claims 1-8.