Video high-definition coding and low-delay transmission method for IPTV (Internet Protocol Television) large-screen short play
By constructing a dual-path bandwidth prediction structure and a CMAF chunk pulling and shaping module, the video encoding and low-latency transmission of IPTV large-screen short dramas are optimized, solving the problems of playback stuttering and switching delay in weak network environments, and achieving a stable and continuous playback experience.
Patent Information
- Application Number
- CN202511522420.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2025-12-09
AI Technical Summary
In IPTV large-screen short drama playback scenarios, existing technologies lack targeted bandwidth prediction and conservative estimation in weak or fluctuating network environments, leading to increased playback stuttering and latency. Furthermore, the lack of refined pre-pull and switching control mechanisms affects the continuity of the user's viewing experience.
A dual-path bandwidth prediction structure consisting of a short-window prediction path and a medium-window prediction path is constructed, and a conservative available bandwidth estimate is generated by fusion. A CMAF chunk pulling and shaping module and a short drama serialization pre-pulling mechanism are configured, and combined with adaptive bitrate control and buffer mode adjustment, the bandwidth determination and data pre-pulling strategy on the player side are optimized.
It effectively reduces the stuttering rate during playback, ensuring that the switching time between episodes is within 0.5 to 0.8 seconds, significantly improving the viewing experience of short dramas on large screens and enhancing the user's sense of immediacy and continuity.
Smart Images

Figure CN121099094A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of low-latency transmission technology. Specifically, it relates to a method for high-definition video encoding and low-latency transmission of short dramas for IPTV large-screen platforms. Background Technology
[0002] In the scenario of playing short dramas on IPTV large screens, high-definition video content places extremely high demands on the real-time performance and stability of the transmission link. Due to the frequent episode switching, the large proportion of opening and closing credits in short dramas, and the significant impact of the first screen loading speed on user experience, it is necessary to comprehensively consider multiple key indicators such as end-to-end latency, first frame rendering time, episode switching response time, and playback stuttering rate throughout the entire process of video encoding, transmission, and playback. At the same time, the uncertainty of the network environment, such as bandwidth fluctuations, increased packet loss rate, and network jitter, also poses higher technical challenges to high-definition video encoding and low-latency transmission.
[0003] However, existing IPTV video transmission technology solutions have the following two shortcomings. First, in weak or fluctuating network environments, there is a lack of targeted bandwidth prediction and conservative estimation mechanisms, which can easily lead to overly aggressive adaptive bitrate selection, resulting in playback stuttering and increased latency. Second, in scenarios where short episodes are switched, there is a lack of refined pre-pull and switching control mechanisms, which often leads to excessively high first frame loading delays when switching between episodes, affecting the continuity of the user's viewing experience. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of existing technologies by providing a method for high-definition video encoding and low-latency transmission of short dramas for IPTV large screens. This method aims to construct a dual-path bandwidth prediction structure consisting of a short-window prediction path and a medium-window prediction path, and then fuse these paths to generate a conservative estimate of available bandwidth. This achieves stable bandwidth determination under weak network conditions, effectively reducing the frequency of stuttering during playback. Furthermore, by configuring a CMAF chunk pulling and shaping module and a short drama consecutive playback pre-pulling mechanism, keyframe data can be pulled in advance during the inter-episode switching phase, ensuring that the cross-episode switching time is controlled within 0.5 to 0.8 seconds. This improves the first-screen loading speed and the smoothness of inter-episode switching, significantly enhancing the overall viewing experience of short dramas on large screens.
[0005] To achieve the above objectives, the present invention provides the following technical solution: A method for high-definition video encoding and low-latency transmission of short dramas for IPTV large screens includes the following: Step S100: Define end-to-end measurement boundaries and time references, configure the acquisition mechanism for real-time status indicators of the network environment, extract high-definition encoded metadata using the encoding end, mark timestamp information, and configure a unified structured running status object.
[0006] Step S200: Configure the short window prediction path and the medium window prediction path, configure the input parameters of the short window prediction path and the medium window prediction path, obtain the prediction results of the short window prediction path and the medium window prediction path, fuse the prediction results to form a conservative available bandwidth estimate, and determine the network level label based on the network environment.
[0007] Step S300: Configure the peak identification module based on high-definition encoding metadata to identify peak bitrate risk periods triggered by IDR frames, peak bitrate risk periods triggered by video buffer verifier buffer backfilling, peak bitrate risk periods caused by sudden changes in shot complexity, and peak bitrate risk periods during cross-episode switching of short dramas, and generate a peak risk list.
[0008] Step S400: Configure the buffer mode adjustment module and set the buffer mode boundary range, configure the adaptive bitrate selection module and the buffer control module, set the switching constraint mechanism, and generate a joint decision output of the buffer target and bitrate.
[0009] Step S500: Configure the CMAFchunk streaming shaping module based on segmentation priority, configure the fine-grained streaming degradation mechanism and trigger dynamic granular adjustment, configure the short drama consecutive playback pre-pull mechanism and weak network first screen optimization strategy, configure the emergency rollback strategy and minimum viewing level guarantee mechanism, and generate the CMAchunk request priority queue and streaming shaping status field.
[0010] Step S600: Configure the playback quality feedback collection module, generate user-side playback experience indicators, configure the online self-calibration rule set, set the trigger mechanism, generate the parameter tuning configuration set, and send it to the execution module.
[0011] As a preferred embodiment of the present invention, step S100 specifically comprises: Step S100.1: Define the end-to-end measurement boundaries and time references, and configure the acquisition mechanism for real-time status indicators of the network environment.
[0012] In the scenario of playing short dramas on large screens for IPTV platforms, quantitative playback control targets are established. These playback control targets include: end-to-end latency threshold, first frame appearance time, episode switching response time within a short drama, and percentage of playback stuttering.
[0013] A network status monitoring module is configured on the player side to collect instantaneous available throughput data, round-trip latency data, network jitter data, packet loss rate data, and CMAFchunk arrival interval data. The sampling period for instantaneous available throughput data is set to 200 milliseconds, and the average value is calculated within a sliding window of about 1 second. The round-trip latency data is calculated by interacting with timestamped probe messages and acknowledgment messages to calculate the round-trip transmission time between the server and the player. The network jitter data is calculated by the standard deviation of the round-trip latency difference of continuous probe messages. The packet loss rate data is calculated within a fixed statistical window of 1 second, showing the ratio of the number of lost data packets to the total number of data packets. The CMAFchunk arrival interval data records the arrival time interval sequence of CMAFchunk data in the player buffer.
[0014] Real-time network environment status indicators are generated based on instantaneous available throughput data, round-trip time data, network jitter data, packet loss rate data, and CMAF chunk arrival interval data.
[0015] Step S100.2 uses the encoding end to extract high-definition encoded metadata, and marks timestamp information and configures a unified structured running status object.
[0016] The key encoding metadata is collected at the encoding end. The key encoding metadata includes: IDR frame interval parameters, video buffer verifier buffer target parameters and HRD buffer target parameters, shot scene complexity labels and high dynamic range marker information.
[0017] The IDR frame interval parameter is: recording the time interval between consecutive InstantaneousDecoderRefresh frames. The video buffer verifier buffer target parameter and HRD buffer target parameter are: recording the VideoBufferingVerifier buffer target and HypotheticalReferenceDecoder buffer target set at the encoding end, respectively. The shot scene complexity label is: marking the complexity of each shot through the pre-encoding image analysis module, including three categories: low complexity, medium complexity, and high complexity. The high dynamic range label information is: recording whether the video content is in HighDynamicRange encoding format.
[0018] High-definition encoded metadata is formed based on IDR frame interval parameters, video buffer verifier buffer target parameters and HRD buffer target parameters, shot scene complexity labels and high dynamic range marker information.
[0019] During the packaging stage, the key encoded metadata will be written into the CMAFchunk header information or bypass signaling channel.
[0020] In response to the highly fragmented content and frequent episode switching characteristics of short dramas, the content provider will inject structural scene boundary timestamp information during the packaging stage. The injected scene boundary timestamp information includes: the start and end timestamps of the opening credits, the start and end timestamps of the advertisement segments, the start and end timestamps of the end credits and the end timestamp of the entire episode, and the timestamp of the first frame of the next episode.
[0021] The collected real-time network environment status indicators, high-definition encoded metadata, and scene boundary timestamp information are integrated and configured into a unified structured runtime status object. The structured runtime status object includes: a set of network status parameters, a set of encoded metadata parameters, and a set of scene boundary information.
[0022] As a preferred embodiment of the present invention, step S200 specifically comprises: Step S200.1: Configure the short window prediction path and the medium window prediction path, configure the input parameters of the short window prediction path and the medium window prediction path, and obtain the prediction results of the short window prediction path and the medium window prediction path.
[0023] Configure the bandwidth prediction module in the player and establish a dual-path bandwidth prediction structure that includes a short-window prediction path and a medium-window prediction path. The input of the bandwidth prediction module is the set of network state parameters in the structured runtime state object.
[0024] The short-window prediction path is as follows: an exponentially weighted moving average is performed on the instantaneous available throughput data within the sliding window of the last 1 second to obtain the 1-second prediction value.
[0025] The mid-window prediction path is as follows: perform least squares time-series regression on the instantaneous available throughput data, network jitter data, and packet loss rate data within the sliding window of the past 3 seconds to obtain the 2-3 second prediction value.
[0026] When the relative fitting error of least squares regression exceeds 5%, a lightweight recurrent neural network model is called to correct the prediction. The lightweight recurrent neural network model adopts a 2-layer structure and 8 hidden nodes, and the single inference time does not exceed 5 milliseconds.
[0027] Step S200.2: Based on the prediction results, perform fusion to form a conservative estimate of available bandwidth, and determine the network level label based on the network environment.
[0028] The prediction results calculated from the short-window and medium-window prediction paths are used to perform confidence interval correction and lower limit fusion, and a conservative estimate of available bandwidth is output.
[0029] The network level label is determined by using the set of network status parameters in the structured operational status object. The determination rules include: stable network level, fluctuating network level, and weak network level.
[0030] The determination of the network stability level requires meeting three conditions: packet loss rate <1%, network jitter <10 milliseconds, and the relative fluctuation of instantaneous available throughput data within the past 2-second window. The determination of the fluctuating network level requires meeting any one of the following conditions: packet loss rate data is between 1% and 3%, network jitter data is between 10 and 25 milliseconds, and the relative decrease in instantaneous available throughput data within the past 2-second window is >15%. The determination of the weak network level requires meeting any one of the following conditions: packet loss rate data ≥3%, network jitter data ≥25 milliseconds, and the relative decrease in instantaneous available throughput data within the past 2-second window is >30%. The update cycle of the network level label is 200 milliseconds, and the network level label submission frequency is once every 2 seconds.
[0031] As a preferred embodiment of the present invention, step S300 specifically comprises: Step S300.1: Configure the peak recognition module based on high-definition encoded metadata.
[0032] Configure a peak recognition module in the player to identify high-risk periods where bitrate may suddenly increase in the future. The input of the peak recognition module is the set of encoding metadata parameters and scene boundary information in the structured running state object. The peak recognition module performs periodic rolling recognition according to a unified time base. The prediction window length is set to 3 seconds and the update period is 200 milliseconds. Step S300.2: Based on the peak identification module, identify the peak risk period, peak risk period, peak risk period and peak risk period of bitrate, and generate a peak risk list.
[0033] As a preferred embodiment of the present invention, step S400 specifically includes: Step S400.1: Configure the buffer mode adjustment module and set the buffer mode boundary range.
[0034] In the player configuration, the buffer mode adjustment module takes the following inputs: network level label, peak risk list, and conservative available bandwidth estimate. The network level label can be set to: stable network level, fluctuating network level, or weak network level.
[0035] Three buffering modes and boundary ranges can be set: low buffering mode, standard buffering mode, and protective buffering mode.
[0036] The low buffering mode, with a buffer target range of 0.35 seconds to 0.60 seconds, is only enabled when the network level label is a stable network level and the peak risk list is not hit within the prediction window of the next 1-3 seconds. The standard buffering mode, with a buffer target range of 0.60 seconds to 1.20 seconds, is used as the default buffering strategy. The protection buffering mode, with a buffer target range of 1.20 seconds to 1.60 seconds, is enabled when the network level label is a weak network level or the peak risk list is hit within the prediction window of the next 1-3 seconds.
[0037] The buffer mode switching adopts hysteresis control, and both upward and downward switching require the switching conditions to be met continuously for no less than 1 second.
[0038] Step S400.2: Configure the adaptive bitrate selection module and the buffer control module, and set the switching constraint mechanism to generate a joint decision output of the buffer target and bitrate.
[0039] The player is configured with an adaptive bitrate control module and a buffer control module. The adaptive bitrate control module selects the bitrate level based on a conservative estimate of available bandwidth and follows constraints including: the average bitrate of the bitrate level satisfies... ,in, It is the average bitrate of the selected bitrate level. This is a conservative estimate of available bandwidth, with a coefficient of 0.85 representing a conservative margin. The bitrate switching frequency is limited to a maximum of once every 2 seconds, based on the player's current buffering duration. If the time interval is less than 0.40 seconds, the bitrate will be increased, and only the current bitrate level can be maintained or the bitrate level can be decreased. Within the stable time window of episode switching, it is prohibited to increase the bitrate level by two consecutive levels, and only one bitrate level can be increased or the current bitrate level can be maintained.
[0040] The player generates a joint decision output based on the network grade label and peak risk list, the overlap of risk periods within a 1–3 second prediction window, a conservative estimate of available bandwidth, the current buffer duration and buffer mode boundary, and the average bitrate of the current bitrate tier. The output includes the target buffer duration. Take the median value according to the selected buffer mode range, and for low buffer mode, take... Standard buffer mode takes Seconds, protection buffer mode take Seconds, adaptive bitrate level In order to satisfy Provided that the conditions for upgrading are not violated, the highest feasible bitrate level should be selected.
[0041] The target buffer duration and adaptive bitrate are written into the playback strategy decision field area of the structured running status object. The buffer control module adjusts the buffer filling strategy according to the target buffer duration, and the streaming scheduling module matches the CMAFchunk pulling rhythm according to the adaptive bitrate.
[0042] As a preferred embodiment of the present invention, step S500 specifically comprises: Step S500.1: Configure the CMAFchunk pull stream shaping module based on segment priority, configure the fine-grained pull stream degradation mechanism, and trigger dynamic granularity adjustment.
[0043] Step S500.2: Configure the short drama serialization pre-pull mechanism and the weak network first screen optimization strategy, configure the emergency rollback strategy and the minimum viewing level guarantee mechanism, and generate the CMAchunk request priority queue and the streaming shaping status field.
[0044] In the scenario of short drama series playback, a pre-playback is performed at the end of the previous episode. The pre-playback includes: when the playback progress is at the end of the episode and the current time is between 1 and 2 seconds away from the timestamp of the first frame of the next episode, the CMAF chunk of the InstantaneousDecoderRefresh frame corresponding to the first frame of the next episode and the following 2 to 3 CMAF chunks are pre-fetched. When the network level label is a weak network level, the pre-fetched CMAF chunk adopts a lower first-screen strategy.
[0045] The timing control for the short drama pre-roll mechanism when switching episodes is set to 0.5 seconds.
[0046] An emergency rollback determination and execution mechanism is set up, which includes: emergency rollback trigger conditions and rollback actions. The emergency rollback determination is: two consecutive CMAF chunk fetch timeouts and a packet loss rate greater than or equal to 5%. The rollback action is: keep the resolution unchanged, reduce the bitrate and turn off high-complexity encoding tools, temporarily switch to protection buffer mode, and set the target buffer duration to 1.40 seconds. If the network level label does not belong to the weak network level for two consecutive seconds and the current buffer duration is greater than or equal to 0.90 seconds, then the normal adaptive strategy is restored.
[0047] Write the CMAchunk request priority queue to the pull control field area of the structured runtime state object.
[0048] As a preferred embodiment of the present invention, step S600 specifically comprises: Step S600.1: Configure the playback quality feedback collection module and generate user-side playback experience metrics.
[0049] Step S600.2: Configure the online self-calibration rule set, set the trigger mechanism, generate the parameter tuning configuration set, and send it to the execution module.
[0050] The playback strategy parameter tuning module takes the user-side playback experience indicators output by the playback quality feedback collection module as input and dynamically adjusts them according to the online self-calibration rule set, including: playback stuttering ratio adjustment rules, first frame presentation time adjustment rules, and short episode switching response adjustment rules.
[0051] The rule for adjusting the percentage of playback stuttering is as follows: when the 95th percentile of the percentage of playback stuttering is greater than 0.5% / hour, the conservative strategy coefficient for bandwidth assessment is increased, and the lower limit of the minimum target buffer time is increased to 0.10 seconds to 0.20 seconds.
[0052] The first frame presentation time adjustment rule is as follows: when the first frame presentation time exceeds the target, the duration threshold of the first segment CMAF chunk of the first screen is adjusted to 300 milliseconds to 500 milliseconds, and the pull priority of the first segment CMAF chunk in the pull stream shaping module is increased by one discrete priority level.
[0053] The short episode switching response adjustment rule is as follows: when the episode switching response time exceeds the target of 0.5 seconds or the 95th percentile exceeds 0.8 seconds, the number of pre-pulled CMAF chunks will be adjusted from 2 to 3, or the episode switching advance will be adjusted to 2 to 3 seconds.
[0054] At the end of each decision cycle, the playback strategy parameter tuning module generates a parameter update set, which includes: bandwidth assessment conservative strategy coefficient, minimum target buffer duration lower limit, first screen first segment CMAF chunk duration threshold, number of pre-pulled CMAF chunks, set switching advance, adaptive bitrate tier upgrade limit parameter, and pull stream shaping granularity adjustment switch.
[0055] The parameter update set will be written into the playback strategy parameter tuning field area of the structured runtime state object and sent to the adaptive bitrate control module, buffer mode adjustment module and CMAFchunk pull stream shaping module. When the encoding end or the packaging end provides a parameter sending interface, it will be sent to the encoding end or the packaging end at the same time.
[0056] Configure a merging and evolution mechanism, which includes: a minimum interval period for parameter tuning of ≥5 seconds, a maximum of 3 consecutive adjustments per parameter, and a parameter tuning freeze period of 20 seconds after exceeding the limit. After each parameter update, an observation period of 10 seconds is entered. When the user-side playback experience indicators continue to improve, the next round of parameter tuning will be allowed. If the user-side playback experience indicators do not improve after the observation period, the system will revert to the previous parameter update set and record the revert status flag.
[0057] After each round of parameter tuning is completed, the playback strategy parameter tuning module writes the currently effective parameter set as a stable parameter set into the strategy evolution record area of the structured running status object. The record content includes: parameter change timestamp, stable parameter set, network level label at the time of recording, user-side playback experience index snapshot, and parameter tuning result mark.
[0058] Compared with the prior art, the beneficial effects of the present invention are: 1. By configuring short-window prediction path and medium-window prediction path, and combining confidence interval correction and lower limit fusion to form a conservative available bandwidth estimate, and triggering protection buffer mode and low bit rate pre-pull strategy when the network level label determines that it is a weak network level, this mechanism effectively solves the problems of frequent playback stuttering and high switching latency in the existing technology in the weak network environment, so that short dramas can still maintain a stable and continuous playback experience in complex network environments.
[0059] 2. Configure the CMAF chunk pulling and shaping module on the player side, and implement a key byte-priority pulling strategy for the first screen loading scenario and cross-episode switching scenario. Combined with the mechanism of pre-pulling Instantaneous Decoder Refresh frames and their subsequent chunks, the average first frame presentation time is controlled within 0.8 seconds, and the response time for switching between short dramas is compressed to the target range of 0.5 seconds. Compared with the shortcomings of existing technologies, such as long first frame waiting time and black screen or obvious pauses during cross-episode switching, this solution can significantly shorten the response time of the first screen and episode switching, and improve the user's sense of immediacy and continuity in watching short dramas.
[0060] 3. The playback quality feedback acquisition module collects experience indicators in real time, such as first frame rendering time, stuttering rate, adaptive bitrate switching frequency, and inter-episode switching response time. Based on the online self-calibration rule set, a parameter tuning configuration set is dynamically generated, forming a closed-loop evolution mechanism for playback strategy parameter tuning. This mechanism can not only automatically optimize bandwidth assessment, buffering targets, and pre-pull strategies based on user feedback, but also avoid frequent parameter oscillations through parameter tuning freeze and observation period settings. This overcomes the problem of fixed and rigid playback strategies in existing technologies that cannot be adaptively adjusted according to actual QoE, thus achieving a self-optimization effect that becomes more stable over time. Attached Figure Description
[0061] Figure 1 This is a flowchart illustrating a method for high-definition video encoding and low-latency transmission of short dramas for IPTV large screens, provided as an embodiment of this application. Detailed Implementation
[0062] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0063] Please see Figure 1 , Figure 1 This application provides a flowchart of a method for high-definition video encoding and low-latency transmission of short dramas for IPTV large screens.
[0064] In this embodiment, a method for high-definition video encoding and low-latency transmission of short dramas for IPTV large screens may include steps S100, S200, S300, S400, S500 and S600. Step S100: Define end-to-end measurement boundaries and time references, configure the acquisition mechanism for real-time status indicators of the network environment, extract high-definition encoded metadata using the encoding end, mark timestamp information, and configure a unified structured running status object.
[0065] Step S200: Configure the short window prediction path and the medium window prediction path, configure the input parameters of the short window prediction path and the medium window prediction path, obtain the prediction results of the short window prediction path and the medium window prediction path, fuse the prediction results to form a conservative available bandwidth estimate, and determine the network level label based on the network environment.
[0066] Step S300: Configure the peak identification module based on high-definition encoding metadata to identify peak bitrate risk periods triggered by IDR frames, peak bitrate risk periods triggered by video buffer verifier buffer backfilling, peak bitrate risk periods caused by sudden changes in shot complexity, and peak bitrate risk periods during cross-episode switching of short dramas, and generate a peak risk list.
[0067] Step S400: Configure the buffer mode adjustment module and set the buffer mode boundary range, configure the adaptive bitrate selection module and the buffer control module, set the switching constraint mechanism, and generate a joint decision output of the buffer target and bitrate.
[0068] Step S500: Configure the CMAFchunk streaming shaping module based on segmentation priority, configure the fine-grained streaming degradation mechanism and trigger dynamic granular adjustment, configure the short drama consecutive playback pre-pull mechanism and weak network first screen optimization strategy, configure the emergency rollback strategy and minimum viewing level guarantee mechanism, and generate the CMAchunk request priority queue and streaming shaping status field.
[0069] Step S600: Configure the playback quality feedback collection module, generate user-side playback experience indicators, configure the online self-calibration rule set, set the trigger mechanism, generate the parameter tuning configuration set, and send it to the execution module.
[0070] In some specific embodiments, step S100 specifically includes: Step S100.1: Define the end-to-end measurement boundaries and time references, and configure the acquisition mechanism for real-time status indicators of the network environment.
[0071] In the scenario of playing short dramas on large screens for IPTV platforms, quantitative playback control targets are established. These playback control targets include: end-to-end latency threshold, first frame appearance time, episode switching response time within a short drama, and percentage of playback stuttering.
[0072] The end-to-end latency threshold is set to 1.2 seconds. End-to-end latency refers to the entire process from the start of encoding processing at the video encoding end to the completion of image rendering by the player. The average threshold for the first frame appearance time is set to 0.8 seconds, and the 95th percentile threshold is set to 1.1 seconds. 95% means that 95% of playback tasks achieve the first frame display within 1.1 seconds. The threshold for the response time of switching between episodes in a short drama is set to 0.5 seconds, and the 95th percentile threshold is set to 0.8 seconds. The percentage of stuttering during playback is set to the 95th percentile threshold of stuttering during playback is set to within 0.5% per hour.
[0073] A network status monitoring module is configured on the player side to collect instantaneous available throughput data, round-trip latency data, network jitter data, packet loss rate data, and CMAF chunk arrival interval data. The sampling period for the instantaneous available throughput data is set to 200 milliseconds, and the average value is calculated within a sliding window of about 1 second. The sampling period is set to 200 milliseconds. While ensuring the real-time performance of sampling, the system's computational load is balanced. When the sampling period is less than 150 milliseconds, the computational occupancy rate on the encoding side increases by about 30%, resulting in accumulated latency. When the sampling period is greater than 250 milliseconds, the response lag of the prediction bandwidth is significant, and the peak error increases by more than 4%.
[0074] A 1-second sliding window is used to smooth short-term bandwidth fluctuations and eliminate the interference of instantaneous changes on the prediction model input. When the window is less than 0.5 seconds, the noise interference is large, which leads to an increase in the average misclassification rate of about 8%. When the window exceeds 1.5 seconds, the response speed decreases, which is not conducive to real-time adjustment in low-latency scenarios. Therefore, 1 second is selected as the balancing window.
[0075] The round-trip latency data is calculated by exchanging timestamped probe messages and acknowledgment messages to determine the round-trip transmission time between the server and the player. The network jitter data is calculated by the standard deviation of the round-trip latency difference between continuous probe messages. The packet loss rate data is calculated within a fixed 1-second statistical window, showing the ratio of the number of lost data packets to the total number of data packets. The CMAFchunk arrival interval data records the arrival time sequence of CMAFchunk data in the player's buffer.
[0076] Real-time network environment status indicators are generated based on instantaneous available throughput data, round-trip time data, network jitter data, packet loss rate data, and CMAF chunk arrival interval data.
[0077] Step S100.2 uses the encoding end to extract high-definition encoded metadata, and marks timestamp information and configures a unified structured running status object.
[0078] The key encoding metadata is collected at the encoding end. The key encoding metadata includes: IDR frame interval parameters, video buffer verifier buffer target parameters and HRD buffer target parameters, shot scene complexity labels and high dynamic range marker information.
[0079] The IDR frame interval parameter is: recording the time interval between consecutive InstantaneousDecoderRefresh frames. The video buffer verifier buffer target parameter and HRD buffer target parameter are: recording the VideoBufferingVerifier buffer target and HypotheticalReferenceDecoder buffer target set at the encoding end, respectively. The shot scene complexity label is: marking the complexity of each shot through the pre-encoding image analysis module, including three categories: low complexity, medium complexity, and high complexity. The high dynamic range label information is: recording whether the video content is in HighDynamicRange encoding format.
[0080] High-definition encoded metadata is formed based on IDR frame interval parameters, video buffer verifier buffer target parameters and HRD buffer target parameters, shot scene complexity labels and high dynamic range marker information.
[0081] During the packaging stage, the key encoded metadata will be written into the CMAFchunk header information or bypass signaling channel.
[0082] In response to the highly fragmented content and frequent episode switching characteristics of short dramas, the content provider will inject structural scene boundary timestamp information during the packaging stage. The injected scene boundary timestamp information includes: the start and end timestamps of the opening credits, the start and end timestamps of the advertisement segments, the start and end timestamps of the end credits and the end timestamp of the entire episode, and the timestamp of the first frame of the next episode.
[0083] The intro start timestamp and intro end timestamp are used to control the first-screen buffering strategy and fast first-frame loading. The ad segment start timestamp and ad segment end timestamp are used to identify ad segments and dynamically adjust the bitrate selection rules. The outro start timestamp and episode end timestamp are used to control the playback ending rhythm and trigger the next episode pre-pull logic. The next episode first frame timestamp is used for advance control of cross-episode streaming to achieve low-latency switching.
[0084] Scene boundary timestamp information is injected through a packaged file structure and identified and stored using the player's parsing module.
[0085] The collected real-time network environment status indicators, high-definition encoded metadata, and scene boundary timestamp information are integrated and configured into a unified structured runtime status object. The structured runtime status object includes: a set of network status parameters, a set of encoded metadata parameters, and a set of scene boundary information.
[0086] In some specific embodiments, step S200 specifically includes: Step S200.1: Configure the short window prediction path and the medium window prediction path, configure the input parameters of the short window prediction path and the medium window prediction path, and obtain the prediction results of the short window prediction path and the medium window prediction path.
[0087] Configure the bandwidth prediction module in the player and establish a dual-path bandwidth prediction structure that includes a short-window prediction path and a medium-window prediction path. The input of the bandwidth prediction module is the set of network state parameters in the structured runtime state object.
[0088] The short-window prediction path is as follows: An exponentially weighted moving average is applied to the instantaneous available throughput data within the nearest 1-second sliding window to obtain the 1-second prediction value. Specifically:
[0089] In the formula: It is the short-window prediction bandwidth value at the current moment. This is the exponentially weighted smoothing coefficient, with a default value of 0.7. It is the instantaneous available throughput data obtained from sampling at the current moment. It is the short-window predicted bandwidth value of the previous moment. The formula expresses the method of smoothly integrating the current instantaneous available throughput data with the historical short-window predicted bandwidth value through an exponentially weighted moving average, which is used to quickly predict the bandwidth availability for the next second.
[0090] The mid-window prediction path is as follows: Least squares temporal regression is performed on the instantaneous available throughput data, network jitter data, and packet loss rate data within the past 3-second sliding window to obtain 2-3 second prediction values. Specifically:
[0091] In the formula: It is about predicting the future. The bandwidth value of the middle window at time 1 second. These are the fitting coefficients of the time-series regression model, dynamically estimated from historical data within the sliding window. This is the current sampling time point. This is the network jitter data at the current sampling time. This is the packet loss rate data at the current sampling time point. The prediction step size ranges from 1 to 3 seconds. The formula expresses the prediction of bandwidth trends within the next 2 to 3 seconds based on a multivariate time series regression method, which combines time, network jitter data, and packet loss rate data.
[0092] When the relative fitting error of least squares regression exceeds 5%, a lightweight recurrent neural network model is called to correct the prediction. The lightweight recurrent neural network model adopts a 2-layer structure and 8 hidden nodes, and the single inference time does not exceed 5 milliseconds.
[0093] The fitting error threshold is set at 5% to balance accuracy and model call frequency. 5% is taken as the trigger boundary to maintain stable processing efficiency while ensuring prediction accuracy. The two-layer structure and eight hidden nodes can achieve better temporal feature fitting effect with limited computing resources. The single inference time is controlled within 5 milliseconds to ensure that real-time inference is completed within the video frame interval without interfering with the encoder's bitrate adjustment process.
[0094] Step S200.2: Based on the prediction results, perform fusion to form a conservative estimate of available bandwidth, and determine the network level label based on the network environment.
[0095] Based on the prediction results calculated from the short-window and medium-window prediction paths, confidence interval correction and lower limit fusion are performed, and a conservative estimate of available bandwidth is output, specifically:
[0096] In the formula: This is a conservative estimate of available bandwidth. It is the short-window prediction bandwidth value. It is the standard deviation of the short-window prediction residuals. It is the mid-window prediction bandwidth value. It is the standard deviation of the mid-window prediction residual. The formula expresses the final conservative available bandwidth estimate obtained by subtracting the standard deviation of the residual from the short-window prediction bandwidth value and the mid-window prediction bandwidth value respectively. The conservative available bandwidth estimate ensures that when the network uncertainty is large, the player uses a safer bandwidth as the basis to avoid stuttering during playback.
[0097] The network level label is determined by using the set of network status parameters in the structured operational status object. The determination rules include: stable network level, fluctuating network level, and weak network level.
[0098] The determination of the network stability level requires meeting three conditions: packet loss rate <1%, network jitter <10 milliseconds, and the relative fluctuation of instantaneous available throughput data within the past 2-second window. The determination of the fluctuating network level requires meeting any one of the following conditions: packet loss rate data is between 1% and 3%, network jitter data is between 10 and 25 milliseconds, and the relative decrease in instantaneous available throughput data within the past 2-second window is >15%. The determination of the weak network level requires meeting any one of the following conditions: packet loss rate data ≥3%, network jitter data ≥25 milliseconds, and the relative decrease in instantaneous available throughput data within the past 2-second window is >30%. The update cycle of the network level label is 200 milliseconds, and the network level label submission frequency is once every 2 seconds.
[0099] In some specific embodiments, step S300 specifically includes: Step S300.1: Configure the peak recognition module based on high-definition encoded metadata.
[0100] Configure a peak recognition module in the player to identify high-risk periods where bitrate may suddenly increase in the future. The input of the peak recognition module is the set of encoding metadata parameters and scene boundary information in the structured running state object. The peak recognition module performs periodic rolling recognition according to a unified time base. The prediction window length is set to 3 seconds and the update period is 200 milliseconds. Step S300.2: Based on the peak identification module, identify the peak risk period, peak risk period, peak risk period and peak risk period of bitrate, and generate a peak risk list.
[0101] The peak identification module identifies peak bitrate risk periods triggered by IDR frames, peak bitrate risk periods triggered by video buffer verifier buffer backfilling, peak bitrate risk periods caused by sudden changes in shot complexity, and peak bitrate risk periods during cross-episode switching in short dramas.
[0102] The bitrate peak risk period triggered by the IDR frame is as follows: Since the InstantaneousDecoderRefresh frame is a key node for decoding and reconstruction in the video stream, its preceding and following regions are usually accompanied by a significant increase in bitrate. Therefore, an identification logic based on the IDR frame interval parameter is constructed to define the peak period triggered by the IDR frame, specifically as follows:
[0103] In the formula: It is the peak bitrate period triggered by the InstantaneousDecoderRefresh frame. It is the time point when the next InstantaneousDecoderRefresh frame appears, which is calculated from the IDR frame interval parameter. The unit of 0.5 is seconds, which means the prediction window range of 0.5 seconds before and after.
[0104] In each prediction cycle, the peak identification module compares the current time in the structured running status object with the time difference of the InstantaneousDecoderRefresh frame. If it is within this range, it is marked as an IDR encoding peak risk period.
[0105] The peak risk period triggered by the video buffer verifier buffer backfill is as follows: When the target parameters of the VideoBufferingVerifier and HypotheticalReferenceDecoder at the encoding end trigger the backfill task, the encoding output rate will experience a short-term surge. The peak risk is determined by obtaining the video buffer verifier buffer filling cycle from the structured runtime state object and combining it with calculations. Specifically:
[0106] In the formula: The peak risk period is triggered by VideoBufferingVerifier buffer backfilling. This is the point in time when the video buffer verifier buffer enters a low-water state. It is the point in time when the video buffer verifier buffer is restored to the target fill level.
[0107] When the buffer fill rate falls below the target lower limit by 10%, backfilling is immediately triggered and that period is marked as a peak risk zone.
[0108] The peak risk period caused by sudden changes in scene complexity is defined as follows: a sudden increase in scene scene complexity will lead to a significant increase in instantaneous bitrate. The scene complexity variation rate is defined as follows:
[0109] In the formula: It is the variation in lens complexity. This represents the complexity level of the scene at the current moment, categorized as follows: high complexity is denoted by 3, medium complexity by 2, and low complexity by 1. It represents the level of shot complexity from the previous moment.
[0110] when If the next frame type is an I-frame or an InstantaneousDecoderRefresh frame, then the next 0.5–1 seconds will be predicted as the period of peak risk of complexity surge.
[0111] The peak bitrate risk period during cross-episode switching in short dramas is as follows: During short drama playback, when the timestamp of the first frame of the next episode is less than or equal to 2 seconds away, and the current playback stage is the end credits stage, potential peak cross-episode switching periods need to be identified in advance. The specific judgment criteria are as follows:
[0112] In the formula: This is the timestamp of the first frame of the next episode. This is the current playback time. "Stage" is a scene stage marker, and its values include: Opening, Content, Ad, and Ending. This indicates a logical AND operation, requiring both conditions to be met simultaneously.
[0113] If the conditions are met, then define the peak risk period for cross-set switching, specifically as follows:
[0114] In the formula: This is the peak risk period triggered by the first frame of an episode. This is the timestamp of the first frame of the next episode. This is the number of CMAF chunks to be pre-fetched; the default value is 3. It is the average duration of a single CMAF chunk, in seconds, and is usually set to 0.5 seconds.
[0115] The results obtained from peak bitrate risk periods triggered by IDR frames, peak bitrate risk periods triggered by video buffering verifier buffer backfilling, peak bitrate risk periods caused by sudden changes in shot complexity, and peak bitrate risk periods during cross-episode switching in short dramas are integrated to form a peak risk list. The peak risk list includes: a list of peak risk periods for InstantaneousDecoderRefresh frames, a list of peak risk periods for VideoBufferingVerifier backfilling, a list of peak risk periods for sudden changes in shot scene complexity, and a list of peak risk periods for cross-episode switching.
[0116] The refresh cycle of the peak risk list is set to 200 milliseconds. Within each refresh cycle, the player merges and deduplicates the four types of time period lists and writes the merged peak risk list into the encoded timing prediction field area of the structured running status object. Based on the start and end times and priority information recorded in the four types of time period lists in the peak risk list, the player executes advance control, streaming shaping and throttling strategies.
[0117] In some specific embodiments, step S400 specifically includes: Step S400.1: Configure the buffer mode adjustment module and set the buffer mode boundary range.
[0118] In the player configuration, the buffer mode adjustment module takes the following inputs: network level label, peak risk list, and conservative available bandwidth estimate. The network level label can be set to: stable network level, fluctuating network level, or weak network level.
[0119] Three buffering modes and boundary ranges can be set: low buffering mode, standard buffering mode, and protective buffering mode.
[0120] The low buffering mode, with a buffer target range of 0.35 seconds to 0.60 seconds, is only enabled when the network level label is a stable network level and the peak risk list is not hit within the prediction window of the next 1-3 seconds. The standard buffering mode, with a buffer target range of 0.60 seconds to 1.20 seconds, is used as the default buffering strategy. The protection buffering mode, with a buffer target range of 1.20 seconds to 1.60 seconds, is enabled when the network level label is a weak network level or the peak risk list is hit within the prediction window of the next 1-3 seconds.
[0121] The buffer mode switching adopts hysteresis control, and both upward and downward switching require the switching conditions to be met continuously for no less than 1 second.
[0122] Step S400.2: Configure the adaptive bitrate selection module and the buffer control module, and set the switching constraint mechanism to generate a joint decision output of the buffer target and bitrate.
[0123] The player is configured with an adaptive bitrate control module and a buffer control module. The adaptive bitrate control module selects the bitrate level based on a conservative estimate of available bandwidth and follows constraints including: the average bitrate of the bitrate level satisfies... ,in, It is the average bitrate of the selected bitrate level. This is a conservative estimate of available bandwidth, with a coefficient of 0.85 representing a conservative margin. The bitrate switching frequency is limited to a maximum of once every 2 seconds, based on the player's current buffering duration. If the time interval is less than 0.40 seconds, the bitrate will be increased, and only the current bitrate level can be maintained or the bitrate level can be decreased. Within the stable time window of episode switching, it is prohibited to increase the bitrate level by two consecutive levels, and only one bitrate level can be increased or the current bitrate level can be maintained.
[0124] The player generates a joint decision output based on the network grade label and peak risk list, the overlap of risk periods within a 1–3 second prediction window, a conservative estimate of available bandwidth, the current buffer duration and buffer mode boundary, and the average bitrate of the current bitrate tier. The output includes the target buffer duration. Take the median value according to the selected buffer mode range, and for low buffer mode, take... Standard buffer mode takes Seconds, protection buffer mode take Seconds, adaptive bitrate level In order to satisfy Provided that the conditions for upgrading are not violated, the highest feasible bitrate level should be selected.
[0125] The target buffer duration and adaptive bitrate are written into the playback strategy decision field area of the structured running status object. The buffer control module adjusts the buffer filling strategy according to the target buffer duration, and the streaming scheduling module matches the CMAFchunk pulling rhythm according to the adaptive bitrate.
[0126] In some specific embodiments, step S500 specifically includes: Step S500.1: Configure the CMAFchunk pull stream shaping module based on segment priority, configure the fine-grained pull stream degradation mechanism, and trigger dynamic granularity adjustment.
[0127] Configure the CMAFchunk streaming shaping module in the player. The CMAFchunk streaming shaping module can dynamically optimize and control the CMAFchunk fetching order, segmentation strategy and streaming rhythm. The input parameters of the CMAFchunk streaming shaping module include: current buffer duration, conservative estimated available bandwidth, peak risk list in the structured running status object, playback scene marker and current playback time position and the timestamp of the first frame of the next episode.
[0128] The CMAFchunk streaming shaping module implements a keyword-first fetching strategy for first-screen loading scenarios and cross-set switching scenarios.
[0129] The first-screen loading scenario is as follows: when the player first enters playback mode and is in the intro stage, the first-screen optimization strategy is activated. This strategy involves: performing segmented fetching of the first CMAF chunk, dividing it into a first half byte segment and a second half byte segment, prioritizing the fetching of the first half byte segment, which contains the complete decoding units necessary for initial decoding and visual presentation. Fixed-speed fetching is enabled, and the peak fetch rate is set to an upper limit. ,in, This is the maximum CMAF chunk fetch rate after the limit. This is a conservative estimate of the available bandwidth at the current moment, and the coefficient 0.95 is the flow control redundancy coefficient.
[0130] The cross-episode switching scenario is as follows: when the playback progress is at the end of the episode and the timestamp of the first frame of the next episode is less than or equal to 2 seconds, the episode switching optimization strategy is activated. The episode switching optimization strategy is as follows: the CMAF chunk of the InstantaneousDecoderRefresh frame corresponding to the first frame of the next episode and the following 2 to 3 CMAF chunks are pre-fetched. If the network level label is a weak network level, the CMAF chunk of the first screen of the next episode that is pre-fetched will preferentially use a lower bitrate.
[0131] Configure a fine-grained streaming degradation mechanism, which includes triggering conditions and degradation actions.
[0132] When any of the triggering conditions is met, the streaming degradation mechanism is triggered. The conditions include: the ratio of the conservative estimated available bandwidth within the next second to the average bitrate of the current bitrate tier is lower than the safety threshold of 1.10, and the current buffer duration is less than or equal to 0.25 seconds. The degradation action is as follows: the duration of subsequent CMAF chunks is adjusted from 0.5 seconds to fine-grained slices of 0.2 to 0.3 seconds, a pre-downgrading strategy is executed, and the adaptive bitrate tier is temporarily reduced by 1 tier. If it is during the episode switching stage, the newly fetched CMAF chunks will uniformly adopt the minimum bitrate tier.
[0133] Step S500.2: Configure the short drama serialization pre-pull mechanism and the weak network first screen optimization strategy, configure the emergency rollback strategy and the minimum viewing level guarantee mechanism, and generate the CMAchunk request priority queue and the streaming shaping status field.
[0134] In the scenario of short drama series playback, a pre-playback is performed at the end of the previous episode. The pre-playback includes: when the playback progress is at the end of the episode and the current time is between 1 and 2 seconds away from the timestamp of the first frame of the next episode, the CMAF chunk of the InstantaneousDecoderRefresh frame corresponding to the first frame of the next episode and the following 2 to 3 CMAF chunks are pre-fetched. When the network level label is a weak network level, the pre-fetched CMAF chunk adopts a lower first-screen strategy.
[0135] The timing control for the short drama pre-roll mechanism when switching episodes is set to 0.5 seconds.
[0136] An emergency rollback determination and execution mechanism is set up, which includes: emergency rollback trigger conditions and rollback actions. The emergency rollback determination is: two consecutive CMAF chunk fetch timeouts and a packet loss rate greater than or equal to 5%. The rollback action is: keep the resolution unchanged, reduce the bitrate and turn off high-complexity encoding tools, temporarily switch to protection buffer mode, and set the target buffer duration to 1.40 seconds. If the network level label does not belong to the weak network level for two consecutive seconds and the current buffer duration is greater than or equal to 0.90 seconds, then the normal adaptive strategy is restored.
[0137] Write the CMAchunk request priority queue to the pull control field area of the structured runtime state object.
[0138] In some specific embodiments, step S600 specifically includes: Step S600.1: Configure the playback quality feedback collection module and generate user-side playback experience metrics.
[0139] The player is configured with a playback quality feedback acquisition module. The playback quality feedback acquisition module collects and quantifies the user-side playback experience indicators and submits them as input to the playback strategy parameter tuning module. The acquisition indicators of the playback quality feedback acquisition module include: first frame presentation time, inter-episode switching response time, playback stuttering ratio, adaptive bitrate level switching frequency, 95th percentile inter-frame jitter, and next episode pre-pull hit rate.
[0140] The first frame presentation time is the time from when the player initiates a playback request to when the image is first decoded and presented. The inter-episode switching response time is the time from when the previous episode ends to when the first frame of the next episode is presented. The playback stuttering percentage is the percentage of cumulative stuttering time to total playback time recorded with a 1-hour statistical period, and the 95th percentile value is output. The adaptive bitrate gradation switching frequency is the number of times the adaptive bitrate gradation is switched per unit time. The 95th percentile inter-frame jitter is the 95th percentile standard deviation of the amplitude of the change in the interval between decoded frames. The next episode pre-pull hit rate is the proportion of pre-pulled CMAF chunks that are directly used for decoding and playback during the episode switching phase without being pulled a second time.
[0141] The playback quality feedback collection module is set to upload feedback every 5 seconds and performs outlier removal.
[0142] Step S600.2: Configure the online self-calibration rule set, set the trigger mechanism, generate the parameter tuning configuration set, and send it to the execution module.
[0143] The playback strategy parameter tuning module takes the user-side playback experience indicators output by the playback quality feedback collection module as input and dynamically adjusts them according to the online self-calibration rule set, including: playback stuttering ratio adjustment rules, first frame presentation time adjustment rules, and short episode switching response adjustment rules.
[0144] The rule for adjusting the percentage of playback stuttering is as follows: when the 95th percentile of the percentage of playback stuttering is greater than 0.5% / hour, the conservative strategy coefficient for bandwidth assessment is increased, and the lower limit of the minimum target buffer time is increased to 0.10 seconds to 0.20 seconds.
[0145] The first frame presentation time adjustment rule is as follows: when the first frame presentation time exceeds the target, the duration threshold of the first segment CMAF chunk of the first screen is adjusted to 300 milliseconds to 500 milliseconds, and the pull priority of the first segment CMAF chunk in the pull stream shaping module is increased by one discrete priority level.
[0146] The short episode switching response adjustment rule is as follows: when the episode switching response time exceeds the target of 0.5 seconds or the 95th percentile exceeds 0.8 seconds, the number of pre-pulled CMAF chunks will be adjusted from 2 to 3, or the episode switching advance will be adjusted to 2 to 3 seconds.
[0147] At the end of each decision cycle, the playback strategy parameter tuning module generates a parameter update set, which includes: bandwidth assessment conservative strategy coefficient, minimum target buffer duration lower limit, first screen first segment CMAF chunk duration threshold, number of pre-pulled CMAF chunks, set switching advance, adaptive bitrate tier upgrade limit parameter, and pull stream shaping granularity adjustment switch.
[0148] The parameter update set will be written into the playback strategy parameter tuning field area of the structured runtime state object and sent to the adaptive bitrate control module, buffer mode adjustment module and CMAFchunk pull stream shaping module. When the encoding end or the packaging end provides a parameter sending interface, it will be sent to the encoding end or the packaging end at the same time.
[0149] Configure a merging and evolution mechanism, which includes: a minimum interval period for parameter tuning of ≥5 seconds, a maximum of 3 consecutive adjustments per parameter, and a parameter tuning freeze period of 20 seconds after exceeding the limit. After each parameter update, an observation period of 10 seconds is entered. When the user-side playback experience indicators continue to improve, the next round of parameter tuning will be allowed. If the user-side playback experience indicators do not improve after the observation period, the system will revert to the previous parameter update set and record the revert status flag.
[0150] After each round of parameter tuning is completed, the playback strategy parameter tuning module writes the currently effective parameter set as a stable parameter set into the strategy evolution record area of the structured running status object. The record content includes: parameter change timestamp, stable parameter set, network level label at the time of recording, user-side playback experience index snapshot, and parameter tuning result mark.
[0151] The strategy evolution record area stores the set of stable parameters from the last 20 rounds for statistical analysis and strategy replay.
[0152] In practical applications, the player first configures a network status monitoring module to periodically collect instantaneous available throughput data, round-trip latency data, network jitter data, packet loss rate data, and CMAF chunk arrival interval data. The sampling period is 200 milliseconds. The encoding end outputs InstantaneousDecoderRefresh frame interval parameters, VideoBufferingVerifier buffer target parameters, HypotheticalReferenceDecoder buffer target parameters, and scene complexity labels. The HighDynamicRange tag information is injected into the packaging end, including the start and end timestamps of the title sequence, the start and end timestamps of the advertisement segments, the start and end timestamps of the end timestamp, the end timestamp of the entire series, and the first frame timestamp of the next episode. All information is uniformly written into a structured runtime state object. The end-to-end latency threshold is set to 1.2 seconds, the average target for the first frame presentation time is 0.8 seconds, and the 95th percentile target is 1.1 seconds. The threshold for the response time between episodes in a short series is 0.5 seconds, and the 95th percentile is 0.8 seconds. The 95th percentile threshold for the percentage of stuttering during playback is within 0.5% per hour.
[0153] Next, the player configures a bandwidth prediction module, establishing short-window and medium-window prediction paths. The short-window prediction path uses an exponentially weighted moving average based on instantaneous available throughput data to predict bandwidth for the next second. The medium-window prediction path uses regression analysis based on instantaneous available throughput data, network jitter data, and packet loss rate data to predict bandwidth for the next 2–3 seconds. When the error exceeds 5%, it is corrected by a lightweight recurrent neural network. The fused results form a conservative available bandwidth estimate R_cons, with an update period of 200 milliseconds. Based on the fluctuation range of packet loss rate data, network jitter data, and instantaneous available throughput data, the network level label is determined as stable network level, fluctuating network level, or weak network level.
[0154] Subsequently, the player is configured with a peak identification module, with a prediction window length set to 3 seconds and an update cycle of 200 milliseconds. The peak identification module identifies four types of bitrate peak risk periods: 0.5 seconds before and after the InstantaneousDecoderRefresh frame; the period from when the VideoBufferingVerifier buffer rate is 10% below the target lower limit until backfilling is complete; the 0.5–1 second period when the scene complexity label mutation level is greater than or equal to 1 and the next frame is a keyframe; and the cross-episode switching period when the playback progress is at the end of the episode and the timestamp of the first frame of the next episode is less than or equal to 2 seconds. The four types of risk periods are merged into a peak risk list and written into the encoding timing prediction field area of the structured runtime object.
[0155] Next, the buffer mode adjustment module takes into account the network level label, peak risk list, and conservative available bandwidth estimate, and outputs the target buffer duration. The low buffer mode range is 0.35–0.60 seconds, the standard buffer mode range is 0.60–1.20 seconds, and the protection buffer mode range is 1.20–1.60 seconds. The adaptive bitrate control module selects an adaptive bitrate level that meets the requirement that the average bitrate is less than or equal to 0.85 times the conservative available bandwidth estimate. It is allowed to increase the bitrate level once every 2 seconds. When the current buffer duration is less than 0.40 seconds, increasing the bitrate level is prohibited. During the end credits stage, when there are 0.5–2 seconds between the timestamp of the first frame of the next episode, it is prohibited to increase the bitrate level by two consecutive levels. The target buffer duration and adaptive bitrate level are written into the playback strategy decision field area of the structured runtime state object.
[0156] Then, the player is configured with a CMAFchunk streaming shaping module. In the first-screen loading scenario, the first CMAFchunk is fetched in segments, with the first half of the byte segment being fetched first. The peak fetch rate is set to a maximum of 0.95 times the conservative estimated available bandwidth. In the cross-episode switching scenario, the CMAFchunk of the InstantaneousDecoderRefresh frame corresponding to the first frame of the next episode and the following 2-3 CMAFchunks are pre-fetched. In weak network conditions, a lower bitrate is used, employing a fine-grained streaming degradation mechanism. Triggered when the conservative available bandwidth estimate is lower than the average bitrate of the current adaptive bitrate tier or the current buffer duration is less than or equal to 0.25 seconds, the CMAF chunk duration is adjusted to 0.2–0.3 seconds and the bitrate is temporarily reduced by one tier. The emergency fallback mechanism is triggered when two consecutive CMAF chunk fetch times out or the packet loss rate is greater than or equal to 5%, forcibly switching to the lowest visible tier, with the target buffer duration set to 1.40 seconds. The recovery condition is that the network level label is not a weak network level for 2 seconds and the current buffer duration is greater than or equal to 0.90 seconds.
[0157] Finally, the player is configured with a playback quality feedback collection module, with a period of 5 seconds, to collect the first frame presentation time, the response time of switching between episodes within a short drama, the percentage of playback stutters, the frequency of adaptive bitrate switching, the 95th percentile inter-frame jitter value, and the hit rate of the next episode pre-pull. The playback strategy parameter tuning module sets a set of rules based on the feedback: when the percentage of playback stutters is greater than 0.5% / hour, the bandwidth evaluation conservative strategy coefficient is increased and the buffer target is increased by 0.1–0.2 seconds; when the first frame presentation time exceeds the average target of 0.8 seconds and the 95th percentile target of 1.1 seconds, the CMAF chunk duration of the first segment of the first screen is shortened to 300–500 milliseconds and the priority is increased; when the response time of switching between episodes within a short drama exceeds the target of 0.5 seconds, the number of pre-pulls is increased to 3, and the advance is set to 2–3 seconds. The parameter update set is written to the playback strategy parameter tuning field area of the structured running state object, the stable parameter set is written to the strategy evolution record area, and 20 rounds of valid parameter sets are retained.
[0158] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for high-definition video encoding and low-latency transmission of short dramas for IPTV large screens, characterized in that, Includes the following steps: S100: Define end-to-end measurement boundaries and time references, configure the acquisition mechanism for real-time status indicators of the network environment, extract high-definition encoded metadata using the encoding end, mark timestamp information, and configure a unified structured running status object; S200: Configure the short window prediction path and the medium window prediction path, configure the input parameters of the short window prediction path and the medium window prediction path, obtain the prediction results of the short window prediction path and the medium window prediction path, fuse the prediction results to form a conservative available bandwidth estimate, and determine the network level label based on the network environment. The S300 uses a peak identification module based on high-definition encoding metadata to identify peak risk periods triggered by IDR frames, peak risk periods triggered by video buffer verifier buffer backfilling, peak risk periods caused by sudden changes in shot complexity, and peak risk periods during cross-episode switching in short dramas, and generates a peak risk list.
2. The method for high-definition video encoding and low-latency transmission of short dramas for IPTV large screens as described in claim 1, characterized in that, Also includes: S400: Configure the buffer mode adjustment module and set the buffer mode boundary range; configure the adaptive bitrate selection module and buffer control module and set the switching constraint mechanism to generate a joint decision output of buffer target and bitrate level. S500, configure a segment-first CMAFchunk streaming shaping module, configure a fine-grained streaming degradation mechanism and trigger dynamic granular adjustment, configure a short drama consecutive playback pre-pull mechanism and a weak network first screen optimization strategy, configure an emergency rollback strategy and a minimum viewing level guarantee mechanism, and generate a CMAchunk request priority queue and streaming shaping status field. The S600 is configured with a playback quality feedback collection module, which generates user-side playback experience metrics, configures an online self-calibration rule set, sets a trigger mechanism, generates a parameter tuning configuration set, and sends it to the execution module.
3. The method for high-definition video encoding and low-latency transmission of short dramas for IPTV large screens as described in claim 1, characterized in that, Specifically, S100 is as follows: S100.1 Define the end-to-end measurement boundaries and time references, and configure the acquisition mechanism for real-time status indicators of the network environment; In the scenario of playing short dramas on large screens for IPTV platforms, quantitative playback control targets are established, including: end-to-end latency threshold, first frame appearance time, episode switching response time in short dramas, and percentage of stuttering during playback. A network status monitoring module is configured on the player side to collect instantaneous available throughput data, round-trip latency data, network jitter data, packet loss rate data, and CMAFchunk arrival interval data. The sampling period for instantaneous available throughput data is set to 200 milliseconds, and the average value is calculated within a sliding window of about 1 second. The round-trip latency data is calculated by interacting with timestamped probe messages and acknowledgment messages to calculate the round-trip transmission time between the server and the player. The network jitter data is calculated by the standard deviation of the round-trip latency difference of continuous probe messages. The packet loss rate data is calculated within a fixed statistical window of 1 second, showing the ratio of the number of lost data packets to the total number of data packets. The CMAFchunk arrival interval data records the arrival time interval sequence of CMAFchunk data in the player buffer. Real-time network environment status indicators are generated based on instantaneous available throughput data, round-trip time data, network jitter data, packet loss rate data, and CMAF chunk arrival interval data.
4. The method for high-definition video encoding and low-latency transmission of short dramas for IPTV large screens as described in claim 3, characterized in that, The S100 further includes: S100.2 uses the encoding end to extract high-definition encoded metadata, and marks timestamp information and configures a unified structured running status object; The key encoding metadata is collected at the encoding end. The key encoding metadata includes: IDR frame interval parameters, video buffer verifier buffer target parameters and HRD buffer target parameters, shot scene complexity labels and high dynamic range marking information; High-definition encoded metadata is formed based on IDR frame interval parameters, video buffer verifier buffer target parameters and HRD buffer target parameters, shot scene complexity labels and high dynamic range marker information; During the packaging stage, the key encoded metadata will be written into the CMAFchunk header information or bypass signaling channel; In response to the highly fragmented content and frequent episode switching characteristics of short dramas, the content provider will inject structural scene boundary timestamp information during the packaging stage; The collected real-time network environment status indicators, high-definition encoded metadata, and scene boundary timestamp information are integrated and configured into a unified structured runtime status object.
5. The method for high-definition video encoding and low-latency transmission of short dramas for IPTV large screens as described in claim 1, characterized in that, Specifically, S200 is as follows: S200.1 Configure the short window prediction path and the medium window prediction path, configure the input parameters of the short window prediction path and the medium window prediction path, and obtain the prediction results of the short window prediction path and the medium window prediction path. Configure the bandwidth prediction module in the player and establish a dual-path bandwidth prediction structure that includes a short-window prediction path and a medium-window prediction path. The input of the bandwidth prediction module is the set of network state parameters in the structured runtime state object. The short-window prediction path is as follows: an exponentially weighted moving average is performed on the instantaneous available throughput data within the sliding window of the last 1 second to obtain the 1-second prediction value; The mid-window prediction path is as follows: perform least squares time-series regression on the instantaneous available throughput data, network jitter data, and packet loss rate data within the sliding window of the last 3 seconds to obtain the 2-3 second prediction value; When the relative fitting error of least squares regression exceeds 5%, a lightweight recurrent neural network model is called to correct the prediction. The lightweight recurrent neural network model adopts a 2-layer structure and 8 hidden nodes, and the single inference time does not exceed 5 milliseconds.
6. The method for high-definition video encoding and low-latency transmission of short dramas for IPTV large screens as described in claim 5, characterized in that, The S200 further includes: S200.
2. Based on the prediction results, a conservative estimate of available bandwidth is formed by fusion, and the network level label is determined based on the network environment; The prediction results calculated from the short window prediction path and the medium window prediction path are used to perform confidence interval correction and lower limit fusion, and a conservative available bandwidth estimate is output. The network level label is determined by using the set of network status parameters in the structured runtime state object. The determination rules include: stable network level, fluctuating network level, and weak network level. The determination of the network stability level requires meeting three conditions: packet loss rate <1%, network jitter <10 milliseconds, and the relative fluctuation of instantaneous available throughput data within the past 2-second window. The determination of the fluctuation network level requires that any one of the following conditions be met: 1) The conditions are: packet loss rate data is in the range of 1% to 3%, network jitter data is in the range of 10 to 25 milliseconds, the relative decrease of instantaneous available throughput data within the past 2-second window is >15%, and the weak network level must meet any of the following conditions; 2) The conditions are: packet loss rate ≥ 3%, network jitter ≥ 25 milliseconds, relative decrease of instantaneous available throughput data within the last 2-second window > 30%, network grade label update cycle is 200 milliseconds, and network grade label submission frequency is once every 2 seconds.
7. The method for high-definition video encoding and low-latency transmission of short dramas for IPTV large screens as described in claim 1, characterized in that, Specifically, S300 is as follows: S300.1, Peak recognition module configured based on high-definition encoded metadata; Configure a peak recognition module in the player to identify high-risk periods where bitrate may suddenly increase in the future. The input of the peak recognition module is the set of encoding metadata parameters and scene boundary information in the structured running state object. The peak recognition module performs periodic rolling recognition according to a unified time base. The prediction window length is set to 3 seconds and the update period is 200 milliseconds. S300.
2. Based on the peak identification module, identify the peak risk period, peak risk period, peak risk period and peak risk period of bitrate, and generate a peak risk list; The peak risk period is identified by the peak identification module as being triggered by IDR frames, triggered by video buffer verifier buffer backfilling, caused by sudden changes in shot complexity, and during cross-episode switching in short dramas. When the buffer fill rate falls below 10% of the target lower limit, backfilling is immediately triggered and that period is marked as a peak risk zone. The peak risk period caused by sudden changes in lens complexity is defined as follows: a sudden increase in the complexity of the lens scene will lead to a significant increase in the instantaneous bit rate, and the lens complexity variation rate is defined as follows; when If the next frame type is an I-frame or an InstantaneousDecoderRefresh frame, then the next 0.5–1 seconds will be predicted as the period of peak risk of complexity surge. The peak bitrate risk period during cross-episode switching of short dramas is as follows: During the playback of short dramas, when the timestamp of the first frame of the next episode is less than or equal to 2 seconds and the current playback stage is the end credits stage, the potential peak cross-episode switching period needs to be identified in advance. If the conditions are met, then define the peak risk period for cross-set switching; The results obtained from the peak bitrate risk periods triggered by IDR frames, the peak bitrate risk periods triggered by video buffer verifier buffer backfilling, the peak bitrate risk periods caused by sudden changes in shot complexity, and the peak bitrate risk periods during cross-episode switching in short dramas are integrated to form a peak risk list. The refresh cycle of the peak risk list is set to 200 milliseconds. Within each refresh cycle, the player merges and deduplicates the four types of time period lists and writes the merged peak risk list into the encoded timing prediction field area of the structured running status object. Based on the start and end times and priority information recorded in the four types of time period lists in the peak risk list, the player executes advance control, streaming shaping and throttling strategies.
8. A method for high-definition video encoding and low-latency transmission of short dramas for IPTV large screens as described in claim 2, characterized in that, Specifically, S400 is: S400.1 Configure the buffer mode adjustment module and set the buffer mode boundary range; In the player configuration, the buffer mode adjustment module is configured. The input to the buffer mode adjustment module is: network level label, peak risk list and conservative available bandwidth estimate. The network level label can be set to: stable network level, fluctuating network level or weak network level. Three buffering modes and boundary ranges are set: low buffering mode, standard buffering mode, and protective buffering mode. The switching of the three buffer modes adopts hysteresis control. Both the upward and downward switching require that the switching conditions be met continuously for no less than 1 second. S400.2 Configure the adaptive bitrate selection module and buffer control module, and set the switching constraint mechanism to generate a joint decision output of buffer target and bitrate level; The player is configured with an adaptive bitrate control module and a buffer control module. The adaptive bitrate control module selects the bitrate level based on a conservative estimate of available bandwidth and follows constraints including: the average bitrate of the bitrate level satisfies... When the player's current buffering time If the time interval is less than 0.40 seconds, the bitrate will be increased, and only the current bitrate level can be maintained or the bitrate level can be decreased. Within the stable time window of episode switching, it is prohibited to increase the bitrate level by two consecutive levels, and only one bitrate level can be increased or the current bitrate level can be maintained. The player generates a joint decision output based on the network grade label and peak risk list, the overlap of risk periods within a 1–3 second prediction window, a conservative estimate of available bandwidth, the current buffer duration and buffer mode boundary, and the average bitrate of the current bitrate tier. The output includes the target buffer duration. Take the median value according to the selected buffer mode range, and for low buffer mode, take... Standard buffer mode takes Seconds, protection buffer mode take Seconds, adaptive bitrate level In order to satisfy Provided that the conditions for upgrading are not violated, the highest feasible bitrate level should be selected.
9. A method for high-definition video encoding and low-latency transmission of short dramas for IPTV large screens as described in claim 2, characterized in that, Specifically, S500 is as follows: S500.1 Configure a segment-first CMAFchunk streaming shaping module, configure a fine-grained streaming degradation mechanism, and trigger dynamic granularity adjustment; Configure the CMAFchunk streaming shaping module in the player to dynamically optimize and control the CMAFchunk fetching order, segmentation strategy, and streaming rhythm. The CMAFchunk streaming shaping module executes a keyword-first fetching strategy for first-screen loading scenarios and cross-set switching scenarios; Configure a fine-grained streaming degradation mechanism, which includes: triggering conditions and degradation actions; When any of the triggering conditions is met, the streaming degradation mechanism is triggered. The conditions include: the ratio of the conservative estimated available bandwidth in the next second to the average bitrate of the current bitrate level is lower than the safety threshold of 1.10, and the current buffer duration is less than or equal to 0.25 seconds. The downgrade action is as follows: the duration of subsequent CMAF chunks is adjusted from 0.5 seconds to fine-grained slices of 0.2 to 0.3 seconds, a pre-downgrade strategy is implemented, and the adaptive bitrate level is temporarily reduced by 1 level. If it is during the episode switching stage, the newly fetched CMAF chunks will uniformly adopt the minimum bitrate level. S500.2, configure the short drama serialization pre-pull mechanism and the weak network first screen optimization strategy, configure the emergency rollback strategy and the minimum viewing level guarantee mechanism, and generate the CMAchunk request priority queue and the streaming shaping status field; In the scenario of short drama series playback, a pre-playback is performed at the end of the previous episode. The pre-playback includes: when the playback progress is at the end of the episode and the current time is between 1 and 2 seconds away from the timestamp of the first frame of the next episode, the CMAF chunk of the InstantaneousDecoderRefresh frame corresponding to the first frame of the next episode and the following 2 to 3 CMAF chunks are pre-fetched. When the network level label is a weak network level, the pre-fetched CMAF chunk adopts a lower first-screen strategy. An emergency rollback determination and execution mechanism is set up, which includes: emergency rollback trigger conditions and rollback actions. The emergency rollback determination is: two consecutive CMAF chunk fetch timeouts and a packet loss rate greater than or equal to 5%. The rollback action is as follows: keep the resolution unchanged, reduce the bitrate and turn off the high-complexity encoding tool, temporarily switch to the protection buffer mode, and set the target buffer duration to 1.40 seconds. If the network level label does not belong to the weak network level within 2 consecutive seconds, and the current buffer duration is greater than or equal to 0.90 seconds, then revert to the normal adaptive strategy.
10. A method for high-definition video encoding and low-latency transmission of short dramas for IPTV large screens as described in claim 2, characterized in that, Specifically, S600 is as follows: S600.1 Configure the playback quality feedback collection module and generate user-side playback experience metrics; The player is configured with a playback quality feedback acquisition module. The playback quality feedback acquisition module collects and quantifies the user-side playback experience indicators and submits them as input to the playback strategy parameter tuning module. The acquisition indicators of the playback quality feedback acquisition module include: first frame presentation time, inter-episode switching response time, playback stuttering ratio, adaptive bitrate level switching frequency, 95th percentile inter-frame jitter, and next episode pre-pull hit rate. S600.2 Configure the online self-calibration rule set, set the trigger mechanism, generate the parameter tuning configuration set, and send it to the execution module; The playback strategy parameter tuning module takes the user-side playback experience indicators output by the playback quality feedback collection module as input, and dynamically adjusts them according to the online self-calibration rule set, including: playback stuttering ratio adjustment rules, first frame presentation time adjustment rules, and short episode switching response adjustment rules. At the end of each decision cycle, the playback strategy parameter tuning module generates a parameter update set, which includes: bandwidth assessment conservative strategy coefficient, minimum target buffer duration lower limit, first screen first segment CMAF chunk duration threshold, number of pre-pulled CMAF chunks, set switching advance, adaptive bitrate tier upgrade limit parameter, and pull stream shaping granularity adjustment switch. The parameter update set will be written into the playback strategy parameter tuning field area of the structured running state object and sent to the adaptive bitrate control module, buffer mode adjustment module and CMAFchunk pull stream shaping module. When the encoding end or the packaging end provides a parameter sending interface, it will be sent to the encoding end or the packaging end at the same time. Configure a merging and evolution mechanism, which includes: a minimum interval period for parameter tuning of ≥5 seconds, a maximum of 3 consecutive adjustments for a single parameter, and a parameter tuning freeze period of 20 seconds after exceeding the limit. After each parameter update, an observation period of 10 seconds is entered. When the user-side playback experience index continues to improve, the next round of parameter tuning will be allowed. If the user-side playback experience index does not improve after the observation period, the system will revert to the previous parameter update set and record the revert status flag. After each round of parameter tuning is completed, the playback strategy parameter tuning module writes the currently effective parameter set as a stable parameter set into the strategy evolution record area of the structured running state object.
Citation Information
Cited By
High-definition wireless video stream data compression method based on adaptive code rate
CN121644809A