Video stream decoding delay prediction and dynamic quality optimization method based on multi-feature fusion
Through the hybrid machine learning model and dynamic video quality optimization strategy of multi-feature fusion, the problem of insufficient decoding delay prediction and video quality optimization in the existing technology is solved, efficient and stable video streaming is achieved, and user experience is improved.
Patent Information
- Application Number
- CN202510326052.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-19
AI Technical Summary
The existing adaptive streaming technology has shortcomings in decoding latency prediction and video quality optimization, resulting in poor user experience, especially in 4K/8K high-resolution scenarios.
Through multi-feature fusion, a hybrid machine learning model is constructed to predict decoding delay, and combined with model prediction control and reinforcement learning strategies, dynamically adjust the CRF value of the video quality encoding parameter to optimize video streaming.
It significantly improves the performance of the video transmission system, accurately predicts decoding delays, stabilizes and optimizes video quality, reduces playback lags and quality fluctuations, and improves user experience.
Smart Images

Figure CN120151531A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of video streaming media transmission, and specifically relates to a dynamic adaptive optimization system for ultra-high definition video transmission. By integrating video coding features, content complexity analysis, and end-side device performance perception, a decoding delay prediction model is constructed, and video quality decisions are jointly optimized in combination with the transmission state. It is applicable to streaming media scenarios that are sensitive to end-to-end delay, such as VR real-time interaction and 4K ultra-high definition live broadcast. Background Art
[0002] With the development of streaming media technology, users' demand for high-quality video content is increasing day by day. Existing adaptive streaming media technologies (such as DASH, HLS) usually adjust video quality based on network bandwidth and buffer status, but ignore the impact of decoding delay on end-to-end delay. However, in practical applications, decoding delay is one of the important factors affecting the user experience. Traditional decoding delay prediction methods usually rely on static look-up tables or simple linear models and cannot adapt to complex dynamic environments (such as network fluctuations, device load changes, etc.). Especially in high-resolution scenarios such as 4K / 8K, the decoding time fluctuates significantly with video content complexity, coding parameters (such as Constant Rate Factor (CRF) value), and device performance, which may lead to playback jitter or quality mutation.
[0003] Defects of the prior art: 1) Coarse-grained decoding delay modeling: Only relying on a static bitrate-delay mapping table, without considering dynamic content features (such as scene switching frequency, object movement speed) and real-time device load fluctuations; 2) Lack of multi-module cooperation: The network transmission optimization and decoding resource allocation are separated, resulting in an end-to-end delay prediction error exceeding 20%; 3) Deterioration of user experience: Frequent quality switches (>5 times / minute) cause visual jitter, and the variance of Video Multimethod Assessment Fusion (VMAF) quality fluctuations exceeds 30%.
[0004] Therefore, there is an urgent need for a method that can combine multiple features, dynamically adjust, and accurately predict decoding delay to optimize the streaming media playback experience. Summary of the Invention
[0005] In order to overcome the deficiencies of the prior art, the present invention proposes a method for predicting video stream decoding delay and dynamically optimizing quality based on multi-feature fusion. By using a machine learning model to predict the decoding time of video blocks at different CRF levels, and combining the transmission delay prediction results, quality selection is optimized under the end-to-end delay constraint, and fluctuations are suppressed.
[0006] The technical solution adopted by the present invention to solve its technical problems is:
[0007] A video stream decoding delay prediction and dynamic quality optimization method based on multi-feature fusion, comprising the following steps:
[0008] Step 1, video stream feature extraction, the process is as follows:
[0009] Step 1.1: Extract video encoding parameters, including CRF, bitrate, group of pictures structure, and frame rate;
[0010] Step 1.2: Content complexity analysis, including parsing motion vectors and calculating the motion intensity index MVI, and analyzing the discrete cosine transform (DCT) texture entropy;
[0011] Step 1.3: Device status monitoring, including collecting the GPU utilization rate U GPU and the video memory bandwidth occupancy rate B mem to construct a device performance index DPI;
[0012] Step 2, video stream decoding delay prediction: Design a hybrid model architecture to predict the decoding delay, that is, adopt a LightGBM sub-model and an LSTM residual compensation sub-model;
[0013] Step 3, video stream quality dynamic decision-making: Combine two methods of an improved model predictive control (MPC) algorithm and a Q-learning reinforcement learning strategy. MPC is used under normal circumstances, considering delay constraints and quality scores, allowing a certain degree of delay overrun; when the network fluctuates violently, switch to Q-learning, and online adjust the decision through reinforcement learning, while restricting the adjustment amplitude to maintain quality stability, and finally determine the sequence output of the CRF value within a rolling window;
[0014] Step 4, video stream smoothing processing: When playing the video, adopt the video quality encoding CRF value determined by the final decision. This encoding parameter reflects the presented video quality; maximize the bitrate R max under the current network to match the optimal picture quality determined by the CRF decision. The video playback adopts double-buffer progressive rendering and bitrate smoothing filtering;
[0015] Step 5, after completing the video stream smoothing processing, transmit the processed video data to a display device for rendering and presentation.
[0016] Furthermore, in the said Step 1.1, CRF is a video quality encoding level that dynamically captures the current video block. The lower the value, the higher the picture quality but the larger the bitrate; the bitrate is the average bitrate of the current video segment calculated based on the encoding parameters; the GOP structure is to extract the I-frame interval, the number of P-frames, and B-frames, and calculate the average GOP length; the frame rate is to parse the number of frames per second (FPS) transmitted from the video stream header.
[0017] Furthermore, in step 1.2, MVI obtains the motion vector components (Δx, Δy) of each macroblock from the decoder interface and calculates the motion amplitude according to the macroblock:
[0018]
[0019] The full-frame MVI value is the weighted average of all macroblocks, and the weight is the proportion of the macroblock area:
[0020]
[0021] where w i is the area weight coefficient of the i-th macroblock, N is the total number of intra-frame macroblocks, Δx i is the motion vector component of the i-th macroblock in the X direction, and Δy i is the motion vector component of the i-th macroblock in the Y direction;
[0022] DCT texture entropy quantifies the image texture complexity through discrete cosine transform DCT, including:
[0023] 8×8 block and DCT transform: Divide the luminance component of the image into multiple 8×8 pixel blocks, perform DCT transform on each block to obtain the corresponding frequency coefficient matrix, denoted as C k , that is, the k-th DCT coefficient;
[0024] Normalized energy distribution calculation: Calculate the total energy of each 8×8 block Normalize the energy proportion of each DCT coefficient:
[0025]
[0026] where p(k) represents the normalized energy distribution probability of the k-th DCT coefficient;
[0027] Calculate the full-frame texture entropy H texture :
[0028]
[0029] Furthermore, in step 1.3, the GPU utilization rate is obtained by real-time obtaining the GPU occupancy rate through the driver API, and the video memory bandwidth occupancy rate is obtained by monitoring the ratio of the video memory read / write bandwidth to the theoretical peak value;
[0030] Construct DPI:
[0031] DPI = 0.7U GPU + 0.3B mem .
[0032] The process of step 2 is as follows:
[0033] Step 2.1: Input the static encoding parameters, namely CRF, bitrate, GOP, and frame rate, into the pre-trained LightGBM model to output the predicted value T of the basic decoding time base ;
[0034] Step 2.2: Input the temporal dynamic features, namely the MVI sequence, DCT texture entropy, and DPI sequence, into the LSTM residual compensation model to predict the decoding time residual correction value ΔT;
[0035] Step 2.3: Calculate the final predicted value T of the decoding delay comprehensively pred :
[0036] T pred = T base + ΔT
[0037] Step 2.4: Adopt an incremental learning strategy. When the single prediction error > 15% or the error > 10% for three consecutive times, trigger the online update mechanism to optimize and update the parameters of the LightGBM model and the LSTM residual compensation model with the recently set number of samples.
[0038] The process of Step 3 is as follows:
[0039] Step 3.1: Design the MPC controller and establish an optimization model with end-to-end delay constraints:
[0040] Rolling window optimization:
[0041]
[0042] Among them, H represents the time length of the optimization window, and RTT represents the network round-trip time (Round-TripTime) in the current environment
[0043] Objective function:
[0044]
[0045] Among them, Q(CRF k ) is the quality score calculated by the VMAF algorithm developed by Netflix when encoding a video with a certain CRF value; CRF k and CRF k-1 are the CRF numerical values of the video quality encoding parameters of the current video block k and the previous video block k-1 respectively; w 1 and w 2 are the weight coefficients; t is the current video playback time progress;
[0046] Optimization with constraint relaxation:
[0047] T trans (CRF k ) + Tpred (CRF k ) ≤ 1.1T max
[0048] Wherein, T trans is the transmission delay when the video quality coding parameter of video block k is CRF under the current bandwidth, and T pred is the predicted decoding processing delay when the video quality coding parameter of video block k is CRF under the current bandwidth; T max is the required delay corresponding to the target frame rate of video block k; 1.1T max allows a 10% delay overrun;
[0049] MPC trigger policy Trigger mpc : When the bandwidth standard deviation in the past 5 seconds < 15% of the average bandwidth, the final continuously used video quality CRF numerical sequence CRF final is the CRF numerical sequence CRF determined by enabling MPC, that is: mpc i.e.:
[0050] Trigger mpc : Bandwidth std5 <15% × Bandwidth mean ,
[0051] CRF final =CRF mpc ;
[0052] Wherein, Bandwidth std5 is the bandwidth standard deviation STD in the past 5 seconds, and Bandwidth mean is the average bandwidth;
[0053] Step 3.2. When the network environment fluctuates violently, switch to the reinforcement learning policy and optimize the decision online through the Q-learning algorithm. The process is as follows:
[0054] State encoding s t : Encode the network bandwidth Bandwidth t , the predicted decoding delay value T pred , the device DPI index, and the historical video coding parameter CRF sequence into a feature vector at the current moment t:
[0055] s t =Concat(Bandwidth t , T pred , DPI t , CRF t-1 )
[0056] Action space constraint at : Limit the CRF adjustment step (-2, -1, 0, +1, +2) to prevent drastic quality fluctuations;
[0057] a t ∈{-2, -1, 0, +1, +2}
[0058] Limit: After adjustment, the CRF needs to be within the range of [15, 45], and if it exceeds, it will be truncated to the nearest boundary value;
[0059] Reward function R: Comprehensively consider video quality, latency, and stability metrics:
[0060] R = Q(CRF t ) - 10×I overdelay - 5×|CRF t - CRF t-1 | + 2×I smooth
[0061] where t is the current video playback time progress; Q(CRF t ) is the quality score calculated by the VMAF algorithm developed by Netflix when encoding a video with a certain CRF value; CRF t and CRF t-1 are the CRF numerical values of the video quality encoding parameters for the current and previous video blocks respectively; CRF t - CRF t-1 represents whether there is a CRF quality switch when detecting the playback of the video; I overdelay is a binary indicator function when the actual latency of playing the current video block exceeds the required latency corresponding to the target frame rate, which is 1 when the latency exceeds the threshold and 0 otherwise; I smooth The stability metric mechanism is triggered when there is no CRF quality switch for 3 consecutive times, encouraging long-term stable decisions, and is a binary indicator function, which is 1 when triggered and 0 otherwise;
[0062] Q-learning trigger strategy Trigger ql : When the standard deviation of the bandwidth in the past 5 seconds is between 15% and 20% of the average bandwidth, trigger the Q-learning strategy, and the sequence of CRF numerical values CRF final of the video quality continuously used is the sequence of CRF numerical values CRF ql decided by enabling Q-learning, that is:
[0063] Trigger ql : 15%×Bandwidth mean <Bandwidth std5 <20%×Bandwidth mean ,
[0064] CRF final = CRF ql ;
[0065] Among them, Bandwidth std5 is the standard deviation STD of the bandwidth in the past 5 seconds, and Bandwidth mean is the average bandwidth;
[0066] Step 3.3: In the rolling time domain window H, when the network condition changes drastically, that is, when the standard deviation of the bandwidth in the past 5 seconds > 20% of the average bandwidth, through the cooperation of MPC and Q-learning, at this time, the outputs of MPC and Q-learning are fused according to the weight, and the CRF value of the video coding parameter is adaptively adjusted, that is:
[0067] Trigger 协同 : Bandwidth std5 > 20% × Bandwidth mean ,
[0068] CRF final = 0.7CRF mpc + 0.3CRF ql .
[0069] Among them, Bandwidth std5 is the standard deviation STD of the bandwidth in the past 5 seconds, and Bandwidth mean is the average bandwidth;
[0070] Duration constraint: When the network state changes from the fluctuating state back to the stable state, it is necessary to maintain the cooperative strategy for at least 3 more seconds before switching back to the MPC strategy, to avoid misjudgment caused by short-term network jitter, resulting in repeated jumps of the CRF value and causing video quality jitter, and to ensure that the decision of strategy switching is based on continuous state changes rather than instantaneous noise. In this way, the video quality is maximized on the premise of controllable delay, taking into account both real-time performance and robustness.
[0071] The process of step 4 is as follows:
[0072] Step 4.1: Double-buffered progressive rendering to ensure smooth picture transition. The process is:
[0073] Step 4.1.1: Piece preloading, and the background thread preloads the next piece into the Back Buffer in advance;
[0074] Step 4.1.2: Use time-varying mixed weights for picture transition, as the picture P finally displayed by the current video player display , to achieve a smooth transition from the old piece to the new piece:
[0075] P display = α t P front +(1 - α t )P back
[0076] where P front is the decoded frame of the currently playing old segment; P back is the decoded frame of the background pre - loaded segment; α decays exponentially with time, α = e -t / τ , and τ is set according to the device refresh rate;
[0077] Step 4.1.3: Dynamic switching trigger: Start pixel - level blending transition when the current segment plays to 50%;
[0078] Step 4.2: Bitrate smoothing filter. Based on the video quality determined by CRF, dynamically adjust the bitrate to adapt to the real - time bandwidth, balancing quality and stability. The process is as follows:
[0079] Step 4.2.1 Dynamic gradient limit:
[0080]
[0081] where σ bandwidth is the standard deviation of the bandwidth in the past 10 seconds;
[0082] When the bandwidth fluctuation in the past 10 seconds is small, allow the maximum single - time bitrate adjustment amplitude to reach 30% to quickly match the available bandwidth and improve video quality; if the bandwidth fluctuates violently, limit the adjustment amplitude within 15% to avoid frequent switching causing buffering;
[0083] Step 4.2.2 Bitrate adjustment, suppressing bitrate mutations and balancing video quality stability and bandwidth utilization:
[0084] R next = γR current +(1 - γ)R target
[0085] where R current is the bitrate at the current moment; R target is the target bitrate (determined by CRF); R next is the bitrate at the next moment; the smoothing factor γ is dynamically adjusted according to the degree of network oscillation:
[0086]
[0087] When the network is stable, the smoothing factor is reduced (γ≈0.7) to make the bit rate quickly approach the target value; when the network is turbulent, the smoothing factor is increased (γ can reach 0.9), the historical bit rate weight is strengthened, and sudden fluctuations are suppressed; through the nonlinear adjustment of the tanh function, the bandwidth standard deviation is mapped to the interval [0,0.4], and nonlinear adjustment of γ between 0.7 and 0.9 is achieved, ensuring that the bit rate adjustment is more robust when the network fluctuates.
[0088] The present invention effectively solves the key problems existing in the existing adaptive streaming media technology by integrating multi-feature perception and intelligent optimization decision-making, and significantly improves the overall performance of the video transmission system.
[0089] The beneficial effects of the present invention are mainly manifested in:
[0090] 1. By combining the dynamic features of video content (such as motion intensity and texture complexity) with the real-time performance data of terminal devices (such as GPU load and video memory usage), a hybrid machine learning model is constructed to more accurately predict the decoding time under different encoding parameters, overcoming the problem that the traditional static table lookup method is not adaptable enough in dynamic scenes;
[0091] 2. Introducing a model predictive control strategy with a constraint relaxation mechanism, while ensuring a strict end-to-end delay upper limit, allowing short-term controllable over-limit in exchange for long-term quality stability, effectively reducing the risk of playback interruption caused by network jitter or sudden changes in device load
[0092] 3. Double-buffered progressive rendering technology is used to achieve seamless transition of different quality slices through time-varying weighted hybrid algorithm, eliminating visual freezes caused by sudden changes in the picture. Combined with dynamic bit rate smoothing control strategy, frequent quality adjustments caused by network fluctuations are suppressed, significantly improving visual coherence. BRIEF DESCRIPTION OF THE DRAWINGS
[0093] Figure 1 This is a system architecture diagram of the present invention;
[0094] Figure 2 A hybrid model diagram for predicting video stream decoding delay according to the present invention;
[0095] Figure 3 This is a diagram of the dynamic decision model for video stream quality according to the present invention. DETAILED DESCRIPTION
[0096] The present invention will be further described below in conjunction with the accompanying drawings.
[0097] Reference Figure 1 , Figure 2 and Figure 3 , a video stream decoding delay prediction and dynamic quality optimization method based on multi-feature fusion, comprising the following steps:
[0098] Step 1. Video stream feature extraction, the process is as follows:
[0099] Step 1.1: Obtain the encoding parameter set in real time through the video decoder interface, including the CRF value, bit rate, GOP structure, and frame rate;
[0100] The CRF is the video quality encoding level for dynamically capturing the current video block. The lower the value, the higher the image quality but the larger the bit rate. In this embodiment, its range is 15 - 45;
[0101] The bit rate is the average bit rate of the current video segment calculated based on the encoding parameters. In this embodiment, its range is 2 - 50 Mbps;
[0102] The GOP structure is to extract parameters such as the I - frame interval, the number of P - frames, and B - frames, and calculate the average GOP length;
[0103] The frame rate is parsed from the video stream header as the number of frames transmitted per second. In this embodiment, its range is 24 - 120 fps;
[0104] Step 1.2: Parse the motion vector data of the video stream. The MVI is to obtain the motion vector components (Δx, Δy) of each macroblock from the decoder interface, and calculate the motion amplitude according to the macroblock:
[0105]
[0106] The full - frame MVI value is the weighted average of all macroblocks, and the weight is the proportion of the macroblock area:
[0107]
[0108] where, w i is the area weight coefficient of the i - th macroblock, N is the total number of intra - frame macroblocks, Δx i is the motion vector component of the i - th macroblock in the X direction, and Δy i is the motion vector component of the i - th macroblock in the Y direction;
[0109] Step 1.3: Perform 8×8 block DCT transformation on the luminance component in YUV format, and calculate the texture entropy value of each block, including:
[0110] 8×8 block and DCT transformation: Divide the luminance component of the image into multiple 8×8 pixel blocks, perform DCT transformation on each block, and obtain the corresponding frequency coefficient matrix, denoted as C k , that is, the k - th DCT coefficient;
[0111] Normalized energy distribution calculation: Calculate the total energy of each 8×8 block Normalize the energy proportion of each DCT coefficient:
[0112]
[0113] Among them, p(k) represents the normalized energy distribution probability of the k-th DCT coefficient;
[0114] Calculate the full-frame texture entropy H texture :
[0115]
[0116] Step 1.4: Monitor the GPU utilization rate U GPU of the terminal device, in %, and the video memory bandwidth occupancy rate B mem of the terminal device, in GB / s, and construct a device performance index:
[0117] DPI = 0.7U GPU + 0.3B mem ;
[0118] Step 2: Video stream decoding delay prediction, refer to Figure 2 , and the process is as follows:
[0119] Step 2.1: Input the static encoding parameters (CRF, bit rate, GOP, frame rate) into the pre-trained LightGBM model, and output the basic decoding time prediction value T base ;
[0120] Step 2.2: Input the temporal dynamic features (MVI sequence, DCT texture entropy, and DPI sequence) into the LSTM residual compensation model to predict the decoding time residual correction value ΔT;
[0121] Step 2.3: Comprehensively calculate the final decoding delay prediction value T pred :
[0122] T pred = T base + ΔT
[0123] Step 2.4: Adopt an incremental learning strategy. When the single prediction error > 15% or the continuous three errors > 10%, trigger the online update mechanism to optimize and update the parameters of the LightGBM model and the LSTM residual compensation model with the latest 100 samples;
[0124] Step 3: Video stream quality dynamic decision-making: Combines two methods of the improved model predictive control (MPC) algorithm and the Q-learning reinforcement learning strategy. MPC is used in the case of a normal and stable network, considering delay constraints and quality scores, allowing a certain delay overrun; when the network fluctuates violently, switch to Q-learning, and online adjust the decision through reinforcement learning while restricting the adjustment range to maintain quality stability. Finally, determine the sequence output of the CRF value within the rolling window, refer to Figure 3, the process is as follows:
[0125] Step 3.1: Design of the MPC controller, establishing an optimization model with end-to-end delay constraints. The specific process is as follows:
[0126] Rolling window optimization:
[0127]
[0128] Among them, H represents the time length of the optimization window, and RTT represents the network round-trip time (Round-Trip Time) in the current environment;
[0129] Objective function:
[0130]
[0131] Among them, Q(CRF k ) is the quality score calculated by the VMAF algorithm developed by Netflix when encoding a video with a certain CRF value; CRF k and CRF k-1 are the CRF numerical values of the video quality encoding parameters of the current video block k and the previous video block k - 1 respectively; w 1 and w 2 are the weight coefficients; t is the current video playback time progress;
[0132] Optimization with constraint relaxation:
[0133] T trans (CRF k ) + T pred (CRF k ) ≤ 1.1T max
[0134] Among them, T trans is the transmission delay when the video quality encoding parameter of video block k is CRF under the current bandwidth, and T pred is the predicted decoding processing delay when the video quality encoding parameter of video block k is CRF under the current bandwidth; T max is the required delay corresponding to the target frame rate of video block k; 1.1T max is to allow a 10% delay overrun;
[0135] MPC trigger strategy Trigger mpc : When the standard deviation of the bandwidth in the past 5 seconds < 15% of the average bandwidth, the continuous video quality CRF numerical sequence CRF final is the CRF numerical sequence CRF mpc decided by enabling MPC, that is:
[0136] Triggermpc : Bandwidth std5 <15%×Bandwidth mean ,
[0137] CRF final =CRF mpc ;
[0138] Among them, Bandwidth std5 is the standard deviation STD of the bandwidth in the past 5 seconds, and Bandwidth mean is the average bandwidth;
[0139] Step 3.2. When the network environment fluctuates violently, switch to the reinforcement learning strategy and optimize the decision online through the Q-learning algorithm. The process is as follows:
[0140] State encoding s t : Encode the network bandwidth Bandwidth t at the current moment t, the predicted value T of the decoding delay pred , the device DPI index, and the historical video encoding parameter CRF sequence into a feature vector:
[0141] s t =Concat(Bandwidth t ,T pred ,DPI t ,CRF t-1 )
[0142] Action space constraint a t : Limit the CRF adjustment step size (-2, -1, 0, +1, +2) to prevent violent quality fluctuations;
[0143] a t ∈{-2, -1, 0, +1, +2}
[0144] Restriction: The adjusted CRF needs to be within the range of [15, 45]. If it exceeds, it will be truncated to the nearest boundary value;
[0145] Reward function R: Comprehensively consider video quality, delay, and stability indicators:
[0146] R=Q(CRF t ) - 10×I overdelay - 5×|CRF t - CRF t-1 | + 2×I smooth
[0147] Among them, t is the current video playback time progress; Q(CRF tis the quality score calculated by the VMAF algorithm developed by Netflix when encoding a video with a certain CRF value; CRF t and CRF t-1 are the CRF numerical values of the video quality encoding parameters of the current video block and the previous video block at the current moment respectively; CRF t -CRF t-1 represents whether a CRF quality switch occurs when detecting the playback of a video; I overdelay is a binary indicator function when the actual delay of playing the current video block exceeds the required delay corresponding to the target frame rate, which is 1 when the delay exceeds the threshold and 0 otherwise; I smooth The stability index mechanism is triggered when there is no CRF quality switch for 3 consecutive times, encouraging long-term stable decisions, and is a binary indicator function, which is 1 when triggered and 0 otherwise.
[0148] Q-learning Trigger Strategy ql : When the standard deviation of the bandwidth in the past 5 seconds is between 15% and 20% of the average bandwidth, the Q-learning strategy is triggered, and the CRF numerical value sequence CRF final is the CRF numerical value sequence obtained by enabling Q-learning decision-making; CRF ql , that is:
[0149] Trigger ql : 15%×Bandwidth mean <Bandwidth std5 <20%×Bandwidth mean ,
[0150] CRF final =CRF ql ;
[0151] Among them, Bandwidth std5 is the standard deviation of the bandwidth STD in the past 5 seconds, and Bandwidth mean is the average bandwidth.
[0152] Step 3.3: In the rolling time domain window H, when the network conditions change violently, that is, when the standard deviation of the bandwidth in the past 5 seconds > 20% of the average bandwidth, through the cooperation of MPC and Q-learning, at this time, the outputs of MPC and Q-learning are fused according to the weights, and the CRF value of the video encoding parameter is adaptively adjusted, that is:
[0153] Trigger 协同 : Bandwidth std5 >20%×Bandwidth mean ,
[0154] CRF final = 0.7CRF mpc + 0.3CRF ql 。
[0155] Among them, Bandwidth std5 is the standard deviation STD of the bandwidth in the past 5 seconds, and Bandwidth mean is the average bandwidth. Duration constraint: When the network state changes back from the fluctuating state to the stable state, the collaborative strategy needs to be maintained for at least another 3 seconds before switching back to the MPC strategy, to avoid misjudgment caused by short-term network jitter, resulting in repeated jumps in the CRF value and causing video quality jitter, and to ensure that the decision of strategy switching is based on continuous state changes rather than instantaneous noise. In this way, the video quality is maximized on the premise of controllable delay, taking into account both real-time performance and robustness.
[0156] Step 4: Video stream smoothing processing. Maximize the bitrate R max under the current network to match the optimal picture quality of the CRF decision. In order to reduce the current video playback stuttering rate, improve the user's subjective picture quality experience, and at the same time avoid the buffer delay caused by sudden changes in the bitrate. Double-buffer progressive rendering and bitrate smoothing filtering are adopted during video playback, and the process is as follows:
[0157] Step 4.1: The double-buffer controller performs progressive rendering, and the specific process is as follows:
[0158] Step 4.1.1: Fragment preloading: The background thread preloads the next fragment into the Back Buffer in advance;
[0159] Step 4.1.2: Use time-varying mixed weights for picture transition as the picture P display finally displayed by the current video player to achieve a smooth transition from the old fragment to the new fragment:
[0160] P display = α t P front +(1 - α t )P back
[0161] Among them, P front is the picture of the old fragment currently being played; P back is the decoded picture of the background preloaded fragment; α decays exponentially with time α = e -t / τ , τ is set according to the device refresh rate, and in this embodiment, it is updated frame by frame for a device with a 60Hz refresh rate;
[0162] In this embodiment, when switching from CRF = 22 to CRF = 24: The Back Buffer preloads the slices of CRF = 24 and completes decoding; when playing to the 50% position of the current slice, start the hybrid rendering, and the alpha value fades from 1 (fully showing the old slice) to 0 (fully showing the new slice); thus avoiding sudden changes in the picture and making the switching process imperceptible to the human eye.
[0163] Step 4.1.3: Dynamic switching trigger: Start the pixel-level hybrid transition when the current slice is played to 50%.
[0164] Step 4.2: The bitrate smoothing filter performs dynamic bitrate smoothing control to adapt to the real-time bandwidth and balance quality and stability. The process is as follows:
[0165] Step 4.2.1: Adaptive adjustment of the bitrate dynamic gradient limit:
[0166]
[0167] Among them, σ bandwidth is the standard deviation of the bandwidth in the past 10 seconds, and the corresponding policy logic is:
[0168] Stable network (low fluctuation): Allow a large increase in the bitrate (from 4Mbps to 5.2Mbps in this embodiment), make full use of the bandwidth to improve the picture quality, and match the high-quality requirements of the CRF decision.
[0169] Turbulent network (high fluctuation): Limit the amplitude of bitrate adjustment (only allow 4Mbps to 4.6Mbps in this embodiment) to avoid buffering caused by frequent switching.
[0170] Step 4.2.2 Bitrate adjustment to suppress sudden changes in the bitrate and balance picture quality stability and bandwidth utilization:
[0171] R next = γR current +(1 - γ)R target
[0172] Among them, R current is the bitrate at the current moment; R target is the target bitrate (determined by the CRF decision); R next is the bitrate at the next moment; the γ smoothing factor is dynamically adjusted according to the degree of network oscillation:
[0173]
[0174] In this embodiment, the corresponding mapping rule: The tanh function compresses the bandwidth standard deviation to [0, 0.4] to ensure that γ ∈ [0.7, 0.9]; for a stable network (σ bandwidth<1Mbps): γ=0.7, quickly approaching the target bitrate, the bitrate smoothing filter can increase the bitrate to the upper limit allowed by CRF, and further optimize the image quality details (such as reducing compression artifacts). Turbulent network (σ bandwidth ≥1Mbps): γ = 0.9, relying on the historical bitrate, keeping the CRF unchanged, maintaining stable image quality, and suppressing sudden fluctuations.
[0175] Step 5: After completing the video stream smoothing process, the processed video data must finally be transmitted to a display device (such as VR, mobile phone, etc.) for rendering and presentation.
[0176] In this embodiment, by predicting the decoding delay of the video stream and dynamically adjusting the video quality coding parameter CRF value, the image quality is prioritized when the network is stable, and the quality is quickly reduced to ensure smoothness when there are drastic fluctuations, thereby achieving a balance between high image quality and low lag. This is suitable for real-time scenarios such as VR real-time interaction and 4K ultra-high-definition live broadcast, and can significantly improve user experience and transmission stability in complex network environments.
[0177] The contents described in the embodiments of this specification are merely an enumeration of the implementation forms of the inventive concept and are for illustrative purposes only. The protection scope of the present invention should not be considered to be limited to the specific forms described in this embodiment, and the protection scope of the present invention also extends to equivalent technical means that can be thought of by ordinary technicians in this field based on the inventive concept.
Claims
1. A method for predicting video stream decoding delay and dynamic quality optimization based on multi-feature fusion, characterized in that: The method comprises the following steps: Step 1: Video stream feature extraction. The process is as follows: Step 1.1: Extract video coding parameters, including CRF, bit rate, picture group structure and frame rate; Step 1.2: Content complexity analysis, including parsing motion vectors and calculating the motion intensity index MVI, and analyzing discrete cosine transform DCT texture entropy; Step 1.3: Device status monitoring, including collecting GPU utilization U GPU and memory bandwidth usage B mem , build the device performance index DPI; Step 2: Predict video stream decoding delay: Design a hybrid model architecture to predict decoding delay, that is, use the LightGBM sub-model and the LSTM residual compensation sub-model; Step 3: Dynamic decision-making on video stream quality: This method combines the improved model predictive control (MPC) algorithm and the Q-learning reinforcement learning strategy. MPC is used in normal situations, taking into account delay constraints and quality scores, and allowing certain delays to exceed the limit. When the network fluctuates violently, it switches to Q-learning, adjusts the decision online through reinforcement learning, and limits the adjustment range to maintain stable quality. Finally, the sequence output of the CRF value is determined within the rolling window. Step 4: Video stream smoothing: The final video quality coding parameter CRF value is used when playing the video. This coding parameter reflects the quality of the presented video. The bit rate R is maximized under the current network. max To match the optimal image quality determined by CRF, video playback uses double-buffered progressive rendering and bitrate smoothing filtering; Step 5: After the video stream smoothing process is completed, the processed video data is transmitted to the display device for rendering and presentation.
2. The method for video stream decoding delay prediction and dynamic quality optimization based on multi-feature fusion as claimed in claim 1, characterized in that: In step 1.1, CRF is the video quality coding level of the current video block captured dynamically. The lower the value, the higher the image quality but the higher the bit rate. The bit rate is based on the real-time calculation of the average bit rate of the current video segment based on the coding parameters. The GOP structure is to extract the I frame interval, P frame, and B frame number parameters to calculate the average GOP length. The frame rate is the number of frames per second (FPS) parsed from the video stream header.
3. The method for video stream decoding delay prediction and dynamic quality optimization based on multi-feature fusion as claimed in claim 1, characterized in that: In step 1.2, MVI obtains the motion vector component (Δx, Δy) of each macroblock from the decoder interface and calculates the motion amplitude by macroblock: The full-frame MVI value is the weighted average of all macroblocks, with the weight being the proportion of the macroblock area: Among them, w i is the area weight coefficient of the i-th macroblock, N is the total number of macroblocks in the frame, Δx i is the motion vector component of the i-th macroblock in the X direction, Δy i is the motion vector component of the i-th macroblock in the Y direction; DCT texture entropy is the quantification of image texture complexity through discrete cosine transform DCT, including: 8×8 block division and DCT transformation: Divide the brightness component of the image into multiple 8×8 pixel blocks, perform DCT transformation on each block, and obtain the corresponding frequency coefficient matrix, denoted as C k , i.e. the kth DCT coefficient; Normalized energy distribution calculation: Calculate the total energy of each 8×8 block Normalize the energy proportion of each DCT coefficient: Where p(k) represents the normalized energy distribution probability of the kth DCT coefficient; Calculate the full frame texture entropy H texture :
4. The method for video stream decoding delay prediction and dynamic quality optimization based on multi-feature fusion as claimed in claim 1, characterized in that: In the step 1.3, the GPU utilization is obtained by obtaining the GPU computing unit occupancy in real time through the driver API, and the video memory bandwidth occupancy is obtained by monitoring the ratio of the video memory read and write bandwidth to the theoretical peak value; Build DPI: DPI=0.7U GPU +0.3B mem 。 5. The method for video stream decoding delay prediction and dynamic quality optimization based on multi-feature fusion according to any one of claims 1 to 4, characterized in that: The process of step 2 is as follows: Step 2.1: Input the static encoding parameters, namely CRF, bit rate, GOP, and frame rate, into the pre-trained LightGBM model and output the basic decoding time prediction value T base ; Step 2.2: Input the temporal dynamic features, i.e., the MVI sequence, DCT texture entropy, and DPI sequence, into the LSTM residual compensation model to predict the decoding time residual correction value ΔT; Step 2.3: Comprehensively calculate the final decoding delay prediction value T pred : T pred =T base +ΔT Step 2.4: Adopting the incremental learning strategy, when the single prediction error is >15% or the error is >10% for three consecutive times, the online update mechanism is triggered to optimize and update the parameters of the LightGBM model and LSTM residual compensation model with the most recent set number of samples.
6. The method for video stream decoding delay prediction and dynamic quality optimization based on multi-feature fusion according to any one of claims 1 to 4, characterized in that: The process of step 3 is as follows: Step 3.1: MPC controller design, build an optimization model with end-to-end delay constraints: Rolling window optimization: Among them, H represents the time length of the optimization window, and RTT represents the network round-trip delay in the current environment; Objective function: Among them, Q(CRF k ) is the quality score calculated by the VMAF algorithm developed by Netflix when encoding a video with a certain CRF value; CRF k and CRF k-1 are the video quality coding parameter CRF values of the current video block k and the previous video block k-1 respectively; w1 and w2 are weight coefficients; t is the current video playback time progress; Optimization with constraint relaxation: T trans (CRF k )+T pred (CRF k )≤1.1T max Among them, T trans is the transmission delay of video block k under the current bandwidth when the video quality coding parameter is CRF, T pred is the prediction decoding processing delay when the video quality coding parameter of video block k is CRF under the current bandwidth; T max is the required delay corresponding to the target frame rate of video block k; 1.1T max A 10% delay overrun is allowed; MPC trigger strategy Trigger mpc : When the bandwidth standard deviation in the past 5 seconds is < 15% of the average bandwidth, the video quality CRF value sequence CRF used continuously final is the CRF numerical sequence CRF decided by enabling MPC mpc ,Right now: Trigger mpc :Bandwidth std5 <15%×Bandwidth mean , CRF final =CRF mpc ; Bandwidth std5 Bandwidth is the standard deviation of bandwidth in the past 5 seconds. mean is the average bandwidth; Step 3.
2. When the network environment fluctuates violently, switch to the reinforcement learning strategy and optimize the decision online through the Q-learning algorithm. The process is as follows: Status Codes t :Bandwidth of the network at the current time t t , decoding delay prediction value T pred , device DPI index and historical video encoding parameter CRF sequence are encoded as feature vectors: s t =Concat(Bandwidth t ,T pred ,DPI t ,CRF t-1 ) Action space constraintsa t : Limit the CRF adjustment step size (-2, -1, 0, +1, +2) to prevent drastic quality fluctuations; a t ∈{-2,-1,0,+1,+2} Restrictions: The adjusted CRF must be in the range of [15,45]. If it exceeds the range, it will be truncated to the nearest boundary value. Reward function R: Comprehensively consider video quality, delay and stability indicators: R=Q(CRF t )-10×I overdelay -5×|CRF t -CRF t-1 |+2×I smooth Among them, t is the current video playback time progress; Q(CRF t ) is the quality score calculated by the VMAF algorithm developed by Netflix when encoding a video with a certain CRF value; CRF t and CRF t-1 are the video quality coding parameter CRF values of the video block at the current moment and the previous moment respectively; CRF t -CRF t-1 Indicates whether CRF quality switching occurs when playing a video; I overdelay When the actual delay of playing the current video block exceeds the required delay corresponding to the target frame rate, it is a binary indicator function, which is 1 when the delay exceeds the threshold, otherwise it is 0; I smooth The stability indicator mechanism is triggered when there is no CRF quality switch for three consecutive times, which encourages long-term stable decision-making. It is a binary indicator function, which is 1 when triggered and 0 otherwise; Q-learning trigger strategy ql : When the bandwidth standard deviation in the past 5 seconds is between 15% and 20% of the average bandwidth, the Q-learning strategy is triggered and the video quality CRF numerical sequence CRF is continuously used final It is the CRF numerical sequence CRF decided by enabling Q-learning ql ,Right now: Trigger ql :15%×Bandwidth mean <Bandwidth std5 <20%×Bandwidth mean , CRF final =CRF ql ; Bandwidth std5 Bandwidth is the standard deviation of bandwidth in the past 5 seconds. mean is the average bandwidth; Step 3.3: In the rolling time domain window H, through the collaboration of MPC and Q-learning, the video quality coding parameter CRF value is adaptively adjusted when the network conditions change, so as to maximize the video quality while ensuring controllable delay, taking into account both real-time and robustness.
7. The method for video stream decoding delay prediction and dynamic quality optimization based on multi-feature fusion according to any one of claims 1 to 4, characterized in that: The process of step 4 is as follows: Step 4.1: Double buffer progressive rendering to ensure smooth transition of the picture. The process is as follows: Step 4.1.1: Segment preloading: the background thread loads the next segment into the Back Buffer in advance; Step 4.1.2: Use time-varying mixing weights to perform screen transition, which is the final screen P displayed by the current video player. display , to achieve a smooth transition from the old shard to the new shard: P display =α t P front +(1-α t )P back Among them, P front The old segment currently being played; P back The decoded picture of the pre-loaded fragment in the background; α decays exponentially with time α = e -t / τ ,τ is set according to the device refresh rate; Step 4.1.3: Dynamic switching trigger: Start pixel-level blending transition when the current segment plays to 50%; Step 4.2: The bitrate smoothing filter dynamically adjusts the bitrate to adapt to the real-time bandwidth based on the image quality determined by CRF, balancing quality and stability. The process is as follows: Step 4.2.1 Dynamic gradient limitation: Among them, σ bandwidth is the bandwidth standard deviation in the past 10 seconds; When the bandwidth fluctuations in the past 10 seconds are small, the bitrate adjustment is allowed to be up to 30% at a time to quickly match the available bandwidth and improve the image quality. If the bandwidth fluctuations are drastic, the adjustment is limited to 15% to avoid buffering caused by frequent switching. Step 4.2.2 Bitrate adjustment, suppressing bitrate mutations, and balancing image quality stability and bandwidth utilization: R next =γR current +(1-γ)R target Among them, R current The current bit rate; R target Target bit rate (determined by CRF); R next The bit rate at the next moment; the γ smoothing factor is dynamically adjusted according to the degree of network oscillation: When the network is stable, the smoothing factor is reduced (γ≈0.7) to make the bit rate quickly approach the target value; when the network is turbulent, the smoothing factor is increased (γ can reach 0.9), the historical bit rate weight is strengthened, and sudden fluctuations are suppressed; through the nonlinear adjustment of the tanh function, the bandwidth standard deviation is mapped to the interval [0,0.4], and nonlinear adjustment of γ between 0.7 and 0.9 is achieved, ensuring that the bit rate adjustment is more robust when the network fluctuates.
Citation Information
Patent Citations
Method and system for assessing real-time transmission quality of video
CN105491403A
Short video code rate adaptive transmission method based on multi-agent reinforcement learning
CN116506626A
Adaptive code rate selection method based on imitation learning and reinforcement learning
CN118573863A
Data video stream adaptive processing system and method based on streaming processing technology
CN118694945A
Network optimization method for low-delay transmission in multimedia live broadcast
CN118827378A
Cited By
Self-adaptive video low-delay hard decoding display method and system
CN121151565A
Network transmission optimization method in video monitoring scene
CN121547560A