Energy-saving video adaptive bit rate optimization method based on deep reinforcement learning
Through the video adaptive bit rate optimization method of deep reinforcement learning and video multi-method fusion evaluation, combined with energy consumption perception model and Actor-Critic architecture, the problem of high energy consumption in video streaming is solved, and a high-quality and low-energy video streaming experience is achieved.
Patent Information
- Application Number
- CN202510157782.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing video streaming methods fail to effectively balance device energy consumption when optimizing video quality, especially on devices with limited resources in IoT environments, resulting in high energy consumption and poor user experience.
Using an energy-saving video adaptive bit rate optimization method based on deep reinforcement learning, combined with video multi-method fusion evaluation and Actor-Critic architecture neural network, the bit rate is dynamically adjusted to optimize the energy efficiency and quality of video streams through energy consumption perception models and improved near-end strategy optimization algorithms.
On the premise of ensuring the quality of user experience, significantly reduce equipment energy consumption, and achieve dual optimization of video streaming quality and energy consumption.
Smart Images

Figure CN120264046A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of streaming media, and particularly to an energy-saving video adaptive bitrate optimization method based on deep reinforcement learning. Background Art
[0002] With the rapid development of the Internet of Things and 5G technologies, video streaming services have gradually become mainstream in various mobile devices and embedded systems. However, limited by the battery capacity of devices, the volatility of the network environment, and the instability of bandwidth, the transmission quality of video streams faces significant challenges.
[0003] Traditional Adaptive Bitrate Streaming (ABR) algorithms usually focus on optimizing video playback quality, smooth switching of bitrates, and buffering time, but consider less about the energy consumption of devices and cannot meet the actual needs of resource-constrained devices in the Internet of Things environment. Existing ABR algorithms can be roughly divided into heuristic methods and learning-based methods. Heuristic algorithms, such as algorithms based on buffer and rate prediction, although can improve the video stream playback quality to a certain extent, often ignore the changes in complex network environments, resulting in frequent switching of video bitrates, which in turn affects the smoothness of video playback and user experience. Learning-based algorithms, especially those based on deep reinforcement learning, such as Pensieve, EAVS, etc., although can better adapt to dynamically changing network environments, these methods usually ignore the balance between the perceived quality of videos and device energy consumption, and often require a large amount of training data and computing resources, and may face problems of insufficient adaptability and unstable effects when deployed in actual environments.
[0004] Meanwhile, with the popularization of video streaming services, users' demands for viewing experience are not limited to video quality itself, and energy efficiency issues have also attracted increasing attention. For mobile devices and embedded systems, the energy consumed during video playback is closely related to users' viewing experience. Summary of the Invention
[0005] The purpose of the present invention is to provide an energy-saving video adaptive bitrate optimization method based on deep reinforcement learning, aiming to solve the problem that existing ABR optimization methods focus on video quality but do not fully consider energy consumption.
[0006] To achieve the above object, the present invention provides an energy-saving video adaptive bitrate optimization method based on deep reinforcement learning, including the following steps:
[0007] Step 1: Prepare video data and network data, where the video data includes the video chunk file size and VMAF value of each bitrate version, and the network data includes the network bandwidth distribution situation within a period of time;
[0008] Step 2: The client simulator based on the DASH standard receives video data and network data as input data, simulates the streaming video playback process, and outputs network statistical information, video chunk statistical information, and system buffer information;
[0009] Step 3: The energy consumption awareness model calculates the energy consumption based on the network statistical information and the video chunk statistical information; the ABR agent receives the energy consumption, network statistical information, video chunk statistical information, and system buffer information as state inputs, and passes them to the policy network and the value network. The policy network outputs the bitrate probability distribution, and the value network outputs the value of the current state;
[0010] Step 4: The ABR agent selects the bitrate with the highest probability according to the probability distribution output by the policy network, interacts with the client simulator to obtain statistical information, calculates the reward based on the statistical information and the reward function formula, and the ABR agent adjusts the policy according to the calculated reward value to maximize the reward;
[0011] Step 5: Train the ABR agent until the reward value converges, improve the bitrate decision-making quality, and optimize the QoE and energy efficiency.
[0012] Optionally, the VMAF value in Step 1 is calculated using libvmaf in the open-source tool library FFmpeg. The video version with the highest bitrate is used as the reference video, and then the VMAF values of other bitrate versions of the video chunks are calculated respectively.
[0013] Optionally, the energy consumption awareness model perceives the energy consumption in two parts, namely the data acquisition energy consumption and the video display energy consumption;
[0014] Data acquisition energy consumption:
[0015]
[0016] where th is the total actual throughput, a and b are constant parameters, and f s is the size of the video chunk;
[0017] Video display energy consumption:
[0018] E l = w·b v + c
[0019] where b v is the video bitrate, and w and c are constant parameters.
[0020] Optionally, the ABR agent extracts the current video playback state information according to the time step t to form the state S t and the state S tIncludes measurements of network dynamics and video players, expressed as:
[0021]
[0022] Among them, and respectively represent the throughput measurement value in past observations and its corresponding download time; and respectively represent the size and VMAF value of the next video chunk at m bitrate levels; b t represents the current buffer size, q t is the VMAF value of the previous video chunk, L t represents the number of remaining video chunks, and g includes two QoE weight parameters for balancing the trade-off between buffering and playback smoothness.
[0023] Optionally, in step 3, the neural network of the Actor-Critic architecture is used to evaluate the network condition and output a selection. Among them, the Actor network corresponds to the policy network and is responsible for generating actions, that is, selecting the bitrate in the current state. The Critic network corresponds to the value network and is responsible for evaluating the value of the current state.
[0024] Optionally, the training process of the ABR agent adopts an improved proximal policy optimization algorithm. The loss function includes the policy update loss and the value function loss, and the Dual-Clip method and adaptive entropy weight are introduced;
[0025] The total loss function of the improved proximal policy optimization algorithm is expressed as:
[0026]
[0027] Optionally, the improved proximal policy optimization algorithm also introduces a dynamically adjusted reward function, and the reward function is defined as:
[0028]
[0029] Among them, i represents the index of the video chunk, and VMAF i represents the perceived quality of the i-th video chunk; the term represents the buffering penalty, where rt i is the buffering time before playing the i-th video chunk; the term represents the smoothness penalty, reflecting the penalty for the sharp change in quality between consecutive video chunks; the term represents the energy consumption penalty; the penalty coefficients w1, w2, and w3 are set to 100, 5, and 20 according to the characteristics of VMAF and the insights of the QoE trend.
[0030] The present invention provides an energy-saving video adaptive bitrate optimization method based on deep reinforcement learning. In the optimization process, a video quality assessment method that combines multiple video methods for evaluation is introduced to effectively evaluate the perceived quality of the video at different bitrates, ensuring that users obtain the most realistic picture experience during viewing. At the same time, an energy consumption perception model is proposed by combining a neural network with the Actor-Critic architecture, and the energy consumption of the device during video playback is estimated based on the energy consumption of video data download and video rendering. In addition, by improving the entropy update strategy of the proximal policy optimization algorithm, multi-dimensional information such as network conditions and user preferences can be more effectively combined to optimize the bitrate decision of the video stream and balance video quality and energy consumption. The present invention significantly reduces the device energy consumption on the premise of ensuring the quality of user experience, thus realizing the dual optimization of the quality and energy consumption of video stream transmission. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0032] Figure 1 It is a schematic flowchart of the steps of an energy-saving video adaptive bitrate optimization method based on deep reinforcement learning of the present invention.
[0033] Figure 2 It is a schematic diagram of the principle of video multi-method fusion evaluation in the present invention.
[0034] Figure 3 It is a schematic diagram of the method framework of the present invention.
[0035] Figure 4 It is a schematic diagram of the neural network structure of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] The following will describe in detail the embodiments of the present invention. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present invention and should not be construed as a limitation of the present invention.
[0037] Please refer to Figure 1 , the present invention provides an energy-saving video adaptive bitrate optimization method based on deep reinforcement learning, including the following steps:
[0038] S1: Prepare video data and network data. The video data includes the video chunk file size and VMAF value for each bitrate version, and the network data includes the network bandwidth distribution over a period of time.
[0039] S2: The client simulator based on the DASH standard receives the video data and network data as input data, simulates the streaming media playback process, and outputs network statistics, video chunk statistics, and system buffer information.
[0040] S3: The energy consumption awareness model calculates the energy consumption based on the network statistics and video chunk statistics. The ABR agent receives the energy consumption, network statistics, video chunk statistics, and system buffer information as state inputs, and passes them to the policy network and value network. The policy network outputs the bitrate probability distribution, and the value network outputs the value of the current state.
[0041] S4: The ABR agent selects the bitrate with the highest probability according to the probability distribution output by the policy network, interacts with the client simulator to obtain statistics, calculates the reward based on the statistics and the reward function formula, and the ABR agent adjusts the policy according to the calculated reward value to maximize the reward.
[0042] S5: Train the ABR agent until the reward value converges, improve the bitrate decision quality, and optimize the QoE and energy efficiency.
[0043] The following is a further description in combination with specific implementation steps and related terms:
[0044] In the present invention, Video Multimethod Assessment Fusion (VMAF) is introduced as a video quality assessment metric. By introducing VMAF, it is possible to more accurately reflect the impact of video bitrate adjustment on the user viewing experience, avoiding relying solely on a single metric of bitrate in traditional methods, thereby improving the quality prediction accuracy of video streaming in different network environments. In implementation, libvmaf in the open-source tool library FFmpeg is used to calculate VMAF, taking the video version with the highest bitrate as the reference video, and then calculating the VMAF values of other bitrate versions of the video chunks respectively. Please refer to Figure 2 , and its working principle includes several key steps:
[0045] 1) First, the input video data is split into frames for frame-by-frame quality analysis. Then, VMAF calculates the basic metrics of the video, including three core features: visual information fidelity, detail loss metric, and motion feature. VIF is used to evaluate the loss of information fidelity of video frames at different scales, reflecting the degree of detail retention in the image; DLM focuses on the loss of image details and redundant distortion, evaluating the fineness of the image; while the motion feature calculates the pixel differences between adjacent frames to measure motion distortion. Through these three basic metrics, VMAF can evaluate video quality from different perspectives.
[0046] 2) After calculating these metrics, VMAF uses the training set (NFLX - TRAIN) to train the support vector machine (SVM) regression model and evaluates the performance of the model through the test set (NFLX - TEST). The SVM model assigns weights to each basic metric to further optimize its ability to predict video quality.
[0047] 3) Finally, VMAF generates a comprehensive score by weighted fusion of the results of each metric. This score, as a prediction of the subjective quality of the video, can accurately reflect the actual experience of the audience when watching the video. Through this multi - method fusion approach, VMAF is more reliable and accurate than a single quality assessment method when evaluating video quality.
[0048] The method of the present invention in step 2 can select the optimal bitrate according to the network condition information, buffer information, and video chunk information provided by the client emulator to balance the quality of user experience and energy consumption. As Figure 3 shown is the framework diagram of the method of the present invention.
[0049] The energy consumption awareness model perceives energy consumption in two parts, namely data acquisition energy consumption and video display energy consumption;
[0050] Data acquisition energy consumption:
[0051]
[0052] where th is the total actual throughput, a and b are constant parameters, and f s is the size of the video chunk; for TCP transmission, in the present invention, a = 210 and b = 28 are set.
[0053] Video display energy consumption:
[0054] E l = w·b v + c
[0055] where b vis the video bitrate, and w and c are constant parameters. According to the suggestions of relevant literature, w = 24.71 and c = 1121.5 are set.
[0056] The energy consumption awareness model calculates the energy consumption using a formula based on the information of the network and the video chunks. The calculated energy consumption, along with the video chunk information and network statistical information, is passed as input to two neural networks: the policy network and the value network (Actor and Critic).
[0057] During the execution of step 3, as Figure 4 shown, the neural network is an Actor-Critic architecture for reinforcement learning, mainly including a policy network and a value network. The two neural networks have the same structure but different outputs. The network uses four one-dimensional convolutional layers (1DConv) to capture temporal dependencies and extract local features from sequential data. Each 1DConv layer is configured with 128 feature channels and a convolutional kernel size of 1, enabling it to independently process four main input features: and In this embodiment, k = 8 is set, which means the network considers the information of the past eight time steps. After the convolutional layers, the network uses four fully connected (FC) layers, each with 128 neurons, to further abstract and integrate the features extracted by the convolutional layers. All convolutional layers and fully connected layers use the ReLU activation function. The network processes the four main input features separately through the convolutional layers and The information that has not been processed by convolution is directly input into a separate fully connected layer, which contains 128 neurons. All feature representations are then concatenated into a comprehensive feature vector, and then passed through a final fully connected layer with 128 neurons to integrate the features from different sources and form a unified representation suitable for subsequent tasks. The Softmax activation function is used in the output layer of the Actor network to generate a probability distribution over the available actions. To ensure numerical stability and prevent extreme probabilities from interfering with the learning process, a Clamp operation is applied after the Softmax function, which limits the probability of each action to the range of [10 -4 , 1 - 10 -4 . The output layer of the Critic network consists of a linear neuron that outputs a scalar value representing the estimated value of the current state.
[0058] In addition, the ABR agent extracts the current video playback state information according to the time step t to form the state S t , and the state S t contains the network dynamics and measurements of the video player, expressed as:
[0059]
[0060] Among them, and respectively represent the throughput measurement values and their corresponding download times in the past k observations; and respectively represent the size and VMAF value of the next video block at m bitrate levels; b t represents the current buffer size, q t is the VMAF value of the previous video block, L t represents the number of remaining video blocks, q includes two QoE weight parameters β and γ, where β, γ ∈ [0, 1], which are used to balance the trade - off between buffering and playback smoothness.
[0061] The action space is defined as the set of available bitrate levels for a given video. The agent selects an action a from the action space A according to the policy π t .
[0062] In step S4, according to the action (bitrate selection) output by the neural network, the video client starts playing the video and monitors the playback state in real - time, including whether buffering occurs, video quality changes, playback smoothness, etc. The emulator calculates the reward signal according to the predefined reward function. The reward signal includes the perceived quality (VMAF) of the video block, the rebuffering time of the video block, whether quality switching occurs between two played blocks, and the energy consumption value. The specific reward function is as follows:
[0063]
[0064] Among them, i represents the index of the video block, VMAF i represents the perceived quality of the i - th video block. The term represents the buffering penalty, where rt i is the buffering time before playing the i - th video block. The term represents the smoothness penalty, which reflects the penalty for a sharp change in quality between consecutive video blocks. The term represents the energy consumption penalty. The penalty coefficients w1, w2, and w3 are set to 100, 5, and 20 according to the characteristics of VMAF and the insights of the QoE trend. The parameters β and γ are dynamic scaling factors, which are uniformly sampled from the interval [0, 1] during initialization and reset regularly.
[0065] In step S5, the calculated reward signal is used as feedback and passed to the reinforcement learning model. The reinforcement learning decision model is trained using an improved proximal policy optimization (PPO) algorithm. The agent continuously adjusts its decision according to the reward signal and continuously optimizes the parameters of the neural network. The ultimate goal is to improve the quality of video playback and optimize energy efficiency simultaneously, providing users with a high - quality and low - energy - consumption video streaming experience.
[0066] Specifically, the improved Proximal Policy Optimization (PPO) algorithm is used to train the ABR agent. The PPO algorithm is one of the most classic deep reinforcement learning algorithms. In this invention, Dual Clip and adaptive entropy weight are introduced to ensure the stability of the training process and can efficiently handle video stream transmission problems under different network conditions. The loss function of the PPO algorithm includes the policy update loss and the value function loss.
[0067] The policy update loss is defined as:
[0068]
[0069] where, π θ represents the current policy, is the old policy, ∈ is the clipping threshold, and in this embodiment, ∈ is set to 0.2. is the advantage function, which is used to measure the relative return of the state-action pair and is defined as:
[0070]
[0071] where, r t represents the immediate reward received at time step t, is the value of the state s t output by the critic network.
[0072] The fixed entropy weight κ cannot adapt to the need for policy exploration during the training process. Therefore, κ is dynamically adjusted, and the formula is as follows:
[0073]
[0074] where, α p is the learning rate of the actor network. We autonomously adjust the entropy weight κ to minimize the gap between the current entropy and the target entropy We set The entropy of the current policy is defined as:
[0075] The policy update loss is updated by minimizing the error of the advantage function and is defined as:
[0076]
[0077] To further improve the training stability, we adopt the Dual Clip strategy to constrain the minimum value of the loss function when the advantage function is negative:
[0078]
[0079] Among them, h is a constant that controls the degree of negative dominance clipping. In this embodiment, h is set to 3.
[0080] The total loss function combines the above-mentioned parts:
[0081]
[0082] In summary, the advantages of the present invention are as follows:
[0083] By fusing multiple methods to evaluate the relationship between video perception quality and energy consumption, and combining factors such as the user's viewing preferences and network bandwidth, the bit rate of the video is dynamically adjusted, thereby achieving a dual optimization of the quality and energy consumption of video stream transmission.
[0084] The above-disclosed is only a preferred embodiment of the present invention. Of course, it cannot be used to limit the scope of the rights of the present invention. Those of ordinary skill in the art can understand the entire or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.
Claims
1. An energy-saving video adaptive bitrate optimization method based on deep reinforcement learning, characterized in that It includes the following steps: Step 1: Prepare video data and network data. The video data includes the video chunk file size and VMAF value of each bitrate version, and the network data includes the network bandwidth distribution over a period of time; Step 2: The client simulator based on the DASH standard receives the video data and network data as input data, simulates the streaming playback process, and outputs network statistics, video chunk statistics, and system buffer information; Step 3: The energy consumption awareness model calculates the energy consumption based on the network statistics and video chunk statistics; The ABR agent receives the energy consumption, network statistics, video chunk statistics, and system buffer information as state inputs, and passes them to the policy network and value network. The policy network outputs the bitrate probability distribution, and the value network outputs the value of the current state; Step 4: The ABR agent selects the bitrate with the highest probability according to the probability distribution output by the policy network, interacts with the client simulator to obtain statistics, calculates the reward based on the statistics and the reward function formula, and the ABR agent adjusts the policy according to the calculated reward value to maximize the reward; Step 5: Train the ABR agent until the reward value converges, improve the bitrate decision quality, and optimize the QoE and energy efficiency.
2. The energy-saving video adaptive bitrate optimization method based on deep reinforcement learning according to claim 1, characterized in that The VMAF value in Step 1 is calculated using libvmaf in the open-source tool library FFmpeg, taking the video version with the highest bitrate as the reference video, and then calculating the VMAF values of other bitrate versions of the video chunks respectively.
3. The energy-saving video adaptive bitrate optimization method based on deep reinforcement learning according to claim 2, characterized in that The energy consumption awareness model perceives the energy consumption in two parts, namely data acquisition energy consumption and video display energy consumption; Data acquisition energy consumption: Among them, th is the total actual throughput, a and b are constant parameters, and f s is the size of the video block; Video display energy consumption: E l = w·b v + c where b v is the video bit rate, and w and c are constant parameters.
4. The energy-saving video adaptive bitrate optimization method based on deep reinforcement learning according to claim 3, characterized in that The ABR agent extracts the current video playback status information according to the time step t to form the state S t , where the state S t includes network dynamics and measurements of the video player, expressed as: wherein, and respectively represent the throughput measurement value in past observations and its corresponding download time; and respectively represent the size and VMAF value of the next video chunk at m bitrate levels; b t represents the current buffer size, q t is the VMAF value of the previous video chunk, L t represents the number of remaining video chunks, and g includes two QoE weight parameters for balancing the trade-off between buffering and playback smoothness.
5. The energy-saving video adaptive bitrate optimization method based on deep reinforcement learning according to claim 4, characterized in that In Step 3, the neural network of the Actor-Critic architecture is used to evaluate the network condition and output a selection. Among them, the Actor network corresponds to the policy network and is responsible for generating actions, that is, selecting the bitrate in the current state. The Critic network corresponds to the value network and is responsible for evaluating the value of the current state.
6. The energy-saving video adaptive bitrate optimization method based on deep reinforcement learning according to claim 5, characterized in that The training process of the ABR agent adopts an improved proximal policy optimization algorithm. The loss function includes the policy update loss and the value function loss, and the Dual-Clip method and adaptive entropy weight are introduced; The total loss function of the improved proximal policy optimization algorithm is expressed as:
7. The energy-saving video adaptive bitrate optimization method based on deep reinforcement learning according to claim 6, characterized in that The improved proximal policy optimization algorithm also introduces a dynamically adjusted reward function, and the reward function is defined as: where i represents the index of the video chunk, and VMAF i represents the perceived quality of the i-th video chunk; the term represents the buffering penalty, where rt i is the buffering time before playing the i-th video chunk; the term represents the smoothness penalty, which reflects the penalty for a sharp change in quality between consecutive video chunks; the term represents the energy consumption penalty; the penalty coefficients w1, w2, and w3 are set to 100, 5, and 20 according to the characteristics of VMAF and insights into the QoE trend.
Citation Information
Cited By
Self-adaptive bit rate control method and system
CN120455745A
Adaptive bit rate control method and system
CN120455745B