Heterogeneous network-oriented video stream adaptive transmission method, system and device and medium

By introducing pre-trained deep reinforcement learning models and dynamic convolutional layers into the ABR algorithm, combined with multiple Critic networks, the problem of unstable performance of existing ABR algorithms in different environments and user groups is solved, and more efficient and robust video streaming is achieved.

CN120034691APending Publication Date: 2025-05-23SHENZHEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510205542.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Existing ABR algorithms cannot provide relatively stable performance in different environments and target different user groups.

Method used

The pre-trained deep reinforcement learning model is adopted to build input data by obtaining the status information of the played video blocks, and the generalization ability and robustness of the model are improved by using dynamic convolutional layers and multiple Critic networks.

Benefits of technology

It alleviates the instability of ABR algorithm training, improves the robustness of strategy evaluation and the generalization ability of the model, and can provide more stable and efficient video streaming in heterogeneous network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034691A_ABST
    Figure CN120034691A_ABST
Patent Text Reader

Abstract

The invention discloses a heterogeneous network-oriented video stream adaptive transmission method, system and device and a medium, and the method comprises the steps: obtaining the state information of a played video block, and constructing input data; inputting the input data into a pre-trained deep reinforcement learning model, wherein the deep reinforcement learning model is used for predicting the bit rate of the next unplayed video block according to the input data; receiving the bit rate of the next unplayed video block generated by the deep reinforcement learning model; and adjusting the bit rate when the next video block is played according to the bit rate of the next video block which is not played. According to the technical scheme provided by the invention, the technical problem that an algorithm in the prior art cannot provide relative stable performance for different user groups in different environments can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video processing technology, and in particular to a method, system, device and medium for adaptive transmission of heterogeneous network video streams. Background Art

[0002] Adaptive Bitrate algorithm (ABR) is an algorithm used in video streaming services, which can dynamically adjust the bit rate of the video according to network conditions. The existing ABR algorithms are generally divided into two types, one is a model-based algorithm and the other is a learning-based algorithm.

[0003] Among them, model-based algorithms are essentially completing a model, or setting thresholds on several indicators and hyperparameters to achieve better results. This will cause certain disadvantages, that is, these solutions are overly dependent on accurate network prediction, precise network modeling and careful parameter adjustment, which may lead to good performance in one network situation, but poor performance in other networks.

[0004] Learning-based algorithms, even with rich historical datasets, have difficulty training a generalized model to cope with different network types. Models trained with multiple network trajectories do not improve adaptability and may even perform worse than models trained with a single network dataset.

[0005] Therefore, neither model-based algorithms nor learning-based algorithms can provide relatively stable performance in different environments and for different user groups. Summary of the invention

[0006] The present invention provides a method, system, device and medium for adaptive transmission of heterogeneous network video streams, aiming to effectively solve the technical problem that algorithms in the prior art cannot provide relatively stable performance in different environments and for different user groups.

[0007] According to a first aspect of the present invention, the present invention provides a method for adaptive transmission of heterogeneous network video streams, including: obtaining status information of played video blocks and constructing input data; inputting the input data into a pre-trained deep reinforcement learning model, wherein the deep reinforcement learning model is used to predict the bit rate of the next unplayed video block based on the input data; receiving the bit rate of the next unplayed video block generated by the deep reinforcement learning model; and adjusting the bit rate of the next video block when it is played based on the bit rate of the next unplayed video block.

[0008] Furthermore, the status information includes: network throughput measurement of k video chunks played, download time of k video chunks played, vector of m available sizes of the next unplayed video chunk, current buffer level, number of remaining video chunks in the video, bit rate of the last video chunk in the downloaded video.

[0009] Furthermore, the training method of the deep reinforcement learning model includes: constructing an A3C algorithm; and modifying the convolution layer of the A3C algorithm to a dynamic convolution layer to improve the generalization ability of the A3C algorithm.

[0010] Furthermore, the dynamic convolution layer includes: attention mechanism, adaptive convolution, batch normalization and activation function.

[0011] Furthermore, the step of generating the dynamic convolution layer includes: obtaining key status information of the video block, the key status information including the bandwidth and cache occupancy of the video block; calculating dynamic weights using the attention mechanism according to the key status information to generate multiple one-dimensional convolution kernel weights; using the adaptive convolution to combine different one-dimensional convolution kernel weights with input features in a weighted manner; performing the batch normalization and the activation function on the combined one-dimensional convolution kernel weights and input features to generate a dynamic convolution layer.

[0012] Furthermore, the A3C algorithm includes a value network and a policy network; the training method of the deep reinforcement learning model also includes: adding a value network to the A3C algorithm so that the value network in the A3C algorithm has N, where N is greater than or equal to 2, and each value network corresponds to a data or environment.

[0013] According to the second aspect of the present invention, the present invention also provides an adaptive transmission system for heterogeneous network video streams, including: a state information acquisition module, used to obtain state information of played video blocks and construct input data; a deep reinforcement learning model module, used to input the input data into a pre-trained deep reinforcement learning model, and the deep reinforcement learning model is used to predict the bit rate of the next unplayed video block based on the input data; a bit rate receiving module, used to receive the bit rate of the next unplayed video block generated by the deep reinforcement learning model; a bit rate adjustment module, used to adjust the bit rate of the next video block when it is played according to the bit rate of the next unplayed video block.

[0014] According to the third aspect of the present invention, the present invention also provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, it implements any one of the above-mentioned methods for adaptive transmission of heterogeneous network video streams.

[0015] According to a fourth aspect of the present invention, the present invention further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, any one of the above-mentioned methods for adaptive transmission of heterogeneous network video streams is implemented.

[0016] According to another aspect of the present invention, the present invention further provides a computer program product for executing any one of the above-mentioned methods for adaptive transmission of heterogeneous network video streams.

[0017] Through one or more of the above embodiments of the present invention, at least the following technical effects can be achieved:

[0018] In the technical solution disclosed in the present invention, the instability of ABR algorithm training can be alleviated and the robustness of strategy evaluation can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The technical solutions and other beneficial effects of the present invention will be made apparent by describing in detail the specific embodiments of the present invention in conjunction with the accompanying drawings.

[0020] Figure 1 (a) is a schematic diagram of network heterogeneity in the diversity of network environments in the prior art;

[0021] Figure 1 (b) is a schematic diagram of user heterogeneity in the diversity of network environments in the prior art

[0022] Figure 2 (a) is the test diagram of Pensieve pre-trained on 3G dataset in the prior art and tested in different network environments;

[0023] Figure 2 (b) is a graph of Pensieve trained on a Hybrid dataset (3G, WIFI, 4G) in the prior art and tested in different network environments;

[0024] Figure 3 A flowchart of a method for adaptively transmitting heterogeneous network video streams provided by an embodiment of the present invention;

[0025] Figure 4 A comparison diagram before and after adding N value networks to the A3C network in the adaptive transmission method for heterogeneous network video streams provided by an embodiment of the present invention;

[0026] Figure 5 A framework diagram of a heterogeneous network video stream adaptive transmission system provided by an embodiment of the present invention;

[0027] Figure 6 A schematic block diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0028] The technical scheme in the embodiment of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiment of the present invention. Obviously, the described embodiment is only a part of the embodiment of the present invention, not all of the embodiments. Based on the embodiment of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0029] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the term "and / or" herein is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " herein, unless otherwise specified, generally indicates that the associated objects before and after are in an "or" relationship.

[0030] Existing ABR algorithms are generally divided into two categories: one is a model-based algorithm and the other is a learning-based algorithm.

[0031] Model-based methods,Model-based methods usually use key features or domain knowledge to build mathematical models,to describe network conditions, such as throughput prediction and,buffer occupancy control, to make appropriate bitrate decisions,for ABR tasks.

[0032] In this intuitive scheme, the client makes presentation decisions based on the measured available network bandwidth. Such approaches require precise bandwidth estimates and suffer from long-term bandwidth fluctuation problems. The throughput can vary greatly over time, resulting in poor ABR performance and, therefore, frequent buffer underruns.

[0033] Buffer-based adaptive video transmission algorithms select the appropriate bitrate by monitoring the client's buffer occupancy, thereby smoothing video playback and avoiding frequent freezes and quality fluctuations. However, such algorithms are prone to QoE (Quality of Experience) degradation and frequent bitrate switching under long-term bandwidth fluctuations or dynamic cross-traffic, resulting in unstable video quality. In addition, existing buffer-based solutions are not adaptable enough in the face of complex heterogeneous network environments, and lack effective responses to network fluctuations, bitrate changes, and stalls, which limits their performance and stability under diverse network conditions.

[0034] Adaptive bitrate algorithms based on combined indicators optimize bitrate selection by comprehensively considering multiple factors such as available bandwidth and buffer occupancy, aiming to improve the stability and quality of video transmission. However, these algorithms have some shortcomings: first, they are usually very sensitive to parameter settings, especially prone to unstable performance under different network conditions; second, although they optimize the quality of a single client, they fail to fully consider fairness among clients, especially in scenarios where multiple clients share bandwidth; finally, despite being able to improve video quality, these algorithms usually do not take into account the user's QoE, which may lead to ignoring the actual needs and satisfaction of users.

[0035] In short, model-based algorithms are essentially completing a model, or setting thresholds on several indicators and hyperparameters to achieve better results. Of course, this also causes certain disadvantages. These solutions rely too much on accurate network prediction, precise network modeling and careful parameter adjustment, which may lead to good performance in one network situation, but poor performance in other networks.

[0036] Based on learning methods, this type of technology defines the rate adaptation decision problem as a Markov decision process, and uses deep reinforcement learning to make real-time decisions on the rate level of video clips. The historical bandwidth and playback status information are used as the input of the deep neural network, and the deep reinforcement learning algorithm based on asynchronous advantage actor-critic is adopted. Through continuous interactive learning with the offline environment, strategies that can improve the quality of user experience can be learned. This type of algorithm contributes to the improvement of user experience quality through a long offline training time. Since then, a lot of work has begun to focus on the shortcomings of learning-based rate adaptation strategies and proposed different coping methods. This includes the introduction of more advanced artificial intelligence technologies, such as imitation learning, self-game reinforcement learning, and meta-reinforcement learning.

[0037] These methods are trained or optimized in an offline setting, i.e., using a fixed network distribution. However, they perform poorly if the online network distribution is different from the training set. And data-driven neural network training, although it performs well, has high execution overhead, large model parameters, and is only suitable for specific scenarios. A model trained for a certain client cannot be generalized to other clients, even if they run in similar environments. Therefore, even with a rich historical dataset, it is difficult to train a generalized model to cope with different network types. Models trained with multiple network trajectories do not improve adaptability, and their performance is even worse than models trained with a single network dataset.

[0038] Regardless of whether it is a model-based algorithm or a learning-based algorithm, in the adaptive video streaming scenario, the system is dynamic and uncertain, and the future state cannot be predicted. Based on experience and testing, it is found that today's Internet network conditions are not only diverse, but also unique. The diversity of the network environment is mainly demonstrated from two perspectives. The first is network heterogeneity, such as Figure 1 (a) On the one hand, through testing on foot, driving, or high-speed rail, we found that different vehicle speeds lead to very different bandwidth distributions; at the same time, according to network type, the bandwidth gap between 3G, 4G, and 5G is even greater. The second perspective is user heterogeneity, such as Figure 1 (b) In the same scenario, some users have very constant bandwidth and can watch videos smoothly, while most users have very unstable bandwidth. This is an inherent problem in the implementation world.

[0039] In addition, current ABR algorithms are usually designed for specific data sets to make ABR decisions, which makes it difficult to always perform well in today's complex network environment, which is reflected in two aspects.

[0040] First, existing DRL (deep reinforcement learning) algorithms are usually task-specific and trained to handle specific network environments independently, which makes it difficult to handle unseen scenarios. Figure 2 (a) The agent is applied to bitrate selection under different network conditions (WiFi, 4G) after training the DRL agent on a 3G network using the Pensieve algorithm. The results show that the agent performs well in the same working environment as the training network, but performs poorly on WiFi and 4G networks, with QoE close to or lower than simple model-based algorithms such as BBA (Buffer-Based Algorithm) and Robust MPC (Robust Model Predictive Control).

[0041] Second, existing DRL (Deep Reinforcement Learning) models trained for a client cannot generalize to other clients, even if they operate in similar environments. Therefore, it is difficult to train a generalized model to cope with different network types, even with rich historical datasets. For example, a DRL model is trained using an augmented hybrid dataset combining 3G, WiFi, and 4G network traces and applied to Figure 2Different network environments in (b). Models trained with multiple network trajectories do not improve adaptability and even perform worse than models trained with a single network dataset. Poor adaptability to mixed datasets may be due to dataset shift: the joint distribution of inputs and outputs differs between training and testing phases. In the example, a DRL model trained to fit a widely distributed (3G+WiFi+4G) data will result in degraded performance when tested only on a relatively narrow distribution.

[0042] Due to the diversity of real-world network environments and the shortcomings of existing methods, there is an urgent need for a dynamic adaptive ABR algorithm that can provide relatively stable performance in different environments and for different user groups. It faces technical challenges such as diverse requirements, dynamic networks, and difficulty in generalizing models.

[0043] Therefore, the present invention provides a method, system, device and medium for adaptive transmission of heterogeneous network video streams, which can combine the knowledge of deep learning dynamic networks and focus on improving the generalization ability. At the same time, data collection and training time, knowledge transfer and generalization ability are also difficulties faced by this research. By solving these problems, this application can bring new breakthroughs to the field of video stream transmission.

[0044] Figure 3 The method for adaptively transmitting heterogeneous network video streams provided by an embodiment of the present invention includes:

[0045] S101, obtaining status information of played video blocks and constructing input data;

[0046] S102, inputting the input data into a pre-trained deep reinforcement learning model, where the deep reinforcement learning model is used to predict the bit rate of the next unplayed video block based on the input data;

[0047] S103, receiving the bit rate of the next unplayed video block generated by the deep reinforcement learning model;

[0048] S104: Adjust the bit rate of the next video block to be played according to the bit rate of the next unplayed video block.

[0049] In this embodiment, after downloading each chunk, the model enters the state as input to the neural network. is the network throughput measurement for the past k video chunks; is the download time of the past k video chunks, which represents the time interval for throughput measurement; is a vector of m available sizes for the next video chunk; bt is the current buffer level; ct is the number of chunks remaining in the video; lt is the bitrate for downloading the last video chunk.

[0050] Strategy: After receiving the input St, the model needs to take an action corresponding to the bit rate of the next video block. The model follows the strategy π θ (s t ,a t )Select an action.

[0051] In some embodiments, the training method of the deep reinforcement learning model includes: constructing an A3C algorithm; and modifying the convolution layer of the A3C algorithm to a dynamic convolution layer to improve the generalization ability of the A3C algorithm.

[0052] The training algorithm used in this embodiment uses the A3C algorithm, which is an advanced asynchronous advantage actor-critic deep reinforcement learning algorithm. The overall process is as follows: when the agent performs an abr (Adaptive Bitrate) task, it first interacts with the environment to generate a new state, and the environment gives a reward. This cycle continues, and the agent and the environment continue to interact to generate more new data. The reinforcement learning algorithm interacts with the environment through a series of action strategies to generate new data, and then uses the new data to modify its own action strategy. After several iterations, the agent will learn the action strategy required to complete the task.

[0053] In addition, the embodiment of the present application also introduces dynamic convolution into the actor network. This is because in the study of adaptive transmission algorithms for video stream bit rates for heterogeneous networks and users, the existing methods have insufficient generalization capabilities under different types of network environments, resulting in unstable transmission efficiency and user experience. In order to solve this problem, the embodiment of the present application introduces a dynamic convolution mechanism and changes the traditional two-dimensional dynamic convolution to a one-dimensional dynamic convolution to enhance the model's adaptability to different network environments and user scenarios.

[0054] Traditional 2D dynamic convolution is widely used in image processing tasks, and can adaptively adjust the convolution kernel weights according to the input features to enhance the expressiveness of the model. However, the key to video bitrate adaptation lies in the modeling of time series data such as bandwidth and cache occupancy, while 2D dynamic convolution is weak in time modeling. In order to better adapt to the characteristics of time series data, this embodiment proposes an optimization solution based on 1D dynamic convolution.

[0055] The dynamic convolution layer includes: attention mechanism, adaptive convolution, batch normalization and activation function.

[0056] This solution improves the generalization ability of the model in the following ways:

[0057] The core idea of ​​1D dynamic convolution is to dynamically adjust the weight of the convolution kernel in the time dimension, so that the model can adaptively adjust the bitrate selection strategy according to different network environments and user conditions. As shown in the figure, its specific structure is as follows:

[0058] Attention Module: Dynamic weights are calculated through the fully connected layer (FC), ReLU activation function (Rectified Linear Unit) and Softmax layer to generate multiple different 1D convolution kernel weights.

[0059] Adaptive Convolution: Combines different 1D convolution kernels with input features in a weighted manner to enhance the dynamic characteristics of the model.

[0060] Batch Normalization (BN) and Activation Function: Ensure the stability of training, prevent gradient disappearance, and improve the convergence speed of the model.

[0061] Therefore, the steps for generating a dynamic convolution layer include: obtaining key status information of the video block, which includes the bandwidth and cache occupancy of the video block; using the attention mechanism to calculate dynamic weights based on the key status information to generate multiple one-dimensional convolution kernel weights; using adaptive convolution to combine different one-dimensional convolution kernel weights with input features in a weighted manner; performing batch normalization and activation functions on the combined one-dimensional convolution kernel weights and input features to generate a dynamic convolution layer.

[0062] By using 1D dynamic convolution, the computational complexity can be reduced: Compared with 2D dynamic convolution, 1D dynamic convolution reduces the computational requirements in the spatial dimension, making the model lighter and able to run efficiently on edge devices or low computing resource environments.

[0063] It can also enhance temporal modeling capabilities: 1D dynamic convolution can effectively capture bit rate change trends, adapt to bandwidth fluctuations, and improve the smoothness of video transmission.

[0064] It can also improve generalization capabilities: Since 1D dynamic convolution only focuses on feature changes in the time dimension, it will not be affected by specific network environments, and can maintain high transmission performance on different types of networks and user devices.

[0065] Among them, the training process of 1D dynamic convolution includes:

[0066] Collect key status information such as bandwidth and cache occupancy rate to construct input data; dynamic convolution calculation and feature extraction and decision making.

[0067] Among them, the dynamic convolution calculation calculates the dynamic weights through the attention mechanism, generates multiple 1D convolution kernels, and then uses the adaptive weighting strategy to select the optimal convolution kernel for convolution calculation.

[0068] Feature extraction and decision-making are input into the policy network (PolicyNetwork) after batch normalization and activation function processing to determine the bit rate selection strategy at the next moment.

[0069] The key to improving generalization ability is to adapt to different input data features through multiple dynamic convolution kernels. Specifically, this embodiment can achieve the following points through multiple convolution kernels in 1D dynamic convolution:

[0070] 1) Adaptability of multiple convolution kernels

[0071] Each convolution kernel can be adaptively adjusted according to the different characteristics of the input data. Through the attention mechanism, the weight of each convolution kernel changes dynamically with each input, allowing the model to flexibly respond to different network environments and user conditions. Different convolution kernels can focus on different types of timing features, so that when faced with heterogeneous network conditions such as bandwidth fluctuations and cache changes, strategies can be adjusted to optimize video transmission quality.

[0072] 2) Enhanced generalization capability of multiple convolution kernels

[0073] By using multiple convolution kernels, the model can select the most suitable convolution operation at each time step to avoid relying on a single feature. This mechanism enables the model to more accurately extract and utilize the temporal features of data when dealing with different types of network environments, thereby improving performance in unknown environments and enhancing generalization capabilities.

[0074] 3). Adaptability to different network conditions

[0075] Since 1D dynamic convolution focuses on the time dimension, through multiple convolution kernels, each convolution kernel may focus on different network states or user behavior patterns (such as bandwidth fluctuations, packet loss, etc.). This adaptive capability helps the model better handle different data distributions, thereby improving stability and adaptability in different network environments.

[0076] Through 1D dynamic convolution, the model can be optimized for different timing features during training, improving the model's adaptability to a variety of network conditions and user devices, thereby improving overall generalization capabilities. The core of this method lies in the adaptive adjustment of the convolution kernel, which enables the model to flexibly select the optimal convolution operation based on the different characteristics of the input data and optimize video transmission performance.

[0077] In addition, the general A3C algorithm includes a value network and a policy network, and in this embodiment, the training method of the deep reinforcement learning model also includes: adding a value network to the A3C algorithm so that the value network in the A3C algorithm has N, where N is greater than or equal to 2, and each value network corresponds to a data or environment.

[0078] After the introduction of 1D dynamic convolution in the embodiment of the present application, although the generalization ability of the model is significantly improved and can better adapt to changes in different network conditions and user devices, it also brings challenges in the training process. Due to the adaptability of dynamic convolution, the convolution kernel is adjusted at each input, and the parameter space of the model becomes more complex, which leads to instability and slow convergence in the training process.

[0079] Especially when using a reinforcement learning framework (such as Actor-Critic), the value estimation of the Critic network is crucial for updating the Actor strategy. However, when training in multiple data sets and network environments, a single Critic is difficult to adapt to all possible network conditions and is prone to overfitting or estimation errors. Therefore, in order to solve this problem, the embodiment of the present application proposes to expand the Actor-Critic to multiple Critic to better cope with changes in different data sets and improve the stability and generalization ability of model training.

[0080] In the training process of the A3C (Asynchronous Advantage Actor-Critic) algorithm, introducing multiple critics to correct the Actor's strategy learning is an effective improvement method. Figure 4 , Figure 4 The left side is the A3C network. Figure 4 The right side shows a schematic diagram of adding multiple critics to the A3C network, especially when training on multiple datasets or in heterogeneous environments. The main advantages of this method are as follows:

[0081] 1) Improve the robustness of strategy evaluation

[0082] Traditional A3C uses a single critic to estimate value, which may be biased in different data distributions or environments. By introducing multiple critics, each critic can focus on different data sets or environments, thereby providing more stable and accurate value assessments and reducing the instability of a single critic due to data bias.

[0083] 2) Enhance generalization ability

[0084] Since each critic learns the value function from different data sets or network environments, the actor can integrate the feedback of multiple critics when updating the strategy, making it more generalizable in a wider range of environments. This is especially important for adaptive rate control tasks in heterogeneous networks or with multiple QoE requirements.

[0085] 3) Alleviate training instability

[0086] In reinforcement learning, the error of the Critic will directly affect the update of the Actor, and the introduction of multiple Critic can reduce the impact of individual Critic errors through voting, weighted averaging or other fusion methods, thereby reducing fluctuations in the training process and improving the convergence speed.

[0087] 4) Adapt to dynamic changes in different network conditions

[0088] Since the network environment for video streaming is complex and changeable, different Critic can independently learn for different network conditions (such as bandwidth fluctuation, packet loss rate, delay, etc.), so that Actors can adapt to different network conditions and choose better bitrate strategies.

[0089] In summary, introducing multiple critics in the A3C algorithm training process not only helps to improve the accuracy and generalization ability of strategy evaluation, but also alleviates training instability and improves training efficiency, making it more suitable for adaptive video transmission tasks in complex heterogeneous network environments.

[0090] In summary, the adaptive transmission method for heterogeneous network video streams provided by this application introduces one-dimensional dynamic convolution into deep reinforcement learning, and enhances the adaptability of the model to different network conditions and improves the generalization of the model through the characteristics of multiple convolution kernels. In addition, in the face of the difficulty in training and optimization caused by the introduction of dynamic convolution, the single critic of the actor-critic framework is expanded to mutil-critic to alleviate training instability and improve the robustness of strategy evaluation.

[0091] The adaptive transmission method for heterogeneous network video streams provided in the embodiment of the present application can be applied to video streaming services according to the technical effects that can be achieved. Specifically, it can be applied to live video and video on demand services such as YouTube and Netflix. During use, the bit rate of the video can be adjusted in real time according to the user's network bandwidth and device performance to ensure a smooth playback experience. Or it can also be applied to online meetings and distance education. Specifically, it can be applied to online meeting platforms such as Zoom and Microsoft Teams, as well as distance education platforms. During use, the clarity and smoothness of video conferencing and online education content in different network environments can be ensured. There are also cloud games, advertising, smart homes and other fields, all of which can apply the adaptive transmission method for heterogeneous network video streams provided in the embodiment of the present application.

[0092] In order to more intuitively demonstrate the effect of the adaptive transmission method for heterogeneous network video streams provided by the embodiment of the present application, this embodiment is also experimentally verified, as follows:

[0093] Experimental setup:

[0094] 1) QoE model

[0095] The core goal of the bitrate adaptation strategy is to improve the QoE of video services. The three key factors affecting QoE are video quality, freeze time, and video quality fluctuation. The QoE evaluation criteria considered in this experiment are as follows:

[0096]

[0097] For a video with N blocks, R n represents the bit rate of the nth block, q(R n ) maps bit rate to user perceived quality, T n Indicates bit rate R n The rebuffering time for downloading the nth video chunk, and the last term penalizes the variation of video quality to promote smoothness. In the experiment, the embodiment of the present application considers q(Rn)=Rn, μ=4.3, and the experimental result is the average QoE of each chunk, that is, the total QoE metric divided by the number of chunks in the video.

[0098] 2) Network bandwidth and corresponding video dataset

[0099] Network: For different network scenarios, the present embodiment uses 4 public data sets collected from real networks (see Table 1 for details) for training and evaluation. The present embodiment divides the first three data sets into a ratio of 8:2 for training set and test set baselines. Oboe is only used for comparison of generalization test.

[0100]

[0101] Table 1: Public datasets used

[0102] Video: The "EnvivioDash3" video from the DASH-246 JavaScript reference client is used. The video is encoded with the H.264 / MPEG-4 codec with bitrates of {300, 750, 1200, 1850, 2850, 4300} kbps. In addition, the video is divided into 48 chunks with a total length of 193 seconds. Therefore, each chunk represents approximately 4 seconds of video playback.

[0103] 3) Comparison of algorithms

[0104] The most representative bitrate adaptation strategies in recent years are selected for performance comparison:

[0105] (1) BB (Buffer-Based Algorithm): Constructs a mapping relationship between the playback buffer size and the bitrate level, and makes a bitrate decision directly based on the buffer size;

[0106] (2) BOLA (Buffer Occupancy based Lyapunov Algorithm): uses the Lyapunov optimization method to balance the video clip bitrate level and the playback buffer size, and selects the maximum average bitrate that minimizes the freeze time;

[0107] (3) RB (Rate-Based Algorithm): The harmonic mean bandwidth is calculated based on the download time and actual file size of the downloaded video clips as the predicted value of future bandwidth, and the maximum bit rate that will not exhaust the buffer area is selected based on the predicted value;

[0108] (4) RobustMPC: With user experience quality as the optimization goal, the harmonic mean of historical bandwidth data is used as the future available bandwidth. The sub-optimization problem is constructed based on model predictive control, and the bitrate of each video segment is determined by the rolling optimization solution step sequence.

[0109] (5) Pensieve: The first rate adaptation strategy using deep reinforcement learning. Using an asynchronous advantage actor-critic algorithm to learn the adaptive strategy directly in the training environment, this work has had a profound impact on the field of rate adaptation.

[0110] 4) Performance comparison

[0111] In order to verify the performance of the algorithm provided in the embodiment of the present application in various network scenarios, the performance of each rate adaptation strategy in multiple scenarios was tested. The average QoE and its various indicators of each strategy were obtained. The algorithm provided in the embodiment of the present application showed significant advantages in improving the user experience quality of video services in the FCC18 and Ghent scenarios. In the FCC18 data set, compared with other strategies, an improvement of 2.0%-20.8% was achieved in the QoE evaluation standard; in the Ghent data set, compared with other strategies, an improvement of 2.2%-14.0% was achieved in the QoE evaluation standard; in the HSDPA data set, compared with Pensieve, the performance was slightly reduced, but compared with the RB strategy, it still achieved an improvement of 118%. The algorithm provided in the embodiment of the present application effectively guarantees the video viewing experience of video users in various scenarios by adopting a data-aware network structure that combines one-dimensional dynamic convolution with mutil-critic.

[0112] The model trained by Pensieve in one scenario has obvious performance degradation in the other two scenarios. In particular, the model trained in the Ghent scenario has a catastrophic application when tested in the HSDPA scenario. However, the model in the embodiment of the present application can achieve a good user video viewing experience in multiple scenarios while ensuring good QoE.

[0113] As mentioned above, models trained with multiple network trajectories will not improve adaptability, and their performance may even be worse than models trained with a single network dataset. In the example, a DRL model trained to fit widely distributed data will result in performance degradation if tested only on a relatively narrow distribution. Both pensieve and the method of this embodiment train DRL models on (HSDPA+Ghent+FCC) data, and the method of this embodiment achieves improvements in QoE indicators in all scenarios.

[0114] Although Pensieve is also trained on multiple data sets, the generalization ability of this method is limited in unseen network environments and scenarios, resulting in a decrease in its performance on out-of-distribution (OOD) data. The method of this embodiment enhances the adaptability of the model to different network conditions and improves the generalization ability of the strategy by introducing 1D dynamic convolution and multi-critic mechanisms. Specifically, the model of this embodiment can better learn the intrinsic structure of the data and maintain a high QoE score in unseen scenarios, showing stronger robustness. This shows that the method of this embodiment has advantages in out-of-distribution generalization, thereby being able to provide a more stable and efficient rate adaptation strategy in a complex and changeable network environment.

[0115] See also Figure 5 The embodiment of the present application also provides an adaptive transmission system for heterogeneous network video streams, including: a state information acquisition module 1, a deep reinforcement learning model module 2, a bit rate receiving module 3 and a bit rate adjustment module 4; the state information acquisition module 1 is used to obtain the state information of the played video block and construct input data; the deep reinforcement learning model module 2 is used to input the input data into a pre-trained deep reinforcement learning model, and the deep reinforcement learning model is used to predict the bit rate of the next unplayed video block based on the input data; the bit rate receiving module 3 is used to receive the bit rate of the next unplayed video block generated by the deep reinforcement learning model; the bit rate adjustment module 4 is used to adjust the bit rate of the next video block when it is played according to the bit rate of the next unplayed video block.

[0116] The adaptive transmission system for heterogeneous network video streams provided in this embodiment can alleviate the instability of ABR algorithm training and improve the robustness of strategy evaluation.

[0117] In some embodiments, the state information includes: a network throughput measurement of k video chunks that have been played, a download time of the k video chunks that have been played, a vector of m available sizes for the next unplayed video chunk, a current buffer level, a number of remaining video chunks in the video, and a bit rate of the last video chunk in the downloaded video.

[0118] In some embodiments, the deep reinforcement learning model module 2 includes an algorithm construction unit and an algorithm correction unit. The algorithm construction unit is used to construct the A3C algorithm; the algorithm correction unit is used to correct the convolution layer of the A3C algorithm to a dynamic convolution layer to improve the generalization ability of the A3C algorithm.

[0119] In some embodiments, the dynamic convolution layer includes: an attention mechanism, an adaptive convolution, a batch normalization, and an activation function.

[0120] In some embodiments, the algorithm correction unit includes: a key state information acquisition subunit, a convolution kernel weight generation subunit, a weighting subunit and an execution subunit; the key state information acquisition subunit is used to obtain the key state information of the video block, and the key state information includes the bandwidth and cache occupancy of the video block; the convolution kernel weight generation subunit is used to calculate the dynamic weights according to the key state information using the attention mechanism to generate multiple one-dimensional convolution kernel weights; the weighting subunit is used to use adaptive convolution to combine different one-dimensional convolution kernel weights with input features in a weighted manner; the execution subunit is used to perform batch normalization and activation function on the combined one-dimensional convolution kernel weights and input features to generate a dynamic convolution layer.

[0121] In some embodiments, the algorithm modification unit further includes: a value network adding subunit, which is used to add a value network to the A3C algorithm so that the value network in the A3C algorithm has N, where N is greater than or equal to 2, and each value network corresponds to a data or environment.

[0122] The present application embodiment provides an electronic device. Figure 6 The electronic device includes: a memory 601, a processor 602, and a computer program stored in the memory 601 and executable on the processor 602. When the processor 602 executes the computer program, the adaptive transmission method for heterogeneous network video streams described above is implemented.

[0123] Furthermore, the electronic device also includes: at least one input device 603 and at least one output device 604 .

[0124] The memory 601 , processor 602 , input device 603 and output device 604 are connected via a bus 605 .

[0125] The input device 603 may be a camera, a touch panel, a physical button or a mouse, etc. The output device 604 may be a display screen.

[0126] The memory 601 may be a high-speed random access memory (RAM) memory or a non-volatile memory, such as a disk memory. The memory 601 is used to store a set of executable program codes, and the processor 602 is coupled to the memory 601 .

[0127] Furthermore, the embodiment of the present application also provides a computer-readable storage medium, which may be disposed in the electronic device in the above embodiments, and the computer-readable storage medium may be the memory 601 in the above embodiments. The computer-readable storage medium stores a computer program, which, when executed by the processor 602, implements the adaptive transmission method for heterogeneous network video streams described in the above method embodiment.

[0128] Furthermore, the computer storable medium may also be a U disk, a mobile hard disk, a read-only memory (ROM), a RAM, a magnetic disk or an optical disk, or other medium that can store program codes.

[0129] An embodiment of the present application also provides a computer program product for executing the adaptive transmission method for heterogeneous network video streams provided in any of the above embodiments.

[0130] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.

[0131] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0132] In addition, each functional module in each embodiment of the present invention may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of software functional modules.

[0133] If the integrated module is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention.

[0134] It should be noted that, for the convenience of description, the aforementioned method embodiments are all described as a series of action combinations, but those skilled in the art should be aware that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.

[0135] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0136] In summary, although the present invention has been disclosed as above in terms of preferred embodiments, the above preferred embodiments are not intended to limit the present invention. A person skilled in the art may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention shall be based on the scope defined in the claims.

Claims

1. A method for adaptive transmission of heterogeneous network video streams, characterized in that: include: Get the status information of the played video blocks and construct the input data; Inputting the input data into a pre-trained deep reinforcement learning model, wherein the deep reinforcement learning model is used to predict the bit rate of the next unplayed video block based on the input data; receiving a bit rate of a next unplayed video chunk generated by the deep reinforcement learning model; The bit rate of the next video block when it is played is adjusted according to the bit rate of the next unplayed video block.

2. The method for adaptive transmission of heterogeneous network video streams according to claim 1, characterized in that: The state information includes: network throughput measurement of k video chunks played, download time of k video chunks played, vector of m available sizes for the next unplayed video chunk, current buffer level, number of remaining video chunks in the video, bit rate of the last video chunk in the downloaded video.

3. The method for adaptive transmission of heterogeneous network video streams according to claim 1, characterized in that: The training method of the deep reinforcement learning model includes: Build the A3C algorithm; The convolution layer of the A3C algorithm is modified to a dynamic convolution layer to improve the generalization ability of the A3C algorithm.

4. The method for adaptive transmission of heterogeneous network video streams as claimed in claim 3, characterized in that: The dynamic convolution layer includes: attention mechanism, adaptive convolution, batch normalization and activation function.

5. The method for adaptively transmitting heterogeneous network video streams as claimed in claim 4, characterized in that: The generation step of the dynamic convolutional layer includes: Acquire key status information of the video block, where the key status information includes bandwidth and cache occupancy of the video block; Calculating dynamic weights using an attention mechanism according to the key state information to generate multiple one-dimensional convolution kernel weights; Using the adaptive convolution to combine different one-dimensional convolution kernel weights with input features in a weighted manner; The batch normalization and the activation function are performed on the combined one-dimensional convolution kernel weights and input features to generate a dynamic convolution layer.

6. The method for adaptive transmission of heterogeneous network video streams as claimed in claim 3, characterized in that: The A3C algorithm includes a value network and a policy network; The training method of the deep reinforcement learning model also includes: A value network is added to the A3C algorithm so that the A3C algorithm has N value networks, where N is greater than or equal to 2, and each value network corresponds to one data or environment.

7. A video stream adaptive transmission system for heterogeneous networks, characterized in that: include: A status information acquisition module is used to obtain status information of played video blocks and construct input data; A deep reinforcement learning model module, used for inputting the input data into a pre-trained deep reinforcement learning model, wherein the deep reinforcement learning model is used for predicting the bit rate of the next unplayed video block based on the input data; A bit rate receiving module, configured to receive a bit rate of a next unplayed video block generated by the deep reinforcement learning model; The bit rate adjustment module is used to adjust the bit rate of the next video block when it is played according to the bit rate of the next unplayed video block.

8. An electronic device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method described in any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method described in any one of claims 1 to 6 is implemented.

10. A computer program product, characterized in that Used to execute the method according to any one of claims 1 to 6.