VR Video Adaptive Bitrate Control Method Based on Multi-Agent Reinforcement Learning

Through the VR video adaptive bit rate control method of multi-intelligent reinforcement learning, the bit rate is dynamically adjusted to adapt to network changes, solving the problem of inflexible bit rate control and insufficient collaboration of multi-intelligent bodies in the existing technology, achieving smoother video playback and higher user experience.

CN120151495BActive Publication Date: 2025-07-11ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510451910.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-11
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The existing VR video adaptive bit rate control method cannot accurately reflect the actual network conditions, resulting in inflexible bit rate adjustment and lag. The multi-agent collaboration mechanism has shortcomings in information sharing and status updates, which affects the accuracy and real-timeness of bit rate selection.

Method used

Using a multi-intelligent reinforcement learning method, by obtaining historical video block transmission data and predicting bandwidth, building network delay feature vectors and buffer states, establishing a bit rate switching impact model, combining multi-layer perceptron structure and Actor network, optimizing the bit rate selection strategy, using multi-agent joint state vectors and timing differential targets to optimize the Critic network, dynamically adjusting the video stream bandwidth share, and achieving optimal bit rate decision.

Benefits of technology

It improves the smoothness and user experience of video transmission, reduces heavy buffering, enhances the system's adaptability and video immersion in different network environments, and improves the overall viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120151495B_ABST
    Figure CN120151495B_ABST
Patent Text Reader

Abstract

The present invention provides a VR video adaptive bitrate control method based on multi-agent reinforcement learning, which relates to the field of VR technology. It includes obtaining historical video block transmission data to predict the bandwidth and evaluating network stability, outputting a bitrate selection probability distribution by an Actor network based on a multi-layer perceptron structure, calculating the structural similarity and rebuffering metrics of video blocks to construct a comprehensive reward function, obtaining local observation information using a multi-agent joint state vector and integrating group information through attention weights, and updating network parameters using a policy gradient method and an Adam optimizer to achieve bitrate optimization. The present invention can effectively improve the VR video transmission quality, reduce the rebuffering probability, ensure the continuity and smoothness of the user viewing experience, and at the same time reduce the bitrate switching frequency and improve the system stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to VR technology, and in particular to a VR video adaptive bitrate control method based on multi-agent reinforcement learning. Background Art

[0002] With the rapid development of virtual reality technology, VR videos are increasingly widely used. However, the streaming transmission of VR videos faces problems such as bandwidth fluctuations, network latency, and buffering, which directly affect the user's viewing experience. To ensure smooth video playback, adaptive bitrate control technology has emerged. This technology dynamically adjusts the video bitrate by real-time monitoring of the network condition and the performance of the user device to adapt to different network environments.

[0003] Existing adaptive bitrate control methods often rely on simple bandwidth prediction models, which cannot accurately reflect the actual network condition, resulting in inflexible bitrate adjustment and easy occurrence of stuttering.

[0004] Many existing technologies lack comprehensiveness in the evaluation of buffer state, and fail to fully consider the balance between video quality and user experience, which may lead to frequent quality fluctuations during the user's viewing process.

[0005] Existing multi-agent cooperation mechanisms have deficiencies in information sharing and state update, and cannot effectively integrate the observation information of each agent, resulting in low overall decision-making efficiency and affecting the accuracy and real-time nature of bitrate selection. Summary of the Invention

[0006] Embodiments of the present invention provide a VR video adaptive bitrate control method based on multi-agent reinforcement learning, which can solve the problems in the existing technology.

[0007] In the first aspect of the embodiments of the present invention,

[0008] A VR video adaptive bitrate control method based on multi-agent reinforcement learning is provided, including:

[0009] Obtaining historical video block transmission data and using predicted bandwidth, constructing a network latency feature vector to evaluate network stability, establishing a bitrate switching influence model based on the predicted bandwidth and buffer state, performing threshold control through a buffer state transition equation, dynamically adjusting in combination with the video stream bandwidth share, and selecting the optimal bitrate based on quality score, switching influence, and network stability;

[0010] Construct a feature vector including buffer state and video quality state based on network delay feature vectors, input the feature vectors into the Actor network with a multi-layer perceptron structure, output the bitrate selection probability distribution through the ReLU activation function and the Softmax function, construct a comprehensive reward function based on video quality, buffer state and switching smoothness, and determine the bitrate level by combining batch sampling from the experience pool;

[0011] Calculate the structural similarity of video blocks to obtain quality mapping values, calculate the rebuffering metrics based on download time and buffer state, evaluate the quality change by combining quality difference and time decay weight, collect viewport position information to calculate smoothness and prediction error, determine the video immersion degree based on the field of view coverage rate and spatial quality distribution, construct a comprehensive reward function from the metrics and normalize it, calculate the cumulative reward based on the discount factor and optimize the weight coefficients;

[0012] Construct a multi-agent joint state vector to obtain local observation information, optimize the mean square error loss of the Critic network using the temporal difference target, update the Actor network by combining policy gradient and entropy regularization, adopt the mechanism of parameter sharing for homogeneous agents and independent update for heterogeneous agents, and integrate the group observation information through attention weights;

[0013] Continuously store the state vector interaction data in the experience pool, extract network features using a multi-layer perceptron, construct an adaptive reward function based on the feature vectors, update the network parameters using the policy gradient method and the Adam optimizer, adjust the reward weights according to the difference in experience quality, and continuously optimize the bitrate selection strategy by sampling from the experience pool.

[0014] Obtain the historical video block transmission data and adopt the predicted bandwidth, construct a network delay feature vector to evaluate network stability, establish a bitrate switching impact model based on the predicted bandwidth and buffer state, perform threshold control through the buffer state transition equation, dynamically adjust in combination with the video stream bandwidth share, and select the optimal bitrate based on quality score, switching impact and network stability, including:

[0015] Obtain the throughput data and download time data of the past k video blocks during video transmission, where the throughput data is the ratio of the size of each video block to the corresponding download time; perform weighted calculation on the throughput data using an exponential decay weight to obtain the weighted predicted bandwidth, where the degree of decay of the exponential decay weight over time is determined by the decay coefficient, and the decay coefficient is used to adjust the influence degree of historical data on bandwidth prediction;

[0016] Construct a network latency feature vector based on the download time data, calculate the variance value of the network latency feature vector as a quantization index for network stability. The larger the variance value, the worse the network stability. Based on the weighted predicted bandwidth, combined with the current buffer size, the predicted size information of the next video block corresponding to different bitrates, and the current bitrate, establish a bitrate switching impact evaluation model, which is used to evaluate the impact degree of bitrate switching on the system performance.

[0017] Construct a buffer state transition equation according to the current buffer size, weighted predicted bandwidth, and video block duration, and set a buffer safety threshold. When the predicted buffer size at the next moment by the buffer state transition equation is less than the buffer safety threshold, select a conservative bitrate based on the weighted predicted bandwidth.

[0018] When detecting multiple video streams, obtain the target bitrate of each video stream, calculate the bandwidth share factor of each video stream, where the bandwidth share factor is the ratio of the target bitrate of the corresponding video stream to the total sum of the target bitrates of all video streams, and adjust the weighted predicted bandwidth based on the bandwidth share factor to obtain the adjusted bandwidth.

[0019] Construct an optimization objective function with the video quality score, the evaluation result of the bitrate switching impact evaluation model, and the variance value. Adjust the weights of each evaluation index by setting a balance factor. Based on the adjusted bandwidth, select the bitrate that makes the optimization objective function obtain the optimal value from the preset bitrate set as the final bitrate decision result.

[0020] Construct a state feature vector including the buffer state and video quality state based on the network latency feature vector, input the feature vector into the Actor network with a multi-layer perceptron structure, and output the bitrate selection probability distribution through the ReLU activation function and Softmax function. Construct a comprehensive reward function based on video quality, buffer state, and switching smoothness, and determine the bitrate level by batch sampling from the experience pool, including:

[0021] Construct a state feature vector, which includes the buffer state and video quality state. The buffer state includes the current buffer size and the buffer safety threshold, and the video quality state includes the current bitrate level and the bitrate switching amplitude. Input the state feature vector into the Actor network, and the Actor network adopts a multi-layer perceptron structure, including a first hidden layer and a second hidden layer, and both the first hidden layer and the second hidden layer adopt the ReLU activation function.

[0022] Construct a comprehensive reward function based on the current bitrate level, the current buffer size, and the bitrate switching amplitude. The comprehensive reward function includes video quality reward, buffer state reward, and switching smoothness reward. Among them, the video quality reward is the weighted value of the ratio of the current bitrate level to the maximum bitrate level. The buffer state reward is the weighted value of the calculation result of the ratio of the current buffer size to the buffer safety threshold. The switching smoothness reward is the normalized weighted value of the ratio of the bitrate switching amplitude to the maximum bitrate level. Calculate the policy gradient according to the comprehensive reward function, and use the policy gradient to iteratively optimize the network parameters of the Actor network through the gradient descent method to obtain the optimized network parameters.

[0023] Construct a transition sample by combining the state feature vector, the action determined based on the bitrate selection probability distribution, the calculation result of the comprehensive reward function, and the next state feature vector, and store it in the experience pool. Perform batch sampling from the experience pool based on a preset sampling size, and update the optimized network parameters in combination with the exploration strategy. Input the updated network parameters into the Actor network, and determine the final bitrate level based on the bitrate selection probability distribution output by the Actor network.

[0024] Calculate the structural similarity of the video block to obtain the quality mapping value, calculate the rebuffering metric based on the download time and the buffer state, evaluate the quality change by combining the quality difference and the time decay weight, collect the viewport position information to calculate the smoothness and the prediction error, determine the video immersion degree based on the field of view coverage rate and the spatial quality distribution, construct a comprehensive reward function for the metrics and normalize it, calculate the cumulative reward based on the discount factor and optimize the weight coefficients, including:

[0025] Obtain the video block bitrate and the video block image data, calculate the structural similarity based on the mean, standard deviation, and covariance of the video block image data, and map the video block bitrate and the structural similarity to obtain the quality mapping value.

[0026] Monitor the video block download process, obtain the download time of each video block and the corresponding remaining buffer duration, accumulate the rebuffering duration when the download time is greater than the remaining buffer duration, and calculate the rebuffering frequency based on the number of rebuffering occurrences and the total playback duration. Obtain the quality mapping values of adjacent video blocks to calculate the quality difference, construct a time decay weight sequence, and use the weighted sum of the quality difference and the time decay weight sequence as the quality change evaluation index.

[0027] Real-time collect the user's viewport position information, extract the azimuth angle and the pitch angle in the user's viewport position information, calculate the Euclidean distance of the azimuth angle and the pitch angle at adjacent moments to obtain the viewport smoothness, and calculate the viewport prediction error based on the deviation distance between the viewport position prediction value and the actual value.

[0028] Calculate the ratio of the area of the field of view region to the total area of the field of view in the video frame to obtain the field of view coverage rate. At the same time, divide the field of view space and assign spatial weights to each region, and use the weighted sum of the quality mapping values of each region and the corresponding spatial weights as the video immersion degree;

[0029] Perform a linear combination of the quality mapping value, re-buffering duration, quality change evaluation index, viewport smoothness, viewport prediction error, and video immersion degree to construct a comprehensive reward function, and perform maximum-minimum normalization processing on the comprehensive reward function to obtain a normalized reward value;

[0030] Calculate the cumulative discounted reward based on a preset discount factor. The cumulative discounted reward is the weighted sum of the normalized reward values at each moment multiplied by the exponential decay coefficient of the corresponding time step; use the gradient ascent method to iteratively optimize the weight coefficients of each evaluation index in the comprehensive reward function. Under the constraint that the sum of the weight coefficients is 1 and each coefficient is non-negative, determine the optimal weight coefficient combination by maximizing the expected value of the cumulative discounted reward.

[0031] Construct a multi-agent joint state vector to obtain local observation information, use the temporal difference target to optimize the mean square error loss of the Critic network, update the Actor network in combination with policy gradient and entropy regularization, adopt the mechanism of parameter sharing for homogeneous agents and independent update for heterogeneous agents, integrate the group observation information through attention weights, and synchronously update the network parameters based on the global experience pool data to achieve collaborative learning, including:

[0032] Construct a multi-agent joint state vector, which contains the state information of all agents, and map the multi-agent joint state vector to the local observation information of each agent through an observation function;

[0033] Based on the local observation information of each agent, independently generate actions using the corresponding policy function, and record the multi-agent joint state vector, actions, the reward value obtained at the current moment, and the information of the transferred next state as a state transition tuple and store it in the experience pool;

[0034] When the amount of data in the experience pool reaches the preset storage threshold, randomly sample batch data from the experience pool, and calculate the temporal difference target based on the reward value in the batch data and the evaluation value of the target network for the next state; use the temporal difference target and the predicted value of the critic network for the current state-action combination to construct a mean square error loss function, and update the parameters of the critic network based on the gradient direction of the mean square error loss function;

[0035] Multiply the action value function of the evaluation state of the critic network by the logarithmic probability of the policy function to calculate the policy gradient, introduce a policy entropy regularization term into the policy gradient, and update the actor network parameters based on the final gradient direction; identify isomorphic agents in the system, adopt a parameter sharing mechanism for isomorphic agents to uniformly update the network parameters, and at the same time keep the parameters of heterogeneous agents independently updated;

[0036] Calculate the attention weights according to the similarity of the observation information between heterogeneous agents, and use the attention weights to weighted integrate the observation information of other agents to obtain an enhanced observation representation containing group information; merge the experience data of all agents into the global experience pool, calculate the overall gradient based on the data in the global experience pool, and synchronously update the global network parameters through the overall gradient to achieve collaborative learning of multiple agents.

[0037] Continuously store the state vector interaction data in the experience pool, use a multi-layer perceptron to extract network features, construct an adaptive reward function based on the feature vectors, adopt a policy gradient method and an Adam optimizer to update the network parameters, adjust the reward weights according to the difference in experience quality, and continuously optimize the bitrate selection strategy by sampling from the experience pool, including:

[0038] Construct an experience pool to store the state vector interaction data, where the state vector interaction data includes state vectors and bitrate decisions; use a multi-layer perceptron to extract network features based on the state vectors, and perform a non-linear transformation on the state vectors through the ReLU activation function to obtain feature vectors; construct an adaptive reward function according to the feature vectors, and the adaptive reward function is obtained by weighted calculating the bitrate quality score, the buffer jitter degree, and the change range of adjacent bitrate decisions, and the weighted calculation uses dynamically adjusted weight coefficients;

[0039] Construct a bitrate selection policy network based on the feature vectors, use a policy gradient method to calculate the parameter gradient of the bitrate selection policy network, and the parameter gradient is obtained through the expected value of the product of the logarithm of the policy function and the state-action value function; introduce an Adam optimizer to update the parameters of the bitrate selection policy network, and the Adam optimizer calculates an adaptive learning rate based on the first-order momentum estimation and the second-order momentum estimation, and uses the adaptive learning rate to update the parameters of the policy network;

[0040] Calculate the difference between the target experience quality and the current experience quality, dynamically adjust the weight coefficients in the adaptive reward function based on the difference, and achieve the adaptive update of the weight coefficients through a preset step size; during the video stream transmission process, continuously randomly sample batch data from the experience pool, and use the batch data to continuously optimize the bitrate selection strategy.

[0041] In the second aspect of the embodiments of the present invention,

[0042] Provide an electronic device, including:

[0043] A processor;

[0044] A memory for storing instructions executable by the processor;

[0045] Wherein, the processor is configured to call the instructions stored in the memory to execute the foregoing method.

[0046] In the third aspect of the embodiments of the present invention,

[0047] A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the foregoing method is implemented.

[0048] The beneficial effects of this application are as follows:

[0049] 1. By obtaining historical video block transmission data and predicting the bandwidth, this method can effectively evaluate network stability, thereby achieving adaptive bitrate control and improving the smoothness of video transmission and the user experience.

[0050] 2. Based on the design of the Actor network with a multi-layer perceptron structure and the comprehensive reward function, the bitrate selection becomes more intelligent, can dynamically adapt to the video quality and buffer status, reduce the re-buffering phenomenon, and improve the continuity of video playback.

[0051] 3. By optimizing the multi-agent joint state vector and the temporal difference target, the learning ability and adaptability of the system are enhanced, enabling a high video immersion and quality to be maintained in different network environments and improving the overall viewing experience. Description of the Drawings

[0052] Figure 1 It is a system diagram of the VR video bitrate adaptive method based on multi-agent reinforcement learning according to the embodiments of the present invention;

[0053] Figure 2 It is a neural network diagram of the VR video bitrate adaptive method based on multi-agent reinforcement learning according to the embodiments of the present invention. Detailed Embodiments

[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0055] The technical solution of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0056] As Figure 1 - Figure 2 shown, the method described in the embodiment of the present invention includes:

[0057] Step 1: Divide the complete 360° panoramic video into multiple segments at fixed intervals, and then each video segment is further divided into multiple view shards. Use the H.265 encoder to encode these shards to generate an H.265 encoded video stream for streaming transmission.

[0058] Step 2: During the video stream transmission process, each agent real-time collects environmental state information , and according to the current state observation , calculates the selection probability of each bitrate level through the actor network, and selects the bitrate of the next video block according to the probability distribution. Apply the selected bitrate to the transmission of the video stream, and store the state transition information in the experience pool.

[0059] Step 3: The bitrate of the video block after the decision of the actor network is transmitted to each user through the wireless network, and the user decodes and plays the target video after receiving it.

[0060] Step 4: According to the state of the transmitted video stream and the feedback information of the user, calculate the reward according to the definition of QoE. The critic network inputs the joint information of the states and actions of all agents, evaluates the value of the current state-action combination, and guides the actor network to make better decisions. When the experience pool meets the update conditions, start the training process, and update the neural network parameters of the agent according to the centralized training method.

[0061] Step 5: During the entire process of video stream transmission, continuously repeat the above steps (2) to (4), and the agent continuously learns and optimizes the bitrate selection strategy until the end of this video.

[0062] Regarding the agent state observation and action selection in Step 2, the specific situation of this embodiment is as follows:

[0063] In this embodiment, the agent continuously observes a series of metrics to comprehensively understand the current network status and user behavior, and observes the state at any time Specifically:

[0064] ;

[0065] Among them, : is the network throughput of the past k video chunks. Network throughput reflects the network's ability to transmit data over a period of time. In this embodiment, the agent monitors , and can understand the current network's busyness and the efficiency of data transmission. The specific judgment executed in this embodiment is: If 's value remains low, it indicates that the network may be congested. When selecting a bitrate, the agent may need to lower the bitrate to avoid video stuttering and ensure that the video can be smoothly transmitted under the current network conditions.

[0066] : is the download time of the past k video chunks. The length of the download time directly affects the smoothness of the user's video viewing. A longer download time may mean higher network latency or insufficient bandwidth. In this embodiment, based on the 's changing trend, the agent can judge whether the network condition is stable. The specific judgment executed in this embodiment is: is increasing continuously, and the agent should tend to select a lower bitrate to reduce the download time and ensure the real-time playback of the video.

[0067] : is a vector that contains the size information of the next video chunk at different bitrates. In this embodiment, based on this information, the agent can predict in advance the impact of different bitrate selections on the data transmission volume. The specific judgment executed in this embodiment is: When the agent considers selecting a higher bitrate, it can view the corresponding video chunk size. If it is found that selecting this high bitrate will cause a significant increase in the data volume of the next video chunk, and the current network bandwidth is limited, the agent will avoid selecting this high bitrate to prevent network congestion caused by excessive data transmission and ensure the stable transmission of the video stream.

[0068] : is the bitrate of the previous video chunk. The bitrate selection of the previous video chunk provides an important reference for the agent's subsequent decisions. In this embodiment, if the bitrate selection of the previous video chunk causes problems in video playback, such as stuttering or poor quality, the agent will adjust its strategy during this decision-making. The specific judgment executed in this embodiment is: If a high bitrate was selected for the previous video chunk but there was stuttering, the bitrate may be lowered this time; conversely, if the low bitrate of the previous video chunk played smoothly but the visual quality was poor, the bitrate may be increased this time.

[0069] : is the current buffer size. The buffer size reflects the remaining space in the client buffer. In this embodiment, when the buffer is small, the agent needs to carefully select the bitrate to avoid quickly filling the buffer due to selecting too high a bitrate, which may cause a re-buffering phenomenon. The specific judgment executed in this embodiment is: If There is insufficient remaining space. The agent should preferentially select a low bitrate to ensure the continuity of video playback and prevent video pauses waiting for data loading due to buffer overflow.

[0070] : is the number of remaining video blocks. The agent has an overall understanding of the progress of video transmission and to a certain extent assists the agent in long-term decision-making and planning.

[0071] : is the bandwidth ratio. In this embodiment, it represents the percentage of bandwidth resources occupied by the current type of video block. In the case of multiple video streams sharing network resources, it determines the bandwidth resources that can be allocated to each agent to ensure fairness among different video streams. The specific judgment executed in this embodiment is: If the current VR 360° video stream already occupies a relatively high proportion of bandwidth and other video streams also have bandwidth requirements, the agent may reduce its bitrate selection to yield some bandwidth to other video streams, achieving a reasonable allocation and fair use of the overall network resources.

[0072] After the agent comprehensively observes the above states it selects an action from 6 pre-set bitrate levels and this action determines the bitrate of the next video block. In this embodiment, respectively corresponding to {300, 750, 1200, 1850, 2850, 4300} Kbps. Each agent uses a policy function to define the exact mapping from state to action, and this function presents as a probability distribution. In this embodiment, a neural network actor with parameters is used to approximate this policy: . During the training process, the agent continuously adjusts the parameters according to the rewards feedback from the environment to optimize the policy. The specific judgment executed is: When the agent selects a bitrate, if the bitrate results in good video stream playback quality and does not cause network problems, the environment will give a higher reward. At this time, the agent will strengthen the tendency to select this bitrate or a similar bitrate, and increase the probability of selecting these bitrates by adjusting the parameters . On the contrary, if the selected bitrate causes video stuttering or a serious decline in quality, the environment gives a lower reward or even a penalty, and the agent will adjust the parameters to reduce the probability of selecting this bitrate. By continuously interacting with the environment and learning, the agent gradually masters the optimal bitrate selection strategy, can flexibly adapt to different network environments and video content requirements, and achieve efficient adaptive bitrate control to provide users with stable and high-quality VR 360° video stream services.

[0073] In the reward mechanism in step 4, the environment gives rewards to the agent according to the QoE definition. In this embodiment, the QoE definition comprehensively considers the visual quality of the video , re-buffering time , quality change , viewport smoothness and immersion , a total of 5 key factors. The specific calculation performed in this embodiment is: when the agent selects a bitrate and causes a change in the state of the video stream, the environment calculates the reward strictly according to the QoE definition based on the new state . The specific definition of QoE is:

[0074] ;

[0075] where is the viewport smoothness penalty; is the immersion of the 360° video; the values are respectively = 1, = 3.

[0076] 1) is the viewport smoothness:

[0077] ;

[0078] Here is the total number of frames of the video; , is the actual viewport center, is the predicted viewport center.

[0079] 2) represents mapping the bitrate of the i-th block of the video to the visually perceived quality by the user. is the sum of the visual quality of all video blocks. In this embodiment, the higher the bitrate, the higher the usually visual quality and the greater the positive contribution to QoE.

[0080] 3) is the pause time caused by the bitrate when downloading the i-th block. In this embodiment, represents the negative impact of the re-buffering time on QoE. The longer the re-buffering time, the lower the QoE. is the re-buffering penalty coefficient, used to adjust the weight of the impact of the re-buffering time on QoE. The value in this embodiment is = 4.3.

[0081] 4) is the quality change between two adjacent video blocks i and i + 1. In this embodiment, It reflects the video quality fluctuations. Frequent quality changes will reduce the QoE.

[0082] If the bitrate selected by the agent results in a significant improvement in the visual quality of the video, along with a short rebuffering time and smooth quality changes, the environment will give a high reward to encourage the agent to continue exploring similar bitrate selection strategies. Conversely, if the bitrate selection leads to frequent quality fluctuations, long rebuffering times, or poor viewport smoothness in the video, the environment will give a low reward or even a penalty, prompting the agent to adjust its strategy. The core goal of each agent is to maximize its expected cumulative discounted reward , where is the set discount factor, which is used to balance the importance of current rewards and future rewards. In this way, when making decisions, the agent will not only focus on the immediate effects brought by the current bitrate selection but also consider the long-term impact on the future video stream transmission quality, thus making a more intelligent and reasonable bitrate selection decision to maximize the overall QoE.

[0083] Initialize the actor and critic network parameters of the agent, and set hyperparameters such as the size of the experience pool, the number of training rounds, and the learning rate. The specific parameter configurations are as follows:

[0084] Figure 2 This is the neural network diagram of a VR 360° video stream bitrate adaptation method based on multi-agent reinforcement learning in the present invention. The number of input state neurons is 9, the number of hidden layers is 2, the number of neurons in each layer is 128, the ReLU activation function is applied, the batch size is 1024, the learning rate of the critic network is 0.01, the learning rate of the actor network is 0.01, the number of steps for policy update is 100, the discount factor is 0.95, the target network update coefficient is 0.005, and the output layer uses the Sigmoid activation function and scales the output to an appropriate range.

[0085] In the simulation environment, use multiple network trace datasets such as FCC18, HSR, Ghent, and Lab for training to ensure that the agent can learn effective bitrate selection strategies under different network conditions.

[0086] 2) Centralized training:

[0087] During the training process, the critic network of each agent will input the joint information of the states and actions of all agents, where and . Through this comprehensive information sharing, the agent can deeply learn the impact of the behaviors of other agents on the overall QoE, and thus better cooperate and compete. The agent actively collects each state transition information and stored in the experience pool. When the size of the experience pool exceeds the set threshold, each agent randomly samples a mini-batch of data from the experience pool. Then, the critic network is updated by minimizing the mean squared error loss, and its loss function is:

[0088] ;

[0089] where represents the Q-value output by the target critic network, with the input being the next state and the action .

[0090] Meanwhile, the actor network is updated through the sampled policy gradient, and its gradient calculation formula is:

[0091] ;

[0092] This training mechanism enables the agent to optimize its own policy in continuous iterations, gradually finding the optimal bitrate selection strategy to adapt to the complex and changing network environment and user requirements.

[0093] 3) Decentralized execution:

[0094] In the actual interaction process of the environment of this embodiment, each actor network makes decisions based on the state information observed by the local agent itself. This decentralized execution method endows the agent with the ability to quickly respond to local state changes. The specific decision-making in this embodiment is: when the local network condition suddenly changes, the agent adjusts the bitrate selection according to the state information it observes, without waiting for feedback from other agents or performing complex global coordination. At the same time, the global cooperation and competition strategies learned during the training process can ensure the overall performance optimization.

[0095] 4) Use multi-step rewards to calculate the TD error of the critic. Change the single-step reward to . Meanwhile, the learner samples state transition information from the buffer according to the priority. These strategies enable the agent to learn from future experiences, accelerate the learning process, and improve the adaptability to complex environments. The specific decision-making in this embodiment is: when the bitrate selection made by the agent at a certain moment may not immediately produce obvious effects at present, but has a positive impact on video quality and network stability in the subsequent several video blocks, the multi-step reward mechanism can accurately capture this situation and give the agent corresponding rewards, guiding the agent to learn this forward-looking strategy. And the experience replay priority ensures that the agent pays more attention to those state transition information that is more valuable to the learning process, further improving the learning efficiency.

[0096] Obtain historical video block transmission data and adopt predicted bandwidth, construct a network delay feature vector to evaluate network stability, establish a bitrate switching impact model based on the predicted bandwidth and buffer state, perform threshold control through the buffer state transition equation, dynamically adjust in combination with the video stream bandwidth share, and select the optimal bitrate based on quality score, switching impact, and network stability;

[0097] Construct a feature vector containing buffer state and video quality state based on the network delay feature vector, input the feature vector into the Actor network with a multi-layer perceptron structure, output the bitrate selection probability distribution through the ReLU activation function and the Softmax function, construct a comprehensive reward function based on video quality, buffer state, and switching smoothness, and determine the bitrate level by combining experience pool batch sampling;

[0098] Calculate the structural similarity of the video block to obtain the quality mapping value, calculate the rebuffering index based on the download time and buffer state, evaluate the quality change by combining the quality difference and the time decay weight, collect the viewport position information to calculate the smoothness and prediction error, determine the video immersion degree based on the field of view coverage rate and spatial quality distribution, construct a comprehensive reward function from the metrics and normalize it, calculate the cumulative reward based on the discount factor and optimize the weight coefficient;

[0099] Construct a multi-agent joint state vector to obtain local observation information, optimize the mean square error loss of the Critic network using the temporal difference target, update the Actor network by combining policy gradient and entropy regularization, adopt the mechanism of homogeneous agent parameter sharing and heterogeneous agent independent update, and integrate the group observation information through the attention weight;

[0100] Continuously store the state vector interaction data in the experience pool, extract network features using a multi-layer perceptron, construct an adaptive reward function based on the feature vector, update the network parameters using the policy gradient method and the Adam optimizer, adjust the reward weight according to the difference in perceived quality, and continuously optimize the bitrate selection strategy by sampling from the experience pool.

[0101] In an optional implementation manner, obtaining historical video block transmission data and adopting predicted bandwidth, constructing a network delay feature vector to evaluate network stability, establishing a bitrate switching impact model based on the predicted bandwidth and buffer state, performing threshold control through the buffer state transition equation, dynamically adjusting in combination with the video stream bandwidth share, and selecting the optimal bitrate based on quality score, switching impact, and network stability includes:

[0102] Obtain the throughput data and download time data of the past k video chunks during video transmission, where the throughput data is the ratio of the size of each video chunk to the corresponding download time; perform weighted calculation on the throughput data using an exponentially decaying weight to obtain a weighted predicted bandwidth, where the degree of decay of the exponentially decaying weight over time is determined by a decay coefficient, and the decay coefficient is used to adjust the influence degree of historical data on bandwidth prediction;

[0103] Construct a network delay feature vector based on the download time data, calculate the variance value of the network delay feature vector as a network stability quantization index, and the larger the variance value, the worse the network stability; based on the weighted predicted bandwidth, combined with the current buffer size, the expected size information of the next video chunk corresponding to different bitrates, and the current bitrate, establish a bitrate switching impact evaluation model, and the bitrate switching impact evaluation model is used to evaluate the impact degree of bitrate switching on system performance;

[0104] Construct a buffer state transition equation according to the current buffer size, weighted predicted bandwidth, and video chunk duration, set a buffer safety threshold, and when the buffer size predicted by the buffer state transition equation at the next moment is less than the buffer safety threshold, select a conservative bitrate based on the weighted predicted bandwidth;

[0105] When detecting the existence of multiple video streams, obtain the target bitrate of each video stream, calculate the bandwidth share factor of each video stream, where the bandwidth share factor is the ratio of the target bitrate of the corresponding video stream to the sum of the target bitrates of all video streams, and adjust the weighted predicted bandwidth based on the bandwidth share factor to obtain an adjusted bandwidth;

[0106] Construct an optimization objective function with the video quality score, the evaluation result of the bitrate switching impact evaluation model, and the variance value, adjust the weights of each evaluation index by setting a balance factor, and based on the adjusted bandwidth, select the bitrate that makes the optimization objective function obtain the optimal value from the preset bitrate set as the final bitrate decision result.

[0107] During video transmission, first, it is necessary to collect the throughput data and download time data of the past k video chunks. The throughput data is calculated by the ratio of the size of each video chunk to the corresponding download time. The download time refers to the time required from the start of the request until the video chunk is completely downloaded. By recording these data, it can provide a basis for subsequent bandwidth prediction.

[0108] After obtaining the historical data, perform weighted calculation on the throughput data using an exponentially decaying weight. The design of the exponentially decaying weight is to make the newer data have a greater impact on bandwidth prediction, while the impact of the older data gradually decreases. The selection of the decay coefficient will directly affect the influence degree of historical data on bandwidth prediction. In this way, a weighted predicted bandwidth value can be obtained to reflect the actual bandwidth situation of the current network.

[0109] Based on the download time data, construct a network latency feature vector. This feature vector contains information in multiple dimensions, such as average download time, maximum download time, and minimum download time, etc. By calculating the variance values of these features, the stability of the network can be quantified. The larger the variance value, the worse the stability of the network, and vice versa.

[0110] After obtaining the weighted predicted bandwidth and the network latency feature vector, combine the current buffer size, the expected size information of the next video block corresponding to different bitrates, and the current bitrate to establish a bitrate switching impact evaluation model. This model is used to evaluate the impact degree of bitrate switching on the system performance and help make decisions on whether bitrate adjustment is needed.

[0111] According to the current buffer size, the weighted predicted bandwidth, and the video block duration, construct a buffer state transition equation. This equation is used to predict the buffer size at the next moment and set a buffer safety threshold. When the predicted buffer size at the next moment is less than the safety threshold, the system will select a conservative bitrate based on the weighted predicted bandwidth to ensure the smoothness of video playback.

[0112] When detecting the existence of multiple video streams, obtain the target bitrate of each video stream and calculate the bandwidth share factor of each video stream. The bandwidth share factor refers to the proportion of the target bitrate of the corresponding video stream in the total target bitrates of all video streams. By calculating the bandwidth share factor, the weighted predicted bandwidth can be adjusted to obtain an adjusted bandwidth value.

[0113] Combine the video quality score, the evaluation result of the bitrate switching impact evaluation model, and the variance value to construct an optimization objective function. By setting a balance factor, the weights of each evaluation index can be adjusted. Based on the adjusted bandwidth, select the bitrate that makes the optimization objective function obtain the optimal value from the preset bitrate set as the final bitrate decision result.

[0114] The solution of this application can:

[0115] Improve the stability of video transmission: By analyzing the network latency feature vector, the stability of the network can be effectively evaluated, so as to take corresponding measures in an unstable network environment to ensure the smoothness of video playback. Optimize the bitrate switching strategy: The established bitrate switching impact evaluation model can accurately evaluate the impact of bitrate switching on the system performance and help the system select the optimal bitrate under different network conditions to enhance the user experience. Dynamically adjust the bandwidth allocation: By calculating the bandwidth share factor, the system can dynamically adjust the bandwidth allocation according to the needs of multiple video streams to ensure that each video stream can obtain sufficient bandwidth and avoid playback jams caused by insufficient bandwidth.

[0116] In an alternative embodiment, a state feature vector including buffer status and video quality status is constructed based on the network latency feature vector. The feature vector is input into the Actor network with a multi-layer perceptron structure, and the bitrate selection probability distribution is output through the ReLU activation function and the Softmax function. A comprehensive reward function is constructed based on video quality, buffer status, and switching smoothness. Determining the bitrate level by batch sampling in combination with the experience pool includes:

[0117] Construct a state feature vector. The state feature vector includes buffer status and video quality status. The buffer status includes the current buffer size and the buffer safety threshold. The video quality status includes the current bitrate level and the bitrate switching amplitude. Input the state feature vector into the Actor network. The Actor network adopts a multi-layer perceptron structure, including a first hidden layer and a second hidden layer. Both the first hidden layer and the second hidden layer adopt the ReLU activation function;

[0118] Construct a comprehensive reward function based on the current bitrate level, the current buffer size, and the bitrate switching amplitude. The comprehensive reward function includes video quality reward, buffer status reward, and switching smoothness reward. Among them, the video quality reward is the weighted value of the ratio of the current bitrate level to the maximum bitrate level. The buffer status reward is the weighted value of the calculation result of the ratio of the current buffer size to the buffer safety threshold. The switching smoothness reward is the normalized weighted value of the ratio of the bitrate switching amplitude to the maximum bitrate level. Calculate the policy gradient according to the comprehensive reward function, and use the policy gradient to iteratively optimize the network parameters of the Actor network by the gradient descent method to obtain the optimized network parameters;

[0119] Construct a transition sample from the state feature vector, the action determined based on the bitrate selection probability distribution, the calculation result of the comprehensive reward function, and the next state feature vector and store it in the experience pool. Perform batch sampling from the experience pool based on a preset sampling size, and update the optimized network parameters in combination with the exploration strategy. Input the updated network parameters into the Actor network, and determine the final bitrate level based on the bitrate selection probability distribution output by the Actor network.

[0120] First, construct a state feature vector. The state feature vector consists of buffer status and video quality status. The buffer status includes the current buffer size and the buffer safety threshold, and the video quality status includes the current bitrate level and the bitrate switching amplitude. By real-time monitoring of network latency and video playback, collect this state information and form a state feature vector.

[0121] Next, the state feature vector is input into the Actor network. The Actor network adopts a multi-layer perceptron structure, including a first hidden layer and a second hidden layer. Each layer uses the ReLU activation function to enhance the network's non-linear expression ability. Through forward propagation, the Actor network converts the state feature vector into a bitrate selection probability distribution.

[0122] Then, a comprehensive reward function is constructed based on the current bitrate level, the current buffer size, and the bitrate switching amplitude. The comprehensive reward function consists of a video quality reward, a buffer state reward, and a switching smoothness reward. The video quality reward is weighted and calculated by the ratio of the current bitrate level to the maximum bitrate level. The buffer state reward is weighted by the ratio of the current buffer size to the buffer safety threshold. The switching smoothness reward is normalized and weighted by the ratio of the bitrate switching amplitude to the maximum bitrate level. The design of the comprehensive reward function aims to balance video quality, buffer state, and switching smoothness to achieve the best user experience.

[0123] After obtaining the comprehensive reward function, the policy gradient is calculated. Using the policy gradient method, the network parameters of the Actor network are iteratively optimized through gradient descent to obtain the optimized network parameters. This process ensures that the network can continuously adjust its policy according to environmental changes to adapt to different network conditions and user needs.

[0124] Next, the state feature vector, the action determined based on the bitrate selection probability distribution, the calculation result of the comprehensive reward function, and the next state feature vector are constructed into a transition sample and stored in the experience pool. The experience pool is used to store historical interaction data for subsequent batch sampling. Based on a preset sampling size, batch sampling is performed from the experience pool, and the optimized network parameters are updated in combination with the exploration strategy. This step ensures that the network can learn and adapt in a diverse environment.

[0125] Finally, the updated network parameters are input into the Actor network, and the final bitrate level is determined based on the bitrate selection probability distribution output by the Actor network. Through this series of steps, the system can dynamically adjust the bitrate of the video stream to achieve a smooth playback experience.

[0126] The solution of this application can:

[0127] Improve the playback quality of the video stream, ensuring that users can obtain a good viewing experience under different network conditions. Dynamically adjust bitrate selection, reduce buffering phenomena, and enhance the continuity and smoothness of the video stream. Through the design of the comprehensive reward function, optimize the balance between video quality, buffer state, and switching smoothness, and enhance user satisfaction.

[0128] In an alternative embodiment, the structural similarity of a video block is calculated to obtain a quality mapping value, a rebuffering metric is calculated based on the download time and buffer status, the quality change is evaluated by combining the quality difference and the time decay weight, the viewport position information is collected to calculate the smoothness and prediction error, the video immersion degree is determined according to the field of view angle coverage rate and the spatial quality distribution, and the metrics are used to construct a comprehensive reward function and normalize it. Based on the discount factor, the cumulative reward is calculated and the weight coefficients are optimized, including:

[0129] Obtain the bitrate of the video block and the video block image data, calculate the structural similarity based on the mean, standard deviation and covariance of the video block image data, and map the video block bitrate and the structural similarity to obtain a quality mapping value;

[0130] Monitor the download process of the video block, obtain the download time of each video block and the remaining duration of the corresponding buffer. When the download time is greater than the remaining buffer duration, the rebuffering duration is cumulatively calculated, and the rebuffering frequency is calculated based on the number of rebuffering occurrences and the total playback duration; Obtain the quality mapping values of adjacent video blocks to calculate the quality difference, construct a time decay weight sequence, and use the weighted sum of the quality difference and the time decay weight sequence as the quality change evaluation index;

[0131] Real-time collect the user's viewport position information, extract the azimuth angle and elevation angle in the user's viewport position information, calculate the Euclidean distance between the azimuth angles and elevation angles at adjacent moments to obtain the viewport smoothness, and calculate the viewport prediction error based on the deviation distance between the viewport position prediction value and the actual value;

[0132] Calculate the ratio of the area of the field of view region in the video frame to the total area of the field of view to obtain the field of view angle coverage rate. At the same time, divide the field of view space and assign spatial weights to each region, and use the weighted sum of the quality mapping values of each region and the corresponding spatial weights as the video immersion degree;

[0133] Linearly combine the quality mapping value, the rebuffering duration, the quality change evaluation index, the viewport smoothness, the viewport prediction error and the video immersion degree to construct a comprehensive reward function, and perform min-max normalization processing on the comprehensive reward function to obtain a normalized reward value;

[0134] Calculate the cumulative discounted reward based on a preset discount factor. The cumulative discounted reward is the weighted sum of the normalized reward values at each moment multiplied by the exponential decay coefficient of the corresponding time step; Use the gradient ascent method to iteratively optimize the weight coefficients of each evaluation index in the comprehensive reward function. Under the constraint that the sum of the weight coefficients is 1 and each coefficient is non-negative, determine the optimal weight coefficient combination by maximizing the expected value of the cumulative discounted reward.

[0135] First, obtain the bitrate and image data of the video block. By analyzing the video block image data, calculate its mean, standard deviation, and covariance to evaluate the structural similarity of the video block. Map the calculation result of the structural similarity to the bitrate of the video block to obtain the corresponding quality mapping value. This process provides the basic data for subsequent quality evaluation.

[0136] Next, monitor the download process of the video block, record the download time and the remaining buffer duration of each video block. When the download time exceeds the remaining buffer duration, start accumulating and calculating the rebuffering duration. By counting the number of rebuffering occurrences and the total playback duration, the rebuffering frequency can be calculated. This metric reflects the smoothness of video playback and has a direct impact on the user experience.

[0137] On this basis, obtain the quality mapping values of adjacent video blocks, calculate their quality differences. Construct a time decay weight sequence, and use it to perform a weighted sum of the quality differences and the time decay weight sequence to form a quality change evaluation metric. This metric can reflect the changing trend of video quality over time and provide a basis for subsequent quality optimization.

[0138] Collect the user's viewport position information in real time, extract the azimuth and elevation angles from it. By calculating the Euclidean distance between the azimuth and elevation angles at adjacent moments, the viewport smoothness is obtained. The viewport prediction error is calculated by comparing the deviation distance between the predicted value and the actual value of the viewport position. These metrics together reflect the immersion and experience smoothness when the user watches the video.

[0139] Furthermore, calculate the ratio of the area of the field of view region in the video frame to the total area of the field of view to obtain the field of view angle coverage rate. At the same time, divide the field of view space and assign spatial weights to each region. Perform a weighted sum of the quality mapping values of each region and the corresponding spatial weights to finally obtain the immersion degree of the video. This process ensures that the impact of the quality of different regions on the overall viewing experience is reasonably considered.

[0140] Finally, perform a linear combination of the quality mapping value, rebuffering duration, quality change evaluation metric, viewport smoothness, viewport prediction error, and video immersion degree to construct a comprehensive reward function. Perform maximum-minimum normalization on the comprehensive reward function to obtain the normalized reward value. Based on the preset discount factor, calculate the cumulative discounted reward. The cumulative discounted reward is the weighted sum of the normalized reward values at each moment and the exponential decay coefficients of the corresponding time steps. Iteratively optimize the weight coefficients of each evaluation metric in the comprehensive reward function through the gradient ascent method to ensure that the sum of the weight coefficients is 1 and each coefficient is non-negative, so as to determine the optimal combination of weight coefficients.

[0141] The solution of this application can:

[0142] Improve the smoothness of video playback, reduce the re-buffering phenomenon during user viewing, and thus enhance the user experience. By real-time monitoring and evaluating video quality changes, the video playback strategy can be adjusted in a timely manner to ensure that users can obtain the best viewing experience in different network environments. Combining user viewport information and spatial quality distribution, optimize the presentation method of video content to enhance the user's immersion and engagement.

[0143] In an alternative implementation, a multi-agent joint state vector is constructed to obtain local observation information, the mean squared error loss of the Critic network is optimized using the temporal difference target, the Actor network is updated by combining policy gradient and entropy regularization, and a homogeneous agent parameter sharing and heterogeneous agent independent update mechanism is adopted. The group observation information is integrated through attention weights, and the network parameters are synchronously updated based on the global experience pool data to achieve collaborative learning, including:

[0144] Construct a multi-agent joint state vector, which contains the state information of all agents, and map the multi-agent joint state vector to the local observation information of each agent through an observation function;

[0145] Based on the local observation information of each agent, an action is independently generated using the corresponding policy function, and the multi-agent joint state vector, action, the reward value obtained at the current moment, and the information of the transferred next state are recorded as a state transition tuple and stored in the experience pool;

[0146] When the amount of data in the experience pool reaches the preset storage threshold, batch data is randomly sampled from the experience pool, and the temporal difference target is calculated based on the reward value in the batch data and the evaluation value of the target network for the next state; a mean squared error loss function is constructed using the temporal difference target and the predicted value of the Critic network for the current state-action combination, and the parameters of the Critic network are updated based on the gradient direction of the mean squared error loss function;

[0147] Multiply the action value function of the evaluation state of the Critic network by the logarithmic probability of the policy function to calculate the policy gradient, introduce a policy entropy regularization term in the policy gradient, and update the parameters of the Actor network based on the final gradient direction; identify the homogeneous agents in the system, adopt a parameter sharing mechanism to uniformly update the network parameters for homogeneous agents, while keeping the parameters of heterogeneous agents independently updated;

[0148] Calculate the attention weights according to the similarity of the observation information between heterogeneous agents, use the attention weights to weight and integrate the observation information of other agents to obtain an enhanced observation representation containing group information; merge the experience data of all agents into the global experience pool, calculate the overall gradient based on the data in the global experience pool, and synchronously update the global network parameters through the overall gradient to achieve collaborative learning of multi-agents.

[0149] First, construct a multi-agent joint state vector, which contains the state information of all agents. By designing an observation function, map the multi-agent joint state vector to the local observation information of each agent. This observation function can be adjusted based on the specific requirements of the agents and environmental characteristics to ensure that each agent can obtain local information relevant to its task.

[0150] Next, based on the local observation information of each agent, independently generate actions using the corresponding policy function. The policy function of each agent can be a deep learning-based model that can output corresponding actions according to the input local observation information. The generated actions, the reward value obtained at the current moment, and the transferred next state information will be recorded as a state transition tuple and stored in the experience pool. The design of the experience pool should consider storage efficiency and data access speed for subsequent training processes.

[0151] When the amount of data in the experience pool reaches the preset storage threshold, randomly sample a batch of data from the experience pool. Based on the reward value in this batch of data and the evaluation value of the target network for the next state, calculate the temporal difference target. The calculation process of the temporal difference target includes evaluating the current state and action combination, combined with the reward value, to form an expectation for the future state.

[0152] Use the temporal difference target and the predicted value of the Critic network for the current state-action combination to construct a mean squared error loss function. The optimization process of this loss function will update the parameters of the Critic network through the backpropagation algorithm to improve its evaluation accuracy of the state-action combination.

[0153] After updating the Critic network, multiply the state-action value function evaluated by the Critic network by the log probability of the policy function to calculate the policy gradient. To enhance the exploration of the policy, introduce a policy entropy regularization term in the policy gradient. Finally, update the parameters of the Actor network based on the calculated gradient direction to improve the decision-making ability of the agent.

[0154] After identifying isomorphic agents in the system, adopt a parameter sharing mechanism to uniformly update the isomorphic agents to ensure the consistency of their network parameters. At the same time, keep the parameters of heterogeneous agents updated independently to adapt to the specific requirements and environments of different agents.

[0155] According to the similarity of the observation information between heterogeneous agents, calculate the attention weights. Use these weights to weighted integrate the observation information of other agents to obtain an enhanced observation representation containing group information. This process can be achieved by designing a suitable attention mechanism to ensure the effective integration of information.

[0156] Finally, the experience data of all agents is merged into the global experience pool. Based on the data in the global experience pool, the overall gradient is calculated, and the global network parameters are synchronously updated through the overall gradient. This process realizes the collaborative learning of multiple agents, ensuring that each agent can share experiences and jointly improve performance.

[0157] The solution of this application can:

[0158] Improve the collaborative working ability of the multi-agent system. By sharing experiences and integrating information, the decision-making ability and adaptability of the agents are enhanced. By introducing the temporal difference objective and the policy entropy regularization term, the learning process of the agents is optimized, enabling them to explore and utilize more effectively in complex environments. The parameter sharing and independent update mechanism ensure the flexibility and efficiency of homogeneous and heterogeneous agents, improving the overall performance and stability of the system.

[0159] In an optional implementation manner, the state vector interaction data is continuously stored in the experience pool, the network features are extracted using a multi-layer perceptron, an adaptive reward function is constructed based on the feature vector, the network parameters are updated using the policy gradient method and the Adam optimizer, and the reward weight is adjusted according to the difference in experience quality. Sampling from the experience pool to continuously optimize the bitrate selection strategy includes:

[0160] Construct an experience pool to store state vector interaction data, where the state vector interaction data includes state vectors and bitrate decisions; extract network features using a multi-layer perceptron based on the state vector, and perform a non-linear transformation on the state vector through the ReLU activation function to obtain a feature vector; construct an adaptive reward function according to the feature vector, and the adaptive reward function is obtained by weighted calculation of the bitrate quality score, the buffer jitter degree, and the change amplitude of adjacent bitrate decisions, and the weighted calculation uses a dynamically adjusted weight coefficient;

[0161] Construct a bitrate selection policy network based on the feature vector, calculate the parameter gradient of the bitrate selection policy network using the policy gradient method, and the parameter gradient is obtained through the product expectation of the logarithm of the policy function and the state-action value function; introduce the Adam optimizer to update the parameters of the bitrate selection policy network, and the Adam optimizer calculates the adaptive learning rate based on the first-order momentum estimate and the second-order momentum estimate, and uses the adaptive learning rate to update the policy network parameters;

[0162] Calculate the difference between the target experience quality and the current experience quality, dynamically adjust the weight coefficient in the adaptive reward function based on the difference, and realize the adaptive update of the weight coefficient through a preset step size; during the video stream transmission process, continuously randomly sample batch data from the experience pool, and use the batch data to continuously optimize the bitrate selection policy.

[0163] In this technical solution, an experience pool is first constructed to store state vector interaction data. The state vector interaction data includes the current state vector and the corresponding bitrate decision. The design of the experience pool allows the system to collect and store a large amount of interaction data at different time points for subsequent learning and optimization.

[0164] Next, a multi-layer perceptron is used to process the state vector to extract network features. The multi-layer perceptron performs a non-linear transformation on the input state vector through a series of hierarchical structures and finally outputs a feature vector. The activation function in this process uses ReLU (Rectified Linear Unit), which can effectively introduce non-linearity and thus improve the expressive power of the model.

[0165] Based on the feature vector, an adaptive reward function is constructed. This reward function forms a comprehensive reward mechanism by performing weighted calculations on the bitrate quality score, buffer jitter degree, and the change amplitude of adjacent bitrate decisions. The weight coefficients are dynamically adjusted during this process to better reflect the current network state and user experience.

[0166] Subsequently, a bitrate selection policy network is constructed based on the feature vector. The parameters of this network are optimized by the policy gradient method. Specifically, the parameter gradient is obtained through the expected value of the product of the logarithm of the policy function and the state-action value function. This process ensures that the policy network can select the optimal bitrate according to the current state and action.

[0167] In terms of parameter update, the Adam optimizer is adopted. This optimizer combines first-order momentum estimation and second-order momentum estimation, calculates an adaptive learning rate, and uses this learning rate to update the parameters of the policy network. In this way, the network can adaptively adjust the learning rate during the training process, thereby improving the convergence speed and stability.

[0168] During the video stream transmission process, the system continuously randomly samples batch data from the experience pool. After each sampling, the difference between the target quality of experience and the current quality of experience is calculated. According to this difference, the weight coefficients in the adaptive reward function are dynamically adjusted. The update of the weight coefficients is achieved through a preset step size, ensuring the flexibility and adaptability of the reward mechanism.

[0169] Finally, by continuously optimizing the bitrate selection policy, the system can achieve a higher user experience and lower latency in video stream transmission. Each optimization is based on the latest experience pool data, ensuring the real-time nature and effectiveness of the policy.

[0170] The solution of this application can:

[0171] It improves the stability and smoothness of video stream transmission and reduces the stuttering phenomenon during users' viewing. Through the adaptive reward mechanism, the system can dynamically adjust the bitrate selection strategy according to the real-time network conditions, thus optimizing the user experience. The design of the experience pool enables the system to continuously learn and improve, enhancing the overall intelligent level and adaptive ability.

[0172] In the second aspect of the embodiments of the present invention,

[0173] A kind of electronic device is provided, including:

[0174] A processor;

[0175] A memory for storing instructions executable by the processor;

[0176] Wherein, the processor is configured to call the instructions stored in the memory to execute the foregoing method.

[0177] In the third aspect of the embodiments of the present invention,

[0178] A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the foregoing method is implemented.

[0179] The present invention can be a method, a device, a system and / or a computer program product. The computer program product can include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are uploaded.

[0180] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An adaptive bitrate control method for VR videos based on multi-agent reinforcement learning, characterized in that Including: Obtain historical video block transmission data and use predicted bandwidth to construct a network latency feature vector to evaluate network stability. Establish a bitrate switching impact model based on the predicted bandwidth and buffer status, perform threshold control through the buffer status transfer equation, dynamically adjust in combination with the video stream bandwidth share, and select the optimal bitrate based on quality score, switching impact, and network stability; Construct a feature vector including buffer status and video quality status based on the network latency feature vector, input the feature vector into the Actor network with a multi-layer perceptron structure, output the bitrate selection probability distribution through the ReLU activation function and Softmax function, construct a comprehensive reward function based on video quality, buffer status, and switching smoothness, and determine the bitrate level by combining experience pool batch sampling; Calculate the structural similarity of the video block to obtain the quality mapping value, calculate the rebuffering index based on the download time and buffer status, evaluate the quality change in combination with quality difference and time decay weight, collect viewport position information to calculate smoothness and prediction error, determine the video immersion degree based on the field of view angle coverage rate and spatial quality distribution, construct a comprehensive reward function from the metrics and normalize it, calculate the cumulative reward based on the discount factor and optimize the weight coefficient; Construct a multi-agent joint state vector to obtain local observation information, optimize the mean square error loss of the Critic network using the temporal difference target, update the Actor network in combination with policy gradient and entropy regularization, adopt the mechanism of homogeneous agent parameter sharing and heterogeneous agent independent update, and integrate group observation information through attention weights; Continuously store the state vector interaction data in the experience pool, extract network features using a multi-layer perceptron, construct an adaptive reward function based on the feature vector, update the network parameters using the policy gradient method and Adam optimizer, adjust the reward weight according to the difference in experience quality, and continuously optimize the bitrate selection strategy by sampling from the experience pool; Obtain the throughput data and download time data of the past k video blocks during video transmission, where the throughput data is the ratio of the size of each video block to the corresponding download time; Perform weighted calculation on the throughput data using an exponentially decaying weight to obtain the weighted predicted bandwidth, where the degree of decay of the exponentially decaying weight over time is determined by the decay coefficient, and the decay coefficient is used to adjust the influence degree of historical data on bandwidth prediction; Construct a network latency feature vector based on the download time data, calculate the variance value of the network latency feature vector as a quantitative index of network stability, and the larger the variance value, the worse the network stability; Based on the weighted predicted bandwidth, in combination with the current buffer size, the expected size information of the next video block corresponding to different bitrates, and the current bitrate, establish a bitrate switching impact evaluation model, which is used to evaluate the impact degree of bitrate switching on system performance; Construct a buffer status transfer equation according to the current buffer size, weighted predicted bandwidth, and video block duration, set a buffer safety threshold, and when the predicted buffer size at the next moment by the buffer status transfer equation is less than the buffer safety threshold, select a conservative bitrate based on the weighted predicted bandwidth; When multiple video streams are detected, obtain the target bitrate of each video stream, calculate the bandwidth share factor for each video stream, where the bandwidth share factor is the ratio of the target bitrate of the corresponding video stream to the sum of the target bitrates of all video streams, and adjust the weighted predicted bandwidth based on the bandwidth share factor to obtain the adjusted bandwidth; Construct an optimization objective function with the video quality score, the evaluation result of the bitrate switching impact evaluation model, and the variance value. Adjust the weights of each evaluation index by setting a balance factor. Based on the adjusted bandwidth, select the bitrate that makes the optimization objective function reach the optimal value from the preset bitrate set as the final bitrate decision result.

2. The method according to claim 1, wherein Construct a state feature vector including the buffer state and the video quality state based on the network delay feature vector. Input the feature vector into the Actor network with a multi-layer perceptron structure. Output the bitrate selection probability distribution through the ReLU activation function and the Softmax function. Construct a comprehensive reward function based on video quality, buffer state, and switching smoothness. Determine the bitrate level by combining batch sampling from the experience pool, including: Construct a state feature vector, which includes the buffer state and the video quality state. The buffer state includes the current buffer size and the buffer safety threshold. The video quality state includes the current bitrate level and the bitrate switching amplitude. Input the state feature vector into the Actor network, and the Actor network adopts a multi-layer perceptron structure, including a first hidden layer and a second hidden layer. Both the first hidden layer and the second hidden layer adopt the ReLU activation function; Construct a comprehensive reward function based on the current bitrate level, the current buffer size, and the bitrate switching amplitude. The comprehensive reward function includes video quality reward, buffer state reward, and switching smoothness reward. Among them, the video quality reward is the weighted value of the ratio of the current bitrate level to the maximum bitrate level. The buffer state reward is the weighted value of the calculation result of the ratio of the current buffer size to the buffer safety threshold. The switching smoothness reward is the normalized weighted value of the ratio of the bitrate switching amplitude to the maximum bitrate level. Calculate the policy gradient according to the comprehensive reward function, and use the policy gradient to iteratively optimize the network parameters of the Actor network through the gradient descent method to obtain the optimized network parameters; Construct a transition sample by combining the state feature vector, the action determined based on the bitrate selection probability distribution, the calculation result of the comprehensive reward function, and the next state feature vector and store it in the experience pool. Perform batch sampling from the experience pool based on the preset sampling size, and update the optimized network parameters in combination with the exploration strategy. Input the updated network parameters into the Actor network, and determine the final bitrate level based on the bitrate selection probability distribution output by the Actor network.

3. The method according to claim 1, characterized in that, Calculate the quality mapping value by calculating the structural similarity of video blocks. Calculate the rebuffering index based on the download time and the buffer state. Evaluate the quality change by combining the quality difference and the time decay weight. Collect the viewport position information to calculate the smoothness and prediction error. Determine the video immersion degree based on the field of view coverage rate and the spatial quality distribution. Construct a comprehensive reward function with the indicators and normalize it. Calculate the cumulative reward based on the discount factor and optimize the weight coefficients, including: Obtain the video block bitrate and video block image data, calculate the structural similarity based on the mean, standard deviation, and covariance of the video block image data, and map the video block bitrate and the structural similarity to obtain a quality mapping value; Monitor the video block download process, obtain the download time of each video block and the corresponding remaining buffer duration. When the download time is greater than the remaining buffer duration, cumulatively calculate the rebuffering duration, and calculate the rebuffering frequency based on the number of rebuffering occurrences and the total playback duration; Obtain the quality differences by calculating the quality mapping values of adjacent video blocks, construct a time decay weight sequence, and use the weighted sum of the quality differences and the time decay weight sequence as the quality change evaluation index; Real-time collect the user's viewport position information, extract the azimuth and pitch angles in the user's viewport position information, calculate the Euclidean distance between the azimuth and pitch angles at adjacent moments to obtain the viewport smoothness, and calculate the viewport prediction error based on the deviation distance between the viewport position prediction value and the actual value; Calculate the ratio of the field-of-view area in the video frame to the total field-of-view area to obtain the field-of-view angle coverage rate. At the same time, divide the field-of-view space and assign spatial weights to each region, and use the weighted sum of the quality mapping values of each region and the corresponding spatial weights as the video immersion degree; Linearly combine the quality mapping value, rebuffering duration, quality change evaluation index, viewport smoothness, viewport prediction error, and video immersion degree to construct a comprehensive reward function, and perform min-max normalization processing on the comprehensive reward function to obtain a normalized reward value; Calculate the cumulative discounted reward based on a preset discount factor. The cumulative discounted reward is the weighted sum of the normalized reward values at each moment multiplied by the exponential decay coefficient of the corresponding time step; Use the gradient ascent method to iteratively optimize the weight coefficients of each evaluation index in the comprehensive reward function. Under the constraint that the sum of the weight coefficients is 1 and each coefficient is non-negative, determine the optimal weight coefficient combination by maximizing the expected value of the cumulative discounted reward.

4. The method according to claim 1, characterized in that, Construct a multi-agent joint state vector to obtain local observation information, use the temporal difference target to optimize the mean square error loss of the Critic network, update the Actor network by combining policy gradient and entropy regularization, adopt the mechanism of homogeneous agent parameter sharing and heterogeneous agent independent update, integrate the group observation information through attention weights, and synchronously update the network parameters based on the global experience pool data to achieve collaborative learning, including: Construct a multi-agent joint state vector, which contains the state information of all agents, and map the multi-agent joint state vector to the local observation information of each agent through an observation function; Based on the local observation information of each agent, independently generate actions using the corresponding policy function, and record the multi-agent joint state vector, actions, the reward value obtained at the current moment, and the transferred next state information as a state transition tuple and store it in the experience pool; When the amount of data in the experience pool reaches the preset storage threshold, randomly sample batch data from the experience pool, and calculate the temporal difference target based on the reward value in the batch data and the evaluation value of the target network for the next state; use the temporal difference target and the predicted value of the critic network for the current state-action combination to construct a mean squared error loss function, and update the parameters of the critic network based on the gradient direction of the mean squared error loss function; Multiply the action value function of the evaluation state of the critic network by the logarithmic probability of the policy function to calculate the policy gradient, introduce a policy entropy regularization term in the policy gradient, and update the parameters of the actor network based on the final gradient direction; identify isomorphic agents in the system, and use a parameter sharing mechanism to uniformly update the network parameters for isomorphic agents, while keeping the parameters of heterogeneous agents updated independently; Calculate the attention weights according to the similarity of the observation information between heterogeneous agents, and use the attention weights to weighted integrate the observation information of other agents to obtain an enhanced observation representation containing group information; merge the experience data of all agents into the global experience pool, calculate the overall gradient based on the data in the global experience pool, and synchronously update the global network parameters through the overall gradient to achieve collaborative learning of multiple agents.

5. The method according to claim 1, characterized in that, Continuously store state vector interaction data in the experience pool, use a multi-layer perceptron to extract network features, construct an adaptive reward function based on the feature vectors, update the network parameters using the policy gradient method and the Adam optimizer, and adjust the reward weights according to the difference in experience quality. Sampling from the experience pool to continuously optimize the bitrate selection strategy includes: Construct an experience pool to store state vector interaction data, where the state vector interaction data includes state vectors and bitrate decisions; use a multi-layer perceptron to extract network features based on the state vectors, and perform a non-linear transformation on the state vectors through the ReLU activation function to obtain feature vectors; construct an adaptive reward function according to the feature vectors, and the adaptive reward function is obtained by weighted calculation of the bitrate quality score, buffer jitter degree, and the change amplitude of adjacent bitrate decisions. The weighted calculation uses dynamically adjusted weight coefficients; Construct a bitrate selection policy network based on the feature vectors, calculate the parameter gradient of the bitrate selection policy network using the policy gradient method, and the parameter gradient is obtained through the expected value of the product of the logarithmic policy function and the state-action value function; introduce the Adam optimizer to update the parameters of the bitrate selection policy network. The Adam optimizer calculates the adaptive learning rate based on the first-order momentum estimation and the second-order momentum estimation, and uses the adaptive learning rate to update the parameters of the policy network; Calculate the difference between the target experience quality and the current experience quality, dynamically adjust the weight coefficients in the adaptive reward function based on the difference, and achieve the adaptive update of the weight coefficients through a preset step size; during the video stream transmission process, continuously randomly sample batch data from the experience pool, and use the batch data to continuously optimize the bitrate selection strategy.

Citation Information

Patent Citations

  • Self-adaptive video stream transmission method and system based on server-free computing

    CN116962414A

  • Audio and video cooperation adaptive code rate control system and method based on reinforcement learning

    CN117939192A