A Bit Rate Adaptive Selection Method and Device
The code rate adaptive method uses a reinforcement learning model to optimize video block rate selection based on network and buffer states, enhancing user experience by aligning with user demands and reducing playback interruptions.
Patent Information
- Application Number
- CN202310085087.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-28
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-01-28
AI Technical Summary
The existing code rate adaptive method based on reinforcement learning cannot adjust the bit rate selection strategy of the video block according to changes in the state space, and cannot feedback user needs, resulting in interruption of video playback and reduced user experience.
By building a state space, using preset reinforcement learning model and neural network, combining network state, video content state and player state, the bit rate copy probability of the video block is calculated, and the model parameters are optimized through the cumulative value of external rewards and internal reward discounts, and the bit rate selection strategy of the video block is dynamically adjusted.
It improves the user's experience quality of video viewing, ensures that the smoothness and quality of video playback meet user needs, and reduces the negative impact of buffer status on user experience.
Smart Images

Figure CN116156228B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and particularly to a bitrate adaptive selection method and device. Background Art
[0002] With the rapid growth of the amount of video stream communication data based on HTTP (HyperText Transfer Protocol), users' demand for the perceived quality of experience of downloading videos is also increasing day by day. Metrics such as buffering time, average playback bitrate, and bitrate switching frequency have become key factors for measuring QoE (Quality of Experience). In a complex Internet video transmission ecosystem, the bitrate adaptive (ABR) algorithm deployed in the HyperText Transfer Protocol server or the client player is crucial for optimizing the user experience. The existing bitrate adaptive methods based on reinforcement learning cannot adjust the bitrate selection strategy of video chunks according to the changes in the state space, resulting in conflicts between task objectives, such as the conflict between high-bitrate video chunks and low network speed. At the same time, the existing bitrate adaptive methods cannot accurately feedback the user's needs, thus causing problems such as video playback interruption and unclear video at the user end, reducing the user experience. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a bitrate adaptive selection method and device, so as to solve the problem that the existing bitrate adaptive methods cannot adjust the bitrate selection strategy of video chunks according to the changes in the state space and cannot feedback the user's needs at the same time.
[0004] One aspect of the present invention provides a bitrate adaptive selection method, and the method includes the following steps:
[0005] The video provider obtains a video source, divides the video source into a plurality of consecutive video chunks according to a set duration, and compresses each video chunk into corresponding multiple bitrate copies according to multiple bitrates; in an initial state, the video provider sends a specified bitrate copy of the first video chunk according to the request of the video requester and transmits it to the buffer space;
[0006] During the transmission process, the network state, video content state, and player state during the transmission of the current video block are constructed into a state space; the network state includes the average throughput of each transmitted video block during the transmission process; the video content state includes the download time set of the current video block and a set number of previous video blocks, the file size set of all bitrate copies of the current video block, and the perceptual quality parameter of the previous video block; the player state includes the duration of all unplayed video blocks in the buffer, the bitrate of the current video block, and the number of remaining untransmitted video blocks in the video source; among them, the average throughput of each transmitted video block during the transmission process, the download time set of the current video block and a set number of previous video blocks, and the file size set of all bitrate copies of the current video block are the vector part of the state, and the rest are the scalar part of the state;
[0007] Construct the available bitrate copies of the next video block into an action space;
[0008] Obtain a preset reinforcement learning model, the preset reinforcement learning model includes a one-dimensional convolutional layer, a first fully connected layer, a second fully connected layer, and a Softmax layer; input the vector part of the state into the one-dimensional convolutional layer to obtain a first feature vector, input the scalar part of the state into the first fully connected layer to obtain a second feature vector, combine the first feature vector and the second feature vector and input them into the second fully connected layer to obtain a third feature vector, input the third feature vector into the Softmax layer and output the probabilities of the available bitrate copies when transmitting the next video block, and select the bitrate copy with the highest probability of the next video block and transmit it to the buffer;
[0009] During the reinforcement learning process, calculate the video quality description score according to the bitrate of the current video block, and calculate the external reward in combination with the playback interruption time and the smoothness of video quality switching; use a preset neural network to extract eigenvalue differences from the state spaces of the previous video block and the current video block respectively as the internal reward; calculate the cumulative value of discounted external rewards according to the external reward, and calculate the cumulative value of discounted internal and external rewards according to the external reward and the internal reward; with maximizing the cumulative value of discounted external rewards as the optimization direction, construct a gradient backpropagation using the cumulative value of discounted external rewards and the cumulative value of discounted internal and external rewards and update the parameters of the preset reinforcement learning model and the preset neural network.
[0010] In some embodiments, when calculating the video quality description score according to the bitrate of the current video block and calculating the external reward in combination with the playback interruption time and the smoothness of video quality switching, the calculation formula of the external reward is:
[0011]
[0012]
[0013]
[0014]
[0015] S m = |Q m -Q m-1 |;
[0016] Wherein, Q m is the video quality description score, and the buffer penalty term T m represents the playback interruption time of the m-th video block, and the buffer penalty term S m represents the video quality switching smoothness of the m-th video block, and μ m and λ are buffer penalty term weight coefficients; d m (x m ) represents the data volume of the m-th video block; c m represents the average throughput when downloading the m-th video block; c(t) represents the time-varying throughput; B m represents the content duration of all video blocks in the buffer; B m+1 represents the content duration of all video blocks in the buffer after the m-th video block is completely downloaded; x m represents the bitrate selected for transmitting the m-th video block, and x m ∈ {x1, x2,..., x q}; M represents the total number of video blocks into which the video source is divided;
[0017] The video quality description score is calculated using the video quality description model VMAF, and the calculation formula is:
[0018] Q m = VMAF(x m ).
[0019] In some embodiments, the method further includes dynamically switching the buffer penalty term weight coefficient μ m , including:
[0020] Define the buffer penalty for the past k video blocks as:
[0021]
[0022] The switching penalty for the past k video blocks is:
[0023]
[0024] Define the buffer penalty ratio as:
[0025]
[0026] Define the weight update factor U m as:
[0027]
[0028] where C is a constant term;
[0029] Define the rate at which video content leaves the buffer as O m , define the change rate of the buffer occupancy as ΔB m , ΔB m The calculation formula of is:
[0030]
[0031] Define the minimum video volume that the buffer should have as B min ;
[0032] When B m < B min it is in a low buffer state, and update the buffer penalty term weight coefficient μ m as follows:
[0033] μ’ m = μ m - U m ;
[0034] When ΔB m < 0 it is in a buffer consumption state, and update the buffer penalty term weight coefficient μ m as follows:
[0035] μ’ m = μ m + U m .
[0036] In some embodiments, calculate the cumulative external reward discount according to the external reward, and the calculation formula is:
[0037]
[0038] where G ex (s t , a t ) represents the cumulative external reward discount, γ i represents the reward discount factor in the (t + i)-th state, represents the external reward corresponding to the (t + i)-th video block, γ l represents the reward discount factor in the (t + l)-th state, V(s t+l ) represents the state value corresponding to the (t + l)-th state, V(s t ) represents the state value corresponding to the t-th state.
[0039] In some embodiments, an internal and external reward discount cumulative value is calculated based on the external reward and the internal reward, and the calculation formula is:
[0040]
[0041] where G ex+in (s t , a t ) represents the internal and external reward discount cumulative value, γ i represents the reward discount factor in the (t + i)-th state, represents the external reward corresponding to the (t + l)-th video block, λ represents the weight coefficient of the buffer penalty term, represents the internal reward corresponding to the state and action of the i-th video block, γ l represents the reward discount factor in the (t + l)-th state, V(s t+l ) represents the state value corresponding to the (t + l)-th state, V(s t ) represents the state value corresponding to the t-th state.
[0042] In some embodiments, with the optimization direction of maximizing the external reward discount cumulative value, a gradient backpropagation is constructed using the external reward discount cumulative value and the internal and external reward discount cumulative value, and the parameters of the preset reinforcement learning model and the preset neural network are updated, including:
[0043] The parameters of the preset reinforcement learning model are updated by constructing a gradient backpropagation using the internal and external reward discount cumulative value, and the parameter update calculation formula is:
[0044]
[0045] where θ represents the parameters of the preset reinforcement learning model before update, θ′ represents the parameters of the preset reinforcement learning model after update, and π θ represents the policy of the preset reinforcement learning model.
[0046] In some embodiments, with the optimization direction of maximizing the external reward discount cumulative value, a gradient backpropagation is constructed using the external reward discount cumulative value and the internal and external reward discount cumulative value, and the parameters of the preset reinforcement learning model and the preset neural network are updated, including:
[0047] The external policy gradient is calculated based on the external reward discount cumulative value, and the calculation formula is:
[0048]
[0049] The policy parameter gradient is calculated, and the calculation formula is:
[0050]
[0051] Update the parameters of the preset neural network, and the calculation formula is:
[0052]
[0053] where represents the external policy gradient, G ex (s t , a t ) represents the cumulative value of the external reward discount, represents the policy parameter gradient, G ex+in (s t , a t ) represents the cumulative value of the internal and external reward discounts, π θ represents the policy of the preset reinforcement learning model, π θ′ represents the updated policy of the preset reinforcement learning model, θ represents the parameters of the preset reinforcement learning model before updating, θ′ represents the parameters of the preset reinforcement learning model after updating, η represents the parameters of the preset neural network before updating, η′ represents the parameters of the preset neural network after updating, β represents the step size parameter, s t represents the state space, a t represents the code rate copy with the largest selection probability, and α represents the step size parameter.
[0054] In some embodiments, using a preset neural network to respectively extract eigenvalue and calculate the difference of the state spaces of the previous video block and the current video block as the internal reward includes:
[0055] Using a preset neural network to respectively extract eigenvalue and calculate the L2 norm of the state spaces of the previous video block and the current video block as the internal reward.
[0056] On the other hand, the present invention also provides an electronic device, including a processor and a memory, wherein computer instructions are stored in the memory, and the processor is configured to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device implements the steps of the above method.
[0057] On the other hand, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above method are implemented.
[0058] The beneficial effects of the present invention are at least:
[0059] The bitrate adaptive selection method and device of the present invention obtain the network state, video content state, and player state of the current video block to construct a state space, obtain each bitrate copy that can be selected for the next video block to construct an action space, input the vector part and scalar part in the state space into a preset reinforcement model respectively to obtain the probabilities of each bitrate copy that can be selected for the next video block, and construct a gradient backpropagation through the cumulative value of external reward discount and the cumulative value of internal and external reward discounts to update the parameters of the preset reinforcement learning model. The bitrate adaptive method based on reinforcement learning can continuously adjust the internal parameters of the preset reinforcement learning model through the change of the current video block state space and internal and external rewards, change the bitrate copy selection strategy of the video block, so that the bitrate copy output for the next video block can obtain higher rewards, thereby improving the quality of experience when users watch videos.
[0060] Further, the video quality score is used to replace the objective mapping of video quality as a component of the external reward, so that the bitrate of the output video block better meets the actual needs of users, thereby improving the user-perceived quality.
[0061] Further, the penalty weight of the video block is dynamically switched according to the state of the buffer, and different bitrate copy selection strategies are adopted according to different buffer states to reduce the negative impact of video block buffering on the user experience.
[0062] Further, a gradient backpropagation is constructed through the cumulative value of external reward discount and the cumulative value of internal and external reward discounts, and the parameters of the preset reinforcement learning model and the preset neural network are continuously updated to obtain a target reinforcement learning model, so that the bitrate copy of the output video block meets the user's needs and improves the quality of user experience.
[0063] Further, the bitrate adaptive method of the present invention can optimize the update of the internal parameter gradient of the preset reinforcement learning model when the task objectives conflict with each other, so that the preset reinforcement learning model can better select the bitrate copy of the video block according to the state space.
[0064] The additional advantages, objectives, and features of the present invention will be partially described below, and will become partially obvious to those of ordinary skill in the art after studying the following text, or can be learned from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the structures specifically pointed out in the specification and the drawings.
[0065] Those skilled in the art will understand that the objectives and advantages that can be achieved by the present invention are not limited to the above specifically described, and the above and other objectives that the present invention can achieve will be more clearly understood according to the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] The accompanying drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention.
[0067] Figure 1 This is the bitrate adaptive selection method according to an embodiment of the present invention.
[0068] Figure 2 This is the internal flowchart of the preset reinforcement learning model according to an embodiment of the present invention. Detailed implementation manners
[0069] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the implementation manners and the accompanying drawings. Here, the illustrative implementation manners of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.
[0070] Here, it should also be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution of the present invention are shown in the drawings, while other details less related to the present invention are omitted.
[0071] It should be emphasized that the term "comprising / including" when used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.
[0072] Here, it should also be noted that if not specifically stated, the term "connection" in this article can not only refer to a direct connection, but also represent an indirect connection with an intermediate.
[0073] In the following, embodiments of the present invention will be described with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.
[0074] The prior art improves the user experience through a bitrate adaptive method based on deep reinforcement learning. The existing bitrate adaptive methods use reinforcement learning methods to learn control strategies. However, there are still many deficiencies in the existing learning-based adaptive bitrate methods.
[0075] Firstly, it is the correlation between the optimization objective and the actual requirements. The mainstream optimization objectives of the current bitrate adaptive methods based on deep reinforcement learning cannot accurately reflect the actual perceived requirements of users.
[0076] Secondly, it is the impact of the buffer state on the quality of the user experience. The existing bitrate adaptive methods based on deep reinforcement learning cannot make different bitrate decisions according to different buffer states.
[0077] Finally, there is the gradient update for policy optimization. There are conflicting relationships among the task objectives of current bitrate adaptation methods based on deep reinforcement learning. For example, there is a contradiction between high-bitrate video chunks and low network speed. When downloading a high-bitrate video chunk copy at low network speed, the download is slow, and the phenomenon of the buffer video chunk being emptied will occur, resulting in video playback interruption problems. The conflict between task objectives will lead to a decline in the performance of the internal parameter gradient update of the reinforcement learning model. Therefore, the present invention provides a bitrate adaptation selection method and device to solve the problem that the existing bitrate adaptation methods cannot feedback the actual perceived needs of users and cannot adjust the selection strategy of video chunk bitrate copies according to the buffer state.
[0078] One aspect of the present invention provides a bitrate adaptation selection method, as Figure 1 shown, the method includes steps S101 to S105:
[0079] S101: The video provider obtains the video source, divides the video source into a plurality of consecutive video chunks according to a set duration, and each video chunk is compressed into corresponding multiple bitrate copies according to multiple bitrates; in the initial state, the video provider sends the specified bitrate copy of the first video chunk according to the request of the video requester and transmits it to the buffer space.
[0080] S102: During the transmission process, the network state, video content state, and player state during the transmission of the current video chunk are constructed into a state space; the network state includes the average throughput during the transmission of each transmitted video chunk; the video content state includes the download time set of the current video chunk and the set number of previous video chunks, the file size set of all bitrate copies of the current video chunk, and the perceived quality parameter of the previous video chunk; the player state includes the duration of all unplayed video chunks in the buffer, the bitrate of the current video chunk, and the number of remaining untransmitted video chunks in the video source; among them, the average throughput during the transmission of each transmitted video chunk, the download time set of the current video chunk and the set number of previous video chunks, and the file size set of all bitrate copies of the current video chunk are the vector part in the state, and the rest are the scalar part in the state.
[0081] S103: Construct the action space with all the bitrate copies that the next video chunk can select.
[0082] S104: Obtain a preset reinforcement learning model, where the preset reinforcement learning model includes a one-dimensional convolutional layer, a first fully-connected layer, a second fully-connected layer, and a Softmax layer; input the vector part in the state into the one-dimensional convolutional layer to obtain a first feature vector, input the scalar part in the state into the first fully-connected layer to obtain a second feature vector, combine the first feature vector and the second feature vector and input them into the second fully-connected layer to obtain a third feature vector, input the third feature vector into the Softmax layer and output the probabilities of each bitrate copy selected when transmitting the next video block, and select the bitrate copy with the highest probability of the next video block and transmit it to the buffer.
[0083] S105: During the reinforcement learning process, calculate the video quality description score according to the bitrate of the current video block, and combine the playback interruption time and the video quality switching smoothness to calculate the external reward; use a preset neural network to extract eigenvalue respectively from the state spaces of the previous video block and the current video block and calculate the difference as the internal reward; calculate the cumulative value of the discounted external reward according to the external reward, and calculate the cumulative value of the discounted internal and external rewards according to the external reward and the internal reward; taking maximizing the cumulative value of the discounted external reward as the optimization direction, use the cumulative value of the discounted external reward and the cumulative value of the discounted internal and external rewards to construct gradient backpropagation and update the parameters of the preset reinforcement learning model and the preset neural network.
[0084] In step S101, each video block is compressed into corresponding multiple bitrate copies at multiple bitrates, where the durations of each video block are the same. Compressing each video block into multiple bitrate copies can provide video blocks with different bitrate copies during the transmission process according to the changes in the network state, buffer state, and player state, so that the video played by the video requester is smooth and clear. During the video playback process, the bitrate of each video block changes constantly. Therefore, the video seen by the video requester can be regarded as a set of several video blocks with different constant bitrates. Among them, the video provider can be a server, and the video requester can be a client. The client requests a video block with a specified bitrate from the server. After receiving the request, the server responds and sends the specified bitrate copy of the video block to the video cache space of the client.
[0085] In step S102, as Figure 2 shown, the network state represents the link state when downloading a video block, which can be expressed as Then the average throughput of each transmitted video block during the transmission process is expressed as The video content state represents the information of the transmitted and to-be-transmitted video blocks that can be obtained when downloading a video block, which can be expressed as Then the set of download times of the previous video block and its previous set number of video blocks is expressed as The set of file sizes of all bitrate copies of the current video block is expressed as The perceptual quality parameter of the previous video block is denoted as v t The player state representation refers to the state of the player buffer when the video block is downloaded, which can be denoted as B t ={b t , x t , ρ t}, then the duration of all unplayed video blocks in the buffer is denoted as b t , the bitrate of the current video block is denoted as xt, and the number of remaining untransmitted video blocks is denoted as ρ t . Where t represents the current video block, l represents the transmitted video blocks, and the remaining untransmitted video blocks refer to the video blocks in the video source that have not been transmitted
[0086] In step S103, all selectable bitrate copies of the next video block form the action space, which is used to subsequently select a suitable bitrate copy from the action space of the next video block as the bitrate copy of the next video block according to the state space of the current video block
[0087] In step S104, the vector part in the state space, that is , after being input into the one-dimensional convolutional layer for feature extraction, the first feature vector is obtained. The vector part in the state space, that is v t , B t , after being input into the first fully connected layer for feature extraction, the second feature vector is obtained. The first feature vector and the second feature vector are combined and input into the second fully connected layer to obtain the third feature vector. The third feature vector is input into the Softmax layer and output to obtain an m-dimensional vector, and this m-dimensional vector represents the probabilities of each bitrate copy of the next video block being selected. The reinforcement learning model can continuously improve the policy of executing actions through interaction with the environment. The purpose of the present invention is to continuously improve the selection policy of the bitrate copies of each video block through the interaction of the preset reinforcement learning model with the state space
[0088] In step S105, multiple factors affecting the user experience quality are obtained and a user experience quality index is constructed, that is, a user experience quality index (QoE) is constructed according to the video quality description score, playback interruption time, and video quality switching smoothness of the current video block, and the user experience quality index is used as the external reward of the preset reinforcement learning model. Among them, the video quality description score of the current video block is used as the scoring item of the external reward, and the playback interruption time and video quality switching smoothness are used as the penalty items of the external reward. The optimization goal of the preset reinforcement learning model is to make the output bitrate copy of the video block obtain a higher external reward
[0089] In some embodiments, when calculating the video quality description score according to the bitrate of the current video block and calculating the external reward in combination with the playback interruption time and video quality switching smoothness, the calculation formula of the external reward is
[0090]
[0091]
[0092]
[0093]
[0094] S m = |Q m - Q m-1 |;
[0095] where Q m is the video quality description score of the m-th video block, and Q m-1 is the video quality description score of the (m - 1)-th video block. The buffer penalty term T m represents the playback interruption time of the m-th video block, and the buffer penalty term S m represents the video quality switching smoothness of the m-th video block. μ m and λ are buffer penalty term weight coefficients; d m (x m ) represents the data volume of the m-th video block; c m represents the average throughput when downloading the m-th video block; c(t) represents the time-varying throughput; B m represents the content duration of all video blocks in the buffer; B m+1 represents the content duration of all video blocks in the buffer after the m-th video block is completely downloaded; x m represents the bitrate selected for transmitting the m-th video block, and x m ∈ {x1, x2,..., x q}; M represents the total number of video blocks into which the video source is divided.
[0096] Furthermore, the video quality description score is calculated using the video quality description model VMAF, and the calculation formula is:
[0097] Q m = VMAF(x m ).
[0098] The present invention uses the video quality description score to replace the objective mapping of video quality to form an external reward, which can more accurately reflect the actual perception of the user on the video quality of the played video. Therefore, the external reward composed of the video quality description score can better adjust the parameter update direction of the preset reinforcement learning model according to the user's perception.
[0099] In some embodiments, the bitrate adaptive selection method of the present invention further includes dynamically switching the buffer penalty term weight coefficient μ m , including:
[0100] Define the buffer penalty for the past k video chunks as:
[0101]
[0102] The switching penalty for the past k video chunks is:
[0103]
[0104] Define the buffer penalty ratio as:
[0105]
[0106] Define the weight update factor U m as:
[0107]
[0108] where C is a constant term.
[0109] Define the rate at which video content leaves the buffer as O m , and define the rate of change of the buffer occupancy as ΔB m , ΔB m The calculation formula for it is:
[0110]
[0111] Define the minimum amount of video that the buffer should have as B min ;
[0112] When B m < B min , it is in a low buffer state, and update the weight coefficient μ m of the buffer penalty term as follows:
[0113] μ’ m = μ m - U m ;
[0114] When the video content in the buffer is less than the minimum amount of video that the buffer should have, it is in a low buffer state, and the weight coefficient of the buffer penalty term is updated to reduce the weight of the buffer penalty term.
[0115] When ΔB m < 0, it is in a buffer consumption state, and update the weight coefficient μ m of the buffer penalty term as follows:
[0116] μ’ m = μ m + U m .
[0117] where μ’ mis the updated weight coefficient of the buffer penalty term. When the change rate of the buffer occupancy rate is less than 0, it indicates the buffer consumption state, and the buffer penalty term weight is increased by updating the buffer penalty term weight coefficient.
[0118] In some other embodiments, when the change rate of the buffer occupancy rate is greater than 0 and the content duration of all video blocks in the buffer is greater than the minimum video volume that the buffer should have, the buffer penalty term weight coefficient remains unchanged.
[0119] During video playback, the buffer state is constantly changing. If video block copies with the same bitrate are continuously transmitted, it will cause problems such as video playback interruption. Therefore, it is necessary to continuously adjust the selection strategy of video block bitrate copies according to the buffer state. The buffer has three states: low buffer state, stable playback state, and buffer consumption state. The low buffer state is 0 < ΔB m , B m <B min , and this state will appear in the initial stage of playback and the stage of re-accumulating the buffer after playback interruption; the stable playback state is 0 < ΔB m , B m ≥B min , and this state is the ideal state and may appear at any stage during playback; the buffer consumption state is ΔB m <0, B m ≥B min , and this state indicates that there is still content in the buffer but it is gradually depleting. This state will appear before an impending playback interruption and at the final stage of playback. Playback interruption can be avoided by changing the bitrate selection strategy. During the entire process of video playback, the buffer state switches among the three states. For the low buffer state, the strategy of the preset reinforcement learning model should tend to select a lower bitrate to minimize the initial playback start time or the time required for recovery after playback interruption; the buffer consumption state means that the video blocks in the buffer are about to be exhausted, and there will be phenomena such as playback interruption or stop. The strategy of the preset reinforcement learning model should tend to select a higher bitrate to fill the buffer with video blocks.
[0120] In some embodiments, a preset neural network is used to extract eigenvalue differences as internal rewards by separately extracting the feature values of the state spaces of the previous video block and the current video block. The calculation formula is:
[0121]
[0122] where s t+1 represents the state space of the previous video block, s t represents the state space of the current video block, represents the eigenvalue of the video block state extracted from the convolutional layer of the preset neural network. The internal reward is used to learn the switching between state spaces and understand the state space where the video block is located.
[0123] In this embodiment, a preset neural network is used to extract eigenvalue from the state spaces of the previous video block and the current video block respectively, and calculate the L2 norm as the internal reward.
[0124] In some other embodiments, after extracting the distribution features of the state spaces of the previous video block and the current video block, the difference between the two state spaces is obtained by calculating the Euclidean distance and used as the internal reward of the preset neural network to iteratively update the preset neural network. The difference between the previous video block and the current video block may be the difference in the network state during video transmission.
[0125] In some embodiments, the cumulative value of the external reward discount is calculated according to the external reward, and the calculation formula is:
[0126]
[0127] where, G ex (s t , a t ) represents the cumulative value of the external reward discount, γ i represents the reward discount factor at the (t + i)-th state, represents the external reward corresponding to the (t + i)-th video block, γ l represents the reward discount factor at the (t + l)-th state, V(s t+l ) represents the state value corresponding to the (t + l)-th state, V(s t ) represents the state value corresponding to the t-th state. The cumulative value of the external reward discount is the sum of the external reward values of all the output video blocks calculated according to the reward discount factors of each video block and the external rewards obtained after outputting the bitrate copies of each video block according to the preset reinforcement learning model.
[0128] In some embodiments, the cumulative value of the internal and external reward discounts is calculated according to the external reward and the internal reward, and the calculation formula is:
[0129]
[0130] where, G ex+in (s t , a t ) represents the cumulative value of the internal and external reward discounts, γ i represents the reward discount factor at the (t + i)-th state, represents the external reward corresponding to the (t + 1)-th video block, λ represents the weight coefficient of the penalty term, represents the internal reward corresponding to the state and action of the i-th video block, γ l represents the reward discount factor at the (t + l)-th state, V(s t+l ) represents the state value corresponding to the (t + l)-th state, V(st ) represents the state value corresponding to the t-th state.
[0131] Furthermore, the gradient backpropagation is constructed by using the cumulative value of internal and external reward discounts to update the parameters of the preset reinforcement learning model. The parameter update calculation formula is:
[0132]
[0133] where θ represents the parameters of the preset reinforcement learning model before update, θ′ represents the parameters of the preset reinforcement learning model after update, and π θ represents the policy of the preset reinforcement learning model.
[0134] The parameters of the preset reinforcement learning model are updated through the cumulative value of internal and external reward discounts to optimize the internal parameters of the preset reinforcement learning model to obtain the target reinforcement learning model. After the preset reinforcement learning model outputs the video block bitrate copy to the state space each time, the state space will assign internal reward values and external reward values according to the state of the video block and feedback them to the preset reinforcement learning model. The preset reinforcement learning model adjusts its internal parameters based on the cumulative value of internal and external reward discounts to adjust the bitrate copy selection policy of the next video block, so that the bitrate copy of the output next video block can obtain a higher reward value. After multiple update iterations, the target reinforcement learning model is obtained. The video block bitrate copy output by the target reinforcement learning model better meets the user's needs and the changes in the state space, thereby improving the user experience quality. The cumulative value of internal and external reward discounts is obtained by accumulating the internal reward values and external reward values of each video block output by the preset reinforcement learning model. The optimization goal of the present invention is:
[0135]
[0136] That is, the cumulative value of external reward discounts for all video blocks in the video source reaches the maximum. M represents the total number of video blocks into which the video source is divided. When the cumulative value of external reward discounts reaches the maximum, a higher user experience quality can be obtained.
[0137] In some embodiments, the gradient backpropagation is constructed by using the cumulative value of external reward discounts and the cumulative value of internal and external reward discounts to update the parameters of the preset reinforcement learning model and the preset neural network, and further includes:
[0138] Calculate the external policy gradient according to the cumulative value of external reward discounts. The calculation formula is:
[0139]
[0140] Calculate the policy parameter gradient. The calculation formula is:
[0141]
[0142] Update the parameters of the preset neural network, and the calculation formula is:
[0143]
[0144] Wherein, represents the external policy gradient, and G ex (s t , a t ) represents the cumulative value of the discounted external reward, represents the policy parameter gradient, and G ex+in (s t , a t ) represents the cumulative value of the discounted internal and external rewards, π θ represents the policy of the preset reinforcement learning model, θ represents the parameters of the preset reinforcement learning model before update, π θ′ represents the policy of the preset reinforcement learning model after update, θ′ represents the parameters of the preset reinforcement learning model after update, η represents the parameters of the preset neural network before update, η′ represents the parameters of the preset neural network after update, β represents the step size parameter, s t represents the state space, a t represents the code rate copy with the largest selection probability, and α represents the step size parameter.
[0145] Calculate the policy parameter gradient of the preset reinforcement learning model by the method of random sampling, calculate the external policy gradient by the importance sampling method, and finally calculate the parameters of the preset reinforcement learning model and the preset neural network by the chain rule. Update the parameters of the preset neural network to iterate the preset neural network, so as to reduce the influence of the network state on the code rate copy selection strategy of the video block, so that the video block can also have better playback quality when the network state is poor, and further promote the preset reinforcement learning model to output a more optimized video block code rate copy selection strategy under the joint action of the internal reward and the external reward.
[0146] On the other hand, the present invention also provides an electronic device, including a processor and a memory, wherein computer instructions are stored in the memory, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device implements the steps of the above method.
[0147] On the other hand, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the above method are implemented.
[0148] In summary, for the bitrate adaptive selection method and device of the present invention, the network state, video content state, and player state of the current video block are obtained and constructed into a state space, and each bitrate copy that can be selected for the next video block is obtained and constructed into an action space. The vector part and scalar part in the state space are respectively input into a preset reinforcement model to obtain the probabilities of each bitrate copy selected for the next video block, and the gradient backpropagation is constructed through the cumulative value of external reward discount and the cumulative value of internal and external reward discounts to update the parameters of the preset reinforcement learning model. The bitrate adaptive method based on reinforcement learning can continuously adjust the internal parameters of the preset reinforcement learning model through the changes in the state space of the current video block and internal and external rewards, change the bitrate copy selection strategy of the video block, so that the bitrate copy of the next video block output can obtain higher rewards, thereby improving the quality of experience when users watch videos.
[0149] Furthermore, the video quality score is used to replace the objective mapping of video quality as a component of the external reward, so that the bitrate of the output video block is more in line with the actual needs of users, thereby improving the user-perceived quality.
[0150] Furthermore, the penalty weight of the video block is dynamically switched according to the state of the buffer, and different bitrate copy selection strategies are adopted according to different buffer states to reduce the negative impact of video block buffering on the user experience.
[0151] Furthermore, the gradient backpropagation is constructed through the cumulative value of external reward discount and the cumulative value of internal and external reward discounts, and the parameters of the preset reinforcement learning model and the preset neural network are continuously updated to obtain the target reinforcement learning model, so that the bitrate copy of the output video block meets the user's needs and improves the quality of user experience.
[0152] Furthermore, the bitrate adaptive method of the present invention can optimize the update of the internal parameter gradient of the preset reinforcement learning model in the case of conflicting task objectives, so that the preset reinforcement learning model can better select the bitrate copy of the video block according to the state space.
[0153] Corresponding to the above method, the present invention also provides a device, which includes a computer device. The computer device includes a processor and a memory. Computer instructions are stored in the memory, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device implements the steps of the method described above.
[0154] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the foregoing edge computing server deployment method are implemented. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, floppy disk, hard disk, removable storage disk, CD-ROM, or any other form of storage medium well-known in the art.
[0155] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it may be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave on a transmission medium or a communication link.
[0156] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.
[0157] In the present invention, the features described and / or illustrated for one embodiment can be used in the same or similar manner in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.
[0158] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A bit rate adaptive selection method, characterized in that The method includes the following steps: Obtain a video source from a video provider, divide the video source into a plurality of consecutive video chunks according to a set duration, and compress each video chunk into corresponding multiple bitrate copies according to multiple bitrates; in an initial state, the video provider sends a specified bitrate copy of the first video chunk according to a request from a video requester, and transmits it to a cache space; During the transmission process, construct a state space from the network state, video content state, and player state during the transmission process of the current video chunk; the network state includes the average throughput of each transmitted video chunk during the transmission process; the video content state includes the download time set of the current video chunk and a set number of previous video chunks, the file size set of all bitrate copies of the current video chunk, and the perceived quality parameter of the previous video chunk; the player state includes the duration of all unplayed video chunks in the buffer, the bitrate of the current video chunk, and the number of remaining untransmitted video chunks in the video source; wherein, the average throughput of each transmitted video chunk during the transmission process, the download time set of the current video chunk and a set number of previous video chunks, and the file size set of all bitrate copies of the current video chunk are the vector part in the state, and the rest are the scalar part in the state; Construct an action space from the bitrate copies that can be selected for the next video chunk; Obtain a preset reinforcement learning model, the preset reinforcement learning model includes a one-dimensional convolutional layer, a first fully connected layer, a second fully connected layer, and a Softmax layer; input the vector part in the state into the one-dimensional convolutional layer to obtain a first feature vector, input the scalar part in the state into the first fully connected layer to obtain a second feature vector, combine the first feature vector and the second feature vector and input them into the second fully connected layer to obtain a third feature vector, input the third feature vector into the Softmax layer and output the probabilities of the bitrate copies selected when transmitting the next video chunk, and select the bitrate copy with the highest probability of the next video chunk and transmit it to the buffer; During the reinforcement learning process, calculate a video quality description score according to the bitrate of the current video chunk, and calculate an external reward in combination with the playback interruption time and video quality switching smoothness; use a preset neural network to extract eigenvalue differences from the state spaces of the previous video chunk and the current video chunk respectively as an internal reward; calculate the cumulative value of the discounted external reward according to the external reward, and calculate the cumulative value of the discounted internal and external rewards according to the external reward and the internal reward; taking maximizing the cumulative value of the discounted external reward as the optimization direction, construct a gradient backpropagation using the cumulative value of the discounted external reward and the cumulative value of the discounted internal and external rewards and update the parameters of the preset reinforcement learning model and the preset neural network; When calculating the video quality description score according to the bitrate of the current video chunk and calculating the external reward in combination with the playback interruption time and video quality switching smoothness, the calculation formula of the external reward is: S m = |Q m -Q m-1 |; Among them, Q m is the video quality description score, and the buffer penalty term T m represents the playback interruption time of the m-th video chunk, and the buffer penalty term S m represents the video quality switching smoothness of the m-th video chunk, μ m and λ are buffer penalty term weight coefficients; d m (x m ) represents the data volume of the m-th video chunk; c m represents the average throughput when downloading the m-th video chunk; c(t) represents the time-varying throughput; B m represents the content duration of all video chunks in the buffer; B m+1 represents the content duration of all video chunks in the buffer after the m-th video chunk is completely downloaded; x m represents the bitrate selected for transmitting the m-th video chunk, x m ∈{x1, x2, …, x q}; M represents the total number of video chunks into which the video source is divided; The video quality description score is calculated using the video quality description model VMAF, and the calculation formula is: Q m = VMAF(x m ); The method further includes dynamically switching the buffer penalty term weight coefficient μ m , including: Define the buffer penalty for the past k video chunks as: The switching penalty for the past k video chunks is: Define the buffer penalty ratio as: Define the weight update factor U m as follows: where C is a constant term; Define the rate at which video content leaves the buffer as O m , and define the rate of change of the buffer occupancy as △B m , △B m The calculation formula for is: Define the minimum video volume that the buffer should have as B min ; When B m <B min is in the low cache state, update the buffer penalty term weight coefficient μ m as follows: μ’ m = μ m - U m ; When △B m <is in the cache consumption state when it is less than 0, update the buffer penalty term weight coefficient μ m as follows: μ’ m = μ m + U m .
2. The bit rate adaptive selection method according to claim 1, wherein Calculate the cumulative value of the external reward discount according to the external reward, and the calculation formula is: Among them, G ex (s t , a t ) represents the cumulative value of the external reward discount, γ i represents the reward discount factor in the (t + i)-th state, represents the external reward corresponding to the (t + i)-th video block, γ l represents the reward discount factor in the (t + l)-th state, V(s t+l ) represents the state value corresponding to the (t + l)-th state, V(s t ) represents the state value corresponding to the t-th state.
3. The bit rate adaptive selection method according to claim 2, wherein Calculate the cumulative value of the internal and external reward discounts according to the external reward and the internal reward, and the calculation formula is: Among them, G ex+in (s t , a t ) represents the cumulative value of the internal and external reward discounts, γ i represents the reward discount factor in the (t + i)-th state, represents the external reward corresponding to the (t + 1)-th video block, λ represents the weight coefficient of the buffer penalty term, represents the internal reward corresponding to the state and action of the i-th video block, γ l represents the reward discount factor in the (t + l)-th state, V(s t+l ) represents the state value corresponding to the (t + l)-th state, V(s t ) represents the state value corresponding to the t-th state.
4. The bit rate adaptive selection method according to claim 3, wherein Taking maximizing the cumulative value of the external reward discount as the optimization direction, use the cumulative value of the external reward discount and the cumulative value of the internal and external reward discounts to construct gradient backpropagation and update the parameters of the preset reinforcement learning model and the preset neural network, including: Use the cumulative value of the internal and external reward discounts to construct gradient backpropagation to update the parameters of the preset reinforcement learning model, and the parameter update calculation formula is: Among them, θ represents the parameters of the preset reinforcement learning model before update, θ′ represents the parameters of the preset reinforcement learning model after update, and π θ represents the policy of the preset reinforcement learning model.
5. The bit rate adaptive selection method according to claim 3, wherein Taking maximizing the cumulative value of the external reward discount as the optimization direction, use the cumulative value of the external reward discount and the cumulative value of the internal and external reward discounts to construct gradient backpropagation and update the parameters of the preset reinforcement learning model and the preset neural network, including: Calculate the external policy gradient according to the cumulative value of the external reward discount, and the calculation formula is: Calculate the policy parameter gradient, and the calculation formula is: Update the parameters of the preset neural network, and the calculation formula is: Among them, represents the external policy gradient, G ex (s t , a t ) represents the cumulative value of the external reward discount, represents the policy parameter gradient, G ex+i (s t , a t ) represents the cumulative value of the internal and external reward discounts, π θ represents the policy of the preset reinforcement learning model, π θ′ represents the updated policy of the preset reinforcement learning model, θ represents the parameters of the preset reinforcement learning model before update, θ′ represents the parameters of the preset reinforcement learning model after update, η represents the parameters of the preset neural network before update, η′ represents the parameters of the preset neural network after update, β represents the step size parameter, s t represents the state space, a t represents the code rate copy with the largest selection probability, and α represents the step size parameter.
6. The bit rate adaptive selection method according to claim 1, wherein Use a preset neural network to extract eigenvalue from the state spaces of the previous video block and the current video block respectively and calculate the difference as the internal reward, including: Use a preset neural network to extract eigenvalue from the state spaces of the previous video block and the current video block respectively and calculate the L2 norm as the internal reward.
7. A bitrate adaptation device based on meta-learning, comprising a processor and a memory, characterized in that, The memory stores computer instructions, and the processor is configured to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device implements the steps of the method according to any one of claims 1 to 6.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.