Robustness adaptive code rate method based on time-frequency domain feature enhancement

By introducing time-frequency domain feature enhancement method into the adaptive code rate algorithm, a dual-branch network model is constructed, which solves the problem of insufficient robustness of existing algorithms under complex network fluctuations, and achieves higher user experience quality and bandwidth utilization.

CN120111287APending Publication Date: 2025-06-06HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510261721.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing adaptive bit rate algorithm based on time domain information cannot balance the relationship between buffer level and video block quality selection when facing complex network fluctuations, resulting in poor robustness.

Method used

A robust adaptive bit rate method based on time-frequency domain feature enhancement is adopted. By building a dual-branch network model, the features of the frequency domain and time domain are extracted respectively, and the long sequence dependencies are extracted in combination with the gating neural unit to generate more accurate bit rate selection.

Benefits of technology

It significantly improves the performance of the adaptive code rate algorithm in complex network environments, improves the user experience quality, reduces lag, improves bandwidth utilization, and enhances the robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120111287A_ABST
    Figure CN120111287A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video processing, in particular to a robust adaptive code rate method based on time-frequency domain feature enhancement. The method comprises the following steps: S1, constructing a video stream transmission model at a server and a client; s2, constructing a double-branch network model; s3, according to the constructed video stream transmission model, acquiring video content features, network features of past blocks and video playing features as input states, and inputting the input states into the network model; s4, selecting a code rate according to a strategy, and calculating an expert action; s5, storing expert actions and states in an experience pool; s6, selecting a batch of training samples from the experience pool for training; s7, updating the strategy network model according to the duration of the block; s8, obtaining a next state according to the state and the code rate, and inputting the next state into the network model; and S9, repeating the steps S4 to S8 until convergence. According to the invention, a flexible system architecture is adopted, and more accurate code rate selection is realized; by combining time domain and frequency domain features, the user experience quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video processing technology, and in particular to a robust adaptive bit rate method based on time-frequency domain feature enhancement. Background Art

[0002] With the explosion of streaming services, streaming video transmission has accounted for 67% of Internet traffic. Traditional adaptive bitrate methods, such as rate-based, buffer-based and MPC technologies, are widely used. Traditional adaptive bitrate algorithms rely solely on the throughput information and buffer occupancy information of video blocks, and their performance is heavily dependent on manual design. They need to be manually adjusted in different usage scenarios or business needs, which severely limits the transmission of high-quality streaming media.

[0003] Recently, ABR algorithms based on deep reinforcement learning (DRL), such as Pensieve and Comyco, have surpassed traditional rule-based methods and achieved higher user quality of experience (QoE) performance. For example, Pensieve uses deep reinforcement learning to select the bitrate of the next video based on the previous state without pre-defined assumptions or manually set rules, achieving excellent user quality of experience (QoE) performance. Comyco introduces imitation learning based on Pensieve, and by replacing the bitrate mapping with VMAF mapping, it is able to select video segments with higher video quality. However, these methods only rely on time domain information, cannot capture global features, and perform poorly in the face of rapid network bandwidth fluctuations.

[0004] Figure 1 The bitrate selection and buffer level changes of Comyco and Pensieve in the face of step-down network bandwidth changes (the blue solid line is the network bandwidth, the green solid line is the Comyco bitrate selection, the green dotted line is the Comyco buffer level change, the red solid line is the Pensieve bitrate selection, and the red dotted line is the Pensieve buffer level change).

[0005] Figure 2 Comyco and Pensieve’s bitrate selection and buffer level changes in the face of periodic network bandwidth changes (the blue solid line is the network bandwidth, the green solid line is the Comyco bitrate selection, the green dotted line is the Comyco buffer level change, the red solid line is the Pensieve bitrate selection, and the red dotted line is the Pensieve buffer level change).

[0006] Temporal imitation learning models (such as Comyco) are highly dependent on expert trajectory examples. Although this approach works well under stable network conditions, it performs poorly in dynamic and complex network environments. Temporal models focus mainly on short-term fluctuations, which can lead to incorrect bitrate decisions and gradually deviate from the expert trajectory, causing distribution shift. This deviation accumulates in subsequent decisions and significantly degrades performance. Figure 1 As shown in Figure 2, Comyco incorrectly selected a bit rate of 4.3 Mbps after 100 seconds, while the bandwidth at that time was insufficient to support such a high bit rate, which eventually led to buffer exhaustion and triggered a continuous buffering event after 200 seconds. Similarly, another mainstream algorithm, Pensieve, also has problems under dynamic network conditions. Figure 1 It is shown that when the network bandwidth continues to decline, Pensieve's bitrate decisions fluctuate frequently, indicating that the time-domain features are overly sensitive to local noise and transient anomalies in the time series, resulting in poor stability of bitrate decisions in dynamic environments. Such fluctuations can cause users to be sensitive to bitrate switching, which can significantly degrade the user quality of experience (QoE). In addition, time-domain models rely on local time series information, making them more reactive than predictive in decision making. This characteristic limits their ability to adapt effectively to rapid bandwidth fluctuations or periodic patterns, resulting in suboptimal bitrate selection.

[0007] When faced with periodic fluctuations in network bandwidth, such as Figure 2 As shown in the figure, neither Comyco nor Pensieve can identify periodic patterns. Comyco adopts an aggressive bitrate selection strategy, resulting in continuous buffering events. Pensieve adopts an overly conservative bitrate selection strategy and fails to fully utilize bandwidth resources. This shows that existing learning-based algorithms cannot balance the relationship between buffer level and video block quality selection when facing complex network fluctuations, that is, they have poor robustness. Summary of the invention

[0008] The present invention provides a robust adaptive bit rate method based on time-frequency domain feature enhancement, aiming to solve the problem that the current adaptive bit rate algorithm using time domain information has poor robustness.

[0009] The present invention provides a robust adaptive bit rate method based on time-frequency domain feature enhancement, comprising the following steps:

[0010] S1. Construct a video streaming transmission model on the server and the client, and in the server, divide a source video into K blocks, each block lasts for L seconds, and each block is encoded into n video blocks with different bit rates.

[0011] S2. Construct a dual-branch network model, which includes a frequency domain network branch and a time domain network branch to extract features in the frequency domain and time domain respectively;

[0012] S3. In the client, according to the video streaming model constructed in step S1, the video content features, the network features of the past blocks, and the video playback features are obtained as the input state S k , then the state S k Input to the network model π θ ;

[0013] S4. According to the strategy π(S k ;θ) Select code rate a k , calculate expert actions

[0014] S5. Expert Action and state S k Stored in experience pool C;

[0015] S6. Select a batch of training samples from the experience pool C Conduct training;

[0016] S7. Update the policy network model π according to the duration L of the block θ ;

[0017] S8. According to S k and a k Get the next state S k+1 , and input into the network model π θ ;

[0018] S9. Repeat steps S4 to S8 until convergence.

[0019] As a further improvement of the present invention, step S2 includes the following steps:

[0020] S21. The time domain network branch includes a network layer, four one-dimensional convolutional layers and three linear layers. The time series information is processed through the convolutional layer and the linear layer. The processed time series information is connected through the network layer and then passed through the gated neural unit to extract the time domain features.

[0021] S22. The frequency domain network branch contains the buffer level of the past block, the VMAF of the previous block, and the number of remaining blocks. The sequence information is transformed into real and imaginary parts through discrete Fourier transform. The features obtained through the real and imaginary two-channel one-dimensional convolution layer are connected and passed through the gated neural unit to extract frequency domain features.

[0022] S23. The frequency domain and time domain features extracted by the two network branches are connected and then used again through the gated neural unit to extract long sequence dependencies. The features are then output through the linear layer, and a probability distribution is generated through Softmax.

[0023] As a further improvement of the present invention, step S3 includes the following steps:

[0024] S31. The size and VMAF of the kth block can be obtained through HTTP request, using D k ={d 1 (R 1 ),d 2 (R 2 ),…,d K (R K )} and Q k = {q 1 (R 1 ),q 2 (R 2 ),…,q K (R K )} means, d k (R k ) indicates the code rate is R k The kth video block size, q k (R k ) is the code rate R k The VMAF score of the k-th video block;

[0025] S32. The coding rate set is expressed as The bit rate of the kth video block is denoted by R k express, Get the download time of the kth video chunk:

[0026]

[0027] Among them, d k (R k ) indicates the code rate is R k The kth video block size, t i and t i+1 Respectively represent the time when downloading starts and ends, c t represents the downlink bandwidth, δ t Indicates the round trip time RTT;

[0028] S33. Video chunks are downloaded to the playback buffer, which contains the video chunks that have been downloaded but not yet viewed; let B(t)∈[0,B max ] represents the occupancy level of the playback buffer at time t, let B k =B(t k) represents the occupancy rate of the playback buffer when the kth block starts to be downloaded. The dynamic change of the buffer is expressed as:

[0029] B k =((B(t k )-τ k ) + +L-δ t ) +

[0030] Among them, (x) + =max{x,0}, if B(t k )-τ k <0, indicating that the buffer is exhausted and a rebuffering event occurs; τ k represents the download time of the kth video block; δ t Indicates the round trip time RTT;

[0031] S34. According to step S32, the throughput of the kth block is obtained

[0032] S35. According to step S31, the video content features are obtained: VMAF of the next block and the size of the next block; according to steps S32 and S33, the video playback features are obtained: download time of the past block, buffer level, VMAF of the previous block, and the number of remaining blocks; therefore:

[0033]

[0034] Where T represents the throughput of the past video chunks, τ represents the download time of the previous few video chunks, q represents the VMAF of the previous video chunk, and r represents the number of remaining video chunks; represents the mapping of discrete Fourier transform, is the discrete Fourier transform of the VMAF selected over the past block.

[0035] As a further improvement of the present invention, step S4 specifically includes:

[0036] According to the strategy π(S k ;θ) Select code rate a k , calculate the expert action:

[0037]

[0038] The expert action is the bitrate combination that maximizes the user experience quality (QOE) of the next N video blocks, where QOE is defined as:

[0039]

[0040] Where N is the total number of video blocks, R nIndicates the bit rate selected for the nth video block, VMAF(R n ) is a function that maps the bitrate to its corresponding VMAF; T n Indicates the duration of the freeze caused by selecting the nth video block; [VMAF(R n+1 )-VMAF(R n )] + VMAF captures the quality improvement when switching from lower video quality to higher video quality, while n+1 )-VMAF(R n )] - It indicates a quality degradation in the opposite direction; the coefficients α, β, γ, and δ are the weights assigned to each item, reflecting the user's preference or aversion to different aspects of the streaming media viewing experience.

[0041] As a further improvement of the present invention, the specific operation process of step S5 is as follows:

[0042] Expert Action and state S k Stored in experience pool C:

[0043] As a further improvement of the present invention, step S7 specifically includes:

[0044] Update the policy network π according to L θ :

[0045]

[0046] where π(s,a;θ) is the network strategy, is the true probability vector of the expert’s actions, H(π(s;θ)) represents the entropy of the policy, and α is a hyperparameter that controls the scope of model exploration.

[0047] As a further improvement of the present invention, the convergence process of step S9 includes:

[0048] Find the optimal strategy by minimizing expected loss

[0049]

[0050] Until Strategy Convergence; among them, represents the expert strategy, T is the strategy set, and d is the strategy The state distribution under Represents the loss of the algorithmic strategy relative to the expert strategy.

[0051] The beneficial effects of the present invention are:

[0052] (1) Improving user experience quality (QoE): This invention significantly improves the performance of the adaptive bitrate algorithm in complex network environments by combining time domain and frequency domain features. This is manifested in higher user viewing quality (QoE), less lag, higher bandwidth utilization, and stronger generalization capabilities, effectively solving the problem of insufficient robustness of traditional algorithms under network fluctuations.

[0053] (2) Flexible system architecture: The system architecture combines convolutional neural networks (CNNs) and gated recurrent units (GRUs) to process time domain and frequency domain features separately, and then fuses the information of the two to achieve more accurate bit rate selection. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is a diagram of the bit rate selection and buffer level change of Comyco and Pensieve in the face of step-down network bandwidth changes in the background technology of the present invention;

[0055] Figure 2 It is a diagram of the code rate selection and buffer level change of Comyco and Pensieve in the face of periodic network bandwidth changes in the background technology of the present invention;

[0056] Figure 3 It is an overall framework diagram of the video stream transmission model constructed on the server and the client of the present invention;

[0057] Figure 4 It is the overall framework diagram of the dual-branch network model constructed by the present invention. DETAILED DESCRIPTION

[0058] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.

[0059] The existing adaptive bitrate algorithm only uses time domain information, and the input in the network model is all time domain information. Studies have shown that compared to time domain information focusing on short-term fluctuations and local noise, frequency domain information has a global perspective and has a certain effect on improving the robustness of the model. However, directly inputting the amplitude and phase of frequency domain information into the network has poor performance or even no convergence. It is necessary to design a separate network branch and feature extraction method for frequency domain information.

[0060] The present invention proposes an enhanced adaptive bit rate method based on time-frequency domain features. Different from the existing methods, the present invention introduces frequency domain information as a new feature dimension, converts time domain information to frequency domain through fast Fourier transform (FFT), and captures the periodic fluctuation and global trend characteristics of network bandwidth. In view of the characteristics of frequency domain information, a "dual-branch network" and "real-virtual feature extraction unit" are specially designed.

[0061] Specific as Figure 3 to Figure 4 As shown, a robust adaptive bit rate method based on time-frequency domain feature enhancement of the present invention comprises the following steps:

[0062] S1. Construct a video streaming transmission model on the server and the client, and in the server, divide a source video into K blocks, each block lasts for L seconds, and each block is encoded into n video blocks with different bit rates.

[0063] S2. Construct a dual-branch network model, which includes a frequency domain network branch and a time domain network branch to extract frequency domain and time domain features respectively.

[0064] Step S2 includes the following steps:

[0065] S21. The time domain network branch includes a network layer, four one-dimensional convolutional layers and three linear layers. The time series information is processed by the convolutional layer and the linear layer. The processed time series information is connected through the network layer and then passed through the gated neural unit to extract the time domain features.

[0066] S22. The frequency domain network branch does not contain the sequence information of the remaining blocks, but contains sequence information such as the buffer level of the past block, the VMAF of the previous block, and the number of remaining blocks. The sequence information such as the buffer level of the past block, the VMAF of the previous block, and the number of remaining blocks are subjected to discrete Fourier transform to obtain the real part and the imaginary part, and the features obtained by the real and imaginary two-channel one-dimensional convolution layer are connected and then passed through the gated neural unit to extract the frequency domain features.

[0067] S23. After the frequency domain and time domain features extracted by the two network branches are connected, the long sequence dependency is extracted again through the gated neural unit. Then the long sequence dependency passes through the linear layer to output the long sequence dependency feature, and the output feature is used to generate a probability distribution through Softmax. The probability distribution is directly related to the bit rate selection and the randomness of the bit rate selection, which affects the overall network performance.

[0068] S3. In the client, according to the video streaming model constructed in step S1, the video content features, the network features of the past blocks, and the video playback features are obtained as the input state S k , then the state S k Input to the network model π θ .

[0069] Step S3 includes the following steps:

[0070] S31. The size and VMAF of the kth block can be obtained through HTTP request, using D k ={d 1 (R 1 ),d 2 (R2 ),…,d K (R K )} and Q k = {q 1 (R 1 ),q 2 (R 2 ),…,q K (R K )}. Among them, d k (R k ) indicates the code rate is R k The kth video block size, q k (R k ) is the code rate R k The VMAF score of the k-th video block.

[0071] S32. The coding rate set is expressed as The bit rate of the kth video block is denoted by R k express, Get the download time of the kth video chunk:

[0072]

[0073] Among them, d k (R k ) indicates the code rate is R k The kth video block size, t i and t i+1 Respectively represent the time when downloading starts and ends, c t represents the downlink bandwidth, δ t Indicates the round trip time RTT.

[0074] S33. Video chunks are downloaded to the playback buffer, which contains the video chunks that have been downloaded but not yet viewed; let B(t)∈[0,B max ] represents the occupancy level of the playback buffer at time t, let B k =B(t k ) represents the occupancy rate of the playback buffer when the kth block starts to be downloaded. The dynamic change of the buffer is expressed as:

[0075] B k =((B(t k )-τ k ) + +L-δ t ) +

[0076] Among them, (x) + = max{x,0}, note that if B(t k )-τ k<0, it means the buffer is exhausted and a rebuffering event occurs; (x) + is a function that takes x when x is greater than 0 and takes 0 when x is less than 0; τ k represents the download time of the kth video block; δ t The round-trip time (RTT) refers to the time from when the sender starts sending data to when the sender receives confirmation from the receiver (the receiver sends confirmation immediately after receiving the data).

[0077] S34. According to step S32, the throughput of the kth block is obtained

[0078] S35. According to step S31, the video content features are obtained: VMAF of the next block and the size of the next block; according to steps S32 and S33, the video playback features are obtained: download time of the past block, buffer level, VMAF of the previous block, and the number of remaining blocks. Therefore:

[0079]

[0080] Where T represents the throughput of the past video chunks, τ represents the download time of the previous few video chunks, q represents the VMAF of the previous video chunk, and r represents the number of remaining video chunks; represents the mapping of discrete Fourier transform, is the discrete Fourier transform of the VMAF selected over the past block.

[0081] S4. According to the strategy π(S k ;θ) Select code rate a k , calculate expert actions

[0082]

[0083] π(S k ; θ) The strategy is that the agent interacts with the environment, performs different actions, obtains feedback (rewards), and continuously updates θ through the reinforcement learning algorithm. θ is the parameter in the network. MPC is used as an expert algorithm. MPC calculates the next action based on the future network state, so an approximate optimal solution can be obtained. Among them, S k is the state, MPC(S k ) is to choose the next action based on the current status (knowing the future network status at this time).

[0084] The expert action is the bitrate combination that maximizes the user experience quality (QOE) of the next N video blocks, where QOE is defined as:

[0085]

[0086] Where N is the total number of video blocks, R n Indicates the bit rate selected for the nth video block, VMAF(R n ) is a function that maps the bitrate to its corresponding VMAF; T n Indicates the duration of the freeze caused by selecting the nth video block; [VMAF(R n+1 )-VMAF(R n )] + VMAF captures the quality improvement when switching from lower video quality to higher video quality, while n+1 )-VMAF(R n )] - It indicates a quality degradation in the opposite direction; the coefficients α, β, γ, and δ are the weights assigned to each item, reflecting the user's preference or aversion to different aspects of the streaming media viewing experience.

[0087] S5. Expert Action and state S k Stored in experience pool C:

[0088] S6. Select a batch of training samples from the experience pool C With the help of expert samples, the model can converge quickly. Here, the imitation learning method is used for training.

[0089] S7. Update the policy network model π according to the duration L of the block θ :

[0090]

[0091] where π(s,a;θ) is the network strategy, is the true probability vector of the expert's actions, H(π(s; θ)) represents the entropy of the strategy, and α is a hyperparameter that controls the scope of model exploration. π(s, a; θ) is the optimal action selection for different states learned by the agent through training, and the strategy π(s, a; θ) is gradually adjusted until it converges to the optimal strategy or an approximate optimal strategy.

[0092] S8. According to S k and a k Get the next state S k+1 , and input into the network model π θ Based on the constructed virtual environment, interact with the environment to get the next state. According to the selected bit rate a k , the environment starts downloading the corresponding video block, and thus obtains the video content features, the network features of the past blocks, and the video playback features, that is, Sk . And by selecting the bit rate a corresponding to the next video block k+1 , in order to obtain the video content features of the next video block, the network features of the previous block, and the video playback features, that is, S k+1 .

[0093] S9. Repeat steps S4 to S8 until convergence.

[0094] The convergence process includes: finding the optimal strategy by minimizing the expected loss As shown below:

[0095]

[0096] Until Strategy Convergence; among them, represents the expert strategy, T is the strategy set, and d is the strategy The state distribution under Represents the loss of the algorithmic strategy relative to the expert strategy.

[0097] The above steps S1 to S9 complete the construction of a video stream transmission system and the training process of an adaptive bit rate strategy.

[0098] The robust adaptive bitrate method based on time-frequency domain feature enhancement of the present invention can be applied to streaming video services. Application: In video on demand (VoD) and real-time streaming services, by adopting the adaptive bitrate algorithm of the present invention, buffering events can be effectively reduced and the user viewing experience can be optimized. Applicable to large-scale video streaming platforms such as YouTube and Netflix. Technical value: It enhances the robustness of video quality under network fluctuations and improves bandwidth utilization.

[0099] The robust adaptive bitrate method based on time-frequency domain feature enhancement of the present invention can also be applied to mobile video streaming. Application mode: It is suitable for video streaming applications on mobile terminals (such as smart phones and tablets). It optimizes the video bitrate through an algorithm based on time-frequency domain features to reduce delays and rebuffering in mobile networks. Technical value: It improves the video viewing experience under 4G / 5G mobile network conditions and improves bandwidth utilization and terminal energy efficiency.

[0100] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.

Claims

1. A robust adaptive bit rate method based on time-frequency domain feature enhancement, characterized in that: The following steps are involved: S1. Construct a video streaming transmission model on the server and the client, and in the server, divide a source video into K blocks, each block lasts for L seconds, and each block is encoded into n video blocks with different bit rates; S2. Construct a dual-branch network model, which includes a frequency domain network branch and a time domain network branch to extract features in the frequency domain and time domain respectively; S3. In the client, according to the video streaming model constructed in step S1, the video content features, the network features of the past blocks, and the video playback features are obtained as the input state S k , then the state S k Input to the network model π θ ; S4. According to the strategy π(S k ;θ) Select code rate a k , calculate expert actions S5. Expert Action and state S k Stored in experience pool C; S6. Select a batch of training samples from the experience pool C Conduct training; S7. Update the policy network model π according to the duration L of the block θ ; S8. According to S k and a k Get the next state S k+1 , and input into the network model π θ ; S9. Repeat steps S4 to S8 until convergence.

2. The robust adaptive bit rate method based on time-frequency domain feature enhancement according to claim 1, characterized in that: The step S2 comprises the following steps: S21. The time domain network branch includes a network layer, four one-dimensional convolutional layers and three linear layers. The time series information is processed through the convolutional layer and the linear layer. The processed time series information is connected through the network layer and then passed through the gated neural unit to extract the time domain features. S22. The frequency domain network branch contains the buffer level of the past block, the VMAF of the previous block, and the number sequence information of the remaining blocks. The sequence information is discrete Fourier transformed to obtain the real part and the imaginary part. The features obtained by the real and imaginary two-channel one-dimensional convolution layer are connected and then passed through the gated neural unit to extract the frequency domain features. S23. The frequency domain and time domain features extracted by the two network branches are connected and then used again through the gated neural unit to extract long sequence dependencies. The features are then output through the linear layer, and a probability distribution is generated through Softmax.

3. The robust adaptive bit rate method based on time-frequency domain feature enhancement according to claim 1, characterized in that: The step S3 comprises the following steps: S31. The size and VMAF of the kth block can be obtained through HTTP request, using D k ={d1(R1),d2(R2),…,d K (R K )} and Q k ={q1(R1),q2(R2),…,q K (R K )} means, d k (R k ) indicates the code rate is R k The kth video block size, q k (R k ) is the code rate R k The VMAF score of the kth video block; S32. The coding rate set is expressed as The bit rate of the kth video block is denoted by R k express, Get the download time of the kth video chunk: Among them, d k (R k ) indicates the code rate is R k The kth video block size, t i and t i+1 Respectively represent the time when downloading starts and ends, c t represents the downlink bandwidth, δ t Indicates the round trip time RTT; S33. Video chunks are downloaded to the playback buffer, which contains the video chunks that have been downloaded but not yet viewed; let B(t)∈[0,B max ] represents the occupancy level of the playback buffer at time t, let B k =B(t k ) represents the occupancy rate of the playback buffer when the kth block starts to be downloaded. The dynamic change of the buffer is expressed as: B k =((B(t k )-t k ) + +L-d t ) + Among them, (x) + =max{x,0}, if B(t k )-τ k <0, indicating that the buffer is exhausted and a rebuffering event occurs; τ k represents the download time of the kth video block; δ t Indicates the round trip time RTT; S34. According to step S32, the throughput of the kth block is obtained S35. According to step S31, the video content features are obtained: VMAF of the next block and the size of the next block; according to steps S32 and S33, the video playback features are obtained: download time of the past block, buffer level, VMAF of the previous block, and the number of remaining blocks; therefore: Where T represents the throughput of the past video chunks, τ represents the download time of the previous few video chunks, q represents the VMAF of the previous video chunk, and r represents the number of remaining video chunks; represents the mapping of discrete Fourier transform, is the discrete Fourier transform of the VMAF selected over the past block.

4. The robust adaptive bit rate method based on time-frequency domain feature enhancement according to claim 1, characterized in that: The step S4 specifically includes: According to the strategy π(S k ;θ) Select code rate a k , calculate the expert action: The expert action is the bitrate combination that maximizes the user experience quality (QOE) of the next N video blocks, where QOE is defined as: Where N is the total number of video blocks, R n Indicates the bit rate selected for the nth video block, VMAF(R n ) is a function that maps the bitrate to its corresponding VMAF; T n Indicates the duration of the freeze caused by selecting the nth video block; [VMAF(R n+1 )-VMAF(R n )] + VMAF captures the quality improvement when switching from lower video quality to higher video quality, while n+1 )-VMAF(R n )] - It indicates a quality degradation in the opposite direction; the coefficients α, β, γ, and δ are the weights assigned to each item, reflecting the user's preference or aversion to different aspects of the streaming media viewing experience.

5. The robust adaptive bit rate method based on time-frequency domain feature enhancement according to claim 1, characterized in that: The specific operation process of step S5 is as follows: Expert Action and state S k Stored in experience pool C:

6. The robust adaptive bit rate method based on time-frequency domain feature enhancement according to claim 1, characterized in that: The step S7 specifically includes: Update the policy network π according to L θ : where π(s,a;θ) is the network strategy, is the true probability vector of the expert’s actions, H(π(s;θ)) represents the entropy of the policy, and α is a hyperparameter that controls the scope of model exploration.

7. The robust adaptive bit rate method based on time-frequency domain feature enhancement according to claim 1, characterized in that: The convergence process of step S9 includes: Find the optimal strategy by minimizing expected loss Until Strategy Convergence; among them, represents the expert strategy, T is the strategy set, and d is the strategy The state distribution under Represents the loss of the algorithmic strategy relative to the expert strategy.

Citation Information

Cited By

  • Video streaming media bit rate self-adaption method and system based on expert guidance

    CN120583080A