Adaptive reward-driven video transport stream control method and device
By deploying an adaptive reward-driven video transmission stream control method in a video transmission environment, and optimizing video code rate using reinforcement learning and user feedback, the problem that the existing technology cannot meet the differentiated needs of users is solved, and efficient resource allocation and excellent user experience of the video system are achieved.
Patent Information
- Application Number
- CN202510051126.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-13
AI Technical Summary
The existing technology cannot effectively meet the differentiated needs of users to control video transmission code rate, especially when facing network heterogeneity and user preference differences, unified optimization goals are difficult to meet the needs of different users.
Adaptive reward-driven video transmission stream control method is adopted, and the code rate control module is deployed in the video transmission environment of the WebRTC framework and the FFmpeg tool. The video bit rate adaptive model is used to select the bit rate according to the transmission layer and application layer indicators of real-time video, and the model parameters are updated through the reward value feedback from users.
The video system is implemented to adjust transmission strategies in a timely manner according to the network environment and user preferences, optimize resource allocation, improve the global service quality of network services, and ensure that users have a good viewing experience.
Smart Images

Figure CN120075486A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video transmission, and in particular, to an adaptive reward-driven video transmission stream control method and device. Background Art
[0002] With the progress of network and communication technologies and the rapid development of the Internet of Things and mobile Internet, video Internet of Things has become a new trend in digital services and industrial development. Different from traditional streaming media services, new video network services such as digital home, smart community, digital village, digital governance, etc. scenarios emphasize more on pushing real-time video streams to the cloud to provide low-latency and real-time interactive video services for end-users.
[0003] In 2024, video traffic accounted for more than 72% of the total Internet traffic. Among them, real-time video streams have become one of the mainstream technologies of network video traffic, accounting for 40.24%. Video Quality of Experience (QoE) is a key factor in measuring the audience's perceived feedback and overall viewing satisfaction. At present, the QoE of real-time video streams significantly depends on the effectiveness of the bitrate control algorithm. An excellent transmission algorithm can quickly perceive network changes, timely correct the transmission method or adjust the video stream bitrate to avoid causing a decline in QoE, thereby providing users with a good viewing experience. In recent years, traditional rule-based algorithms (such as GCC, BBR, etc.) have been rapidly transformed into learning-based artificial intelligence algorithms. However, in the context of industrial video transmission, network heterogeneity has brought huge challenges to learning-based bitrate control algorithms. On the one hand, in order to ensure the model effect in different scenarios, a large number of hyperparameters and model tests and adjustments are required, and the implementation cost has increased significantly; on the other hand, due to the differences in user groups, the preferences for videos among different groups are significantly different, and a unified optimization goal cannot well meet the different user needs. Therefore, there is an urgent need for a new video transmission stream control scheme. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide an adaptive reward-driven video transmission stream control method and device to eliminate or improve one or more defects existing in the prior art and solve the problem that the prior art cannot meet the differential requirements of users for video transmission bitrate control.
[0005] One aspect of the present invention provides an adaptive reward-driven video transmission stream control method, which is deployed and runs in a bitrate control module in a video transmission environment based on the WebRTC framework and FFmpeg tool. The method includes the following steps:
[0006] Obtain the transport layer and application layer metrics of the real-time video, and the metrics include: round-trip delay, packet loss rate, throughput, jitter, and original bitrate;
[0007] Construct the metrics of the real-time video for a preset duration before the current time as the state space, and construct multiple optional bitrates as the action space; based on a video bitrate adaptation model pre-trained by reinforcement learning, use the parameters of the state space as the input and output the bitrate for the next time period selected in the action space; the bitrate is used to encode the real-time video and transmit it to the user side through the network;
[0008] Receive the reward value feedback from the user side, and maximize the reward value to update the parameters of the video bitrate adaptation model; the reward value is obtained by weighting the model score according to the real user evaluation interval time; the model score is the score of the pre-trained reward model for the real-time video received by the user side, and the reward model is a generator trained based on adversarial learning for fitting the real user score; the reward value is obtained by exponentially weighting the historical reward values of multiple past cycles, and linearly scale the reward value outside the set value range.
[0009] In some embodiments, the video bitrate adaptation model is pre-trained based on historical data, including the following steps:
[0010] Slice the historical videos received by the user side, and obtain the metrics corresponding to each slice, and construct them into a real-time video quality data set;
[0011] Construct the metrics of each slice in the real-time video quality data set and the preset duration before it as the state space, and construct multiple optional bitrates as the action space; obtain a preset recurrent neural network as the video bitrate adaptation model, use the parameters of the state space as the input and output the bitrate for the next time period selected in the action space; calculate the total reward value of the selected bitrate using a preset rule, update the parameters of the video bitrate adaptation model and the evaluation network by maximizing the total reward value, and use it as the video bitrate adaptation model;
[0012] The calculation formula of the total reward value is:
[0013] Reward=(B - a×loss) / RTT
[0014] Wherein, B represents the bitrate, a is the weight coefficient, loss represents the packet loss rate, and RTT represents the round-trip delay.
[0015] In some embodiments, the pre-training steps of the reward model include:
[0016] Obtain a training sample set, which contains the state parameters of multiple video slices and the artificial scores based on human feedback corresponding to each video slice;
[0017] A generator is constructed based on a recurrent neural network. The generator takes the state parameters of the video slices as input and outputs prediction scores that mimic human feedback. A discriminator is constructed based on a convolutional neural network, which takes the artificial scores or the prediction scores as input and outputs a judgment result on whether it belongs to human feedback. The generator and the discriminator are alternately trained in the form of adversarial learning using the training sample set, and an adversarial loss is constructed to update the parameters of the generator and the discriminator. The updated generator is used as the reward model.
[0018] In some embodiments, the adversarial loss is composed of a generator loss and a discriminator loss;
[0019] The calculation formula for the generator loss is:
[0020]
[0021] where s represents the state parameters of the video slices input to the generator, G(S) represents the probability of outputting each of the prediction scores, x represents the prediction scores, and D(x) represents the confidence in judging whether it belongs to human feedback;
[0022] The calculation formula for the discriminator loss is:
[0023]
[0024] where α and β are weight coefficients, x represents the input to the discriminator, y represents the label, and R h represents the artificial scores.
[0025] In some embodiments, the reward value is obtained by weighting the model scores with reference to the time interval of real user evaluations, including:
[0026] Obtain the time interval T of real user evaluations, and obtain the timestamps of each of the model scores generated by the reward model; then the calculation formula for the reward value is:
[0027]
[0028] where R i represents the i-th model score, t i represents the timestamp of the i-th model score, R a represents the reward value, and n represents the number of model scores.
[0029] In some embodiments, the reward value is obtained by exponentially weighting the historical reward values of multiple past cycles, and the calculation formula is:
[0030]
[0031] Among them, represents the reward value obtained through exponential weighted calculation, and R pi represents the historical reward value of the i-th cycle, and m represents the total number of cycles.
[0032] In some embodiments, the method further includes:
[0033] Performing fractional interval unification correction on the manual scoring, and the expression is:
[0034]
[0035] Among them, represents the median of the given interval, up represents the upper limit of the given interval, low represents the lower limit of the given interval, and R h represents the manual scoring of the current user, and Mode(R H ) represents the median of the manual scoring generated by multiple users; Max(R H ) represents the maximum value in the manual scoring, and Min(R H ) represents the minimum value in the manual scoring.
[0036] On the other hand, the present invention also provides an adaptive reward-driven video transmission stream control device, including a processor, a memory, and a computer program / instruction stored on the memory. The processor is used to execute the computer program / instruction, and when the computer program / instruction is executed, the device implements the steps of the above method.
[0037] On the other hand, the present invention also provides a computer-readable storage medium, on which a computer program / instruction is stored. When the computer program / instruction is executed by a processor, the steps of the above method are implemented.
[0038] On the other hand, the present invention also provides a computer program product, including a computer program / instruction. When the computer program / instruction is executed by a processor, the steps of the above method are implemented.
[0039] The beneficial effects of the present invention are at least:
[0040] For the adaptive reward-driven video transmission stream control method and device of the present invention, in the context of industrial video transmission, the video bitrate adaptation model introduces the form of reinforcement learning to select the bitrate according to the index data of real-time video in the past multiple cycles. The reward model fits the real user's evaluation of the real-time video to establish a reward value to guide the update and optimization of the video bitrate adaptation model. This enables the video system to adjust the transmission strategy in a timely manner according to the network environment and user preferences. While maintaining a good viewing experience, it optimizes resource allocation and improves the global service quality of the network service.
[0041] Additional advantages, objects, and features of the present invention will be partly set forth in the description which follows, and will partly become obvious to those of ordinary skill in the art upon examination of the following, or may be learned by practice of the present invention. The objects and other advantages of the present invention may be realized and attained by the structure particularly pointed out in the specification and the drawings.
[0042] Those skilled in the art will understand that the objects and advantages that can be achieved by the present invention are not limited to those specifically described above, and the above and other objects that the present invention can achieve will be more clearly understood according to the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention. In the drawings:
[0044] Figure 1 is a schematic flow chart of an adaptive reward-driven video transmission stream control method according to an embodiment of the present invention.
[0045] Figure 2 is a schematic logic diagram of parameter update of a video bitrate adaptation model based on reinforcement learning in an adaptive reward-driven video transmission stream control method according to an embodiment of the present invention.
[0046] Figure 3 is a schematic logic diagram of training a reward model based on adversarial learning in an adaptive reward-driven video transmission stream control method according to an embodiment of the present invention.
[0047] Figure 4 is a schematic framework diagram of an adaptive reward-driven video transmission stream control method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] To make the objects, technical solutions, and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with the embodiments and the drawings. Herein, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.
[0049] Here, it should also be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution of the present invention are shown in the drawings, and other details less related to the present invention are omitted.
[0050] It should be emphasized that the term "comprising / including" when used herein refers to the presence of features, elements, steps, or components, but does not exclude the presence or addition of one or more other features, elements, steps, or components.
[0051] Existing real-time video bitrate control algorithms can be divided into two categories: (i) traditional methods based on heuristic rules, which use manually written state transition rules to regulate the bitrate; (ii) machine learning-based algorithms, which use the network state change information within a period of time as input to infer the optimal encoding bitrate for the next moment under the current network conditions. Although the heuristic method has high operating efficiency, it is often unable to meet the QoE requirements of users due to fixed rules and robustness limitations. Therefore, various methods that directly optimize QoE have been proposed, and learning-based algorithms have achieved good results. However, due to the generalization limitations of neural networks, when the distribution of online input metrics changes, the performance of offline models is often unsatisfactory. Although some studies (including PCC Vivace, OnRL, Loki, etc.) have made the model adapt to the new network environment by introducing online learning methods, since the loss function / reward function they use is a fixed formula rule, using the same guiding function in highly heterogeneous online scenarios is doomed to affect the online performance of the model and cannot obtain the optimal solution.
[0052] On the other hand, although the formula rules adopted by current adaptive transmission algorithms comprehensively consider metrics such as throughput, packet loss, latency, and bitrate, this technology metric-based control method usually lacks consideration of human perception and subjective differences, does not always correctly reflect the actual feelings of users, and cannot adapt to different viewing preferences of users. Resources cannot be personalized and allocated on demand, which greatly limits the overall quality of QoE of network services. Therefore, designing an adaptive transmission algorithm based on perception is of great significance for real-time video services.
[0053] The present invention provides an adaptive reward-driven video transmission stream control method, which is deployed and runs in a bitrate control module in a video transmission environment based on the WebRTC framework and FFmpeg tools, as Figure 1 shown. This method includes the following steps S101 to S103:
[0054] Step S101: Obtain the transport layer and application layer metrics of the real-time video, and the metrics include: round-trip delay, packet loss rate, throughput, jitter, and original bitrate.
[0055] Step S102: Construct the metrics of the real-time video within a set duration before the current time into a state space, and construct multiple optional bitrates into an action space; based on a video bitrate adaptation model pre-trained by reinforcement learning, use the parameters of the state space as input and output the bitrate for the next time period selected in the action space; the bitrate is used to encode the real-time video and transmit it to the user side through the network.
[0056] Step S103: Receive the reward value feedback from the client, and maximize the reward value to update the parameters of the video bitrate adaptation model; the reward value is obtained by weighting the model score with reference to the real user evaluation interval time; the model score is the score of the pre-trained reward model for the real-time video received by the client, and the reward model is a generator trained based on adversarial learning for fitting the real user score; the reward value is obtained by exponentially weighting the historical reward values of multiple past cycles, and linearly scales the reward values outside the set value range.
[0057] WebRTC (Web Real-Time Communication) is an open-source framework and technical standard aimed at enabling real-time audio and video communication and data transmission between browsers and mobile devices through standardized APIs without relying on plugins or additional software support. WebRTC is jointly developed by the World Wide Web Consortium (W3C) and the Internet Engineering Task Force (IETF), supports peer-to-peer connections, ensures low-latency and high-quality real-time communication, and is widely used in fields such as video conferencing, online education, real-time chat, streaming media, gaming, and the Internet of Things. Its main functions include: Audio and video capture and processing. WebRTC can directly capture audio and video data from devices (such as cameras, microphones, etc.) and compress them through built-in codecs (such as VP8, VP9, H.264, etc.). Real-time transmission and network optimization. It solves the network penetration problem through NAT (Network Address Translation) and STUN / TURN server technologies, supports reliable data transmission, and can guarantee communication even in complex network environments. Data Channel. In addition to audio and video communication, WebRTC also provides the function of real-time transmission of any data, which is suitable for scenarios such as text, file sharing, or online collaboration. Encryption and security. It encrypts audio and video data using SRTP (Secure Real-Time Transport Protocol) and also supports DTLS (Datagram Transport Layer Security Protocol) to ensure the security of the communication process.
[0058] The peer-to-peer architecture of WebRTC allows applications to directly establish connections between users, thus significantly reducing latency and server load. Developers can build real-time communication applications based on WebRTC without server forwarding, and at the same time reduce operating costs through a distributed architecture.
[0059] FFmpeg is a powerful and widely used multimedia processing tool, which is a cross-platform and open-source solution for processing audio, video, and multimedia file recording, conversion, and streaming. It consists of a set of libraries and command-line tools, supporting almost all mainstream audio and video formats as well as codecs, such as H.264, H.265, VP8, VP9, AAC, MP3, etc. FFmpeg provides rich functions, including file format conversion, audio and video encoding and decoding, video editing, audio processing, subtitle embedding, filter application, screen recording, live streaming, etc., and can meet various complex multimedia processing requirements.
[0060] In this application, WebRTC is used as the base, and the Ffmpeg tool is used as the encoding module to encode and transmit real-time video. The bitrate control module selects the corresponding bitrate for video transmission according to the metric status of the real-time video.
[0061] In step S101, the round-trip time (RTT) and packet loss rate are important metrics at the transport layer, which directly affect the stability and continuity of video transmission. If the RTT is too high or the packet loss rate is large, it may be necessary to reduce the bitrate to reduce the impact of data loss and latency. Throughput is also an important transport layer metric, which reflects the actual bandwidth capacity of the network. When the network bandwidth is insufficient, appropriately reducing the bitrate can avoid network congestion and thus improve the efficiency of video transmission. Jitter is also a factor that cannot be ignored. It will cause inconsistent arrival times of video frames and affect the smoothness of playback. By adjusting the buffering strategy and bitrate, the problems caused by jitter can be effectively alleviated. At the application layer, the frame rate and inter-frame delay are equally important. A high frame rate usually means a smoother video experience, but if the inter-frame delay is too large, it may cause playback stuttering. Therefore, when selecting the bitrate, it is necessary to balance the frame rate and inter-frame delay to ensure the smoothness and real-time performance of the video. In addition, the original bitrate itself is also a key metric, which determines the clarity and data volume of the video. Dynamically adjusting the bitrate according to the network conditions and application requirements can achieve the best transmission effect. In some other embodiments, other metrics can also be selected according to the actual application scenario to meet the requirements.
[0062] In step S102, reinforcement learning is introduced to select the bitrate during the video transmission process. First, key metrics (such as round-trip delay, packet loss rate, throughput, jitter, original bitrate, frame rate, and inter-frame delay, etc.) included in a period of time before the current time of the real-time video are constructed into a state space, and these metrics can comprehensively reflect the real-time state of the transmission system. Then, multiple optional bitrate values (such as 1000 kbps, 2000 kbps, etc.) are constructed into an action space, and these bitrates are different encoding parameters that the video encoder can choose. Based on these definitions, a reinforcement learning model is designed. Through pre-training on historical data, it can learn the dynamic characteristics and optimal decision-making strategies of the system from the input of the state space.
[0063] In practical applications, the working process of the model is as follows: Whenever the system needs to select a bitrate for the next time period, the model takes the parameters in the state space as input, and the trained neural network outputs an optimal bitrate action. This action is the best bitrate selection evaluated by the model in the current state and is used to encode the real-time video. After encoding, the video is transmitted to the user side through the network, and the feedback information at the user side (such as user evaluation, playback stuttering situation, video quality score, etc.) will be recorded and used to further update the parameters of the reinforcement learning model to continuously optimize its decision-making ability.
[0064] In this way, the video bitrate adaptive model can dynamically adapt to changes in the network environment, balance the relationship between video quality and smoothness, reduce stuttering, and improve the user experience. At the same time, through offline pre-training, the model can converge quickly and provide stable and reliable performance when going online. This method not only effectively solves the limitations of traditional fixed bitrate or rule-based strategies but also significantly improves the overall efficiency and user satisfaction of the system through the adaptability of reinforcement learning.
[0065] In some embodiments, the video bitrate adaptive model is pre-trained based on historical data. As Figure 2 shown, it includes the following steps S201 - S202:
[0066] Step S201: Slice the historical videos received at the user side, obtain the metrics corresponding to each slice, and construct them into a real-time video quality dataset;
[0067] Step S202: Construct the metrics of each slice in the real-time video quality dataset and the metrics of the set duration before it into a state space, and construct multiple optional bitrates into an action space; obtain a preset recurrent neural network as the video bitrate adaptive model, take the parameters of the state space as input, and output the bitrate for the next time period selected in the action space; calculate the total reward value of the selected bitrate using a preset rule, update the parameters of the video bitrate adaptive model and the evaluation network by maximizing the total reward value, and use it as the video bitrate adaptive model.
[0068] The calculation formula for the total reward value is as follows:
[0069] Reward=(B - a×loss) / RTT
[0070] Wherein, B represents the bit rate, a is the weight coefficient, loss represents the packet loss rate, and RTT represents the round-trip delay.
[0071] In step S103, the reward value feedback from the user side is introduced to further optimize the parameter update of the model, so as to maximize the user experience and improve the system performance. Specifically, the generation and use of the reward value involve multiple key links, aiming to construct an accurate and dynamic evaluation mechanism to guide the learning process of the video bitrate adaptation model.
[0072] When the user side plays the video, a feedback message will be generated based on the user's feelings, that is, an evaluation of the video quality, indicating the impact of the currently selected bit rate on the user experience. Based on the idea of reinforcement learning from human feedback (RLHF), this evaluation can be used as a reward value to feedback to the parameter update and optimization of the video bitrate adaptation model based on reinforcement learning. Since the evaluation process of humans is inefficient and highly subjective and is affected by the external environment, in this application, a pre-trained reward model is used to fit the real user ratings. This model is trained through generative adversarial learning (GAN). The reward model includes a generator and a discriminator, where the generator fits the distribution of real user ratings, and the discriminator judges the authenticity of the ratings. Through the optimization of adversarial learning, the generator can give a video rating close to the real user perception as the model rating.
[0073] The reward model can take multiple metric parameters of the real-time video received by the receiving end as inputs and simulate human senses for scoring to achieve fast, objective, and efficient scoring. The receiving end continuously scores the received real-time video according to slices as the reward value. Since the feedback granularity of real human scoring (usually 2s) is much lower than the millisecond-level scoring generated by the reward model (usually 50ms to 200ms), in order to avoid problems such as dilution of network state fluctuations and inability to capture the panoramic changes of the network condition caused by the direct averaging method, the model score is weighted according to the real user evaluation interval time to obtain the reward value.
[0074] In some embodiments, the reward value is obtained by weighting the model score with reference to the real user evaluation interval time, including: obtaining the real user evaluation interval time T, and obtaining the timestamps of each of the model scores generated by the reward model; then the calculation formula for the reward value is:
[0075]
[0076] Among them, R i represents the i-th model score, and t i represents the timestamp of the i-th model score, and R a represents the reward value, and n represents the number of model scores.
[0077] The reward value not only refers to the current user feedback but also combines the historical reward values of multiple past cycles for exponential weighting. In this way, the system can smooth the reward fluctuations, capture the changing trends of the user experience on a long time scale, and at the same time quickly respond to recent changes.
[0078] In some embodiments, the reward value is obtained by exponential weighting calculation from the historical reward values of multiple past cycles, and the calculation formula is:
[0079]
[0080] Among them, represents the reward value obtained by exponential weighting calculation, and R pi represents the historical reward value of the i-th cycle, and m represents the total number of cycles.
[0081] To avoid the interference of the reward value being too large or too small on the learning process, an effective range of the reward value is set (such as [0, 1] or [-1, 1]). For the reward value outside the range, a linear scaling method is adopted to map it to the set range, so as to maintain the stability and controllability of the reward signal.
[0082] In some embodiments, the method further includes:
[0083] Performing a unified correction of the score interval for the manual score, and the expression is:
[0084]
[0085] Among them, represents the median of the given interval, up represents the upper limit of the given interval, low represents the lower limit of the given interval, and R h represents the manual score of the current user, and Mode(R H ) represents the median of the manual scores generated by multiple users; Max(R H ) represents the maximum value in the manual scores, and Min(R H ) represents the minimum value in the manual scores.
[0086] In some embodiments, as Figure 3 shown, the pre-training steps of the reward model include S301 to S302:
[0087] Step S301: Obtain a training sample set, where the training sample set includes the state parameters of multiple video slices and the human feedback-based artificial scores corresponding to each video slice.
[0088] Step S302: Construct a generator based on a recurrent neural network. The generator takes the state parameters of video slices as inputs and outputs predicted scores that mimic human feedback. Construct a discriminator based on a convolutional neural network, which takes artificial scores or predicted scores as inputs and outputs a judgment result on whether it belongs to human feedback. Use the training sample set to alternately train the generator and the discriminator in the form of adversarial learning, construct an adversarial loss to update the parameters of the generator and the discriminator, and use the updated generator as a reward model.
[0089] In some embodiments, the adversarial loss is composed of a generator loss and a discriminator loss;
[0090] The calculation formula for the generator loss is:
[0091]
[0092] where s represents the state parameters of the video slices input to the generator, G(S) represents the probability of outputting each predicted score, x represents the predicted score, and D(x) represents the confidence level of judging whether it belongs to human feedback;
[0093] The calculation formula for the discriminator loss is:
[0094]
[0095] where α and β are weight coefficients, x represents the input to the discriminator, y represents the label, and R h represents the artificial score.
[0096] On the other hand, the present invention also provides an adaptive reward-driven video transmission stream control device, including a processor, a memory, and a computer program / instructions stored in the memory. The processor is used to execute the computer program / instructions, and when the computer program / instructions are executed, the device implements the steps of the above method.
[0097] On the other hand, the present invention also provides a computer-readable storage medium, on which computer program / instructions are stored. When the computer program / instructions are executed by a processor, the steps of the above method are implemented.
[0098] On the other hand, the present invention also provides a computer program product, including computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the above method are implemented.
[0099] The present invention is described below in conjunction with a specific embodiment:
[0100] This embodiment provides an adaptive reward-driven video transmission flow control method. In a real-time video session, both parties capture images, encode them, and send data packets to each other. Due to fluctuations in network bandwidth, the sender's bitrate control algorithm will dynamically adjust its sending bitrate to control the amount of data transmitted, thereby avoiding network congestion and meeting the user's demand for video quality at the same time.
[0101] In the training framework of this embodiment, the input of the reinforcement learning model is the core metrics in the video transmission process, including throughput, bitrate, transmission delay, etc. The reinforcement learning model inputs a sequence of state metrics for a period of time, generates a predicted bitrate, and adjusts the transmission load. The adaptive reward model (RM) generates a score for the decision-making effect of the reinforcement learning model, guides the decision-making of the neural network model, and the goal is to maximize the cumulative reward of the reinforcement learning, that is, to achieve the best viewing experience. This embodiment uses a human feedback-based RM to replace the formula reward and supports the online training of the reinforcement learning model, which greatly guarantees the generalization ability of the model under various network conditions and introduces the ability of personalized adaptation and low-cost adjustment to the bitrate control algorithm. Specifically, only by collecting the transmission content and transmission metrics in a specific scenario and adding user feedback annotations, the RM effect can be quickly adjusted offline to achieve the effect of optimizing specific network scenarios and personalizing specific users.
[0102] This embodiment designs a viewing-based scoring fitting model (i.e., the reward model, RM) for real-time video transmission. The data base of the scheme is a real-time video QoE data set captured in a real scenario and scored by real people, and the training framework is GAN.
[0103] During the training process, this embodiment attempts and compares the performance of human feedback reward models based on linear layers, CNNs, RNNs, and GANs, and finds that the RNN can deduce the feedback trend through the state sequence, but cannot identify the changes in human feedback caused by slight fluctuations in the network state. Using the CNN model is prone to overfitting the details of network state changes, resulting in suboptimal performance. Finally, by introducing the GAN method to synergize the capabilities of the two models, the best fitting effect is obtained. Specifically, the RNN is used as the generator to imitate human feedback and output a reward value that the discriminator cannot distinguish; the CNN is used as the discriminator to evaluate whether the reward value corresponding to the state sequence is human feedback and output a confidence level. The advantage of the RNN model is that it can quickly determine the approximate range of human feedback according to the network state. Subsequently, it gradually refines its output through the evaluation of the discriminator to approach the actual human feedback.
[0104] Among them, the loss function of the generator is composed of the difference between the probability distributions generated by the generator and the discriminator. The discriminator provides confidence for each reward value, thereby generating the probability distribution of the rewards. The formula is as follows, where G and D represent the generator and the discriminator respectively:
[0105] The calculation formula for the generator loss is:
[0106]
[0107] Among them, s represents the state parameter of the video slice input to the generator, G(S) represents the probability of outputting each predicted score, x represents the predicted score, and D(x) represents the confidence in judging whether it belongs to human feedback;
[0108] The input of the discriminator can be divided into two parts: actual human feedback and the result of the generator. The discriminator should be able to distinguish between them. In particular, when the input is human feedback, the output should be as close to 1 as possible, and when the input is the result of the generator, the output should be as close to 0 as possible. The calculation formula for the discriminator loss is:
[0109]
[0110] Among them, α and β are weight coefficients, x represents the input of the discriminator, y represents the label, and R h represents the manual score. The first item calculates the difference between the output result and the label, and the second item is to make the model output different results for human feedback and the result of the generator, imposing a penalty to make the confidence of the result of the generator close to 0.
[0111] In order to introduce a personalized bitrate control algorithm based on real perception in a real-time transmission framework, it is a necessary requirement to collect and construct a real perception dataset. In this embodiment, based on the widely used WebRTC framework in the industry, a set of real-time video transmission and scoring collection tools are constructed in combination with the FFmpeg tool.
[0112] Specifically, in this embodiment, WebRTC is used as the base, and the following functions are introduced: (1) The input source supports FFmpeg access, which solves problems such as single sampling specifications caused by camera capture, lack of high-definition scenes, and monotonous video content, ensuring uniform sample distribution in the dataset to improve the generalization ability of the model. (2) Precise metric collection module: The storage granularity of the built-in metric calculation module in WebRTC is too coarse. On the other hand, the metrics need to be aligned with video slices to ensure the corresponding relationship. Therefore, the metric collection module is reformed to be able to collect fine-grained transport layer and application layer metrics, including round-trip delay, packet loss rate, throughput, jitter, bit rate, frame rate, inter-frame delay, etc. On this basis, video capture is performed at the receiving end, and long video slices are sliced, the transmission metrics corresponding to the video content are obtained, played, and user scores are collected to construct an original real-time video QoE dataset. The overall process is as Figure 4 shown.
[0113] After the original dataset is collected, since the feedback granularity of human scoring (usually 2s) is much lower than the millisecond-level scoring generated by the reward model (usually 50ms to 200ms). To avoid problems such as diluting network state fluctuations and being unable to capture the panoramic changes of network conditions caused by the direct averaging method, a weighted method is used to align the two scores, thereby providing a more accurate description of the network state sequence. Among them, R i represents the i-th model score, t i represents the timestamp of the i-th model score, R a represents the reward value, and n represents the number of model scores.
[0114] On the other hand, it is observed that user scores are not only affected by the current video quality but also by the difference from the historical video quality. After viewers are exposed to low-bitrate videos for a period of time, when switching to a video with a slightly higher bitrate, they tend to give higher feedback. For this characteristic, an index is set in this embodiment to monitor the playback quality within a period of time. The index approximates the historical quality of the video by the formula-based reward in the past n cycles, expressed as where, represents the reward value obtained by exponential weighted calculation, R pi represents the historical reward value in the i-th cycle, and m represents the total number of cycles. Subsequently, for the reward values outside the range, linear scaling is performed to make it more in line with the real perception.
[0115] Finally, for data from different users, the value of the user (the median will be used as based on sorting) is used to unify the score intervals. The specific rules are as follows:
[0116]
[0117] Among them, represents the median value of the given interval, up represents the upper limit of the given interval, low represents the lower limit of the given interval, and R h represents the manual score of the current user, and Mode(R H ) represents the median of the manual scores generated by multiple users; Max(R H ) represents the maximum value in the manual scores, and Min(R H ) represents the minimum value in the manual scores.
[0118] This embodiment utilizes the emerging RLHF technology to implement an adaptive reward-driven video transmission flow control method and device. Specifically, by constructing and training a scoring model that can accurately map the true visual experience through transmission indicators, guiding the reinforcement learning model for bitrate control, when the network properties change (such as Wi-Fi - Wired), there is no need for high-cost algorithm testing and hyperparameter fine-tuning, thus ensuring the transmission effect of the algorithm during industrial deployment. More importantly, it can be personalized adjusted to the viewing preferences of users through the collected user feedback, further enhancing the visual experience and improving the resource utilization efficiency. Introducing it into the reinforcement learning method with better effects in the industry, the comprehensive test shows that the performance of each index is equal to or better than the fine-tuning formula, and it is significantly improved under weak network conditions. The maximum reduction rate of the algorithm jitter rate exceeds 40% (OnRL), and the frame rate, jitter, and bitrate indicators are more stable (Loki&OnRL).
[0119] Adaptive reward-driven video transmission flow control method. The solution proposed in this embodiment does not require a large number of experiments and adjustments for the hyperparameter settings of the reward function in different network environments. Instead, it utilizes the rapidly emerging LLM fine-tuning technology RLHF, taking the true visual experience feedback as the reward function to guide the algorithm training. In different network scenarios, people's QoE evaluation criteria for videos are relatively unified, and the technical difficulty and cost of collecting visual experience data are much lower than hyperparameter fine-tuning, making it easy for industrial deployment and in-depth optimization.
[0120] Visual experience reward function fitting model based on generative adversarial network (GAN) and human scoring dataset. This embodiment proposes to use neural networks to simulate human feedback to ensure the training efficiency of the reinforcement learning model fine-tuning and control the cost of manual annotation.
[0121] This reward model combines a convolutional neural network (CNN) and a recurrent neural network (RNN) using GAN. Among them, the RNN is responsible for generating scores, and the CNN is used as a discriminator to output the confidence of the score results. The two achieve a more accurate fit through an adversarial manner and also have a certain ability to identify dirty data.
[0122] Data processing and dataset construction method for real-time video scene in terms of real-person QoE scoring. Different from the on-demand scenario, real-time video does not segment the video, and the metric granularity is usually at the millisecond level (50 ms to 200 ms). When watching, it is closer to an image, which is contrary to human perception of video. On the other hand, the relative change in video quality will also affect the perception score (typically, after a low-quality video, a high-definition video will tend to get a higher score). In data processing, the embodiments also take this into account. In view of the above problems, the embodiments propose corresponding data processing methods to make the dataset more conducive to model training.
[0123] Correspondingly, the present invention also provides a device / system, which includes a computer device. The computer device includes a processor and a memory. Computer instructions are stored in the memory, and the processor is configured to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device / system implements the steps of the method described above.
[0124] The embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the foregoing edge computing server deployment method. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the art.
[0125] In summary, for the adaptive reward-driven video transmission stream control method and device of the present invention, in the context of industrial video transmission, the video bitrate adaptation model introduces the form of reinforcement learning to select the bitrate according to the metric data of real-time video in multiple past cycles. By fitting the real user's evaluation of real-time video through a reward model to establish a reward value to guide the update and optimization of the video bitrate adaptation model. This enables the video system to adjust the transmission strategy in a timely manner according to the network environment and user preferences. While maintaining a good viewing experience, it optimizes resource allocation and improves the global service quality of the network service.
[0126] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in connection with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement it in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave on a transmission medium or a communication link.
[0127] It should be clear that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present invention is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.
[0128] In the present invention, the features described and / or illustrated for one embodiment can be used in the same or similar manner in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.
[0129] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and variations can be made to the embodiments of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An adaptive reward-driven video transmission stream control method, characterized in that: The method is deployed in a bit rate control module and runs in a video transmission environment based on a WebRTC framework and an FFmpeg tool. The method comprises the following steps: Obtaining transport layer and application layer indicators of real-time video, including round-trip delay, packet loss rate, throughput, jitter, and original bit rate; The indicator of the set time length of the real-time video before the current time is constructed as a state space, and multiple optional bit rates are constructed as an action space; based on the video bit rate adaptive model pre-trained by reinforcement learning, the parameters of the state space are used as input and the bit rate of the next time period selected in the action space is output; the bit rate is used to encode the real-time video and transmit it to the user end through the network; Receive the reward value fed back by the user end, and maximize the reward value to update the parameters of the video bitrate adaptive model; the reward value is obtained by weighting the model score with reference to the real user evaluation interval; the model score is the score of the real-time video received by the user end by the pre-trained reward model, and the reward model is a generator for fitting the user's real score obtained through adversarial learning training; the reward value is obtained by exponentially weighted calculation of historical reward values of multiple periods in the past, and the reward value outside the set value range is linearly scaled.
2. The adaptive reward-driven video transmission stream control method according to claim 1, characterized in that: The video bitrate adaptation model is pre-trained based on historical data, including the following steps: Slicing the historical video received by the user terminal, and obtaining the indicator corresponding to each slice to construct a real-time video quality data set; Constructing each slice in the real-time video quality data set and the indicator of the previously set duration thereof as a state space, and constructing a plurality of optional bit rates as an action space; Obtaining a preset recurrent neural network as a video bit rate adaptive model, taking the parameters of the state space as input and outputting the bit rate of the next time period selected in the action space; The total reward value of the selected bit rate is calculated using a preset rule, and the parameters of the video bit rate adaptation model and the evaluation network are updated by maximizing the total reward value, and the parameters are used as the video bit rate adaptation model; The calculation formula of the total reward value is: Reward=(Ba×loss) / RTT Among them, B represents the bit rate, a is the weight coefficient, loss represents the packet loss rate, and RTT represents the round-trip delay.
3. The adaptive reward-driven video transmission stream control method according to claim 1, characterized in that: The pre-training steps of the reward model include: Acquire a training sample set, wherein the training sample set includes state parameters of multiple video slices and a manual score based on human feedback corresponding to each of the video slices; A generator is constructed based on a recurrent neural network, the generator takes the state parameters of the video slice as input and outputs a predicted score that imitates human feedback; a discriminator is constructed based on a convolutional neural network, takes the manual score or the predicted score as input, and outputs a judgment result of whether it belongs to human feedback; the generator and the discriminator are alternately trained in the form of adversarial learning using the training sample set, an adversarial loss is constructed to update the parameters of the generator and the discriminator, and the updated generator is used as the reward model.
4. The adaptive reward-driven video transmission stream control method according to claim 3, characterized in that: The adversarial loss consists of a generator loss and a discriminator loss; The generator loss calculation formula is: Wherein, s represents the state parameter of the video slice input by the generator, G(S) represents the probability of outputting each predicted score, x represents the predicted score, and D(x) represents the confidence level of judging whether it belongs to human feedback; The discriminator loss is calculated as: Among them, α and β are weight coefficients, x represents the input of the discriminator, y represents the label, and R h Represents the manual rating.
5. The adaptive reward-driven video transmission stream control method according to claim 1, characterized in that: The reward value is obtained by weighting the model score with reference to the real user evaluation interval, including: Get the real user evaluation interval T, and get the timestamp of the reward model generating the model score; then the reward value is calculated as: Among them, R i represents the i-th model score, t i represents the timestamp of the i-th model score, R a represents the reward value, and n represents the number of model scores.
6. The adaptive reward-driven video transmission stream control method according to claim 1, characterized in that: The reward value is calculated by exponentially weighting the historical reward values of the past multiple cycles, and the calculation formula is: in, represents the reward value obtained by exponential weighting calculation, R pi represents the historical reward value of the i-th cycle, and m represents the total number of cycles.
7. The adaptive reward-driven video transmission stream control method according to claim 3, characterized in that: The method further comprises: The manual scoring is corrected to a uniform score interval, and the expression is: in, represents the median of a given interval, up represents the upper limit of the given interval, low represents the lower limit of the given interval, R h Indicates the manual rating of the current user, Mode(R H ) represents the median of the manual ratings generated by multiple users; Max(R H ) represents the maximum value among the manual scores, Min(R H ) represents the minimum value among the manual scores.
8. An adaptive reward-driven video transmission stream control device, comprising a processor, a memory, and a computer program / instruction stored in the memory, characterized in that: The processor is used to execute the computer program / instructions. When the computer program / instructions are executed, the device implements the steps of the method according to any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method as claimed in any one of claims 1 to 7 are implemented.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Video stream transmission control method and device, equipment and medium
CN117336521A
Reward model training method and system based on human feedback reinforcement learning
CN118095402A
Model training and application method and device for content generation and storage medium
CN118796989A
Method and apparatus for transmitting video data
KR102392383B1
Cited By
Network adaptive video conference transmission optimization method and system
CN120358347A
Network-adaptive video conferencing transmission optimization method and system
CN120358347B