Streaming media adaptive transmission method and system based on reinforcement learning, and electronic equipment

Through online reinforcement learning, the method of optimizing real-time streaming media transmission has been solved, and the problem that the existing technology is difficult to adapt to the changes in modern Internet networks has been achieved, and a higher quality video transmission experience has been achieved.

CN119996388APending Publication Date: 2025-05-13ZHEJIANG UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510102364.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing real-time streaming media transmission methods are difficult to adapt to network quality changes in modern Internet networks that are highly heterogeneous, diversified and dynamically changing, making it difficult to ensure the quality of video transmission.

Method used

Adaptive real-time streaming media transmission method based on online reinforcement learning is adopted, and the bit rate, resolution, frame rate, and redundant encoding algorithm type, media packet number and redundant packet number of videos are optimized through deep reinforcement learning algorithms, and dynamically adjust it according to the real-time network status and video quality evaluation scores.

Benefits of technology

It significantly improves the user experience quality of audio and video communication, can learn and capture the impact of various network states on system performance more comprehensively, adapt to complex network environments, and improve the stability and quality of video transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996388A_ABST
    Figure CN119996388A_ABST
Patent Text Reader

Abstract

The invention discloses a streaming media adaptive transmission method and system based on reinforcement learning and electronic equipment, and the method comprises the steps: obtaining current streaming media state data, obtaining and sending streaming media control information through an online reinforcement learning model, then obtaining latest streaming media state data, and calculating a model award to continuously update the model; adopting a federated learning method to aggregate the models of all streaming media clients to obtain a new global model; the system comprises a sending end module, a receiving end module, a forward error correction module, a network state detection module, a video quality evaluation module, a model aggregation module and a reinforcement learning agent module used for adaptively adjusting audio and video coding and forward error correction algorithm parameters. The method is suitable for various heterogeneous weak network environments, can effectively deal with the problems of occasionality, complexity and the like of the network, realizes self-adaptive transmission of audios and videos, and remarkably improves the user experience quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a real-time streaming media adaptive transmission method, system and electronic equipment based on online reinforcement learning, and belongs to the technical field of adaptive streaming media transmission. Background Art

[0002] In recent years, with the rapid development of 5G networks, rich media, especially audio and video streaming, has become an important part of daily life and multiple industries, among which the main traffic is concentrated in streaming services such as live broadcast and video conferencing. Whether it is live broadcast on social media, online classroom of distance education, or remote conference and telemedicine consultation of enterprises, real-time streaming transmission plays a vital role. Correspondingly, the data transmission environment has become more complex, showing the characteristics of heterogeneity, diversification and dynamic change. On the one hand, 5G technology has provided new development opportunities for audio and video transmission with the advantages of high bandwidth, wide connection and low latency; on the other hand, weak network environment still exists widely. In high-speed mobile scenarios, frequent access point switching, signal attenuation in underground space, high latency and network congestion make the network status more unstable and unpredictable. In the process of streaming data transmission such as high-definition video, complex weak network environment, such as increased network delay (RTT), network jitter, packet loss and network congestion, often make it difficult to ensure transmission quality. Traditional transmission methods are difficult to meet the requirements of high-quality transmission and adaptability to weak network environments at the same time, which easily leads to problems such as video data packet loss, disorder and delay fluctuation, significantly reducing the quality of user video viewing experience. Therefore, it is very necessary to develop an efficient adaptive video streaming transmission technology that takes into account the current network conditions, video quality and other conditions to improve the user viewing experience.

[0003] In order to improve the user viewing experience, traditional interactive real-time streaming communication applications generally adopt rule-based models, such as the GCC congestion control algorithm, forward error correction algorithm, and some video bitrate adaptive algorithms used in the commonly used WebRTC framework. The packet protocol is usually based on the standard RTP protocol, and the underlying UDP protocol is used. Taking the GCC congestion control algorithm as an example, its core idea is to control the sending rate by predicting the available bandwidth. It will be combined with the bandwidth estimated by both the sender and the receiver for comprehensive calculation. The bandwidth estimation of the sender mainly depends on the packet loss rate, and the bandwidth estimation of the receiver depends on the change of delay. Combining the results of both ends, it is determined by fixed rules whether it is currently in normal state, overload state or underload state, and then the bit rate is adjusted according to fixed rules. However, the rule-based model cannot adapt to the highly heterogeneous, diverse and dynamically changing modern Internet networks today, and cannot quickly and sensitively detect changes in network quality, so it cannot achieve good video effects.

[0004] Reinforcement learning is a branch of machine learning, which refers to the process in which an intelligent system learns the mapping from environment to behavior through trial and error to obtain the maximum reward. In reinforcement learning, the agent maximizes the cumulative reward signal by trying different actions. This process is usually modeled as a Markov decision process (MDP), in which the agent observes the state of the environment at each time step, takes actions to affect the environment, and obtains rewards. The process is as follows: Figure 1 In recent years, reinforcement learning methods have been proposed to improve the QOE of streaming media transmission, such as the Concerto algorithm and the OnRL algorithm. However, the above algorithms all use bit rate as the target action. Practice shows that simply changing the bit rate of the video cannot achieve good results in network environments such as resource constraints and network packet loss. Summary of the invention

[0005] The invention aims to overcome the shortcomings of the above-mentioned existing real-time streaming media transmission methods and proposes a real-time streaming media adaptive transmission method, system and electronic device based on online reinforcement learning to improve the user experience quality.

[0006] The present invention adopts the following technical solution:

[0007] A real-time streaming media adaptive transmission method based on online reinforcement learning comprises the following steps:

[0008] Step 1: Expose two software interfaces at the sending end of the real-time streaming media transmission system, which are used to obtain the current system status and modify the parameters of video encoding and forward error correction, respectively. The current system status includes round-trip delay, actual sending bit rate, target sending bit rate, network packet loss rate, network delay gradient change, receiving bit rate, actual packet loss rate after forward error correction recovery, and video quality evaluation score at the receiving end; the parameters of video encoding and forward error correction include forward error correction redundant coding algorithm type, forward error correction algorithm media packet number, forward error correction algorithm redundant packet number, video encoding bit rate, video resolution, and video frame rate;

[0009] Step 2: Use a deep reinforcement learning algorithm to optimize the video bit rate, resolution, frame rate decision, and forward error correction redundant coding algorithm type, media packet number, and redundant packet number decision, including the following steps:

[0010] Step 2.1: Represent the state space;

[0011] Design multi-dimensional features and design the state space into an eight-dimensional vector, which represents round-trip delay, actual transmission bit rate, target transmission bit rate, network packet loss rate, network delay gradient change, reception bit rate, actual packet loss rate after forward error correction recovery, and the video quality evaluation score at the receiving end;

[0012] Step 2.2: Represent the action space;

[0013] Design multiple discrete action spaces, involving six actions, including forward error correction redundant coding algorithm type, forward error correction algorithm media packet number, forward error correction algorithm redundant packet number, video encoding bit rate, video resolution, and video frame rate; extract features from the state space through multiple fully connected layers in the deep reinforcement learning algorithm, and then design six output layers, each of which uses a normalization function to calculate the probability distribution of the corresponding action;

[0014] Step 2.3: Design reward function;

[0015] After the state and action design, further design in state s t Next take an action a t Rewards t ,Specifically, the reward function is shown in Formula 1,

[0016]

[0017] in,

[0018] r t : reward value;

[0019] q recv : The bit rate measured by the receiving part, α is the weight value of this item;

[0020] q pre-recv : The previous bit rate measured by the receiving part, β is the weight value of the corresponding bit rate change term;

[0021] l: The network packet loss rate measured by the receiving part, δ is the weight value of this item;

[0022] l pre : The previous state network packet loss rate measured by the receiving part, ζ is the weight value corresponding to the change of packet loss rate;

[0023] T: forward error correction decoding time measured by the receiving part, ε is the weight value of this item;

[0024] D: The network delay gradient change measured by the receiving part, is the weight value of this item;

[0025] f: The network packet loss rate after forward error correction recovery measured by the receiving part, η is the corresponding forward error correction recovery rate.

[0026] The weight value corresponding to the reduced packet loss rate after recovery processing;

[0027] s: video quality evaluation score, γ is the weight value of this item;

[0028] In this reward function design, positive rewards are given to the bit rate, packet loss rate, video quality evaluation, and packet loss rate recovered by forward error correction, and penalties are given to the packet loss rate, forward error correction decoding time, network delay gradient change, and bit rate change;

[0029] Step 2.4: Train the deep reinforcement learning model;

[0030] The interface for obtaining the current state in step 1 is used to obtain eight parameters in the current real-time streaming media adaptive transmission system, namely, the round-trip delay, the actual sending bit rate, the target sending bit rate, the network packet loss rate, the network delay gradient change, the receiving bit rate, the actual packet loss rate after forward error correction recovery, and the video quality evaluation score at the receiving end, to generate the state space s t , input into the reinforcement learning model, and get the corresponding action a t , according to the action value, use the interface for modifying the video encoding and forward error correction parameters in step 1 to modify the video encoding and forward error correction parameters of the transmitter, and then use the interface for obtaining the current state to obtain the new state s t+1 , the specific reward value r is calculated according to the reward function t , the above-obtained {s t ,a t ,r t ,s t+1} is stored in a buffer as an experience sample. When the amount of data in the data buffer exceeds a preset threshold, a gradient strategy that maximizes the cumulative reward is executed to train the online reinforcement learning model and perform model parameter update processing on the online reinforcement learning model;

[0031] Step 3: Perform model aggregation in the central server, conduct federated learning between the streaming media client and the central server, integrate the training results of K media clients, and aggregate a more universal model. In the aggregation process, the average aggregation method can be used. The specific process is: the streaming media client sends the locally trained model to the central server. After the central server receives the model parameters sent by all clients, it performs weighted average of all model parameters with the same weight value to generate a new global model. In addition, the personalized aggregation method can be used to generate a model that better meets the user's network conditions. In the aggregation process, the model parameters of the streaming media client can be given priority, and the weight of the model parameters of the client can be increased to make it significantly higher than the weight of other streaming media client models, so as to obtain a model that better meets the needs of the client itself.

[0032] For the model aggregation described in step 3, the model aggregation equation is:

[0033]

[0034] in,

[0035] Represents the model parameters after aggregation;

[0036] λ k represents the weight of the client k model;

[0037] W k,i,j represents the jth parameter of the i-th layer in the network model of the k-th client;

[0038] For the new global model, the present invention adopts an average aggregation method, and for each streaming client, Generate a model with the average experience of all users;

[0039] For the client personalized model, the present invention adopts a personalized aggregation method, giving priority to the weight of the corresponding client itself, and setting λ k =p, p∈[0,1], And m≠k.

[0040] The receiving end in the streaming media transmission is responsible for the video quality evaluation score of the receiving end described in step 1. It decodes and renders the received data, and identifies and evaluates the video quality. The evaluation parameters mainly include: video freeze, video frame skipping, decoding failure, and set corresponding weights for each evaluation parameter, and then perform weighted average to obtain the final result. The evaluation formula for video quality evaluation is as follows:

[0041] f(x)=γ1x1+γ2x2+γ3x3 (3)

[0042] in,

[0043] f(x) is the video quality score result. The higher the score, the higher the video quality.

[0044] x1: quantized value of video freeze, γ1: weight corresponding to video freeze;

[0045] x2: quantized value of video frame skipping, γ2: weight corresponding to video frame skipping;

[0046] x3: quantized value of the video mosaic situation, γ3: weight corresponding to the video mosaic situation;

[0047] And the weight values ​​corresponding to the above evaluation parameters can be set manually.

[0048] Furthermore, the online reinforcement learning method used is a proximal policy optimization algorithm, in which the network output is a multi-dimensional discrete action space. A multi-discrete method is adopted to add multiple heads of the same dimension to the output layer, and a mask is used to handle the problem of inconsistent action dimensions.

[0049] A second aspect of the present invention relates to a real-time streaming media transmission system that implements the present invention and adopts the above-mentioned real-time streaming media adaptive transmission method based on online reinforcement learning, comprising:

[0050] The sending end module is used to collect and encode streaming media data from PC and mobile terminals, unpack the encoded large data stream, encapsulate data packets using the RTP protocol, and send data to the receiving end module using the UDP protocol, thus realizing the complete process of data collection, encoding, and sending;

[0051] The receiving module mainly includes receiving UDP data packets sent by the sending end, parsing RTP data packets, caching and sorting data packets, splicing streaming media raw stream data, decoding and rendering streaming media data, realizing the complete process of data reception, data packet buffering and splicing, decoding and playback;

[0052] The forward error correction module mainly caches the input RTP data packets and selects an appropriate number of media packets for forward error correction processing according to the coding algorithm type and coding redundancy parameter information of the module to generate an appropriate amount of redundant data;

[0053] The network status detection module mainly detects and collects the current network packet loss rate, round-trip delay, network delay gradient, actual sending bit rate of the sender, and actual receiving bit rate of the receiver;

[0054] The video quality evaluation module is used to evaluate the quality of the video obtained after decoding by the receiving end to obtain a video quality evaluation score. The evaluation indicators include video freeze, video frame skipping, and video mosaic.

[0055] Model aggregation module: Due to the high heterogeneity and complexity of real networks, different users have different network states in different scenarios. The network model of a single user cannot adapt to all network states. According to the method described in step 3 of claim 1, the present invention adopts an average aggregation method to generate a new global model for new clients to use, and updates all client model parameters to a new global model at fixed intervals. At the same time, considering that the network state of each client will not change too much in a short period of time, a personalized aggregation method is adopted to generate a personalized model to meet the usage of different users.

[0056] The reinforcement learning agent module uses the network packet loss rate, round-trip delay, network delay gradient, actual sending bit rate of the sender, receiving bit rate of the receiver, packet loss rate after forward error correction recovery, specified sending bit rate of the sender, and evaluation score obtained by the video quality evaluation module detected by the network status detection module as state input into the reinforcement learning network model. Then, according to the result information output by the reinforcement learning agent, it is mapped to the encoding bit rate, frame rate, resolution, forward error correction algorithm type, and number of forward error correction redundant packets of the sender, and the sending parameters of the sender are modified. The network status detection module is used to detect the network information after the modification of the sending parameters, and the reward function is used to generate a reward value to continuously optimize the reinforcement learning agent.

[0057] The sending end module includes a streaming media data acquisition device and an encoding device. The streaming media data acquisition device can receive parameters to modify video resolution, audio sampling rate, and video frame rate parameters; the streaming media encoding device can receive parameters to modify the encoding bit rate.

[0058] In the forward error correction module, different forward error correction algorithms and different numbers of media packets and redundant packets can be selected. Among them, the forward error correction coding algorithm supports three types of check codes: Reed-Solomon code, ladder-type low-density parity check (LDPC) code, and two-dimensional parity check matrix code.

[0059] In the network status detection module, network status detection is performed on both the sending and receiving sides. The sending side detection mainly detects the actual sending bit rate and round-trip delay of the sending side; the receiving side detection mainly detects the network delay gradient, the receiving bit rate, the packet loss rate, and the packet loss rate after error correction processing.

[0060] A third aspect of the present invention provides an electronic device, comprising:

[0061] A memory, characterized in that it is used to store an implementation program of the real-time streaming media transmission system based on online reinforcement learning in the above claims;

[0062] One or more processors, configured to run the program stored in the memory to execute the real-time streaming media transmission system of claim 5 and train the real-time streaming media transmission method based on online reinforcement learning of claim 1;

[0063] Microphone, used for collecting audio data in the above-mentioned real-time streaming media transmission system, must support three sampling rates: 8kHZ, 44.1kHZ, and 48kHZ;

[0064] A loudspeaker, used for playing audio data in the above-mentioned real-time streaming media transmission system;

[0065] The camera is used to collect video data in the above-mentioned real-time streaming media transmission system. It needs to support three resolutions of 640×480, 1280×720, and 1920×1080, and three frame rates of 15fps, 24fps, and 30fps.

[0066] It can be seen from the above description of the present invention that, compared with the prior art, the present invention has the following beneficial effects:

[0067] First, the present invention proposes a real-time streaming media adaptive transmission method based on online reinforcement learning. The rule-based model is difficult to adapt to the highly changing network status of the modern Internet, while the present invention is suitable for heterogeneous, diversified and dynamically changing network environments, and significantly improves the user experience quality of audio and video communications.

[0068] Second, the reinforcement learning model in the present invention takes eight network parameters as state inputs, including round-trip delay, actual sending bit rate, target sending bit rate, network packet loss rate, network delay gradient change, receiving bit rate, actual packet loss rate after forward error correction recovery, and video quality evaluation score at the receiving end. By collecting multi-dimensional network states as state inputs for reinforcement learning, the model can more comprehensively learn and capture the impact of various network states on system performance.

[0069] Third, the reinforcement learning model in the present invention adopts multiple discrete action spaces, forward error correction redundant coding algorithm type, forward error correction algorithm media packet number, forward error correction algorithm redundant packet number, video encoding bit rate, video resolution, and video frame rate. Through the coordinated optimization of six actions, it effectively solves the problem that simply adjusting the bit rate is difficult to cope with complex network environments such as resource constraints and random packet loss.

[0070] Fourth, the present invention adopts the method of "offline pre-training, online adjustment, and model aggregation". The reinforcement learning model is pre-trained in a simulated environment through a network simulation tool. The pre-trained model is stored in a central server. In real-time streaming media communication applications, the program first obtains the latest model from the central server, and continues model training during the communication process. The training results are then aggregated to the central server to generate the latest global model and the client's personalized model. This method effectively solves the problem of cold start of online models. At the same time, through federated learning, it can not only achieve global optimization, but also provide each client with a personalized model that is more in line with the current network conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 It is an overall framework diagram of the streaming media transmission system of the present invention.

[0072] Figure 2 It is a principle block diagram of the reinforcement learning network model of the present invention.

[0073] Figure 3 It is a schematic diagram of the model aggregation federated reinforcement learning framework of the present invention.

[0074] Figure 4 It is a block diagram of the working principle of the video quality evaluation module of the present invention.

[0075] Figure 5 It is a flow chart of the data processing method of the present invention.

[0076] Figure 6 It is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0077] In order to make the purpose and technical solution of the present invention clearly and completely described, and the advantages more clearly understood. It should be understood that the specific embodiments described herein are part of the embodiments of the present invention, rather than all of the embodiments, and are only used to explain the embodiments of the present invention, and are not used to limit the embodiments of the present invention. All other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0078] The embodiment of the present invention adopts an online reinforcement learning model. According to the real-time streaming media status data of the streaming media client, streaming media control information for guiding the streaming media to perform encoding and forward error correction processing is obtained, so as to realize online feedback control processing of the streaming media client, thereby improving the video quality of the streaming media. At the same time, the online reinforcement learning model is also trained online according to the real-time streaming media status data, so that the online reinforcement learning model can adapt to the real-time changing network environment.

[0079] The above-mentioned reinforcement learning model is first trained extensively in a simulated network environment to obtain a pre-trained basic model, which is then uploaded to the central server. Each streaming client obtains the latest reinforcement learning network model from the central server and uses the reinforcement learning model to adjust the streaming parameters of the sender. At this time, the training data used by the reinforcement learning model is the streaming status data of the client in the real call scenario.

[0080] Furthermore, the present invention also proposes a scheme for aggregating multiple online reinforcement learning models, adopts a federated learning method, integrates the training results of all streaming clients, and aggregates a more universal reinforcement learning model. At the same time, the aggregation method adopts two methods: average aggregation and personalized aggregation. Average aggregation is to weighted average the model parameters of all streaming clients with the same weight, and then use the new model parameters to update the global model in the central server. This method can obtain a more universal model that integrates the experience of all client models; and personalized aggregation is that, considering that the network status of the corresponding streaming client will not change much over a period of time, the model weight of the streaming client is increased during the aggregation process, making it significantly higher than the weight of other streaming clients, so that the obtained aggregated model parameters are more in line with their own needs, realizing personalized features, and at the same time, the client model parameters are updated using the global model after a period of time.

[0081] The technical solution of the present invention is further illustrated by some specific embodiments below.

[0082] Embodiment 1

[0083] Reference Figures 1 to 5 The present invention provides a real-time streaming media adaptive transmission method based on online reinforcement learning, comprising the following steps:

[0084] Step 1: Expose two software interfaces at the sending end of the real-time streaming media transmission system, which are used to obtain the current system status and modify the parameters of video encoding and forward error correction, respectively. The current system status includes round-trip delay, actual sending bit rate, target sending bit rate, network packet loss rate, network delay gradient change, receiving bit rate, actual packet loss rate after forward error correction recovery, and video quality evaluation score at the receiving end; the parameters of video encoding and forward error correction include forward error correction redundant coding algorithm type, forward error correction algorithm media packet number, forward error correction algorithm redundant packet number, video encoding bit rate, video resolution, and video frame rate;

[0085] Step 2: Use a deep reinforcement learning algorithm to optimize the video bit rate, resolution, frame rate decision, and forward error correction redundant coding algorithm type, media packet number, and redundant packet number decision, including the following steps:

[0086] Step 2.1: State space representation;

[0087] Design multi-dimensional features and design the state space into an eight-dimensional vector, which represents round-trip delay, actual transmission bit rate, target transmission bit rate, network packet loss rate, network delay gradient change, reception bit rate, actual packet loss rate after forward error correction recovery, and the video quality evaluation score at the receiving end;

[0088] Step 2.2: Action space representation;

[0089] Design multiple discrete action spaces, involving six actions, including forward error correction redundant coding algorithm type, forward error correction algorithm media packet number, forward error correction algorithm redundant packet number, video encoding bit rate, video resolution, and video frame rate; extract features from the state space through multiple fully connected layers in the deep reinforcement learning algorithm, and then design six output layers, each of which uses a normalization function to calculate the probability distribution of the corresponding action;

[0090] Step 2.3: Reward function design;

[0091] After the state and action design, further design in state s t Next take an action a t Rewards t Specifically, the reward function is shown in Formula 1, where positive rewards are given to the network bit rate, packet loss rate, and the difference between the video quality evaluation and the packet loss rate after error correction, and penalties are given to the packet loss rate, forward error correction decoding time, network delay gradient change, and bit rate change. These indicators are weighted by corresponding weights. After calculating the rewards for each time period, the cumulative reward R is calculated according to the following Formula 4 t :

[0092]

[0093] Step 2.4, deep reinforcement learning model training

[0094] The present invention selects the proximal strategy optimization algorithm for training. The detailed process of the reinforcement learning training of the present invention is as follows:

[0095] S1, set the learning rate, experience pool size, and randomly initialize the strategy network and value network;

[0096] S2, obtain State: S0 from the real-time streaming media transmission system, where State represents state information and S0 represents the initial state;

[0097] S3. According to the current state, select action a through the Actor network t , where a t This includes actions such as the bit rate mentioned above;

[0098] S4, the real-time streaming media transmission system executes action a t , get the downward state s t+1 And calculate the reward r t ;

[0099] S5. t , a t , r t ,s​t+1 >Store historical experience in the playback pool and update the current environment;

[0100] S6, repeat steps S3, S4, and S5 until the amount of data in the experience pool exceeds the size of the experience pool;

[0101] S7. When the amount of data in the experience pool exceeds the size of the experience pool, N experience samples are randomly selected from it;

[0102] S8. Calculate the advantage function based on the value function of the critic network and update the actor network by maximizing the objective function of PPO.

[0103] S9, repeat the above S3, S4, S5, S6, S7, S8 until the model converges;

[0104] Step 3: Upload the trained model to the central server as the basic model for subsequent clients to use.

[0105] Step 4: The streaming client obtains the latest model from the central server and deploys it in the corresponding device. When making a real-time streaming call, step 2 is repeated, and the obtained model is used for action selection and model training. After a period of time, the latest model is uploaded to the central server for model aggregation and obtaining the latest personalized model. The global model is obtained from the central server at a specified time to update the local model.

[0106] Embodiment 2

[0107] Reference Figures 1 to 5 The present invention provides a real-time streaming media adaptive transmission system based on online reinforcement learning. Specifically, the system includes a sending end module, a receiving end module, a forward error correction module, a network status detection module, a video quality evaluation module, a model aggregation module and a reinforcement learning agent module.

[0108] The sending end module is used to collect and encode streaming media data from PC and mobile terminals, unpack the encoded large data stream, encapsulate data packets using the RTP protocol, and send data to the receiving end module using the UDP protocol, thus realizing the complete process of data collection, encoding, and sending;

[0109] The receiving module mainly includes receiving UDP data packets sent by the sending end, parsing RTP data packets, caching and sorting data packets, splicing streaming media raw stream data, decoding and rendering streaming media data, realizing the complete process of data reception, data packet buffering and splicing, decoding and playback;

[0110] The forward error correction module mainly caches the input RTP data packets and selects an appropriate number of media packets for forward error correction processing according to the coding algorithm type and coding redundancy parameter information of the module to generate an appropriate amount of redundant data;

[0111] The network status detection module mainly detects and collects the current network packet loss rate, round-trip delay, network delay gradient, actual sending bit rate of the sender, and actual receiving bit rate of the receiver;

[0112] The video quality evaluation module is used to evaluate the quality of the video obtained after decoding by the receiving end to obtain a video quality evaluation score. The evaluation indicators include video freeze, video frame skipping, and video mosaic.

[0113] Model aggregation module: Due to the high heterogeneity and complexity of real networks, different users have different network states in different scenarios. The network model of a single user cannot adapt to all network states. According to the method described in step 3 of claim 1, the present invention adopts an average aggregation method to generate a new global model for new clients to use, and updates all client model parameters to a new global model at fixed intervals. At the same time, considering that the network state of each client will not change too much in a short period of time, a personalized aggregation method is adopted to generate a personalized model to meet the usage of different users.

[0114] The reinforcement learning agent module uses the network packet loss rate, round-trip delay, network delay gradient, actual sending bit rate of the sender, receiving bit rate of the receiver, packet loss rate after forward error correction recovery, specified sending bit rate of the sender, and evaluation score obtained by the video quality evaluation module detected by the network status detection module as state input into the reinforcement learning network model. Then, according to the result information output by the reinforcement learning agent, it is mapped to the encoding bit rate, frame rate, resolution, forward error correction algorithm type, and number of forward error correction redundant packets of the sender, and the sending parameters of the sender are modified. The network status detection module is used to detect the network information after the modification of the sending parameters, and the reward function is used to generate a reward value to continuously optimize the reinforcement learning agent.

[0115] Embodiment 3

[0116] This embodiment relates to an electronic device, which is used to implement a real-time streaming media adaptive transmission method based on online reinforcement learning described in Embodiment 1, such as Figure 6 As shown, it is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention, which specifically includes:

[0117] The memory is used to store the implementation program of the above-mentioned real-time streaming media transmission system based on online reinforcement learning. At the same time, it can also store various other data, such as reinforcement learning model files and system pictures. The memory is implemented by a combination of volatile storage devices and non-volatile storage devices. Among them, the volatile storage device is mainly used for high-speed computing and processing when the system is running, and can use dynamic random access memory (DRAM) and static random access memory (SRAM); non-volatile storage devices are mainly used for permanent storage of programs and data, and can use solid-state drives, mechanical hard drives, and optical disks.

[0118] One or more processors are used to run the program stored in the memory, to implement the above-mentioned real-time streaming media transmission system, and to train a real-time streaming media transmission algorithm based on online reinforcement learning.

[0119] The microphone is used to collect audio data in the above-mentioned real-time streaming media transmission system. It converts the vibration signal in the sound into an electronic signal that can be processed, stored or transmitted. After being encoded by the system, it is sent to other devices using a network card. In this embodiment, the microphone needs to support three sampling rates: 8kHZ, 44.1kHZ, and 48kHZ;

[0120] The loudspeaker is used for playing audio data in the above-mentioned real-time streaming media transmission system. It converts electrical signals into sound waves and generates sound by vibrating the air. In this example, various types of loudspeakers can be used, such as electric loudspeakers and piezoelectric loudspeakers.

[0121] The camera is used to collect video data in the above-mentioned real-time streaming media transmission system. It is an optical image capture device that can convert optical images into digital signals through photoelectric conversion. In this example, the camera needs to support three resolutions of 640×480, 1280×720, and 1920×1080, and three frame rates of 15fps, 24fps, and 30fps.

[0122] Furthermore, as shown in the figure, the electronic device may also include a network card and a display.

[0123] The network card is configured to facilitate wired or wireless communication between the electronic device and other electronic devices, including, for example, 2G, 3G, 4G / LTE, 5G, and WiFi.

[0124] The display is mainly a screen, and a liquid crystal display or a touch screen display can be used.

[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A real-time streaming media adaptive transmission method based on online reinforcement learning, characterized in that: The following steps are involved: Step 1: Expose two software interfaces at the sending end of the real-time streaming media transmission system, which are used to obtain the current system status and modify the parameters of video encoding and forward error correction, respectively. The current system status includes round-trip delay, actual sending bit rate, target sending bit rate, network packet loss rate, network delay gradient change, receiving bit rate, actual packet loss rate after forward error correction recovery, and video quality evaluation score at the receiving end; the parameters of video encoding and forward error correction include forward error correction redundant coding algorithm type, forward error correction algorithm media packet number, forward error correction algorithm redundant packet number, video encoding bit rate, video resolution, and video frame rate; Step 2: Use a deep reinforcement learning algorithm to optimize the video bit rate, resolution, frame rate decision, and forward error correction redundant coding algorithm type, media packet number, and redundant packet number decision, including the following steps: Step 2.1: Represent the state space; Design multi-dimensional features and design the state space into an eight-dimensional vector, which represents round-trip delay, actual transmission bit rate, target transmission bit rate, network packet loss rate, network delay gradient change, reception bit rate, actual packet loss rate after forward error correction recovery, and the video quality evaluation score at the receiving end; Step 2.2: Represent the action space; Design multiple discrete action spaces, involving six actions, including forward error correction redundant coding algorithm type, forward error correction algorithm media packet number, forward error correction algorithm redundant packet number, video encoding bit rate, video resolution, and video frame rate; extract features from the state space through multiple fully connected layers in the deep reinforcement learning algorithm, and then design six output layers, each of which uses a normalization function to calculate the probability distribution of the corresponding action; Step 2.3: Design reward function; After the state and action design, further design in state s t Next take an action a t Rewards t ,Specifically, the reward function is shown in Formula 1, in, r t : reward value; q recv : The bit rate measured by the receiving part, α is the weight value of this item; q pre-recv : The previous bit rate measured by the receiving part, β is the weight value of the corresponding bit rate change term; l: The network packet loss rate measured by the receiving part, δ is the weight value of this item; l pre : The previous state network packet loss rate measured by the receiving part, ζ is the weight value corresponding to the change of packet loss rate; T: forward error correction decoding time measured by the receiving part, ε is the weight value of this item; D: The network delay gradient change measured by the receiving part, is the weight value of this item; f: the network packet loss rate after forward error correction recovery processing measured by the receiving part, η is the weight value corresponding to the reduced packet loss rate after forward error correction recovery processing; s: video quality evaluation score, γ is the weight value of this item; In this reward function design, positive rewards are given to the bit rate, packet loss rate, video quality evaluation, and packet loss rate recovered by forward error correction, and penalties are given to the packet loss rate, forward error correction decoding time, network delay gradient change, and bit rate change; Step 2.4: Train the deep reinforcement learning model; The interface for obtaining the current state in step 1 is used to obtain eight parameters in the current real-time streaming media adaptive transmission system, namely, the round-trip delay, the actual sending bit rate, the target sending bit rate, the network packet loss rate, the network delay gradient change, the receiving bit rate, the actual packet loss rate after forward error correction recovery, and the video quality evaluation score at the receiving end, to generate the state space s t , input into the reinforcement learning model, and get the corresponding action a t , according to the action value, use the interface for modifying the video encoding and forward error correction parameters in step 1 to modify the video encoding and forward error correction parameters of the transmitter, and then use the interface for obtaining the current state to obtain the new state s t+1 , the specific reward value r is calculated according to the reward function t , the above-obtained {s t ,a t ,r t ,s t+1 } is stored in a buffer as an experience sample. When the amount of data in the data buffer exceeds a preset threshold, a gradient strategy that maximizes the cumulative reward is executed to train the online reinforcement learning model and perform model parameter update processing on the online reinforcement learning model; Step 3: Perform model aggregation in the central server, conduct federated learning between the streaming media client and the central server, integrate the training results of K media clients, and aggregate a more universal model; in the aggregation process, adopt the average aggregation method, and the specific process is: the streaming media client sends the locally trained model to the central server. After the central server receives the model parameters sent by all clients, it performs weighted average of all model parameters with the same weight value to generate a new global model; in addition, the personalized aggregation method can also be used to generate a model that better meets the user's network conditions. In the aggregation process, the model parameters of the streaming media client can be given priority, and the weight of the model parameters of the client can be increased to make it significantly higher than the weight of other streaming media client models, so as to obtain a model that better meets the needs of the client itself.

2. The real-time streaming media transmission method based on online reinforcement learning as claimed in claim 1, characterized in that: The model aggregation described in step 3 has the following model aggregation equation: in, Represents the model parameters after aggregation; λ k Represents the weight of the client k model; W k,i,j represents the jth parameter of the i-th layer in the network model of the k-th client; The new global model adopts the average aggregation method. For each streaming client, Generate a model with the average experience of all users; For the client personalized model, the present invention adopts a personalized aggregation method, giving priority to the weight of the corresponding client itself, and setting λ k =p, p∈[0,1], And m≠k.

3. The real-time streaming media adaptive transmission method based on online reinforcement learning as claimed in claim 1, characterized in that: The receiving end in the streaming media transmission is responsible for the video quality evaluation score of the receiving end described in step 1. It decodes and renders the received data, and identifies and evaluates the video quality. The evaluation parameters include: video freeze, video frame skipping, decoding failure, and set corresponding weights for each evaluation parameter, and then perform weighted average to obtain the final result. The evaluation formula for video quality evaluation is as follows: f(x)=γ1x1+γ2x2+γ3x3 (3) in, f(x) is the video quality score result. The higher the score, the higher the video quality. x1: quantized value of video freeze, γ1: weight corresponding to video freeze; x2: quantized value of video frame skipping, γ2: weight corresponding to video frame skipping; x3: quantized value of the video mosaic situation, γ3: weight corresponding to the video mosaic situation; And the weight values ​​corresponding to the above evaluation parameters can be set manually.

4. The real-time streaming media adaptive transmission method based on online reinforcement learning as claimed in claim 1, characterized in that: The online reinforcement learning method used is the proximal policy optimization algorithm.

5. A real-time streaming media transmission system that implements the real-time streaming media adaptive transmission method based on online reinforcement learning as described in claims 1 to 4, characterized in that: include: The sending end module is used to collect and encode streaming media data from PC and mobile terminals, unpack the encoded large data stream, encapsulate data packets using the RTP protocol, and send data to the receiving end module using the UDP protocol, thus realizing the complete process of data collection, encoding, and sending; The receiving module mainly includes receiving UDP data packets sent by the sending end, parsing RTP data packets, caching and sorting data packets, splicing streaming media raw stream data, decoding and rendering streaming media data, realizing the complete process of data reception, data packet buffering and splicing, decoding and playback; The forward error correction module mainly caches the input RTP data packets and selects an appropriate number of media packets for forward error correction processing according to the coding algorithm type and coding redundancy parameter information of the module to generate an appropriate amount of redundant data; The network status detection module mainly detects and collects the current network packet loss rate, round-trip delay, network delay gradient, actual sending bit rate of the sender, and actual receiving bit rate of the receiver; The video quality evaluation module is used to evaluate the quality of the video obtained after decoding by the receiving end to obtain a video quality evaluation score. The evaluation indicators include video freeze, video frame skipping, and video mosaic. Model aggregation module: Due to the high heterogeneity and complexity of real networks, different users have different network states in different scenarios. The network model of a single user cannot adapt to all network states. According to the method described in step 3 of claim 1, the present invention adopts an average aggregation method to generate a new global model for new clients to use, and updates all client model parameters to a new global model at fixed intervals. At the same time, considering that the network state of each client will not change too much in a short period of time, a personalized aggregation method is adopted to generate a personalized model to meet the usage of different users. The reinforcement learning agent module uses the network packet loss rate, round-trip delay, network delay gradient, actual sending bit rate of the sender, receiving bit rate of the receiver, packet loss rate after forward error correction recovery, specified sending bit rate of the sender, and evaluation score obtained by the video quality evaluation module detected by the network status detection module as state input into the reinforcement learning network model. Then, according to the result information output by the reinforcement learning agent, it is mapped to the encoding bit rate, frame rate, resolution, forward error correction algorithm type, and number of forward error correction redundant packets of the sender, and the sending parameters of the sender are modified. The network status detection module is used to detect the network information after the modification of the sending parameters, and the reward function is used to generate a reward value to continuously optimize the reinforcement learning agent.

6. The real-time streaming media transmission system according to claim 5, characterized in that: The transmitting end module includes a streaming media data acquisition device and an encoding device, and the streaming media data acquisition device can receive parameters to modify video resolution, audio sampling rate, and video frame rate parameters; The streaming media encoding device can receive parameters to modify the encoding bit rate.

7. The real-time streaming media transmission system according to claim 5, characterized in that: The forward error correction module selects different forward error correction algorithms and different numbers of media packets and redundant packets, wherein the forward error correction coding algorithm supports three types of check codes: Reed-Solomon code, ladder-type low-density parity check (LDPC) code, and two-dimensional parity check matrix code.

8. The real-time streaming media transmission system according to claim 5, characterized in that: The network status detection module performs network status detection on both the sending end and the receiving end. The sending end detection mainly detects the actual sending bit rate and round-trip delay of the sending end; the receiving end detection mainly detects the network delay gradient, the receiving bit rate, the packet loss rate, and the packet loss rate after error correction processing.

9. The real-time streaming media transmission system according to claim 5, characterized in that: The streaming media encoding parameters of the transmitting end part may also include key frame intervals; the transmitting end and the receiving end also include an automatic feedback retransmission mechanism.

10. An electronic device, characterized in that include: A memory, characterized in that it is used to store an implementation program of the real-time streaming media transmission system based on online reinforcement learning in the above claims; One or more processors, configured to run the program stored in the memory to execute the real-time streaming media transmission system according to claim 5 and implement the real-time streaming media transmission method based on online reinforcement learning according to claim 1; Microphone, used for collecting audio data in the above-mentioned real-time streaming media transmission system, must support three sampling rates: 8kHZ, 44.1kHZ, and 48kHZ; A loudspeaker, used for playing audio data in the above-mentioned real-time streaming media transmission system; The camera is used to collect video data in the above-mentioned real-time streaming media transmission system. It needs to support three resolutions of 640×480, 1280×720, and 1920×1080, and three frame rates of 15fps, 24fps, and 30fps.

Citation Information

Cited By

  • Reinforced learning method and system for personalized optimization of traditional Chinese medicine perception parameters

    CN122266666A

  • An Immersive Application Optimization System and Method Based on Intelligent Agents

    CN122578900A