Strong interaction video stream transmission quality optimization method, device, controller and system

By predicting the network packet loss rate and packet loss aggregation, and using the random forest model to decide the redundancy of video frames for forward error correction encoding, the problem of inability to effectively balance the quality and redundancy cost of strong interactive video streaming in the prior art is solved, and efficient video streaming is achieved.

CN120050451APending Publication Date: 2025-05-27BEIJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510030705.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing forward error correction (FEC) algorithms cannot effectively balance transmission quality and redundancy costs in strong interactive video streams, especially the problem of wasting redundant data in large frames and unable to effectively recover packet loss in small frames.

Method used

By predicting the network packet loss rate and packet loss aggregation in the next streaming protocol feedback cycle based on the packet network feedback data of the historical streaming protocol feedback cycle, the network packet loss rate and packet loss aggregation in the next streaming protocol feedback cycle is used to determine the redundancy of video frames frame by frame, and forward error correction encoding is performed.

Benefits of technology

It accurately adapts to the size difference and network dynamics of different video frames in strong interactive video streams, ensures the interaction quality of video streams, and greatly reduces the overhead of redundant bandwidth, and balances transmission quality and redundant costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050451A_ABST
    Figure CN120050451A_ABST
Patent Text Reader

Abstract

The invention provides a strong interaction video stream transmission quality optimization method, device, controller and system. The method comprises the following steps: a slow module determines a network packet loss probability prediction value and a packet loss aggregation prediction value of strong interaction video stream data in a next streaming media protocol feedback period; the fast module adopts a random forest model to decide the fine-grained redundancy corresponding to each video frame coding period frame by frame based on the network packet loss probability prediction value, the packet loss aggregation prediction value and the length of each video frame to be transmitted by the network corresponding to the strong interaction video stream data; and performing forward error correction coding on the data packet of each video frame based on the redundancy to obtain the data packet added with redundancy protection corresponding to each video frame. According to the method, the device and the system, the size difference and the network dynamics of different video frames of strong interaction video streams such as cloud games can be accurately self-adapted, the interaction quality of the strong interaction video streams can be ensured, the redundant bandwidth overhead can be greatly reduced, and the transmission quality and the redundant cost of the strong interaction video streams can be balanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of video data processing, and in particular to a method, device, controller, and system for optimizing the transmission quality of strongly interactive video streams. Background Art

[0002] A strongly interactive video stream refers to an interactive real-time video stream that always maintains an ultra-low latency transmission requirement. For example, cloud game video data, VR headset video data, and live video data between users all belong to strongly interactive video stream data. Taking the cloud game scenario as an example, cloud gaming, as an increasingly popular strongly interactive video stream application, involves transmitting high-quality strongly interactive video streams from a remote server to various devices, enabling games to run on all devices connected by users, especially mobile devices with limited computing and storage capabilities. With the rapid development of strongly interactive video stream applications, the quality and stability of video streams are crucial for players' gaming experiences. To ensure the interactive quality (QoE) of strongly interactive video streams, it is necessary to maintain ultra-low latency from user requests to responses, usually less than 100 milliseconds.

[0003] Existing interactive video streaming technologies face challenges in ensuring video quality and controlling costs. Currently, forward error correction (FEC) technology is commonly used in video streaming to pre-increase packet redundancy before packet transmission so that the complete transmission of packets can still be satisfied when packet loss occurs during transmission, thereby improving the quality and interactive latency stability of videos.

[0004] However, although existing FEC technologies are widely used to reduce transmission latency, most existing FEC algorithms are coarse-grained and only set the same redundancy for all video frames based on network conditions, without further adaptive adjustment according to different video frame size situations. This results in excessive waste of redundant data in large frames by FEC, while in small frames, it is unable to effectively recover lost packets, leading to resource waste or unstable video latency and affecting the interactive quality of videos.

[0005] Based on this, there is an urgent need to design a method that can balance the transmission quality of strongly interactive video streams and redundancy costs in strongly interactive streaming media such as cloud games. Summary of the Invention

[0006] In view of this, embodiments of this application provide a method, device, controller, and system for optimizing the transmission quality of strongly interactive video streams to eliminate or improve one or more defects existing in the prior art.

[0007] One aspect of this application provides a method for optimizing the transmission quality of strongly interactive video streams, including:

[0008] Determine the predicted value of the network packet loss rate and the predicted value of packet loss aggregation of the strong interaction video stream data in the next streaming protocol feedback cycle according to the packet network feedback data corresponding to each historical streaming protocol feedback cycle and the current streaming protocol feedback cycle respectively;

[0009] Based on the predicted value of the network packet loss rate and the predicted value of packet loss aggregation corresponding to the next streaming protocol feedback cycle of the strong interaction video stream data, and the frame lengths of each video frame to be network-transmitted corresponding to the strong interaction video stream data, use a random forest model to make a decision on the redundancy of each video frame corresponding to each video frame coding cycle frame by frame, so as to perform forward error correction coding on each data packet corresponding to each video frame based on the redundancy of each video frame, and obtain each transmission quality-optimized data packet corresponding to each video frame.

[0010] In some embodiments of the present application, the method of using a random forest model to make a decision on the redundancy of each video frame corresponding to each video frame coding cycle frame by frame based on the predicted value of the network packet loss rate and the predicted value of packet loss aggregation corresponding to the next streaming protocol feedback cycle of the strong interaction video stream data, and the frame lengths of each video frame to be network-transmitted corresponding to the strong interaction video stream data, so as to perform forward error correction coding on each data packet corresponding to each video frame based on the redundancy of each video frame, and obtain each transmission quality-optimized data packet corresponding to each video frame, includes:

[0011] Generate a data set corresponding to each video frame based on the predicted value of the network packet loss rate and the predicted value of packet loss aggregation corresponding to the next streaming protocol feedback cycle of the strong interaction video stream data, and the frame lengths of each video frame to be network-transmitted corresponding to the strong interaction video stream data, wherein each data set includes the frame length of one video frame and the predicted value of the network packet loss rate and the predicted value of packet loss aggregation corresponding to the strong interaction video stream data in the next streaming protocol feedback cycle;

[0012] Input each data set into the random forest model in sequence, so that the random forest model outputs the redundancy of each video frame correspondingly;

[0013] Transmit the redundancy output by the random forest model each time to the forward error correction encoder, so that the forward error correction encoder performs forward error correction coding on each data packet corresponding to the video frame according to the received redundancy of the video frame, and obtains each transmission quality-optimized data packet corresponding to the video frame, so as to perform network transmission on each transmission quality-optimized data packet corresponding to the video frame respectively.

[0014] In some embodiments of the present application, determining the predicted value of the network packet loss rate and the predicted value of packet loss clustering of the strong interaction video stream data in the next streaming protocol feedback period according to the packet network feedback data corresponding to each historical streaming protocol feedback period and the current streaming protocol feedback period respectively includes:

[0015] Receiving the packet network feedback data corresponding to the strong interaction video stream data in the current streaming protocol feedback period;

[0016] Constructing a current network feedback sequence, a network packet loss rate sequence, and a packet loss clustering sequence respectively, where the network feedback sequence includes: the packet network feedback data of the strong interaction video stream data in the current streaming protocol feedback period and the packet network feedback data of the strong interaction video stream data pre-acquired in each historical streaming protocol feedback period; the network packet loss rate sequence includes: the predicted values of the network packet loss rate corresponding to each historical streaming protocol feedback period and the current streaming protocol feedback period of the strong interaction video stream data pre-acquired; the packet loss clustering sequence includes: the predicted values of packet loss clustering corresponding to each historical streaming protocol feedback period and the current streaming protocol feedback period of the strong interaction video stream data pre-acquired;

[0017] Inputting the network feedback sequence, the network packet loss rate sequence, and the packet loss clustering sequence into a convolutional neural network, so that the convolutional neural network correspondingly outputs the predicted value of the network packet loss rate and the predicted value of packet loss clustering of the strong interaction video stream data in the next streaming protocol feedback period.

[0018] In some embodiments of the present application, inputting the network feedback sequence, the network packet loss rate sequence, and the packet loss clustering sequence into a convolutional neural network, so that the convolutional neural network correspondingly outputs the predicted value of the network packet loss rate and the predicted value of packet loss clustering of the strong interaction video stream data in the next streaming protocol feedback period includes:

[0019] Inputting the network feedback sequence, the network packet loss rate sequence, and the packet loss clustering sequence into a convolutional neural network, so that the one-dimensional convolutional block in the convolutional neural network extracts spatial features from the network feedback sequence and outputs the spatial feature information corresponding to the network feedback sequence, and then enabling the fully connected layer in the convolutional neural network to correspondingly output the predicted value of the network packet loss rate and the predicted value of packet loss clustering of the strong interaction video stream data in the next streaming protocol feedback period according to the spatial feature information, the network packet loss rate sequence, and the packet loss clustering sequence.

[0020] In some embodiments of the present application, before determining the predicted value of the network packet loss rate and the predicted value of packet loss clustering of the strong interaction video stream data in the next streaming protocol feedback cycle according to the packet network feedback data corresponding to each historical streaming protocol feedback cycle and the current streaming protocol feedback cycle, it further includes:

[0021] Training the convolutional neural network based on a preset loss function and historical training data, where the historical training data includes: a historical network feedback sequence corresponding to historical strong interaction video stream data, a historical network packet loss rate sequence, and a historical packet loss clustering sequence; wherein, the historical network feedback sequence includes: the packet network feedback data corresponding to the historical strong interaction video stream data in each historical streaming protocol feedback cycle; the historical network packet loss rate sequence includes: the predicted values of the network packet loss rate corresponding to the historical strong interaction video stream data in each historical streaming protocol feedback cycle; the historical packet loss clustering sequence includes: the predicted values of packet loss clustering corresponding to the historical strong interaction video stream data in each historical streaming protocol feedback cycle.

[0022] In some embodiments of the present application, before using a random forest model to make a frame-by-frame decision on the redundancy of each video frame in each video frame coding cycle according to the predicted value of the network packet loss rate and the predicted value of packet loss clustering corresponding to the strong interaction video stream data in the next streaming protocol feedback cycle, and the frame lengths of each video frame to be network-transmitted corresponding to the strong interaction video stream data, it further includes:

[0023] Training the random forest model with the mean squared error as the evaluation index of the random forest model based on the predicted value of the network packet loss rate and the predicted value of packet loss clustering corresponding to the historical strong interaction video stream data, and the frame lengths of each video frame corresponding to the historical strong interaction video stream data.

[0024] The second aspect of the present application provides a frame-level forward error correction device, including:

[0025] A slow module, configured to determine the predicted value of the network packet loss rate and the predicted value of packet loss clustering of the strong interaction video stream data in the next streaming protocol feedback cycle according to the packet network feedback data corresponding to each historical streaming protocol feedback cycle and the current streaming protocol feedback cycle;

[0026] A fast module, which is used to predict the network packet loss rate and packet loss aggregation prediction value corresponding to the next streaming media protocol feedback cycle based on the strong interactive video stream data, and the frame length of each video frame to be network-transmitted corresponding to the strong interactive video stream data. The random forest model is used to make a decision on the redundancy of each video frame corresponding to each video frame coding cycle frame by frame, so as to perform forward error correction coding on each data packet corresponding to each video frame based on the redundancy of each video frame, and obtain each transmission quality optimized data packet corresponding to each video frame.

[0027] The third aspect of the present application provides a forward error correction controller, including: a frame-level forward error correction device and a forward error correction encoder connected by communication;

[0028] The frame-level forward error correction device is used to execute the strong interactive video stream transmission quality optimization method described in the first aspect, and transmit the redundancy output by the random forest model each time to the forward error correction encoder;

[0029] The forward error correction encoder is used to perform forward error correction coding on each data packet corresponding to the video frame according to the redundancy of the received video frame, and obtain each transmission quality optimized data packet corresponding to the video frame.

[0030] The fourth aspect of the present application provides a strong interactive video stream transmission system, including: a data sending end device and a data receiving end device;

[0031] The data sending end device is provided with a video encoder, the forward error correction controller described in the third aspect, and a data transmitter that are connected by communication in sequence, and a network monitor that is communicatively connected to the forward error correction controller;

[0032] The video encoder is used to encode the target video data to obtain corresponding strong interactive video stream data;

[0033] The forward error correction controller is further used to transmit each transmission quality optimized data packet corresponding to the obtained video frame to the data transmitter;

[0034] The data transmitter is used to transmit each transmission quality optimized data packet corresponding to the video frame to the data receiving end device through the network.

[0035] The fifth aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the strong interactive video stream transmission quality optimization method described in the first aspect.

[0036] The sixth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for optimizing the transmission quality of a strong interaction video stream described in the first aspect.

[0037] The seventh aspect of the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the method for optimizing the transmission quality of a strong interaction video stream described in the first aspect.

[0038] For the method for optimizing the transmission quality of a strong interaction video stream provided by the present application, according to the packet network feedback data corresponding to each historical streaming media protocol feedback period and the current streaming media protocol feedback period respectively, the predicted value of the network packet loss rate and the predicted value of packet loss aggregation in the next streaming media protocol feedback period for the strong interaction video stream data are determined; based on the predicted value of the network packet loss rate and the predicted value of packet loss aggregation corresponding to the next streaming media protocol feedback period for the strong interaction video stream data, and the frame length of each video frame to be network-transmitted corresponding to the strong interaction video stream data, a random forest model is used to make a decision on the redundancy of each video frame for each video frame coding period respectively, so as to perform forward error correction coding on each packet corresponding to each video frame respectively based on the redundancy of each video frame, and obtain each transmission quality-optimized packet corresponding to each video frame. It can accurately adapt to the size difference and network dynamics between different video frames of strong interaction video streams such as cloud games, can guarantee the interaction quality QoE of the strong interaction video stream, and can greatly reduce the redundant bandwidth overhead, so as to balance the transmission quality and redundant cost of the strong interaction video stream.

[0039] The additional advantages, objectives, and features of the present application will be partially elaborated in the following description, and will become partially obvious to those of ordinary skill in the art after studying the following text, or can be learned from the practice of the present application. The objectives and other advantages of the present application can be achieved and obtained through the structure specifically pointed out in the description and the drawings.

[0040] Those skilled in the art will understand that the objectives and advantages that can be achieved by the present application are not limited to the above specifically described, and the above and other objectives that the present application can achieve will be more clearly understood according to the following detailed description. Description of the Drawings

[0041] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application, but do not limit the present application. The components in the drawings are not drawn to scale, but are only for showing the principle of the present application. For the convenience of showing and describing some parts of the present application, the corresponding parts in the drawings may be enlarged, that is, they may become larger relative to other components in the exemplary device actually manufactured according to the present application. In the drawings:

[0042] Figure 1 This is the first schematic flowchart of the method for optimizing the transmission quality of strong interaction video streams in an embodiment of the present application.

[0043] Figure 2 This is the second schematic flowchart of the method for optimizing the transmission quality of strong interaction video streams in an embodiment of the present application.

[0044] Figure 3 This is the schematic diagram of the fast module design details in an application example of the present application.

[0045] Figure 4 This is the third schematic flowchart of the method for optimizing the transmission quality of strong interaction video streams in an embodiment of the present application.

[0046] Figure 5 This is the schematic diagram of the slow module design details in an application example of the present application.

[0047] Figure 6 This is the schematic structural diagram of the frame-level forward error correction device in an embodiment of the present application.

[0048] Figure 7 This is the schematic structural diagram of the forward error correction controller in an embodiment of the present application.

[0049] Figure 8 This is the schematic structural diagram of the strong interaction video stream transmission system in an embodiment of the present application.

[0050] Figure 9 This is the schematic diagram of the cloud game system deployment in an application example of the present application.

[0051] Figure 10 This is the schematic diagram of executing the Tooth algorithm in the cloud game system in an application example of the present application. Detailed implementation manners

[0052] To make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below in combination with the implementation manners and the drawings. Herein, the illustrative implementation manners of the present application and their descriptions are used to explain the present application, but do not limit the present application.

[0053] Herein, it should also be noted that in order to avoid obscuring the present application due to unnecessary details, only the structures and / or processing steps closely related to the solution of the present application are shown in the drawings, while other details less related to the present application are omitted.

[0054] It should be emphasized that the term "including / containing" when used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.

[0055] Here, it should also be noted that, without special specification, the term "connection" in this text can not only refer to direct connection, but also indirect connection with intermediate substances.

[0056] In the following, embodiments of the present application will be described with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.

[0057] It should be noted that in the art, a large frame can refer to a video frame containing more than 10 data packets, and a small frame can refer to a video frame containing no more than 10 data packets. For example, if the number of data packets in video frame A is 25, then video frame A is determined to be a large frame; if the number of data packets in video frame B is 5, then video frame B is determined to be a small frame. Of course, the above division of large and small frames mentioned in the present application can also be set according to other data packet values, not limited to 10 data packets, and can be specifically set according to actual application needs.

[0058] In the field of real-time video, to ensure the quality of video transmission, recovering lost data packets is one of the key issues, and forward error correction (FEC) technology is a commonly used means. However, these algorithms have obvious defects:

[0059] (1) Most of the existing FEC algorithms are coarse-grained and apply a unified redundancy on almost all video frames. This setting method does not fully consider the size difference of video frames in cloud games. Since the frame sizes in cloud games vary greatly, the redundancy requirements for frames of different lengths are significantly different. At the same time, the LRIF (defined as the ratio of the lost data volume per frame to the total data volume per frame) of different frames also varies. However, the existing algorithms ignore these factors, resulting in too much redundancy added in large frames, causing bandwidth waste; while in small frames, due to insufficient redundancy, lost data packets cannot be effectively recovered, which in turn causes frequent lags and seriously affects the user's gaming experience.

[0060] (2) The network environment is complex and dynamically changing, and cloud game streaming media needs to adjust strategies according to the real-time state of the network. However, the existing algorithms are insufficient in this regard and it is difficult to dynamically determine the appropriate redundancy in real time according to the network. They only make decisions based on rough historical network packet loss information. In fact, rough network packet loss patterns, such as packet loss rates, cannot reflect the data loss situation of each video frame at a fine-grained level. Therefore, the existing algorithms are even more unable to effectively ensure video QoE and reduce redundant bandwidth overhead.

[0061] Based on this, in order to solve the problems existing in the existing forward error correction coding method for video streams, such as the inability to balance the transmission quality of strong interaction video streams and the redundancy cost in strong interaction streaming media such as cloud games, the embodiments of the present application respectively provide a method for optimizing the transmission quality of strong interaction video streams, a frame-granularity forward error correction device for executing the method for optimizing the transmission quality of strong interaction video streams, a forward error correction controller, a strong interaction video stream transmission system, an electronic device, a computer-readable storage medium, and a computer program product, which can effectively guarantee video QoE and reduce redundant bandwidth overhead.

[0062] Specifically, it will be described in detail through the following embodiments.

[0063] Based on this, the embodiments of the present application provide a method for optimizing the transmission quality of strong interaction video streams that can be implemented by a frame-granularity forward error correction device. Refer to Figure 1 , and the method for optimizing the transmission quality of strong interaction video streams specifically includes the following contents:

[0064] Step 100: Determine the predicted value of the network packet loss rate and the predicted value of packet loss clustering of the strong interaction video stream data in the next streaming media protocol feedback cycle according to the packet network feedback data corresponding to each historical streaming media protocol feedback cycle and the current streaming media protocol feedback cycle respectively.

[0065] It can be understood that the streaming media protocol is the RTCP (Real-Time Control Protocol) protocol. RTCP provides data distribution quality feedback information, which is part of the function of RTP as a transport protocol and it involves flow control and congestion control of other transport protocols. Correspondingly, the streaming media protocol feedback cycle can be abbreviated as the RTCP feedback cycle.

[0066] In one or more embodiments of the present application, the historical streaming media protocol feedback cycle, the current streaming media protocol feedback cycle, and the next streaming media protocol feedback cycle refer to the streaming media protocol feedback cycles at different times, and the historical streaming media protocol feedback cycle and the next streaming media protocol feedback cycle refer to the streaming media protocol feedback cycles before and after the current streaming media protocol feedback cycle respectively, with the current streaming media protocol feedback cycle as the reference. That is to say, in the execution stage of step 100, the next streaming media protocol feedback cycle refers to the future streaming media protocol feedback cycle of this stage.

[0067] Among them, after the data receiving end receives the video data packets from the remote end, it will periodically send data packet network feedback data (i.e., RTPFB feedback packets), recording the time of receiving each data packet and the sending time at the receiving data end, etc. That is to say, the data packet network feedback data corresponding to a streaming media protocol feedback cycle received by the data sending end needs to include the arrival time of each data packet previously transmitted over the network to the data receiving end, as well as information such as whether it has arrived or been lost.

[0068] It should be noted that the streaming media protocol feedback cycle is 100 ms (milliseconds). However, when the duration of the streaming media protocol feedback cycle changes or there are other possibilities of streaming media protocol feedback cycles greater than 16.7 ms in the future, the method for optimizing the transmission quality of strong interactive video streams provided in this application is still applicable.

[0069] In addition, it should be emphasized that the above step 100 predicts the predicted value of the network packet loss rate and the predicted value of packet loss aggregation in the next streaming media protocol feedback cycle in units of the streaming media protocol feedback cycle. It does not need to be executed frame by frame for each video frame, but only runs when new RTCP network feedback is received to reduce overhead. It can effectively reduce the cost required to obtain the predicted data on the premise of providing effective and reliable predicted data for the following step 200.

[0070] Step 200: Based on the predicted value of the network packet loss rate and the predicted value of packet loss aggregation in the next streaming media protocol feedback cycle corresponding to the strong interactive video stream data, and the frame length of each video frame corresponding to the strong interactive video stream data to be transmitted over the network, use the random forest model to make a decision on the redundancy of each video frame corresponding to each video frame coding cycle frame by frame, so as to perform forward error correction coding on each data packet corresponding to each video frame based on the redundancy of each video frame, and obtain each transmission quality optimized data packet corresponding to each video frame.

[0071] In step 200, the redundancy of the video frame refers to the fine-grained redundancy at the video frame level.

[0072] It can be understood that compared with the execution frequency in units of the streaming media protocol feedback cycle in step 100, step 200 is executed in units of the video frame coding cycle. That is, for each video frame corresponding to the strong interactive video stream data to be transmitted over the network, the random forest model is used to obtain the redundancy of each video frame respectively, so as to ensure accurate prediction of future network conditions for various historical network packet loss patterns and ensure the provision of key information in a compact coding cycle.

[0073] Furthermore, step 200 works during each video frame encoding by using a very lightweight machine learning model, namely the random forest model, which can ensure that a suitable redundancy decision for the current frame is determined within an extremely short time, meeting the high real-time requirements of strong interactive video streams.

[0074] The process of using the random forest model to make a decision on the redundancy of each video frame corresponding to each video frame encoding period in step 100 and step 200 above can be called the frame-level forward error correction algorithm (which can also be called the frame-level FEC algorithm or the frame-level dual-module FEC algorithm). This frame-level forward error correction algorithm can be abbreviated as: Tooth. That is to say, by first proposing the Tooth algorithm in this application, it is possible to predict the future network packet loss pattern based on historical network information and make a real-time decision on the redundancy of each frame based on a low-overhead machine learning algorithm.

[0075] As can be seen from the above description, the method for optimizing the transmission quality of strong interactive video streams provided by the embodiments of this application can accurately adapt to the size differences and network dynamics between different video frames of strong interactive video streams such as cloud games, ensure the interactive quality of strong interactive video streams, and can significantly reduce the redundant bandwidth overhead to balance the transmission quality and redundant cost of strong interactive video streams.

[0076] To further improve the application effectiveness and reliability of using the random forest model to make a decision on the redundancy of each video frame corresponding to each video frame encoding period, in a method for optimizing the transmission quality of strong interactive video streams provided by the embodiments of this application, refer to Figure 2 Step 200 in the method for optimizing the transmission quality of strong interactive video streams specifically includes the following content:

[0077] Step 210: Generate a dataset corresponding to each video frame based on the predicted value of the network packet loss rate and the predicted value of packet loss aggregation corresponding to the next streaming media protocol feedback period of the strong interactive video stream data, and the frame length of each video frame to be network-transmitted corresponding to the strong interactive video stream data, where each dataset includes the frame length of one video frame and the predicted value of the network packet loss rate and the predicted value of packet loss aggregation corresponding to the next streaming media protocol feedback period of the corresponding strong interactive video stream data.

[0078] Step 220: Input each dataset into the random forest model in sequence so that the random forest model outputs the redundancy of each video frame correspondingly.

[0079] Specifically, in the ensemble learning based on multiple decision trees, on the one hand, it can effectively handle high-cardinality samples. Each redundant decision tree will be constructed by randomly selecting training samples, thus avoiding the overfitting problem and ensuring sufficient generalization ability. On the other hand, each redundant decision tree will split its nodes based on different input features during the establishment process, so as to be able to learn the influence degree of each feature on the redundancy of each frame. Specifically, refer to Figure 3 , first, partition the data set according to the predicted network packet loss rate lr (which can also be written as lr f )、the predicted packet loss clustering la (which can also be written as la f ) and the frame length fl in the next streaming media protocol feedback cycle to obtain each subset, and at the same time ensure that the elements in each subset are evenly distributed to avoid the situation that a large amount of data is concentrated near a certain value. Subsequently, adjust the features carried by each subset. The explicit features carried by each subset are divided into (lr, fl), (la, fl) and (lr, la, fl). This strategy prevents the decision tree from overly relying on a certain feature. The final redundant decision of the RF model is the average value of the outputs of all decision trees. The combination of multiple decision trees can help the fast module learn the non-linear relationship between lr, la, fl and the redundancy of each frame, and autonomously evaluate the importance of each input feature, and finally significantly improve the generalization ability of the fast module in different network conditions and game environments.

[0080] In one or more embodiments of the present application, the predicted network packet loss rate can also be referred to as the estimated future packet loss rate, and the predicted packet loss clustering can also be referred to as the estimated future packet loss clustering degree. It can be understood that the predicted packet loss clustering value is greater than 0, and there is a positive correlation between the packet loss clustering degree and the predicted packet loss clustering value.

[0081] Step 230: Transmit the redundancy output by the random forest model each time to the forward error correction encoder, so that the forward error correction encoder performs forward error correction encoding on each data packet corresponding to the video frame according to the redundancy of the received video frame, and obtains each transmission quality optimized data packet corresponding to the video frame, so as to perform network transmission on each transmission quality optimized data packet corresponding to the video frame respectively.

[0082] In one or more embodiments of the present application, the forward error correction encoder can be written as the FEC encoder.

[0083] In order to further improve the application effectiveness and reliability of determining the predicted network packet loss rate and the predicted packet loss clustering value in the next streaming media protocol feedback cycle for strong interactive video stream data, in a method for optimizing the transmission quality of strong interactive video stream provided in the embodiments of the present application, refer to Figure 2 , step 100 in the method for optimizing the transmission quality of strong interactive video stream specifically includes the following content:

[0084] Step 110: Receive the packet network feedback data corresponding to the strong interactive video stream data within the current streaming media protocol feedback cycle.

[0085] Step 120: Construct the current network feedback sequence, network packet loss rate sequence, and packet loss clustering sequence respectively. Among them, the network feedback sequence includes: the packet network feedback data of the strong interactive video stream data within the current streaming media protocol feedback cycle and the packet network feedback data of the pre-acquired strong interactive video stream data corresponding to each historical streaming media protocol feedback cycle; the network packet loss rate sequence includes: the predicted network packet loss rate values of the pre-acquired strong interactive video stream data corresponding to each historical streaming media protocol feedback cycle and the current streaming media protocol feedback cycle; the packet loss clustering sequence includes: the predicted packet loss clustering values of the pre-acquired strong interactive video stream data corresponding to each historical streaming media protocol feedback cycle and the current streaming media protocol feedback cycle.

[0086] In an example, if the current streaming media protocol feedback cycle is the nth streaming media protocol feedback cycle, then the network packet loss rate sequence contains lr 1 ,lr 2 ,…lr n ,where lr 1 ,lr 2 ,…lr n-1 are respectively the predicted network packet loss rate values corresponding to n - 1 historical streaming media protocol feedback cycles; lr n is the predicted network packet loss rate value corresponding to the current streaming media protocol feedback cycle; and the packet loss clustering sequence contains la 1 ,la 2 ,…la n ,where la 1 ,la 2 ,…la n-1 are respectively the predicted packet loss clustering values corresponding to n - 1 historical streaming media protocol feedback cycles; la n is the predicted packet loss clustering value corresponding to the current streaming media protocol feedback cycle.

[0087] Step 130: Input the network feedback sequence, the network packet loss rate sequence, and the packet loss clustering sequence into a convolutional neural network, so that the convolutional neural network correspondingly outputs the predicted network packet loss rate value and the predicted packet loss clustering value of the strong interactive video stream data in the next streaming media protocol feedback cycle.

[0088] To further improve the application effectiveness and reliability of using a convolutional neural network to obtain the predicted values of network packet loss rate and packet loss clustering in the next streaming media protocol feedback cycle for the strong interaction video stream data, in a method for optimizing the transmission quality of a strong interaction video stream provided in an embodiment of the present application, refer to Figure 4 Step 130 in the method for optimizing the transmission quality of the strong interaction video stream specifically includes the following content:

[0089] Step 131: Input the network feedback sequence, the network packet loss rate sequence, and the packet loss clustering sequence into the convolutional neural network, so that the one-dimensional convolutional block in the convolutional neural network extracts spatial features from the network feedback sequence and outputs the spatial feature information corresponding to the network feedback sequence. Then, the fully connected layer in the convolutional neural network outputs the predicted value of the network packet loss rate and the predicted value of packet loss clustering for the strong interaction video stream data in the next streaming media protocol feedback cycle according to the spatial feature information, the network packet loss rate sequence, and the packet loss clustering sequence.

[0090] Specifically, refer to Figure 5 , first, the one-dimensional convolutional block (1D-CNN) processes the spatial distribution characteristics of historical packet loss events (i.e., the packet loss clustering sequence). The one-dimensional convolutional block consists of two cascaded one-dimensional convolutional layers Convl and Conv2. Each convolutional layer uses the ReLU activation function and the pooling layer (Max-pool) to extract more representative and key spatial feature information. Subsequently, the processed spatial feature information is combined with other input information (i.e., the network packet loss rate sequence and the packet loss clustering sequence ) and fed into the cascaded fully connected layer together. At this stage, the predicted value of the network packet loss rate lr (which can also be written as lr f ) and the predicted value of packet loss clustering la (which can also be written as la f ) for the next streaming media protocol feedback cycle are estimated.

[0091] To further improve the application effectiveness and reliability of using the convolutional neural network, in a method for optimizing the transmission quality of a strong interaction video stream provided in an embodiment of the present application, refer to Figure 4 , before step 100 in the method for optimizing the transmission quality of the strong interaction video stream, it specifically includes the following content:

[0092] Step 010: Train the convolutional neural network based on a preset loss function and historical training data, where the historical training data includes: a historical network feedback sequence corresponding to historical strong interaction video stream data, a historical network packet loss rate sequence, and a historical packet loss clustering sequence; wherein, the historical network feedback sequence includes: packet network feedback data corresponding to the historical strong interaction video stream data in respective historical streaming media protocol feedback cycles; the historical network packet loss rate sequence includes: predicted network packet loss rate values corresponding to the historical strong interaction video stream data in respective historical streaming media protocol feedback cycles; the historical packet loss clustering sequence includes: predicted packet loss clustering value columns corresponding to the historical strong interaction video stream data in respective historical streaming media protocol feedback cycles.

[0093] It should be noted that for the loss function, an overestimated sum will lead Tooth to add too much redundancy and increase the bandwidth cost. However, underestimating the future sum may result in insufficient redundancy, thus causing stuttering, which is more unacceptable than wasting bandwidth. Therefore, a larger loss is set for the underestimation case.

[0094] Therefore, the loss function L in step 131 of the embodiment of the present application S is shown in the following formula (1):

[0095]

[0096] In formula (1), represents the true predicted network packet loss rate value. is the true predicted packet loss clustering value. The first part on the left side of the multiplication sign is to calculate the estimated lr f and la f and their respective corresponding true values and The difference between them. For the second part on the right side of the multiplication sign, if the estimated value is greater than the true value, is a negative value, and the loss L S can be appropriately reduced; if the predicted value is less than the true value, the loss L S will increase. β is a constant to avoid L S from being 0.

[0097] In order to further improve the application effectiveness and reliability of the random forest model, in a method for optimizing the transmission quality of strong interaction video streams provided in the embodiment of the present application, refer to Figure 4 , before step 200 in the method for optimizing the transmission quality of strong interaction video streams, the following specific content is further included:

[0098] Step 020: Based on the predicted network packet loss rate value and packet loss aggregation prediction value corresponding to the historical strong interaction video stream data, and the frame length of each video frame corresponding to the historical strong interaction video stream data, train a random forest model with the mean squared error as the evaluation index of the random forest model.

[0099] From a software perspective, the present application also provides a frame-level forward error correction device for executing all or part of the strong interaction video stream transmission quality optimization method. Refer to Figure 6 , the frame-level forward error correction device specifically includes the following content:

[0100] Slow module 10, configured to determine the predicted network packet loss rate value and packet loss aggregation prediction value of the strong interaction video stream data in the next streaming media protocol feedback cycle according to the packet network feedback data corresponding to each historical streaming media protocol feedback cycle and the current streaming media protocol feedback cycle.

[0101] Fast module 20, configured to, based on the predicted network packet loss rate value and packet loss aggregation prediction value corresponding to the strong interaction video stream data in the next streaming media protocol feedback cycle, and the frame length of each video frame to be network-transmitted corresponding to the strong interaction video stream data, use a random forest model to make a decision on the redundancy of each video frame for each video frame coding cycle, and perform forward error correction coding on each data packet corresponding to each video frame based on the redundancy of each video frame, so as to obtain each transmission quality-optimized data packet corresponding to each video frame.

[0102] It can be understood that the frame-level forward error correction device can also be simply referred to as the Tooth device.

[0103] The embodiment of the frame-level forward error correction device provided by the present application can specifically be used to execute the processing flow of the embodiment of the strong interaction video stream transmission quality optimization method in the above embodiment, and its functions will not be elaborated here. Reference can be made to the detailed description of the embodiment of the strong interaction video stream transmission quality optimization method above.

[0104] The part of the frame-level forward error correction device for optimizing the strong interaction video stream transmission quality can be implemented in the server or completed in the client device. Specifically, it can be selected according to the processing capacity of the client device and the limitations of the user usage scenario. The present application does not limit this. If all operations are completed in the client device, the client device may further include a processor for specifically processing the strong interaction video stream transmission quality optimization.

[0105] The above-mentioned client device may have a communication module (i.e., communication unit), which can communicate with a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side, and in other implementation scenarios, it may also include a server of an intermediate platform, such as a server of a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, or a server cluster composed of multiple servers, or a server structure of a distributed device.

[0106] Any suitable network protocol can be used for communication between the above-mentioned server and the client device, including network protocols that have not been developed as of the filing date of this application. The network protocol may, for example, include TCP / IP protocol, UDP / IP protocol, HTTP protocol, HTTPS protocol, etc. Of course, the network protocol may also, for example, include the RPC protocol (Remote Procedure Call Protocol) and the REST protocol (Representational State Transfer) used on top of the above-mentioned protocols.

[0107] That is to say, the core of the embodiment of this application lies in using the frame-level FEC technology to accurately and adaptively handle the differences in the sizes of cloud game video frames and network dynamics, ensuring the quality of video stream interaction (QoE) and significantly reducing the redundant bandwidth overhead at the same time.

[0108] The technical problem solved by this application is mainly the adaptation of the redundant algorithm to network dynamics. The key problems of existing FEC algorithms are as follows: (1) In real-time video stream communication, the available bandwidth of the network is constantly changing, and the video stream coding bitrate will also change accordingly, ultimately leading to changes in the frame length (number of data packets). This results in a huge difference in the data loss rate per frame (LRIF). The coarse-grained FEC algorithm ignores the impact of frame length on redundancy, which will cause waste of bandwidth for large frames and inability to recover small frames, affecting the quality of video stream interaction and causing waste of bandwidth resources. (2) In the actual network environment, the change of the packet loss pattern is extremely complex, involving the distribution of packet loss events in time and space. The traditional network packet loss pattern only considers the overall packet loss rate of the network and cannot reflect the details of the change of the network packet loss pattern. Due to the lack of in-depth understanding of the details of the packet loss pattern, the FEC decision based on this cannot effectively balance video QoE and redundant costs.

[0109] Therefore, the embodiment of this application proposes a frame-level dual-module FEC algorithm - Tooth to address the above problems, which has two major innovations: (1) A fast module that predicts the future network packet loss pattern based on historical network information. (2) A fast module that makes real-time decisions on the redundancy of each frame based on a low-overhead machine learning algorithm.

[0110] As can be seen from the above description, the frame-level forward error correction device provided by the embodiments of the present application can accurately adapt to the size differences and network dynamics between different video frames of strong interaction video streams such as cloud games, can ensure the interaction quality of strong interaction video streams, and can greatly reduce the redundant bandwidth overhead to balance the transmission quality and redundant cost of strong interaction video streams.

[0111] Based on the embodiments of the above-mentioned method for optimizing the transmission quality of strong interaction video streams and / or the embodiments of the frame-level forward error correction device, the present application further provides an embodiment of a forward error correction controller. Refer to Figure 7 The forward error correction controller (which can be abbreviated as the FEC controller) specifically includes the following contents:

[0112] A frame-level forward error correction device 1 and a forward error correction encoder 2 that are communicatively connected;

[0113] The frame-level forward error correction device 1 is used to execute the method for optimizing the transmission quality of strong interaction video streams described in the foregoing embodiments, and transmit the redundancy degree output by the random forest model each time to the forward error correction encoder 2;

[0114] The forward error correction encoder 2 is used to perform forward error correction encoding on each data packet corresponding to the video frame according to the redundancy degree of the received video frame, and obtain each data packet after optimizing the transmission quality corresponding to the video frame.

[0115] Furthermore, the present application further provides a strong interaction video stream transmission system including the forward error correction controller. Refer to Figure 8 The strong interaction video stream transmission system specifically includes the following contents:

[0116] A data sending end device and a data receiving end device;

[0117] The data sending end device is provided with a video encoder 01, the forward error correction controller 02 provided in the foregoing embodiments, and a data sender 03 that are communicatively connected in sequence, and a network monitor 04 that is communicatively connected with the forward error correction controller 02;

[0118] The video encoder 01 is used to encode the target video data to obtain corresponding strong interaction video stream data;

[0119] The forward error correction controller 02 is further used to obtain the packet network feedback data corresponding to the current streaming media protocol feedback period transmitted by the network monitor, and transmit each data packet after optimizing the transmission quality corresponding to the video frame to the data sender 03;

[0120] The data transmitter 03 is used to transmit each packet with optimized transmission quality corresponding to the video frame to the data receiving end device via the network.

[0121] To better understand the QoE of strong interaction video streams in the real world, the inventors of this application conducted large-scale real-world measurements covering 66,128 players in 22 cities around the world. The measurement results reveal that the downstream transmission of game video content (i.e., from the remote server to the terminal device) is the main bottleneck affecting the interactive QoE.

[0122] Based on this, to further illustrate the above embodiments, this application also provides a specific application example of a method for optimizing the transmission quality of strong interaction video streams. To comprehensively evaluate the performance of different FEC methods in the strong interaction video stream scenario, the application example of this application was deployed and evaluated on a commercial cloud game video stream system on a large scale. In this application example, see Figure 9 , taking the cloud game server as an example of the data sending end device, taking the cloud game client as an example of the data receiving end device, and taking the cloud game system as an example of the strong interaction video stream transmission system, the method for optimizing the transmission quality of the strong interaction video stream is described in detail.

[0123] The specific contents of the two major innovations of this application example are as follows:

[0124] (1) A slow module that predicts future network packet loss patterns based on historical network information. This application example proposes to use a compressed neural network model to replace the coarse-grained FEC that only considers the loss rate, considering packet loss clustering to achieve accurate prediction. The so-called slow module runs once per network RTCP feedback cycle. For strong interaction video stream applications, this module considers historical packet loss patterns, learns network volatility, and estimates future network packet loss patterns (packet loss rate and packet loss clustering) packet by packet.

[0125] (2) A fast module that makes real-time decisions on the redundancy of each frame based on a low-overhead machine learning algorithm. This application example proposes to use a lightweight random forest method to replace the traditional coarse-grained redundancy coding strategy during redundancy coding. Specifically, the fast module runs at the video frame coding cycle (16.7 ms when FPS = 60), defines the non-linear and discontinuous mapping between LRIF-related factors and the redundancy of each frame as a regression problem, and uses a lightweight random forest method to model the mapping relationship. At the same time, it works in cooperation with the slow module, fully considering frame size differences and network dynamics, effectively determining the appropriate FEC redundancy for each frame, and achieving a balance between video QoE and redundancy cost.

[0126] Specifically, the cloud game system mainly consists of three parts: the cloud game server, the Internet forwarding node, and the client device.

[0127] (i) The cloud game server uses its powerful computing power to run the game engine, which converts the real-time game images into data packets through the video encoder and FEC encoder, and then sends them to the Internet through the data generator;

[0128] (ii) Internet forwarding nodes transmit the video stream data generated by the cloud game server to the game client in real time, which covers a variety of real and complex network environment conditions, including different network bandwidth conditions, packet loss patterns, and network latency levels;

[0129] (iii) The client device is responsible for receiving the video stream transmitted from the server, and has the function of recording various relevant parameters during the entire video playback process. These recorded parameters will provide detailed data support for subsequent in-depth and detailed analysis work, and help to fully understand the various situations during the video stream transmission process.

[0130] Based on this cloud gaming system, the designers of this application have carried out a series of targeted measurements. On the one hand, this application example has conducted an in-depth study of the transmission characteristics of highly interactive video streams, and analyzed key transmission factors such as video stream interaction delay bottlenecks, video frame size, and network packet loss patterns. On the other hand, the designers of this application have fully deployed multiple existing FEC methods, and accurately and meticulously analyzed their transmission performance in highly interactive video stream scenarios in a real network environment.

[0131] Based on the above system framework, the specific description of the Tooth algorithm (frame granularity FEC algorithm) used in the strong interactive video stream transmission quality optimization method provided in the specific application example of this application is as follows:

[0132] Tooth is a frame-granularity FEC algorithm designed for highly interactive video streams such as cloud games. The algorithm composition and operation process are as follows: Figure 10 As shown in the figure, during the video stream transmission process, Tooth runs the fast module and the slow module to make decisions. Due to the fluctuation of network bandwidth, the sender will accurately consider the video frame length and network dynamic factors, and adaptively set the appropriate FEC coding redundancy for each video frame to achieve lower redundancy overhead, thereby improving video QoE.

[0133] In the training framework of this application example, it is divided into two modules. The slow module uses a neural network model to learn network fluctuations to predict future packet loss patterns. The input is data related to historical packet loss patterns during video transmission, including network packet loss sequences and packet loss sequences within multiple past RTCP (Real-Time Transport Control Protocol) feedback cycles. It does not need to execute for each video frame and only runs when new RTCP network feedback is received to reduce overhead. Finally, it generates the predicted network packet loss rate and packet loss clustering. The fast module uses a very lightweight machine learning model. The input is the video frame length and the future network packet loss pattern obtained from the slow module, and it determines the redundancy of each frame frame by frame. This application example uses a neural network model, ensuring that the model can make accurate predictions about future network conditions for various historical network packet loss patterns and providing key information for the fast module in a compact encoding cycle. At the same time, a lightweight random forest is used, which works during each video frame encoding, ensuring that a suitable redundancy decision can be determined for the current frame within an extremely short time and meeting the high real-time requirements of strong interactive video streams. Specifically:

[0134] (1) This application example designs a slow module for real-time video transmission to predict future network packet loss patterns based on historical network packet loss information. The task of the slow module is to learn network fluctuations and estimate the future network packet loss rate lr f and packet loss clustering la f , and update them to the fast module. Obviously, estimating packet loss clustering is very challenging because, in addition to its own intrinsic characteristics, it may also be affected by the future network packet loss rate. Existing simple methods, such as the arithmetic mean of heuristic observations, will bring serious biases to the judgment of per-frame redundancy. Therefore, this application example introduces a CNN neural network model to construct it, as Figure 5 shown.

[0135] When processing this information, first, the spatial distribution characteristics of historical packet loss events of packets are processed through a one-dimensional convolutional 1D-CNN block, which consists of two cascaded one-dimensional convolutional layers Conv1 and Conv2. Each convolutional layer uses a ReLU activation function and a pooling layer to extract more representative and key spatial feature information. Subsequently, the processed spatial feature information is combined with other input information (i.e., the network packet loss rate sequence and the packet loss clustering sequence ) and fed into a cascaded fully connected network. At this stage, the slow module estimates the future network packet loss rate lr f and packet loss clustering la fFor the estimation, this application example emphasizes that for the loss function, an overestimated sum will lead Tooth to add too much redundancy, increasing the bandwidth cost. However, underestimating the future sum may result in insufficient redundancy, leading to stuttering, which is more unacceptable than wasting bandwidth. Therefore, a larger loss is set for the underestimation case. Finally, the loss function shown in the aforementioned formula (1) is adopted for the slow module among them.

[0136] (2) Since the input has high cardinality, the feature space essentially exhibits variable and complex characteristics, and each feature has different non-linear effects on the redundancy per frame. This application example designs a fast module based on the Random Forest (RF) model for real-time video transmission, and its core structure is multiple LRIF decision trees, as Figure 3 shown.

[0137] Based on the ensemble learning of multiple decision trees, on the one hand, it can effectively handle high-cardinality samples. Each redundant decision tree will be constructed by randomly selecting training samples, thus avoiding the overfitting problem and ensuring sufficient generalization ability; on the other hand, each redundant decision tree will split its nodes based on different input features during the establishment process, so as to be able to learn the influence degree of each feature on the redundancy per frame. As Figure 3 shown, first, the dataset is partitioned according to lr, la, and the frame length fl, while ensuring that the elements in each dataset are evenly distributed to avoid the situation where a large amount of data is concentrated near a certain value. Subsequently, the features carried by each subset are adjusted. The explicit features carried by each subset are divided into (lr, fl), (la, fl), and (lr, la, fl). This strategy prevents the decision tree from relying too much on a certain feature. In addition, the Mean Squared Error (MSE) is selected as the criterion for the RF model to train the fast module, and the final redundant decision of the RF model is the average value of the outputs of all decision trees. The combination of multiple decision trees can help the fast module learn the non-linear relationship between lr, la, fl, and the redundancy per frame, and autonomously evaluate the importance of each input feature, ultimately significantly improving the generalization ability of the fast module in different network conditions and game environments.

[0138] That is to say, this application example of the present application proposes a frame-granularity FEC flow control algorithm. By precisely considering the frame size difference and the network dynamic packet loss pattern, a unique algorithm and strategy are adopted to determine the redundancy per frame, so as to achieve the best balance between the QoE of the strong-interaction video stream and the redundant bandwidth overhead. This application example also proposes a dual-module FEC encoding technology. The slow module estimates the future network packet loss pattern; the fast module determines the redundancy per frame according to the received information of the slow module and the frame size. The two modules work together to effectively cope with the network dynamic changes while reducing the computational and inference time overhead.

[0139] Based on this, the frame-granularity FEC technical solution proposed in this application example has important application value. In a commercial strong-interaction video stream environment, compared with the existing technologies, this application example significantly reduces the interaction pause frequency by 40.2% - 85.2%, greatly improving the video fluency. At the same time, the video bit rate is increased by 11.4% - 29.2%, and the video quality is significantly improved. In addition, the bandwidth cost is reduced by 54.9% - 75.0%, effectively improving the utilization efficiency of network resources. In various environments such as different network types (such as WiFi, 4G, 5G), game types (such as 2D games and 3D games), Internet service providers (ISPs), and the distance from the game server to the terminal device (cross-city and in-city sessions), this application example can achieve excellent QoE, indicating that this application example has good generality and adaptability, providing strong technical support for the development of interactive video streams.

[0140] The embodiment of this application also provides an electronic device, which may include a processor, a memory, a receiver, and a transmitter. The processor is used to execute the method for optimizing the transmission quality of strong-interaction video streams mentioned in the above embodiment. The processor and the memory may be connected through a bus or other means. Taking the connection through the bus as an example, the receiver can be connected to the processor and the memory in a wired or wireless manner.

[0141] The processor may be a Central Processing Unit (CPU). The processor may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. chips, or combinations of the above types of chips.

[0142] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the method for optimizing the transmission quality of strong-interaction video streams in the embodiment of this application. The processor executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions, and modules stored in the memory, that is, implementing the method for optimizing the transmission quality of strong-interaction video streams in the above method embodiment.

[0143] The memory may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data created by the processor and the like. In addition, the memory may include high-speed random access memory and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0144] The one or more modules are stored in the memory and, when executed by the processor, implement the method for optimizing the quality of strong interaction video stream transmission in the embodiments.

[0145] In some embodiments of the present application, the user equipment may include a processor, a memory, and a transceiver unit. The transceiver unit may include a receiver and a transmitter. The processor, the memory, the receiver, and the transmitter may be connected through a bus system. The memory is used to store computer instructions, and the processor is used to execute the computer instructions stored in the memory to control the transceiver unit to transmit and receive signals.

[0146] As an implementation manner, the functions of the receiver and the transmitter in the present application may be implemented by considering a transceiver circuit or a dedicated chip for transceiver. The processor may be implemented by considering a dedicated processing chip, a processing circuit, or a general-purpose chip.

[0147] As another implementation manner, it may be considered to use a general-purpose computer to implement the server provided in the embodiments of the present application. That is, the program codes for implementing the functions of the processor, the receiver, and the transmitter are stored in the memory, and the general-purpose processor implements the functions of the processor, the receiver, and the transmitter by executing the codes in the memory.

[0148] The embodiments of the present application further provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the foregoing method for optimizing the quality of strong interaction video stream transmission are implemented. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium well-known in the technical field.

[0149] The embodiments of the present application further provide a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the foregoing method for optimizing the quality of strong interaction video stream transmission are implemented.

[0150] The embodiment of the present application further provides a computer program product, including a computer program, which when executed by a processor, implements the steps of the foregoing method for optimizing the transmission quality of strong interaction video streams.

[0151] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in connection with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present application are programs or code segments used to execute the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave on a transmission medium or a communication link.

[0152] It should be clear that the present application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, the detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.

[0153] In the present application, the features described and / or illustrated for one embodiment can be used in the same or similar manner in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.

[0154] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and variations can be made to the embodiments of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A method for optimizing the transmission quality of highly interactive video streams, characterized in that: include: Determine a network packet loss rate prediction value and a packet loss aggregation prediction value of the strongly interactive video stream data in the next streaming protocol feedback cycle according to the network feedback data of the data packets corresponding to each historical streaming protocol feedback cycle and the current streaming protocol feedback cycle; Based on the predicted value of the network packet loss rate and the predicted value of the packet loss aggregation corresponding to the strongly interactive video stream data in the next streaming media protocol feedback cycle, and the frame lengths of each video frame to be transmitted over the network corresponding to the strongly interactive video stream data, a random forest model is used to decide the redundancy of the video frames corresponding to each video frame encoding cycle frame by frame, so as to perform forward error correction encoding on each data packet corresponding to each video frame based on the redundancy of each video frame, so as to obtain each data packet with optimized transmission quality corresponding to each video frame.

2. The method for optimizing the transmission quality of highly interactive video streams according to claim 1, characterized in that: The method adopts a random forest model to decide the redundancy of each video frame corresponding to each video frame encoding period frame by frame based on the network packet loss rate prediction value and packet loss aggregation prediction value corresponding to the next streaming media protocol feedback cycle of the strongly interactive video stream data, and the frame length of each video frame to be transmitted over the network corresponding to the strongly interactive video stream data, so as to perform forward error correction encoding on each data packet corresponding to each video frame based on the redundancy of each video frame, and obtain each data packet after transmission quality optimization corresponding to each video frame, including: Based on the network packet loss rate prediction value and packet loss aggregation prediction value corresponding to the strongly interactive video stream data in the next streaming media protocol feedback cycle, and the frame lengths of the respective video frames to be transmitted over the network corresponding to the strongly interactive video stream data, a data set corresponding to each of the video frames is generated, wherein each of the data sets contains a frame length of the video frame and the corresponding network packet loss rate prediction value and packet loss aggregation prediction value corresponding to the strongly interactive video stream data in the next streaming media protocol feedback cycle; Inputting each of the data sets into the random forest model in sequence, so that the random forest model outputs the redundancy of each of the video frames accordingly; The redundancy outputted by the random forest model each time is transmitted to a forward error correction encoder, so that the forward error correction encoder performs forward error correction encoding on each data packet corresponding to the video frame according to the redundancy of the received video frame, and obtains each data packet after transmission quality optimization corresponding to the video frame, so as to respectively transmit each data packet after transmission quality optimization corresponding to the video frame through the network.

3. The method for optimizing the transmission quality of strongly interactive video streams according to claim 1, characterized in that: Determining the network packet loss rate prediction value and the packet loss aggregation prediction value of the strong interactive video stream data in the next streaming protocol feedback cycle according to the data packet network feedback data corresponding to each historical streaming protocol feedback cycle and the current streaming protocol feedback cycle includes: Receive data packet network feedback data corresponding to the strongly interactive video stream data within a current streaming media protocol feedback cycle; The current network feedback sequence, network packet loss rate sequence and packet loss aggregation sequence are respectively constructed, wherein the network feedback sequence includes: the data packet network feedback data of the strong interactive video stream data in the current streaming media protocol feedback cycle and the data packet network feedback data of the pre-acquired strong interactive video stream data corresponding to each historical streaming media protocol feedback cycle; the network packet loss rate sequence includes: the network packet loss rate prediction values ​​of the pre-acquired strong interactive video stream data corresponding to each historical streaming media protocol feedback cycle and the current streaming media protocol feedback cycle; the packet loss aggregation sequence includes: the packet loss aggregation prediction values ​​of the pre-acquired strong interactive video stream data corresponding to each historical streaming media protocol feedback cycle and the current streaming media protocol feedback cycle; The network feedback sequence, the network packet loss rate sequence and the packet loss aggregation sequence are input into a convolutional neural network so that the convolutional neural network correspondingly outputs a network packet loss rate prediction value and a packet loss aggregation prediction value of the strongly interactive video stream data in the next streaming media protocol feedback cycle.

4. The method for optimizing the transmission quality of strongly interactive video streams according to claim 3, characterized in that: The step of inputting the network feedback sequence, the network packet loss rate sequence, and the packet loss aggregation sequence into a convolutional neural network so that the convolutional neural network correspondingly outputs a network packet loss rate prediction value and a packet loss aggregation prediction value of the strongly interactive video stream data in the next streaming media protocol feedback cycle comprises: The network feedback sequence, the network packet loss rate sequence and the packet loss aggregation sequence are input into a convolutional neural network, so that the one-dimensional convolution block in the convolutional neural network extracts spatial features of the network feedback sequence and outputs spatial feature information corresponding to the network feedback sequence, and then the fully connected layer in the convolutional neural network outputs the network packet loss rate prediction value and the packet loss aggregation prediction value of the strongly interactive video stream data in the next streaming media protocol feedback cycle according to the spatial feature information, the network packet loss rate sequence and the packet loss aggregation sequence.

5. The method for optimizing the transmission quality of strongly interactive video streams according to claim 3, characterized in that: Before determining the network packet loss rate prediction value and the packet loss aggregation prediction value of the strong interactive video stream data in the next streaming protocol feedback cycle according to the data packet network feedback data corresponding to each historical streaming protocol feedback cycle and the current streaming protocol feedback cycle, the method further includes: The convolutional neural network is trained based on a preset loss function and historical training data, wherein the historical training data includes: a historical network feedback sequence, a historical network packet loss rate sequence and a historical packet loss aggregation sequence corresponding to historical strongly interactive video stream data; wherein the historical network feedback sequence includes: network feedback data of data packets corresponding to each historical streaming media protocol feedback cycle of the historical strongly interactive video stream data; the historical network packet loss rate sequence includes: network packet loss rate prediction values ​​corresponding to each historical streaming media protocol feedback cycle of the historical strongly interactive video stream data; the historical packet loss aggregation sequence includes: packet loss aggregation prediction value columns corresponding to each historical streaming media protocol feedback cycle of the historical strongly interactive video stream data.

6. The method for optimizing transmission quality of highly interactive video streams according to claim 1, characterized in that: Before using a random forest model to decide the redundancy of the video frames corresponding to each video frame encoding cycle frame by frame based on the network packet loss rate prediction value and the packet loss aggregation prediction value corresponding to the next streaming media protocol feedback cycle of the strongly interactive video stream data, and the frame lengths of the respective video frames to be transmitted over the network corresponding to the strongly interactive video stream data, the method further includes: Based on the network packet loss rate prediction value and packet loss aggregation prediction value corresponding to the historical strong interactive video stream data, and the frame lengths of each video frame corresponding to the historical strong interactive video stream data, the random forest model is trained using the mean square error as the evaluation indicator of the random forest model.

7. A frame granularity forward error correction device, characterized in that: include: The slow module is used to determine the network packet loss rate prediction value and packet loss aggregation prediction value of the strong interactive video stream data in the next streaming protocol feedback cycle according to the network feedback data of the data packets corresponding to each historical streaming protocol feedback cycle and the current streaming protocol feedback cycle; The fast module is used to use a random forest model to decide the redundancy of each video frame corresponding to each video frame encoding cycle frame by frame based on the network packet loss rate prediction value and packet loss aggregation prediction value corresponding to the next streaming media protocol feedback cycle of the strongly interactive video stream data, and the frame length of each video frame to be transmitted over the network corresponding to the strongly interactive video stream data, so as to perform forward error correction encoding on each data packet corresponding to each video frame based on the redundancy of each video frame, so as to obtain each data packet corresponding to each video frame after transmission quality optimization.

8. A forward error correction controller, characterized in that: include: A frame granularity forward error correction device and a forward error correction encoder for the communication connection; The frame granularity forward error correction device is used to execute the strong interactive video stream transmission quality optimization method according to any one of claims 1 to 6, and transmit the redundancy output by the random forest model each time to the forward error correction encoder; The forward error correction encoder is used to perform forward error correction encoding on each data packet corresponding to the received video frame according to the redundancy of the received video frame, so as to obtain each data packet after transmission quality optimization corresponding to the video frame.

9. A highly interactive video streaming transmission system, characterized in that: include: Data sending end device and data receiving end device; The data transmitting end device is provided with a video encoder, a forward error correction controller as claimed in claim 8 and a data transmitter which are communicatively connected in sequence, and a network monitor which is communicatively connected to the forward error correction controller; The video encoder is used to encode the target video data to obtain corresponding strongly interactive video stream data; The forward error correction controller is also used to transmit each data packet after transmission quality optimization corresponding to the obtained video frame to the data transmitter; The data transmitter is used to transmit each data packet after transmission quality optimization corresponding to the video frame to the data receiving end device via the network.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for optimizing the transmission quality of highly interactive video streams as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Wireless low-delay signal and data transmission method based on hierarchical redundancy mechanism

    CN108880754A

  • Deep learning adaptive forward error correction transmission method based on video frame

    CN117856973A

  • Adaptive real-time video transmission FEC (Forward Error Correction) control system and method based on feedback of receiver

    CN117979033A

  • Placement of repair packets in multimedia data streams

    WO2022269642A1

Cited By

  • Video processing method and device and electronic equipment

    CN120416539A

  • Forward error correction control method and device for real-time video stream and computer equipment

    CN121356735A

  • Video quality defect intelligent detection system based on CV model and big data analysis

    CN121924289A