Video stream data processing method and device, computer device and readable storage medium
Patent Information
- Application Number
- CN202611033404.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-13
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2046-07-13
AI Technical Summary
这种方式无法根据网络状态进行精准预测和提前调整
[0027]The aforementioned video stream data processing method, apparatus, computer equipment, and computer-readable storage medium predict future network states based on historical network state data to obtain network state prediction data. The acquired terminal device performance data and network state prediction data are configured together as state dimension parameters. A perceptual constraint function is generated based on the action dimension parameters. This perceptual constraint function is used to adjust the near-end strategy to optimize the final output action decision of the network. An effect reward function is generated based on the video stream's quality and bitrate metrics. The near-end strategy optimizes the network output video stream transmission strategy based on the state dimension parameters, the effect reward function, and the perceptual constraint function. The video stream transmission strategy is used to adjust the encoding bitrate and quality parameters of the video stream during transmission. This video stream transmission strategy is an adjustment strategy based on network state prediction data and terminal device performance data, which can optimize the encoding bitrate and quality parameters of the video stream according to the network state and terminal device performance data, thereby improving the transmission quality and stability of real-time video communication scenarios.
Smart Images

Figure CN122534247B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video transmission technology, and in particular to a video stream data processing method, apparatus, computer equipment, and readable storage medium. Background Technology
[0002] WebRTC (Web Real-Time communication), as an open real-time communication technology, has been widely used in various real-time audio and video communication scenarios. From video conferencing, online education, live interactive streaming, and cloud gaming, efficient transmission of video data is crucial. However, under conditions of fluctuating network environments, the video encoding and transmission process faces even greater challenges.
[0003] Existing technologies adapt to changes in network conditions by dynamically adjusting the video bitrate based on monitoring network status or predictions of network status. However, this approach sacrifices video transmission quality in complex network environments and can also impact transmission efficiency to some extent.
[0004] In addition, existing technologies can reduce the impact of network fluctuations on transmission efficiency by coupling video coding and transmission strategies. However, this approach cannot accurately predict and adjust based on network conditions in advance.
[0005] In summary, existing technologies struggle to adaptively adjust video bitrate and transmission strategies based on network conditions, making it impossible to balance the quality and stability of real-time video communication. Summary of the Invention
[0006] Therefore, it is necessary to provide a video stream data processing method, apparatus, computer equipment, and readable storage medium that can improve the quality and stability of real-time video communication in response to the above-mentioned technical problems.
[0007] Firstly, this application provides a video stream data processing method, including:
[0008] Network state prediction data is obtained by predicting the future network state based on historical network state data.
[0009] The acquired device performance data of the terminal device and the network state prediction data are jointly configured as the state dimension parameters of the near-end policy optimization network, and an effect reward function is generated based on the quality index and bitrate index of the video stream. The state dimension parameters are used to describe the state space of the near-end policy optimization network, and the effect reward function is used to constrain the optimization effect of the action dimension parameters in the near-end policy optimization network. The terminal device includes the sending end and / or receiving end of the video stream.
[0010] A perceptual constraint function is generated based on the action dimension parameters. The perceptual constraint function is used to adjust the action decision output by the proximal policy optimization network.
[0011] Based on the state dimension parameters, the effect reward function, and the perception constraint function, the network output video stream transmission strategy is optimized through the near-end strategy. The video stream transmission strategy is used to adjust the encoding bitrate and quality parameters of the video stream when transmitting the video stream.
[0012] Secondly, this application also provides a video stream data processing apparatus, comprising:
[0013] The network state prediction module is used to predict the future network state based on historical network state data, and obtain network state prediction data.
[0014] The parameter configuration module is used to configure the acquired device performance data of the terminal device and the network state prediction data together as the state dimension parameters of the near-end policy optimization network, and generate an effect reward function based on the quality index and bitrate index of the video stream. The state dimension parameters are used to describe the state space of the near-end policy optimization network, and the effect reward function is used to constrain the optimization effect of the action dimension parameters in the near-end policy optimization network. The terminal device includes the sending end and / or receiving end of the video stream.
[0015] The action dimension setting module is used to generate a perceptual constraint function based on the action dimension parameters. The perceptual constraint function is used to adjust the action decision output by the proximal policy optimization network.
[0016] The transmission strategy output module is used to optimize the network output video stream transmission strategy according to the state dimension parameters, the effect reward function, and the perception constraint function through the near-end strategy. The video stream transmission strategy is used to adjust the encoding bitrate and quality parameters of the video stream when transmitting the video stream.
[0017] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0018] Network state prediction data is obtained by predicting the future network state based on historical network state data.
[0019] The acquired device performance data of the terminal device and the network state prediction data are configured together as state dimension parameters, and an effect reward function is generated based on the quality index and bitrate index of the video stream. The state dimension parameters are used to describe the state space of the near-end policy optimization network, and the effect reward function is used to constrain the optimization effect of the action dimension parameters in the near-end policy optimization network. The terminal device includes the sending end and / or receiving end of the video stream.
[0020] This is used to generate a perceptual constraint function based on the action dimension parameters, and the perceptual constraint function is used to adjust the action decision output by the proximal policy optimization network.
[0021] Based on the state dimension parameters, the effect reward function, and the perception constraint function, the network output video stream transmission strategy is optimized through the near-end strategy. The video stream transmission strategy is used to adjust the encoding bitrate and quality parameters of the video stream when transmitting the video stream.
[0022] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0023] Network state prediction data is obtained by predicting the future network state based on historical network state data.
[0024] The acquired device performance data of the terminal device and the network state prediction data are configured together as state dimension parameters, and an effect reward function is generated based on the quality index and bitrate index of the video stream. The state dimension parameters are used to describe the state space of the near-end policy optimization network, and the effect reward function is used to constrain the optimization effect of the action dimension parameters in the near-end policy optimization network. The terminal device includes the sending end and / or receiving end of the video stream.
[0025] A perceptual constraint function is generated based on the action dimension parameters. The perceptual constraint function is used to adjust the action decision output by the proximal policy optimization network.
[0026] Based on the state dimension parameters, the effect reward function, and the perception constraint function, the network output video stream transmission strategy is optimized through the near-end strategy. The video stream transmission strategy is used to adjust the encoding bitrate and quality parameters of the video stream when transmitting the video stream.
[0027] The aforementioned video stream data processing method, apparatus, computer equipment, and computer-readable storage medium predict future network states based on historical network state data to obtain network state prediction data. The acquired terminal device performance data and network state prediction data are configured together as state dimension parameters. A perceptual constraint function is generated based on the action dimension parameters. This perceptual constraint function is used to adjust the near-end strategy to optimize the final output action decision of the network. An effect reward function is generated based on the video stream's quality and bitrate metrics. The near-end strategy optimizes the network output video stream transmission strategy based on the state dimension parameters, the effect reward function, and the perceptual constraint function. The video stream transmission strategy is used to adjust the encoding bitrate and quality parameters of the video stream during transmission. This video stream transmission strategy is an adjustment strategy based on network state prediction data and terminal device performance data, which can optimize the encoding bitrate and quality parameters of the video stream according to the network state and terminal device performance data, thereby improving the transmission quality and stability of real-time video communication scenarios. Attached Figure Description
[0028] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0029] Figure 1 This is an application environment diagram of a video stream data processing method in one embodiment;
[0030] Figure 2 This is a flowchart illustrating a video stream data processing method in one embodiment;
[0031] Figure 3 This is a flowchart illustrating a video stream data processing method in another embodiment;
[0032] Figure 4 This is a flowchart illustrating a video stream data processing method in another embodiment;
[0033] Figure 5 This is a flowchart illustrating a video stream data processing method in another embodiment;
[0034] Figure 6 This is a structural block diagram of a video stream data processing device in one embodiment;
[0035] Figure 7 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0036] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0037] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.
[0038] The video stream data processing method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or located on the cloud or other network servers. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, drones, low-altitude aircraft, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc. Head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0039] In this embodiment, WebRTC (Web Real-Time Communication) is used in real-time video communication scenarios. As an open real-time communication technology standard, it is widely used in various real-time audio and video communication scenarios and can realize real-time video communication between browsers.
[0040] Real-time video communication is a process of exchanging real-time video data between at least two communication nodes over a network. A communication node refers to a terminal device with network access and display capabilities. The sending end is the source end that initiates or provides the video data stream, and the receiving end is the destination end device that receives and displays the video data stream.
[0041] In different real-time video communication scenarios, the sending terminal and the receiving terminal can interchange or coexist based on the interaction logic of the communication scenario. That is, there is a behavior of the terminal device transmitting video stream data to other terminal devices, and there is also a behavior of receiving video stream data sent by other terminal devices.
[0042] In one exemplary embodiment, such as Figure 2 As shown, a video stream data processing method is provided, which can be applied to... Figure 1 Taking the server in the example, the explanation includes the following steps S202 to S206. Wherein:
[0043] Step S202: Based on historical network state data, predict the future network state to obtain network state prediction data.
[0044] Specifically, historical network status data of terminal devices is collected, and future network status is predicted based on the historical network status data to obtain network status prediction data.
[0045] Terminal equipment refers to devices that exchange real-time video data over a network. Terminal equipment has network access capabilities and can collect, encode, and transmit video data to other terminal equipment.
[0046] Network state prediction data is a predicted value that forecasts the network state for a future time period. Specifically, the network state prediction data is set to be the average value of network state values within 500ms, including the average bandwidth within 500ms as the predicted bandwidth, the average latency within 500ms as the predicted latency, and the average packet loss rate within 500ms as the predicted packet loss rate.
[0047] Historical network status data refers to the network status parameters of terminal devices over a certain period of time.
[0048] When making predictions, a training dataset is first constructed, with the collection window set to the historical network state data of the past 30 seconds. For each timestamp, with one second as a time step, the following features are extracted:
[0049] Bandwidth sequence: Extracts bandwidth values over the past 30 time steps; Delay sequence: Extracts delay values over the past 30 time steps; Packet loss rate sequence: Packet loss rate over the past 30 time steps.
[0050] In addition to acquiring historical network status data, it also acquires current video frame characteristics and device performance parameters of the terminal device.
[0051] Among them, video frame features are used to characterize the quality parameters during video stream transmission. Quality parameters include video frame resolution, frame rate, and complexity metrics.
[0052] Complexity metrics can be calculated based on the resolution, frame rate, bit rate, and quantization parameters of the video stream. The higher the video quality, the higher the complexity.
[0053] The device performance parameters of a terminal device are used to characterize the device performance of the terminal device, including the CPU utilization, GPU utilization, and memory usage.
[0054] As an alternative implementation, the future network state is predicted using a Transformer encoder-decoder-based model architecture.
[0055] The Transformer model architecture consists of three layers: Input Feature Layer: This layer integrates multi-dimensional input features, including network state history parameters, video frame features, and device performance parameters. Transformer Encoder: This layer employs a 6-layer stacked structure, with each layer containing a multi-head self-attention mechanism and a feedforward neural network. The self-attention mechanism captures the dependencies between any positions of the input features. Transformer Decoder: This layer also employs a 6-layer stacked structure, using a cross-attention mechanism to obtain contextual information from the encoder output, enabling prediction of future network states and generating network state prediction data.
[0056] In the input feature layer, the collected multi-dimensional input features are transformed into a vector representation of a unified dimension through an embedding layer. This embedded vector representation is then input to the Transformer encoder, where it undergoes six layers of encoding to obtain a feature representation with global dependency information. The Transformer decoder receives the feature representation output from the Transformer encoder and uses it as context information. Simultaneously, it receives the vector representation output from the embedding layer and, through a cross-attention mechanism, focuses on information in the feature representation related to the current prediction data, generating a prediction sequence of future network states—i.e., network state prediction data. This network state prediction data includes state prediction data for a certain future time period, as well as the quality and bitrate features of the current video frame. The video frame's quality features correspond to resolution, frame rate, and quantization parameters, while the video frame's bitrate features correspond to the current video encoding bitrate.
[0057] Step S204: The acquired device performance data and network state prediction data of the terminal device are configured together as the state dimension parameters of the near-end policy optimization network, and an effect reward function is generated based on the quality index and bitrate index of the video stream. The state dimension parameters are used to describe the state space of the near-end policy optimization network, and the effect reward function is used to constrain the optimization effect of the near-end policy optimization network. The terminal device includes the sending end and / or receiving end of the video stream.
[0058] Specifically, the Proximal Policy Optimization Network is a deep learning network based on PPO (Proximal Policy Optimization), which improves decision-making ability while strictly controlling the magnitude of each update.
[0059] The agent in the proximal policy optimization network perceives the external environment through the state space and maps the perception results to each specific output in the action space.
[0060] The state space refers to the set of all possible information that an agent observes from the environment, and defines the input information for the agent's decision-making.
[0061] In this embodiment, the device performance data and network state prediction data of the terminal device are configured as state dimension parameters of the state space, representing the set of all possible information observed by the agent from the environment.
[0062] The network status prediction data includes predicted bandwidth, predicted latency, current bitrate, current frame rate, and current resolution. Device performance data includes CPU load, GPU load, and memory usage.
[0063] Predicted bandwidth is the average bandwidth over the next 500ms, and predicted latency is the average latency over the next 500ms. Current bitrate refers to the video encoding bitrate at the current moment, current frame rate refers to the video frame rate at the current moment, and current resolution refers to the video resolution at the current moment.
[0064] CPU load refers to the CPU utilization rate of a terminal device, GPU load refers to the GPU utilization rate of a terminal device, and memory usage refers to the memory utilization rate of a terminal device.
[0065] Video stream quality metrics characterize the quality of a video stream, and are reflected by the video stream's resolution and frame rate. Video stream bitrate metrics characterize the video stream's encoding bitrate, and are reflected by the video stream's bitrate.
[0066] The effect reward function is an external scalar signal used in near-end policy optimization networks to evaluate the quality of an agent's behavior. In each decision step, the agent performs an action based on the parameters of each state dimension in the current state space and returns a reward value based on the effect reward function. The quality and bitrate metrics of the video stream affect the magnitude of the reward value.
[0067] In this context, terminal equipment refers to the sending end of the video stream and / or the receiving end of the video stream.
[0068] The video encoding bitrate is a core conversion parameter between the network connection status and the video content quality. It converts limited network bandwidth into the amount of video data that can be transmitted per unit time, and directly determines the level of detail of the image that the receiving end can reproduce.
[0069] In real-time video communication, the sending and receiving ends of a video stream can be the same terminal device that simultaneously acts as both the sender and receiver; alternatively, one terminal device can act as the sender, and the other as the receiver. Understandably, in unidirectional transmission scenarios such as live streaming, online education, and cloud gaming, one terminal device can only act as the sender, and the other as the receiver. In bidirectional interactive scenarios such as video conferencing, the same terminal device can be both the sender and receiver of the video stream.
[0070] The video stream referred to in this embodiment is a continuous data sequence formed by packaging multiple video frames in chronological order, with each video frame having a strict temporal relationship.
[0071] Step S206: Optimize the network output video stream transmission strategy through a near-end strategy based on the state dimension parameters, effect reward function, and perception constraint function. The video stream transmission strategy is used to adjust the encoding bitrate and quality parameters of the video stream when transmitting the video stream.
[0072] Specifically, video streaming strategies are used to adjust the encoding bitrate and quality parameters of the video stream. The encoding bitrate refers to the amount of video data that the video stream can transmit per unit of time. This amount of video data is usually related to quality parameters, which include the frame rate and resolution of the current video stream.
[0073] The video stream is compressed based on quality parameters and encoding bitrate, and the compressed video stream is transmitted to the receiving end.
[0074] State dimension parameters, as input information to the near-end policy optimization network, play an important role in constraining the encoding bitrate and quality parameters of the video stream.
[0075] The impact of device performance data on the coding rate is reflected in the following aspects:
[0076] When CPU load / GPU load > 80%, the increase in encoding bitrate is limited, with a maximum increase of only 10%.
[0077] When CPU load / GPU load is less than 30%, increase the encoding bitrate increase, with a maximum allowable increase of 30%.
[0078] When memory usage exceeds 85%, the encoding bitrate will be forcibly reduced, with a maximum reduction of 10%.
[0079] The impact of device performance data on frame rate is reflected in:
[0080] When CPU / GPU load exceeds 70%, the frame rate selection of the video stream is limited, with a maximum allowed 30fps.
[0081] When CPU load / GPU load > 85%, force a reduction in the video stream frame rate, for example, to 24fps or 15fps.
[0082] When memory usage is greater than 80%, prioritize lower frame rates.
[0083] When the CPU load or GPU load is greater than or less than a certain value, the corresponding increase in the encoding bitrate and the selection of the video stream frame rate will be adjusted.
[0084] The impact of equipment performance data on quantitative parameters is reflected in:
[0085] The higher the CPU or GPU load, the higher the QP value; the QP value is positively correlated with the device load represented by the device performance data.
[0086] The quantization parameter (QP) is a core control parameter in video coding, directly determining the compression strength and output image quality. A smaller QP value retains more detail and produces better image quality, but at a higher bitrate; a larger QP value results in greater loss of detail and a blurrier image, but at a lower bitrate.
[0087] The system acquires the current device performance data of the terminal device, generates a video stream transmission strategy based on network status prediction data, and adjusts the quality parameters and encoding bitrate of the video stream through the video stream transmission strategy to ensure the quality and stability of the video stream transmission process.
[0088] In this embodiment, a model based on the Transformer architecture predicts the network state and outputs network state prediction data. This data is then used to obtain device performance data from the terminal device. The device performance data and network state prediction data are used as state dimension parameters in the state space of the near-end policy optimization network. The agent of the near-end policy optimization network perceives the current environment based on these state dimension parameters and optimizes and rewards the perception results through an effect reward function. This effect reward function is designed based on the quality and bitrate metrics of the video stream, enabling the near-end policy network to optimize the adjustment effect of the bitrate and quality metrics of the video stream. By predicting the network state, it can quickly respond to changes in the network state, thereby rapidly responding to the video stream transmission strategy and ensuring the quality and stability of the video stream transmission.
[0089] In one exemplary embodiment, such as Figure 3 As shown, step S206 includes steps S302 to S306. Wherein:
[0090] Step S302: Based on network state prediction data and device performance data, configure multiple state dimension parameters for the state space of the near-end learning policy network.
[0091] Step S304: Generate the effect reward function of the near-end policy optimization network based on the quality index and bit rate index in the state parameters.
[0092] Step S306: Based on multiple state dimension parameters and the effect reward function, adjust the convergence action of the near-end policy optimization network in the action space to obtain the video stream transmission strategy. The video stream transmission strategy includes multiple action dimension parameters, which are used to describe the adjustment actions of the near-end policy optimization network on the encoding bitrate and quality parameters.
[0093] Specifically, the state space of the near-end policy optimization network has multiple state dimension parameters. Table 1 explains the state dimension parameters.
[0094] Quality metrics refer to parameters that measure the quality of a video stream. These metrics include frame rate, resolution, and quantization parameters. Bitrate, on the other hand, measures the amount of data transmitted per unit of video data.
[0095] Under a given coding standard, video quality is positively correlated with coding bitrate, but there is a trade-off between the two due to bandwidth constraints.
[0096] Action space is the set of all possible operations or decisions that an agent can perform in a given state, defining the output information of the agent influencing the environment.
[0097] Introducing network state prediction data and device performance data into the near-end policy optimization network's action decision-making process can balance video transmission quality and terminal device load, avoiding terminal device overload and improving video transmission stability. Table 1 shows the value range of each state dimension parameter.
[0098] Table 1. Value range of state dimension parameters
[0099] State Dimension State dimension parameters Range of values S1 Predicted bandwidth 0-100Mbps S2 Prediction delay 0-500ms S3 Current bitrate 0.5-20Mbps S4 Current frame rate 15fps / 24fps / 30fps / 60fps S5 Current resolution 480p / 720p / 1080p / 4K S6 CPU load 0-100% S7 GPU load 0-100% S8 Memory usage 0-100%
[0100] The action space is configured with multiple action dimension parameters, including four dimensions. The action type of each dimension parameter can be either a continuous action or a discrete action. Table 2 illustrates each action dimension parameter.
[0101] Video streaming strategies include bitrate, frame rate, resolution, and quantization parameters. Frame rate, resolution, and quantization parameters represent the quality metrics of the video stream, while bitrate represents the bitrate of the video stream.
[0102] Table 2. Range of values for dynamic dimension parameters
[0103] Action dimension Action dimension parameters range of motion Action type A1 Bitrate -30%~+30% Continuous action A2 Frame rate 15fps / 24fps / 30fps / 60fps Discrete Actions A3 resolution 480p / 720p / 1080p / 4K Discrete Actions A4 Quantization parameters QP20~QP50 Continuous action
[0104] Continuous and discrete actions are used to represent how the near-end policy optimization network adjusts transmission parameters.
[0105] Discrete actions refer to selecting one as the output from a finite number of countable options, while continuous actions refer to outputting a real number within a range.
[0106] For example, in the frame rate action dimension parameter, you can only choose one from 15fps, 24fps, 30fps, and 60fps. If you choose 15fps, it means that the frame rate of the video stream is 15fps.
[0107] In the bitrate action dimension parameter, a bitrate value can be output within the range of -30% to 30%. This bitrate value represents the adjustment value for the encoding bitrate of the video stream.
[0108] For example, an output of -20% indicates a 20% reduction in the encoding bitrate, while an output of 10% indicates a 10% increase in the encoding bitrate. The adjusted bitrate is based on the current encoding bitrate.
[0109] For example, the proximal policy optimization network uses a two-layer fully connected neural network. The input layer includes 8-dimensional state parameters. The hidden layer consists of 256 neurons that use the ReLU activation function. This function extracts higher-level output features by non-linearly activating the 8-dimensional state parameters of the input layer through weighted summation. The output layer is designed with two action modes: discrete actions and continuous actions.
[0110] In this embodiment, device performance data and network status prediction data are introduced into the near-end policy optimization network to achieve intelligent matching between video encoding and terminal device capabilities. During the policy adjustment process of the near-end policy optimization network, an effect reward function is designed by combining quality indicators and bitrate indicators. Through weight allocation, a balance of multi-objective optimization is achieved, which can adapt to the application requirements of video streaming transmission scenarios.
[0111] As an optional implementation, step S306 includes steps S312 to S316. Wherein:
[0112] Step S312: Generate a perception constraint function based on the action dimension parameters. The perception constraint function is used to adjust the final output action decision.
[0113] Step S314: Generate an effect reward function based on the QoE dimension, resource consumption dimension, and delay penalty dimension.
[0114] Step S316: Adjust the policy expression of the near-end policy optimization network through the perception constraint function and the effect reward function to obtain the video stream transmission policy.
[0115] Specifically, the perceptual constraint function represents the adjustment of the action dimension parameters of the final output of the action space for the proximal policy optimization network when optimizing the policy based on the state dimension parameters.
[0116] The perception constraint function can be expressed by the following formula:
[0117]
[0118] Where Device_Load represents the weighted sum of device performance data, which can be calculated using the following formula:
[0119] Device_Load = CPU load * 0.4 + GPU load * 0.4 + Memory usage * 0.2
[0120] This indicates the adjusted bitrate. This indicates the original bitrate. This represents the equipment influence coefficient, with a value range of 0.5-1.0.
[0121] Among these, the video streaming transmission strategy also needs to balance the load capacity of the terminal device with the complexity of the video stream.
[0122] The load capacity level of terminal devices is distinguished by the comprehensive value obtained after weighted summation of device performance data. Different load capacity levels can handle different levels of complexity. Table 3 lists the comprehensive value range corresponding to each load capacity level of terminal devices.
[0123] Table 3 Capability Levels of Terminal Equipment
[0124] Load capacity level Comprehensive value range Level 1 <30% Level 2 30%-60% Level 3 60%-80% Level 4 ≥80%
[0125] Among them, the adjustment range of each motion dimension parameter in the motion space is related to the capability level of the terminal device, that is, the adjustment of each motion dimension parameter in the motion space is constrained based on the capability level of the terminal device.
[0126] Table 4 illustrates the constraint mechanism of each action dimension parameter in the action space.
[0127] Table 4. Constraint Mechanisms for Parameters of Each Action Dimension in Action Space
[0128] Load capacity level Maximum bitrate Maximum frame rate Maximum resolution QP minimum value Level 1 20Mbps 60fps 4K QP20 Level 2 10 Mbps 30 fps 1080p QP24 Level 3 5 Mbps 24 fps 720p QP28 Level 4 2 Mbps 15 fps 480p QP32
[0129] When encoding video streams, the capability level of the terminal device affects the values of various indicators in the final video stream transmission strategy.
[0130] The performance reward function optimizes the near-end policy optimization network from three dimensions: Quality of Experience (QoE), resource consumption, and latency penalty.
[0131] The effect reward function is expressed by the formula:
[0132]
[0133] in, As a reward value, This is the weight for the QoE dimension, with a value of 0.5; This is the weight for the resource consumption dimension, with a value of 0.3. The weight for the delayed penalty dimension is 0.2.
[0134] In the performance reward function, the user experience quality score has the highest weight. In real-time video communication scenarios, user experience is the most important consideration.
[0135] The formula for calculating the user experience quality score is as follows:
[0136]
[0137] Among them, the user experience quality score It is calculated using the following three parameter values.
[0138] The normalized score of PSNR (Peak Signal-to-Noise Ratio). =Current frame rate / Target frame rate =1 - Percentage of time spent lagging.
[0139] PNSR calculates the pixel-level differences between the transmitted video and the original video through mathematical calculations, quantization compression, and other methods to evaluate image quality.
[0140] The stuttering time percentage is the proportion of the viewing time during which the player cannot output new frames normally due to buffer underload. It is used to measure the smoothness of video streaming and playback.
[0141] The formula for calculating resource consumption is as follows:
[0142] Resource_Cost=0.5*Normalized_CPU_Load+0.3*Normalized_GPU_Load+0.2*Normalized_Memory_Usage
[0143] Resource consumption (Resource_Cost) is calculated based on CPU load, GPU load, and memory usage.
[0144] Normalized_CPU_Load = Current CPU load / Maximum CPU load.
[0145] Normalized_GPU_Load = Current GPU load / Maximum GPU load
[0146] Normalized_Memory_Usage = Current memory usage / Maximum memory capacity.
[0147] The formula for calculating the delay penalty is as follows:
[0148] Delay_Penalty = max(0, (Current_Delay - Target_Delay) / Target_Delay)
[0149] Here, Current_Delay represents the current delay between terminals. Target_Delay represents the target delay, which is typically 150ms in real-time communication scenarios.
[0150] In this embodiment, device performance data and network status prediction data are introduced into the near-end policy optimization network to achieve intelligent matching between video encoding and terminal device capabilities. An effect reward function is designed based on three dimensions: QoE, resource consumption, and latency penalty. A perceptual constraint function is generated based on device performance data. Based on the effect reward function and the perceptual constraint function, a video stream transmission strategy is output.
[0151] In one exemplary implementation, such as Figure 4 As shown, after step S306, the video stream data processing method further includes steps S402 to S404, wherein:
[0152] Step S402: Obtain real-time network status data and adjust the congestion window size according to the relationship between the network status prediction data and the window adjustment increment. The window adjustment increment is the product of the real-time network status data and the weight value.
[0153] Step S404: Transmit the video stream processed by the video stream transmission strategy using the adjusted congestion window size.
[0154] Specifically, real-time network status data refers to network status data during video stream transmission. The Congestion Window (CMND) is one of the core mechanisms of TCP (Transmission Control Protocol) for network congestion control. The congestion parameter is a state variable maintained by the sender, which determines the maximum amount of data the sender can inject into the network before receiving an acknowledgment (ACK) from the receiver.
[0155] The formula for calculating the congestion window is as follows:
[0156]
[0157] in, This indicates the adjusted new congestion window size. This indicates the current congestion window size. The adjustment factor has a range of 0.1-0.5. Indicates the predicted bandwidth. Given the current bandwidth. The congestion window adjustment strategy, determined based on the relationship between network state prediction data and the window adjustment increment, is as follows:
[0158] Increase the size of the congestion window when the predicted bandwidth is greater than the current bandwidth * 1.2.
[0159] When the predicted bandwidth is less than 0.8 times the current bandwidth, reduce the size of the congestion window.
[0160] Limit the growth of the congestion window when the prediction delay is greater than 150ms.
[0161] The window adjustment increment is the product of the current bandwidth and the weight value, and the specific value of the weight value will be adjusted based on the specific value of the predicted bandwidth.
[0162] During video stream transmission, the size of the congestion window is adjusted based on the difference between the predicted network status data and the real-time network status data, and the video stream is transmitted based on the size of the congestion window.
[0163] In this embodiment, by predicting the network state, network state prediction data is obtained. Based on the network state prediction data, the size of the congestion window is adjusted, which can adapt to network changes in advance, reduce transmission jitter, and improve the stability of video stream transmission.
[0164] Furthermore, after step S404, the video stream data processing method further includes steps S412 to S414, including:
[0165] Step S412: Obtain real-time round-trip time data and historical round-trip time data. Calculate round-trip time trend data based on the real-time and historical round-trip time data. The round-trip time trend data is used to characterize changes in network status.
[0166] Step S414: Based on the value of the round-trip time trend data, adjust the congestion window size using the corresponding adjustment strategy.
[0167] Specifically, Round-Trip Time (RTT) is a latency metric in network communication. RTT is defined as the total time it takes for a data packet to travel from the sender to the receiver, and then back to the sender with an acknowledgment (ACK) packet. The unit of RTT is usually milliseconds (ms).
[0168] Real-time round-trip time data refers to the round-trip time data sent to the video stream at the current time of transmission. Historical round-trip time data refers to the round-trip time data of video stream transmissions in past time periods.
[0169] The formula for calculating round-trip time trend data is as follows:
[0170]
[0171] in, For round-trip time trend data, For real-time round-trip time data, This is historical data on round-trip times.
[0172] if >0.1,
[0173] if <-0.1,
[0174] In other words, if the current round-trip time increases compared to historical data, the congestion window size is reduced; if the current round-trip time decreases compared to historical data, the congestion window size is increased. This method dynamically adjusts the congestion window size based on the current network communication quality to adjust the amount of data sent.
[0175] In this embodiment, the round-trip time can characterize the congestion signal on the transmission path and has self-stability. By adjusting the congestion window through the round-trip time, more timely flow control can be achieved.
[0176] In one exemplary implementation, the video streaming strategy includes a layered encoding transmission strategy, which divides the video stream into multiple video layers and transmits at least one video layer, such as... Figure 5As shown, the video stream data processing method further includes steps S502 to S504, wherein:
[0177] Step S502: Determine the available bandwidth based on the predicted bandwidth data included in the network status prediction data and the transmission overhead of the video transmission process.
[0178] Step S504: Determine the number of video layers to be transmitted based on the relationship between the available bandwidth and the sum of the video layer bitrates of multiple video layers.
[0179] Specifically, Scalable Video Coding (SVC) refers to packaging a video stream into a "core layer" plus several "enhancement layers".
[0180] The "core layer" is the basic layer, ensuring the basic quality of the video stream and requiring priority transmission. At least one "enhancement layer" is included, used to improve video quality and selectively transmitted based on network conditions.
[0181] Available bandwidth is the maximum effective transmission rate that the current transmission link can stably carry. Transmission overhead refers to the network bandwidth and computing resources consumed by additional auxiliary data to ensure that the valid data of the video stream can arrive correctly and reliably from the sending end to the receiving end.
[0182] For example, a layered coding transmission strategy divides the video stream into a base layer, enhancement layer 1, and enhancement layer 2.
[0183] Specifically, the transmission layering and number of layers are dynamically determined by using the predicted bandwidth and device performance data from the network status prediction data.
[0184]
[0185] in, Indicates available bandwidth. Indicates the predicted bandwidth. This indicates the transmission overhead.
[0186] when Then the transmission consists of the base layer, enhancement layer 1, and enhancement layer 2.
[0187] when Then the transmission base layer and enhancement layer 1 are transmitted.
[0188] when Then only the base layer is transmitted.
[0189] in, Indicates the bitrate of the base layer. This indicates the bitrate of enhancement layer 1. This indicates the bitrate of enhancement layer 2.
[0190] As an optional implementation, forward error correction (FEC) is implemented by adding redundant data at the transmitting end during transmission, enabling the receiving end to correct a certain degree of data errors without retransmission.
[0191] As an optional implementation, a layered coding transmission strategy is used to apply differentiated FEC protection to each layer that needs to be transmitted.
[0192] Table 5. FEC Strategy for Layered Transmission
[0193] Video layer FEC redundancy Priority base layer 0.3-0.5 high Enhancement layer 1 0.1-0.2 middle Enhancement layer 2 0.05-0.1 Low
[0194] Based on the information in Table 5, different redundancy levels are set for the priority of each video layer.
[0195] The FEC redundancy is dynamically adjusted based on the predicted packet loss rate. The dynamic adjustment strategy can be expressed by the following formula:
[0196]
[0197] in, This represents redundancy, with a value ranging from 0 to 0.5. This is the redundancy coefficient, with a value ranging from 1.5 to 2.0. To predict packet loss rate.
[0198] As an optional implementation, the FEC redundancy adjustment strategy is optimized by predicting latency, as expressed by the following formula:
[0199] when > The FEC redundancy is then adjusted based on the following formula:
[0200]
[0201] in, To predict packet loss rate, The delay threshold is set to 150ms. This is the delay penalty coefficient, with a value ranging from 0.5 to 1.0.
[0202] In this embodiment, based on the layered coding transmission strategy and FEC redundancy, the video layer transmission strategy during video transmission is adjusted. When network congestion occurs, the basic layer of the video stream is transmitted first. At the same time, the FEC redundancy is adjusted based on network status prediction data, which minimizes bandwidth consumption while ensuring reliability and optimizes resource allocation.
[0203] Furthermore, the packet loss rate of the video stream during transmission is obtained, and the layered encoding transmission strategy is adjusted based on the packet loss rate.
[0204] Specifically, a transmission feedback mechanism is set up, and the layered encoding transmission strategy is adjusted based on the transmission effect returned by the transmission feedback mechanism. When the packet loss rate is greater than a preset threshold, the layered encoding transmission strategy is executed.
[0205] As an optional implementation, the parameters of the transmission feedback mechanism also include actual bandwidth, packet loss rate, latency rate, and queue length. When the packet loss rate is greater than 0.1%, a layered coding transmission strategy is implemented.
[0206] By forming a closed-loop optimization through a transmission feedback mechanism, the overall performance of video transmission can be improved.
[0207] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.
[0208] Based on the same inventive concept, this application also provides a video stream data processing apparatus for implementing the video stream data processing method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more video stream data processing apparatus embodiments provided below can be found in the limitations of the video stream data processing method described above, and will not be repeated here.
[0209] In one exemplary embodiment, such as Figure 6 As shown, a video stream data processing device is provided, including: a network state prediction module 610, a parameter configuration module 620, an action dimension setting module 630, and a transmission strategy output module 640, wherein:
[0210] The network state prediction module 610 is used to predict the future network state based on historical network state data, and obtain network state prediction data.
[0211] The parameter configuration module 620 is used to configure the acquired device performance data and network state prediction data of the terminal device as state dimension parameters, and generate an effect reward function based on the quality index and bitrate index of the video stream. The state dimension parameters are used to describe the state space of the near-end policy optimization network, and the effect reward function is used to constrain the optimization effect of the near-end policy optimization network. The terminal device includes the sending end and / or receiving end of the video stream.
[0212] The action dimension setting module 630 is used to generate a perceptual constraint function based on the action dimension parameters. The perceptual constraint function is used to adjust the action decision output by the proximal policy optimization network.
[0213] The transmission strategy output module 640 is used to optimize the network output video stream transmission strategy according to the state dimension parameters and the effect reward function through the near-end strategy. The video stream transmission strategy is used to adjust the encoding bitrate and quality parameters of the video stream when transmitting the video stream.
[0214] In one exemplary embodiment, the transmission policy output module 640 includes the following units:
[0215] The state dimension setting unit is used to configure multiple state dimension parameters for the state space of the near-end learning policy network based on network state prediction data and device performance data.
[0216] The reward function generation unit is used to generate a reward function for the near-end policy optimization network based on the quality and bitrate metrics in the state parameters.
[0217] The transmission strategy generation unit is used to adjust the convergence action of the near-end policy optimization network in the action space based on multiple state dimension parameters and the effect reward function to obtain the video stream transmission strategy. The video stream transmission strategy includes multiple action dimension parameters, which are used to describe the adjustment actions of the near-end policy optimization network on the coding bitrate and quality parameters.
[0218] The transmission policy generation unit includes the following sub-units:
[0219] The objective function generation subunit generates an effect reward function based on the QoE, resource consumption, and latency penalty dimensions. The policy adjustment subunit adjusts the policy expression of the near-end policy optimization network using the perceptual constraint function and the effect reward function to obtain the video stream transmission policy.
[0220] The video stream data processing device also includes the following modules:
[0221] The first module for window adjustment is used to acquire real-time network status data and adjust the congestion window size based on the relationship between the predicted network status data and the window adjustment increment. The window adjustment increment is the product of the real-time network status data and the weight value.
[0222] The video streaming module is used to transmit video streams processed by the video streaming strategy using an adjusted congestion window size.
[0223] The round-trip time calculation module is used to acquire real-time and historical round-trip time data, and calculate round-trip time trend data based on the real-time and historical round-trip time data. The round-trip time trend data is used to characterize the changes in network status.
[0224] The second module for window adjustment is used to adjust the congestion window size based on the value of round-trip time trend data and the corresponding adjustment strategy.
[0225] The available bandwidth determination module is used to determine the available bandwidth based on the predicted bandwidth data included in the network status prediction data and the transmission overhead of the video transmission process. The video stream transmission strategy includes a layered coding transmission strategy, which is used to divide the video stream into multiple video layers and transmit at least one video layer.
[0226] The transmission quantity determination module is used to determine the transmission quantity of a video layer based on the relationship between the available bandwidth and the sum of the bitrates of multiple video layers.
[0227] The transmission strategy adjustment module is used to obtain the packet loss rate of the video stream during transmission and adjust the layered encoding transmission strategy according to the packet loss rate.
[0228] Each module in the aforementioned video stream data processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0229] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 7As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and databases. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media to run. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a video stream data processing method.
[0230] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0231] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0232] Network state prediction data is obtained by predicting the future state of the network based on historical network state data.
[0233] The acquired device performance data and network state prediction data of the terminal devices are configured as state dimension parameters, and an effect reward function is generated based on the quality index and bitrate index of the video stream. The state dimension parameters are used to describe the state space of the near-end policy optimization network, and the effect reward function is used to constrain the optimization effect of the near-end policy optimization network. The terminal devices include the sender and / or receiver of the video stream.
[0234] Based on the state dimension parameters and the effect reward function, the network output video stream transmission strategy is optimized through a near-end strategy. The video stream transmission strategy is used to adjust the encoding bitrate and quality parameters of the video stream during transmission.
[0235] Based on network state prediction data and device performance data, multiple state dimension parameters are configured for the state space of the near-end learning policy network.
[0236] Generate a reward function for the near-end policy optimization network based on the quality and bitrate metrics in the state parameters.
[0237] Based on multiple state dimension parameters and the effect reward function, the convergence action of the near-end policy optimization network in the action space is adjusted to obtain the video stream transmission strategy. The video stream transmission strategy includes multiple action dimension parameters, which are used to describe the adjustment actions of the near-end policy optimization network on the encoding bitrate and quality parameters.
[0238] A perception constraint function is generated based on the action dimension parameters. This perception constraint function is used to adjust the final output action decision.
[0239] The effect reward function is generated based on the QoE dimension, resource consumption dimension, and delay penalty dimension.
[0240] By adjusting the policy expression of the near-end policy optimization network using the perception constraint function and the effect reward function, the video stream transmission policy is obtained.
[0241] The system acquires real-time network status data and adjusts the congestion window size based on the relationship between the predicted network status data and the window adjustment increment. The window adjustment increment is the product of the real-time network status data and the weight value.
[0242] The video stream, processed by the video streaming strategy, is transmitted using an adjusted congestion window size.
[0243] Real-time and historical round-trip time data are obtained. Based on the real-time and historical round-trip time data, round-trip time trend data is calculated. The round-trip time trend data is used to characterize the changes in network status.
[0244] Based on the numerical values of round-trip time trend data, the congestion window size is adjusted using corresponding adjustment strategies.
[0245] In one exemplary embodiment, the video stream transmission strategy includes a layered encoding transmission strategy, which is used to divide the video stream into multiple video layers and transmit at least one video layer. The video stream data processing method further includes:
[0246] Available bandwidth is determined based on the predicted bandwidth data included in the network status prediction data and the transmission overhead of the video transmission process.
[0247] The number of video layers to be transmitted is determined based on the relationship between the available bandwidth and the sum of the bitrates of the multiple video layers.
[0248] Furthermore, the packet loss rate of the video stream during transmission is obtained, and the layered encoding transmission strategy is adjusted based on the packet loss rate.
[0249] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0250] Network state prediction data is obtained by predicting the future state of the network based on historical network state data.
[0251] The acquired device performance data and network state prediction data of the terminal devices are configured as state dimension parameters, and an effect reward function is generated based on the quality index and bitrate index of the video stream. The state dimension parameters are used to describe the state space of the near-end policy optimization network, and the effect reward function is used to constrain the optimization effect of the near-end policy optimization network. The terminal devices include the sender and / or receiver of the video stream.
[0252] Based on the state dimension parameters and the effect reward function, the network output video stream transmission strategy is optimized through a near-end strategy. The video stream transmission strategy is used to adjust the encoding bitrate and quality parameters of the video stream during transmission.
[0253] Based on network state prediction data and device performance data, multiple state dimension parameters are configured for the state space of the near-end learning policy network.
[0254] Generate a reward function for the near-end policy optimization network based on the QoE dimension, resource consumption dimension, and latency penalty dimension.
[0255] Based on multiple state dimension parameters and the effect reward function, the convergence action of the near-end policy optimization network in the action space is adjusted to obtain the video stream transmission strategy. The video stream transmission strategy includes multiple action dimension parameters, which are used to describe the adjustment actions of the near-end policy optimization network on the encoding bitrate and quality parameters.
[0256] Multiple action dimension parameters are set based on the video stream's quality and bitrate metrics.
[0257] Generate an optimization objective function based on quality and bitrate metrics according to the action dimension parameters.
[0258] By optimizing the objective function and the effect reward function to adjust the policy expression of the near-end policy optimization network, a video stream transmission policy is obtained.
[0259] The system acquires real-time network status data and adjusts the congestion window size based on the relationship between the predicted network status data and the window adjustment increment. The window adjustment increment is the product of the real-time network status data and the weight value.
[0260] The video stream, processed by the video streaming strategy, is transmitted using an adjusted congestion window size.
[0261] Real-time and historical round-trip time data are obtained. Based on the real-time and historical round-trip time data, round-trip time trend data is calculated. The round-trip time trend data is used to characterize the changes in network status.
[0262] Based on the numerical values of round-trip time trend data, the congestion window size is adjusted using corresponding adjustment strategies.
[0263] Available bandwidth is determined based on the predicted bandwidth data included in the network status prediction data and the transmission overhead of the video transmission process.
[0264] The number of video layers to be transmitted is determined based on the relationship between the available bandwidth and the sum of the bitrates of the multiple video layers.
[0265] Furthermore, the packet loss rate of the video stream during transmission is obtained, and the layered encoding transmission strategy is adjusted based on the packet loss rate.
[0266] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps:
[0267] Network state prediction data is obtained by predicting the future state of the network based on historical network state data.
[0268] The acquired device performance data and network state prediction data of the terminal devices are configured as state dimension parameters, and an effect reward function is generated based on the quality index and bitrate index of the video stream. The state dimension parameters are used to describe the state space of the near-end policy optimization network, and the effect reward function is used to constrain the optimization effect of the near-end policy optimization network. The terminal devices include the sender and / or receiver of the video stream.
[0269] Based on the state dimension parameters and the effect reward function, the network output video stream transmission strategy is optimized through a near-end strategy. The video stream transmission strategy is used to adjust the encoding bitrate and quality parameters of the video stream during transmission.
[0270] Based on network state prediction data and device performance data, multiple state dimension parameters are configured for the state space of the near-end learning policy network.
[0271] Generate a reward function for the near-end policy optimization network based on the QoE dimension, resource consumption dimension, and latency penalty dimension.
[0272] Based on multiple state dimension parameters and the effect reward function, the convergence action of the near-end policy optimization network in the action space is adjusted to obtain the video stream transmission strategy. The video stream transmission strategy includes multiple action dimension parameters, which are used to describe the adjustment actions of the near-end policy optimization network on the encoding bitrate and quality parameters.
[0273] Multiple action dimension parameters are set based on the video stream's quality and bitrate metrics.
[0274] Generate an optimization objective function based on quality and bitrate metrics according to the action dimension parameters.
[0275] By optimizing the objective function and the effect reward function to adjust the policy expression of the near-end policy optimization network, a video stream transmission policy is obtained.
[0276] The system acquires real-time network status data and adjusts the congestion window size based on the relationship between the predicted network status data and the window adjustment increment. The window adjustment increment is the product of the real-time network status data and the weight value.
[0277] The video stream, processed by the video streaming strategy, is transmitted using an adjusted congestion window size.
[0278] Real-time and historical round-trip time data are obtained. Based on the real-time and historical round-trip time data, round-trip time trend data is calculated. The round-trip time trend data is used to characterize the changes in network status.
[0279] Based on the numerical values of round-trip time trend data, the congestion window size is adjusted using corresponding adjustment strategies.
[0280] In one exemplary embodiment, the video stream transmission strategy includes a layered encoding transmission strategy, which is used to divide the video stream into multiple video layers and transmit at least one video layer. The video stream data processing method further includes:
[0281] Available bandwidth is determined based on the predicted bandwidth data included in the network status prediction data and the transmission overhead of the video transmission process.
[0282] The number of video layers to be transmitted is determined based on the relationship between the available bandwidth and the sum of the bitrates of the multiple video layers.
[0283] Furthermore, the packet loss rate of the video stream during transmission is obtained, and the layered encoding transmission strategy is adjusted based on the packet loss rate.
[0284] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0285] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0286] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A video stream data processing method, characterized in that, The method includes: Network state prediction data is obtained by predicting the future network state based on historical network state data. The acquired device performance data of the terminal device and the network state prediction data are jointly configured as the state dimension parameters of the near-end policy optimization network, and an effect reward function is generated based on the quality index and bitrate index of the video stream. The state dimension parameters are used to describe the state space of the near-end policy optimization network, and the effect reward function is used to constrain the optimization effect of the action dimension parameters in the near-end policy optimization network. The terminal device includes the sending end and / or receiving end of the video stream. A perceptual constraint function is generated based on the action dimension parameters. The perceptual constraint function is used to adjust the action decision output by the proximal policy optimization network. Based on the state dimension parameters, the effect reward function, and the perception constraint function, the network output video stream transmission strategy is optimized through the near-end strategy. The video stream transmission strategy is used to adjust the encoding bitrate and quality parameters of the video stream when transmitting the video stream.
2. The method according to claim 1, characterized in that, The step of optimizing the network output video stream transmission strategy based on the state dimension parameters and the effect reward function using the near-end strategy includes: Based on the network state prediction data and the device performance data, multiple state dimension parameters are configured for the state space of the near-end learning policy network. The effect reward function of the near-end policy optimization network is generated based on the quality index and bitrate index in the state dimension parameters. The video stream transmission strategy is obtained by adjusting the convergence action of the near-end policy optimization network in the action space based on multiple state dimension parameters and the effect reward function. The video stream transmission strategy includes multiple action dimension parameters, which are used to describe the adjustment actions of the near-end policy optimization network on the coding bitrate and quality parameters.
3. The method according to claim 2, characterized in that, The step of optimizing the network output video stream transmission strategy based on the state dimension parameters, the effect reward function, and the perceptual constraint function through the near-end strategy includes: The effect reward function is generated based on the QoE dimension, resource consumption dimension, and latency penalty dimension. The video stream transmission strategy is obtained by adjusting the policy expression of the near-end policy optimization network through the perceptual constraint function and the effect reward function.
4. The method according to claim 1, characterized in that, After optimizing the network output video stream transmission strategy according to the state dimension parameters and the effect reward function through the near-end strategy, the method further includes: Acquire real-time network status data and adjust the congestion window size according to the relationship between the predicted network status data and the window adjustment increment. The window adjustment increment is the product of the real-time network status data and the weight value. The video stream processed by the video stream transmission strategy is transmitted using the adjusted congestion window size.
5. The method according to claim 4, characterized in that, After transmitting the video stream processed by the video stream transmission strategy using the adjusted congestion window size, the method further includes: Real-time and historical round-trip time data are acquired, and round-trip time trend data is calculated based on the real-time and historical round-trip time data. The round-trip time trend data is used to characterize the changes in network status. Based on the values of the round-trip time trend data, the congestion window size is adjusted using a corresponding adjustment strategy.
6. The method according to claim 1, characterized in that, The video stream transmission strategy includes a layered coding transmission strategy, which is used to divide the video stream into multiple video layers and transmit at least one enhancement layer. The method further includes: The available bandwidth is determined based on the predicted bandwidth data included in the network status prediction data and the transmission overhead of the video transmission process; The number of video layers to be transmitted is determined based on the relationship between the available bandwidth and the sum of the video layer bitrates of the multiple video layers.
7. The method according to claim 6, characterized in that, The method further includes: The packet loss rate of the video stream during transmission is obtained, and the layered encoding transmission strategy is adjusted according to the packet loss rate.
8. A video stream data processing device, characterized in that, The device includes: The network state prediction module is used to predict the future network state based on historical network state data, and obtain network state prediction data. The parameter configuration module is used to configure the acquired device performance data of the terminal device and the network state prediction data together as the state dimension parameters of the near-end policy optimization network, and generate an effect reward function based on the quality index and bitrate index of the video stream. The state dimension parameters are used to describe the state space of the near-end policy optimization network, and the effect reward function is used to constrain the optimization effect of the action dimension parameters in the near-end policy optimization network. The terminal device includes the sending end and / or receiving end of the video stream. The action dimension setting module is used to generate a perceptual constraint function based on the action dimension parameters. The perceptual constraint function is used to adjust the action decision output by the proximal policy optimization network. The transmission strategy output module is used to optimize the network output video stream transmission strategy according to the state dimension parameters, the effect reward function, and the perception constraint function through the near-end strategy. The video stream transmission strategy is used to adjust the encoding bitrate and quality parameters of the video stream when transmitting the video stream.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Distributed data encryption and verification system and method based on block chain
CN120710777A
Method and system for efficient streaming video dynamic rate adaptation
US20120005365A1