Reinforcement Learning Based on Rate Control

The reinforcement learning-based rate control method addresses the inefficiencies of existing methods in handling screen content by dynamically adjusting encoding parameters, resulting in improved visual quality and reduced computational overhead for real-time communication.

JP7687633B2Active Publication Date: 2025-06-03MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022581327
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2020-06-30
Publication Date
2025-06-03
Estimated Expiration
2040-06-30

AI Technical Summary

Technical Problem

Existing rate control methods for video encoding are ineffective for screen content in real-time communication due to its complex and sudden changes, leading to suboptimal quality of experience (QOE).

Method used

A reinforcement learning-based rate control method that determines encoding parameters for a video encoder by associating the encoding state with the encoding of a video unit, allowing for dynamic adjustment of encoding parameters to achieve better QOE.

Benefits of technology

The reinforcement learning-based rate control method achieves improved visual quality with reduced computational overhead and efficiently handles sudden changes in screen content, enhancing the overall quality of experience in real-time communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007687633000007
    Figure 0007687633000007
  • Figure 0007687633000008
    Figure 0007687633000008
  • Figure 0007687633000009
    Figure 0007687633000009
Patent Text Reader

Abstract

An implementation of the subject matter described herein provides a solution for rate control based on reinforcement learning. In this solution, a coding state of a video encoder is determined, and the coding state is associated with encoding of a first video unit by the video encoder. Coding parameters for rate control in the video encoder are determined based on the coding state of the video encoder using a reinforcement learning model. A second video unit different from the first video unit is coded based on the coding parameters. In this way, it is possible to achieve better quality of experience (QOE) for real-time communication with reduced computational overhead.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001]

[0001] In real-time communication (RTC), a common requirement is screen sharing with various users. For example, a participant may need to present his or her desktop screen to other participants in a multi-user video conference. In this situation, the technical goal is to provide a better quality of experience (QOE), which is often determined by various factors such as visual quality, drop rate, transmission delay, etc. Rate control plays an important role in achieving this goal by determining the encoding parameters of the video encoder to achieve the target bitrate.

[0002]

[0002] Existing rate control methods are mainly designed for videos with natural scenes. However, unlike natural videos, which often involve smooth content movement, screen content is usually combined with complex sudden changes or static scenes. Due to this unique movement characteristic, existing rate control methods designed for natural videos cannot function well for screen content.

Background Art

[0003]

[0003] According to the implementation of the subject matter described herein, a solution for rate control based on reinforcement learning is provided. In this solution, the encoding state of a video encoder is determined, and the encoding state is associated with the encoding of a first video unit by the video encoder. Encoding parameters related to the rate control of the video encoder are determined by a reinforcement learning model based on the encoding state of the video encoder. A second video unit different from the first video unit is encoded based on the encoding parameters. The reinforcement learning model is configured to receive the encoding states of one or more video units and determine encoding parameters to be used in another video unit. The encoding state has a limited state dimension and is capable of achieving better QOE for real-time communication with reduced computational overhead.

[0004]

[0004] This summary is provided to introduce, in a simplified form, a selection of concepts that are further described in the detailed description below. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

Brief Description of the Drawings

[0005]

[0005] Through a more detailed description of several implementations of the subject matter described herein in the accompanying drawings, the above and other objects, as well as the features and advantages of the subject matter described herein, will become more apparent.

Figure 1

[0006] FIG. 1 shows a block diagram of a computing device capable of implementing various implementations of the subject matter described herein.

Figure 2

[0007] FIG. 2 shows a block diagram of a reinforcement learning module according to an implementation of the subject matter described herein.

Figure 3

[0008] Figure 3 shows an example of an agent for use in a reinforcement learning module by implementation of the subject matter described in this case.

Figure 4

[0009] Figure 4 shows a flowchart of a reinforcement learning-based rate control method by implementation of the subject matter described in this case.

[0010] Throughout the drawings, the same or similar reference numerals represent the same or similar elements.

Embodiments for Carrying Out the Invention

[0006]

[0011] The subject matter described in this case will be described below with reference to several exemplary implementations. It should be understood that these implementations are discussed only for the purpose of enabling those skilled in the art to better understand and implement the subject matter described in this case, and do not imply any limitation on the scope of the subject matter.

[0007]

[0012] As used in this case, the terms "comprising" and its variations should be read as open terms meaning "including, but not limited to". The term "based on" should be read as "at least partially based on". The terms "an implementation" and "implementations" should be read as "at least one implementation". The term "another implementation" should be read as "at least one other implementation". The terms "first", "second", etc. may refer to different or the same objects. Other definitions, whether explicit or implicit, may be included below.

[0008]

[0013] FIG. 1 shows a block diagram of an arithmetic device 100 in which various implementations of the subject matter described herein can be implemented. It will be understood that the arithmetic device 100 shown in FIG. 1 is for illustrative purposes only and does not imply any limitation in any way to the functions and scope of the implementation of the subject matter described herein. As shown in FIG. 1, the arithmetic device 100 includes a general-purpose arithmetic device 100. The components of the arithmetic device 100 may include, but are not limited to, one or more processors or processing units 110, a memory 120, a storage device 130, one or more communication units 140, one or more input devices 150, and one or more output devices 160.

[0009]

[0014] In some implementations, the arithmetic device 100 may be implemented as any user terminal or server terminal having arithmetic capabilities. The server terminal may be a server, a large-scale arithmetic device, etc., and may be provided by a service provider. The user terminal may be, for example, a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet, an Internet node, a communication device, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / video camera, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination thereof, including accessories and peripheral devices of these devices, etc., and may be any type of mobile terminal, fixed terminal, or portable terminal, or a combination thereof. It is envisioned that the arithmetic device 100 can support any type of interface for the user (e.g., a "wearable" circuit, etc.).

[0010]

[0015] The processing unit 110 may be a physical or virtual processor and can execute various processes based on programs stored in the memory 120. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the arithmetic device 100. The processing unit 110 is also referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.

[0011]

[0016] The arithmetic device 100 typically includes various computer storage media. Such media can be any media accessible by the arithmetic device 100, including but not limited to volatile and non-volatile media, or removable and non-removable media. The memory 120 can be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory), or any combination thereof. The storage device 130 can be any removable or non-removable media and may include machine-readable media such as memory, flash memory drives, magnetic disks, or other media that can be used to store information and / or data and can be accessed within the arithmetic device 100.

[0012]

[0017] The arithmetic device 100 may further include additional removable / non-removable, volatile / non-volatile memory media. Although not shown in FIG. 1, it is possible to provide a magnetic disk drive for reading from and / or writing to a removable non-volatile magnetic disk, and an optical disk drive for reading from and / or writing to a removable non-volatile optical disk. In such a case, each drive may be connected to a bus (not shown) via one or more data media interfaces.

[0013]

[0018] The communication unit 140 communicates with another computing device via a communication medium. Further, the functions of the components within the computing device 100 can be realized by a single computing cluster or multiple computing machines capable of communicating via a communication connection. Accordingly, the computing device 100 can operate in a networked environment using a logical connection to one or more other servers, networked personal computers (PCs), or more general network nodes.

[0014]

[0019] The input device 150 may be one or more of various input devices such as a mouse, keyboard, tracking ball, voice input device, etc. The output device 160 may be one or more of various output devices such as a display, loudspeaker, printer, etc. The communication unit 140 enables the computing device apparatus 100 to communicate further with one or more external devices (not shown) such as a storage device and a display device, with one or more devices that enable a user to interact with the computing device 100, or with any device that enables the computing device 100 to communicate with one or more other computing devices as needed (e.g., a network card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).

[0015]

[0020] In some implementations, as an alternative to being integrated into a single device, some or all of the components of the computing device 100 may be located in a cloud computing architecture. In a cloud computing architecture, the components may be remotely provided and may cooperate to implement the functions described in the subject matter described herein. In some implementations, cloud computing provides computing, software, data access, and storage services, but the end user is not required to be aware of the physical location or configuration of the system or hardware that provides these services. In various implementations, cloud computing uses appropriate protocols to provide services over a wide area network (e.g., the Internet, etc.). For example, a cloud computing provider provides an application on a wide area network that can be accessed via a web browser or any other computing component. The software or components of the cloud computing architecture, and the corresponding data, may be stored on a remote server. The computing resources in a cloud computing environment may be consolidated or distributed in locations within a remote data center. The cloud computing infrastructure operates as a single access point for the user but may provide services via a shared data center. Thus, the cloud computing architecture may be used to provide the components and functions described herein from a remote service provider. Alternatively, these may be provided from a conventional server or may be installed directly or otherwise on the client device.

[0016]

[0021] The computing device 100 may be used to perform reinforcement learning-based rate control for the implementation of the subject matter described herein. The memory 120 may include one or more reinforcement learning modules 122 having one or more program instructions. These modules are accessible and executable by the processing unit 110 to perform the functions of the various implementations described herein. For example, the input device 150 may provide video or a series of frames of the environment of the computing device 100 to the reinforcement learning module 122 to enable a video conferencing application, while the processing unit 110 and / or the memory 120 may provide at least a portion of the screen content to the reinforcement learning module 122 to enable the execution of a screen content sharing application. Multimedia content may be encoded by the reinforcement learning module 122 to achieve rate control with good QOE.

[0017]

[0022] Referring now to FIG. 2, a block diagram of a reinforcement learning module 200 according to the implementation described herein is shown. The reinforcement learning module 200 may be implemented, for example, as the reinforcement learning module 122 within the computing device 100. The reinforcement learning module 200 includes an encoder 204, which is configured to encode multimedia content from other components of the computing device 100, such as the processing unit 110, the memory 120, the storage 130, the input device 150, and / or the like. For example, the input device 150 may provide one or more frames of video to the reinforcement learning module 200, while the processing unit 110 and / or the memory 120 may provide at least a portion of the screen content to the reinforcement learning module 200. For example, the encoder 204 may be a video encoder, particularly a video encoder optimized for screen content from the computing device 100.

[0018]

[0023] Encoding parameters related to rate control, such as quantization parameter (QP) or lambda, control the granularity of compression of video units, such as frames, blocks, or macroblocks within a frame. A large value means that there will be higher quantization, more compression, and lower quality. A lower value means the opposite. Therefore, it is possible to achieve good QOE by performing rate control to adjust the encoding parameters of the encoder, such as the quantization parameter or lambda. Here, references are made to the quantization parameter and lambda, but it should be noted that the quantization parameter and lambda are provided for illustrative purposes, and any other appropriate encoding parameters associated with rate control may be adjusted or controlled.

[0019]

[0024] As shown in FIG. 2, the reinforcement learning module 200 can include an agent 202 configured to make a decision to control the encoding parameters of the encoder 204. In some implementations, the agent 202 may employ a reinforcement learning model implemented by a neural network, such as a recurrent neural network.

[0020]

[0025] Next, the encoded bitstream is output to the transmission buffer. The encoder 204 may also include such a transmission buffer (not shown) to implement the bitstream transmission process. After encoding, the bitstream of the latest encoded video unit is stored or added to the transmission buffer. During transmission, the bitstream stored in the transmission buffer is transmitted to one or more receivers via one or more channels at a certain bandwidth, and the transmitted bitstream is removed from the buffer along with the transmission at that bandwidth. The state of the transmission buffer is in a constantly changing process due to the incoming and outgoing bitstreams entering and leaving the transmission buffer.

[0021]

[0026] At each time step t, agent 202 observes the encoded state s of encoder 204 t . The encoded state s at time step t t may be determined based on the encoding of at least the video unit at time step t - 1. Based on this input information, agent 202 performs an estimation and outputs an action a t . The action a t indicates how finely encoder 204 should compress the video unit at time step t. The action a t may be the encoding parameter of encoder 204 for rate control, such as the quantization parameter (QP), or may be mapped to the encoding parameter of encoder 204. After obtaining the encoding parameter, encoder 204 can start encoding the video unit (e.g., a screen content frame). The encoding of the video unit at time step t will be used to update the encoded state s of agent 202 at time step t + 1 t+1 . It should be understood that the reinforcement learning module 200 may be applicable to any other suitable multimedia application other than real - time screen content sharing

[0022]

[0027] Rather than relying on traditional manual rules, a reinforcement learning-based solution implemented for the subject matter described herein controls encoding parameters based on the encoder's encoding state through the agent's actions, achieving better visual quality with negligible drop rate changes. The encoder's encoding state has a limited state space, thus enabling the determination of encoding parameters to be made with reduced computational overhead and improved efficiency. In particular, when a sudden change in the scene occurs in the screen content, a well-trained reinforcement learning model can update the encoding parameters very quickly to achieve better QOE, which is particularly beneficial for screen content sharing in real-time communication. The reinforcement learning-based architecture is not limited to any codec and can work with various codecs, such as H.264, HEVC, and AV1.

[0023]

[0028] In some implementations, to assist the agent 202 of the reinforcement learning module 200 in making accurate and reliable decisions, the encoding state at time step t as an input to the agent 202 can include a number of elements for representing the encoding state from various perspectives. For example, the video unit may be a frame, and the encoding state s t may include a state representing the outcome of encoding at least the frame at time step t - 1; the state of the transmission buffer at time step t; and a state associated with the network situation at time step t for transmitting the encoded frame.

[0024]

[0029] For example, the results related to encoding at least the frame at time step t-1 may further include the results related to encoding frames before time step t-1, such as the frame at time step t-2. In one example, the results may include the encoding parameters (e.g., QP or lambda) of the encoded frame at time step t-1, and the size of the encoded frame at time step t-1. If a frame is missing, the encoding parameters of the frame encoded at time step t-1 may be set to a predetermined value such as zero. In one example, the frame size at time step t-1 may be expressed by the frame size ratio of the frame, which is defined by the ratio of the frame size to the average target frame size. In other words, the frame size at time step t-1 may be normalized by the average target frame size. For example, the frame size may be expressed by the bitstream size of the frame, the average target frame size may represent the average of the target number of bits in the frame, and may also be calculated by dividing the target bit rate by the frame rate. The target bit rate represents the target number of bits to be transmitted, and the frame rate represents the frequency or rate at which the frames are transmitted. Both the target bit rate and the frame rate can be determined from the video encoder.

[0025]

[0030] In one example, the state of the transmission buffer may include the usage of the buffer, such as the ratio of the occupied space to the maximum space of the buffer; the remaining space of the buffer measured within the frame; or a combination thereof. The remaining space of the buffer measured within the frame may be calculated by dividing the remaining space of the buffer by the average target frame size. This value describes the usage of the buffer from another aspect where the influence of the frame rate is considered.

[0026]

[0031] In one example, the state related to the network situation includes the target bits per pixel (BPP). This parameter is defined by the number of bits used by each pixel, and the target bit rate may also be calculated by dividing the target bit rate by the number of pixels in a frame per unit time. The number of pixels in a frame and the target bit rate may be determined, for example, from a video encoder.

[0027]

[0032] In some implementations, the encoding state described above is with respect to a frame, and the reinforcement learning module 200 makes decisions on a frame-by-frame basis. In other implementations, the reinforcement learning module 200 may be applied or adapted to any other suitable video unit for compression or encoding. For example, the reinforcement learning module may make decisions at the block level, such as at the macroblock (H.264), coding tree unit (HEVC), superblock (AV1), etc. Therefore, the encoding state s used as input to the agent 202 t may include a state representing the result of encoding at least one block at time step t-1, the state of the transmission buffer at time step t, and a state associated with the network situation at time step t for transmitting the encoded block.

[0028]

[0033] For example, the result regarding encoding at least one block may include the result regarding encoding one or more neighboring blocks. The neighboring blocks may include blocks that are spatially to the left, right, above, and / or below the block being processed. Encoding of spatially neighboring blocks may be performed at time step t-1 or other preceding time steps. The result of encoding spatially neighboring blocks may be stored in storage, and the result of encoding spatially neighboring blocks may be retrieved from storage. Additionally or alternatively, the neighboring blocks may include one or more corresponding blocks in a preceding frame, which are also referred to as temporally neighboring blocks. The result of encoding temporally neighboring blocks can be stored in storage and retrieved therefrom.

[0029]

[0034] In one example, the result may include encoding parameters of at least one encoded block, such as QP or lambda, and the size of at least one encoded block. For example, the size of the encoded block may be expressed by a block size ratio defined by the ratio of the size of the encoded block to the average target block size. In other words, the block size can be normalized by the average target block size. For example, the block size may be expressed by the bitstream size for encoding the block, the average target block size may represent the average of the target number of bits in the block, and the target bit rate may also be calculated by dividing the number of blocks transmitted per unit time.

[0030]

[0035] In one example, the state of the transmission buffer may include how the buffer is used, such as the ratio of the occupied space to the maximum space of the buffer, the remaining space of the buffer measured within the block, or a combination thereof. The remaining space of the buffer measured within the block may be calculated by dividing the remaining space of the buffer by the average target block size.

[0031]

[0036] In one example, the state related to the network state includes the target bits per pixel (BPP). This parameter is defined by the number of bits used by the pixel and can be calculated in the same way as the implementation related to the frame.

[0032]

[0037] The encoding state is described in terms of encoding parameters such as quantization parameters or lambda. Note that the encoding state may be applied to any other appropriate encoding parameter associated with rate control used by the encoder.

[0033]

[0038] Referring again to FIG. 2, the action a output by the agent 202 t can control the encoding quality of the encoder 204. For example, the action a determined by the agent 202 t may be normalized and within the range of 0 to 1. In some implementations, the action can be mapped to a QP that the encoder can understand. For example, the mapping may be implemented by the following formula:

Equation

[0039] Here, QP max and QP min represent the maximum and minimum QPs respectively, and QP currepresents the QP that will be used for encoding by the encoder 204. Although this mapping function is illustrated as a linear function, it should be understood that any other suitable function can be used instead. Smaller QP values cause the encoder to perform compression in a more delicate manner and obtain a higher reconstruction quality. However, the sacrifice is generating a larger encoded bitstream. An overly large bitstream can easily overshoot the buffer, and accordingly, frames may be dropped (for example, in the case of frame-level rate control). On the other hand, larger QP values employ coarser encoding, but a smaller encoded bitstream will be generated.

[0034]

[0040] In some further implementations, the encoding parameter may be implemented as lambda. The action a output by the agent 202 t can be mapped to a lambda that the encoder can understand. For example, the mapping may be implemented by the following formula:

[0035]

Number

[0041] Here, lambda max and lambda min represent the maximum and minimum lambdas respectively, and lambda cur represents the lambda that will be used by the encoder 204. This mapping function behaves linearly in the logarithmic domain of lambda. In addition to or instead of the above-described mapping function, any other suitable function may be used for mapping instead. Lower lambda values control encoding in a more delicate manner and obtain a higher reconstruction quality. However, there is a possibility of generating a larger encoded bitstream, and the buffer may be easily overshot, while higher lambda values employ coarser encoding, but a smaller encoded bitstream will be generated.

[0036]

[0042] Continuing to refer to FIG. 2, in the training of the reinforcement learning module 200, it is necessary to evaluate how good the actions taken by the agent 202 are. For this purpose, after the encoder 204 finishes encoding each video unit with the action a t , the reward r t is provided. When the agent 202 has obtained a certain number of training samples, the agent 202 can update its policy based on the reward r t . The agent 202 can be trained to converge in a direction that maximizes the accumulated reward. To obtain a better QOE, one or more factors reflecting the QOE can be incorporated into the reward. For example, the reward r t imposes a penalty on buffer overshoot and is set to increase when the encoding parameters result in higher visual quality. For example, the visual quality increases as the quantization parameter or lambda decreases.

[0037]

[0043] In one example, the reward r t may be calculated as follows:

[0038]

Equation

[0044] Here, a is a constant factor, b is a negative number, r base is the basic reward, Bandwidth cur represents the bandwidth of the channel that transmits the bitstream at time step t, Bandwidth max represents the maximum bandwidth, and r final represents the final reward.

[0039]

[0045] The basic reward r baseis calculated by Equation (3). For example, higher visual quality may result in higher QOE, especially in the scenario of screen content sharing. Therefore, in order to achieve higher visual quality, it is desirable to use a smaller QP or lambda, and as shown in Equation (3), as the current quantization parameter QP cur decreases, the reward increases. However, very small QP values may result in a large bitstream size, which can easily cause buffer overshoot, and as a result, frames may be dropped with respect to frame-level rate control. Therefore, in the case of buffer overshoot, the reward is set as a negative number (i.e., b). Setting a negative number as a penalty is used to train agent 202 to avoid buffer overshoot.

[0040]

[0046] r base After calculating, for example, as shown in Equation (4), the final reward r final can be obtained by scaling r base . The scaling factor is related to the ratio of the bandwidth at time step t to the maximum bandwidth. When the bandwidth at time point t is large, the reward r t is scaled to a larger value, and the penalty is also larger when buffer overshoot occurs. It is more aggressive to pursue better visual quality under wide bandwidth conditions, while buffer overshoot is more serious. It should be noted that without departing from the spirit of the implementation described herein, any other appropriate function may be used instead to calculate the reward.

[0041]

[0047] In some implementations, the Proximal Policy Optimization (PPO) algorithm is used with the reward r tIt may be adopted to train the agent 202 based on this. PPO is implemented based on the actor-critic architecture, which includes an actor network for the actor and a critic network for the critic. The actor operates as the agent 202. The input to the actor network is the encoded state, and the output of the actor network is the action. The actor network is configured to estimate the policy π θ (a t |s t ), where θ represents the policy parameter (e.g., the weights in the actor network), and a t , s t represent the action and the encoded state at time step t, respectively. The critic's critic network is set to evaluate how good the encoded state st is and only operates during the training process.

[0042]

[0048] In the PPO algorithm, as follows, the policy loss L policy may be used to update the actor, and the value loss L value may be used to update the critic:

[0043]

Number

[0049] Here, the value loss is calculated as the square of

[0044]

Number

[0045]

Number

[0046]

[0050] The encoded state used in the reinforcement learning module 200 enables the use of a lightweight network architecture for the agent 202 and a lightweight network architecture for training the agent 202. For example, the neural network implementing the agent 202 may include one or more input fully connected layers configured to extract features from the encoded state s t and the extracted features may be provided to one or more recurrent neural networks to extract temporal features or correlations from the features. Then, the features may be, for example, the action a tTo make a decision to generate, it may be provided to one or more output fully connected layers. The recurrent neural network may be, for example, a gated recurrent unit (GRU) or a long short-term memory (LSTM). The neural network has a lightweight but efficient architecture and can meet the requirements of real-time applications, especially the requirements of screen content coding (SCC).

[0047]

[0051] Figure 3 shows an example of a neural network 300 for training an agent 202 in accordance with the implementation of the subject matter described herein. The neural network 300 includes an actor network 302 and a critic network 304. The actor network 302 and the critic network 304 may share a common network module to reduce the parameters to be optimized. In this example, the input passes through two fully connected (FC) layers and is converted into a feature vector. Although a leaky Rectified Linear Unit (RELU) is shown in Figure 3, it should be understood that any other suitable activation function may be used in the network.

[0048]

[0052] Considering that rate control is a time - series problem, two gated recurrent units (GRUs) are introduced to further extract features in combination with historical information. It should be understood that any other suitable recurrent neural network can be used as well. After the GRUs, the actor and critic networks start to have their respective network modules. Both the actor and the critic reduce the dimension of the feature vector using fully - connected (FC) layers. Finally, both networks generate their respective outputs using one FC layer, and the actor network uses a sigmoid layer to normalize the range of actions to [0, 1]. It should be understood that any other suitable activation function may be used instead of the sigmoid function.

[0049]

[0053] The neural network 300 has a lightweight but efficient architecture to meet the requirements of real - time applications. Regarding screen content coding (SCC), the reinforcement - learning - based solution can achieve better visual quality with negligible drop - rate variations compared to traditional rule - based rate - control methods. In particular, this method can bring about a quality improvement extremely quickly after a rapid scene change occurs in screen content. The reinforcement - learning - network - based architecture is not limited to any codec and can cooperate with various different codecs, such as H.264, HEVC, and AV1.

[0050]

[0054] FIG. 4 shows a flowchart of a reinforcement - learning - based rate - control method 400 according to the implementation of the subject matter described in this case. The method 400 may be implemented by the computing device 100, for example, by the reinforcement - learning module 122 within the computing device 100. Also, the method 400 may be implemented by any other device, a cluster of devices, or a distributed parallel system similar to the computing device 100. For the purpose of explanation, the method 400 is described with reference to FIG. 1.

[0051]

[0055] In block 402, the computing device 100 determines the encoding state of the video encoder. The encoding state may be associated with encoding a first video unit by the video encoder. The video encoder may be configured to encode screen content for real-time communication. For example, the video encoder may be encoder 204 as shown in FIG. 2 and within the reinforcement learning module 200. The encoding state associated with encoding the first video unit includes: a state representing at least the result of encoding the first video unit; the state of a buffer configured to buffer the video unit encoded by the video encoder before transmission; and a state associated with the network situation for transmitting the encoded video unit. The video unit may include a frame, a block, or a macroblock in a frame. In some implementations, the result of encoding the first video unit includes the encoding parameters of the encoded first video unit and the size of the encoded first video unit, the state of the buffer includes the usage of the buffer, and the state associated with the network situation includes the target bits per pixel. In some implementations, the usage of the buffer includes at least one of: the ratio of the occupied space to the maximum space of the buffer; and the remaining space of the buffer measured in the video unit.

[0052]

[0056] In block 404, the computing device 100 determines, by means of a reinforcement learning model, encoding parameters related to the rate control of the video encoder based on the encoding state of the video encoder. The encoding parameters may be quantization parameters or lambda. In some implementations, the encoding parameters are determined based on the action output by the agent based on the encoding state of the video encoder. The agent may include a neural network implementing the reinforcement learning model, and the action output by the agent is mapped to the encoding parameters.

[0053]

[0057] In block 406, the computing device 100 encodes a second video unit different from the first video unit based on the encoding parameters. The first video unit may be the first frame, and the second video unit may be the second frame following the first frame. Alternatively, the first video unit may be the first block, and the second video unit may be a neighboring second block, for example, a spatially neighboring block or a temporally neighboring block.

[0054]

[0058] In some implementations, the reinforcement learning model is trained by a step of determining a reward for the encoding parameters based on the encoding of the second video unit, where the reward is set to impose a penalty on buffer overshoot and increase when the encoding parameters result in higher visual quality.

[0055]

[0059] In some implementations, the step of determining the reward includes: determining the base reward in such a way that the base reward has a negative value when there is a buffer overshoot, and the base reward is proportional to the encoding parameter with a negative coefficient when there is no buffer overshoot; and scaling the base reward by a scaling factor to obtain the reward, where the scaling factor is based on the ratio of the bandwidth associated with encoding the second video unit to the maximum bandwidth of the transmission channel. For example, the reward may be calculated based on equations (3) and (4).

[0056]

[0060] In some implementations, the reinforcement learning model is further trained by: determining an action associated with the encoding parameter based on the encoding state of the video encoder; determining an evaluation value of the encoding state regarding encoding the second video unit; determining a value loss based on the reward and the evaluation value; determining a policy loss based on the action; and updating the reinforcement learning model based on the value loss and the policy loss.

[0057]

[0061] In some implementations, the agent includes a neural network, and the neural network includes: at least one input fully connected layer configured to extract features from the encoding state; at least one recurrent neural network coupled to receive the extracted features; and at least one output fully connected layer configured to determine the action of the agent.

[0058]

[0062] In some implementations, the neural network is trained based on an actor-critic architecture, where the actor is configured to generate an action based on the encoding state, and the critic is configured to generate an evaluation value regarding the encoding state; and the actor and the critic share a common part by the neural network including at least one input fully connected layer and at least one recurrent neural network.

[0059]

[0063] Some exemplary implementations of the subject matter described herein are listed below.

[0060]

[0064] In a first aspect, the subject matter described herein provides a method executable by a computer. The method includes: determining an encoding state of a video encoder, the encoding state being associated with encoding a first video unit by the video encoder; determining, by a reinforcement learning model, encoding parameters regarding rate control of the video encoder based on the encoding state of the video encoder; and encoding a second video unit different from the first video unit based on the encoding parameters.

[0061]

[0065] In some implementations, the step of determining the encoding parameters includes: determining, by a reinforcement learning model, an action based on the encoding state of the video encoder; and mapping the action to the encoding parameters.

[0062]

[0066] In some implementations, the second video unit follows the first video unit, and the encoding state associated with encoding the first video unit includes: a state representing at least an outcome regarding encoding the first video unit; a state of a buffer configured to buffer video units encoded by the video encoder before transmission; and a state associated with a network situation for transmitting the encoded video units.

[0063]

[0067] In some implementations, the outcome regarding encoding the first video unit includes the encoding parameters of the encoded first video unit and the size of the encoded first video unit, the state of the buffer includes how the buffer is used, and the state associated with the network situation includes a target bits per pixel.

[0064]

[0068] In some implementations, the usage of the buffer includes at least one of: the ratio of the occupied space to the maximum space of the buffer; and the remaining space of the buffer measured in the video unit.

[0065]

[0069] In some implementations, the reinforcement learning model is trained by a step of determining a reward for encoding parameters based on the encoding of a second video unit, where the reward imposes a penalty for buffer overshoot and increases when the encoding parameters result in higher visual quality.

[0066]

[0070] In some implementations, the step of determining the reward includes: determining a basic reward in such a way that the basic reward has a negative value when buffer overshoot occurs and the basic reward is proportional to the encoding parameters with a negative coefficient when buffer overshoot does not occur; and scaling the basic reward by a scaling factor to obtain the reward, where the scaling factor is based on the ratio of the bandwidth associated with encoding the second video unit to the maximum bandwidth of the transmission channel.

[0067]

[0071] In some implementations, the reinforcement learning model is further trained by a step of determining an action associated with the encoding parameters based on the encoding state of the video encoder; determining an evaluation value of the encoding state for encoding the second video unit; determining a value loss based on the reward and the evaluation value; determining a policy loss based on the action; and updating the reinforcement learning model based on the value loss and the policy loss.

[0068]

[0072] In some implementations, the reinforcement learning model includes an agent's neural network, which includes: at least one input fully-connected layer configured to extract features from an encoded state; at least one recurrent neural network coupled to receive the extracted features; and at least one output fully-connected layer configured to determine the agent's actions.

[0069]

[0073] In some implementations, the neural network is trained based on an actor-critic architecture, where the actor is configured to generate actions based on an encoded state, the critic is configured to generate an evaluation value for the encoded state; and the actor and the critic share a common portion by a neural network including at least one input fully-connected layer and at least one recurrent neural network.

[0074] In some implementations, the encoding parameters include at least one of quantization parameters and lambda parameters.

[0070]

[0075] In some implementations, the video encoder is configured to encode screen content for real-time communication.

[0071]

[0076] In a second aspect, the subject matter described herein provides an electronic device. The electronic device includes a processing unit and a memory coupled to the processing unit and storing instructions that, when executed by the processing unit, cause the electronic device to perform any of the steps of the method.

[0072]

[0077] In a third aspect, the subject matter described herein provides a computer program product tangibly stored on a computer storage medium and including machine-executable instructions that, when executed by a device, cause the device to perform the method according to the aspect in the first aspect. The computer storage medium may be a non-transitory computer storage medium.

[0073]

[0078] In a fourth aspect, a non-transitory computer storage medium storing machine-executable instructions is provided, and when the instructions are executed by a device, the device is caused to execute the method according to the aspect of the first aspect.

[0074]

[0079] The functions described herein can be executed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chip systems (SOCs), complex programmable logic devices (CPLDs), and the like.

[0075]

[0080] The program code for performing the methods of the subject matter described herein may be written in any combination of one or more programming languages. The program code may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing apparatus, such that the program code, when executed by the processor or controller, causes the functions / operations shown in the flowchart and / or block diagram to be executed. The program code may be executed entirely or partially on the machine, or may be executed as a stand-alone software package partially on the machine, partially on a remote machine, or entirely on a remote machine or server.

[0076]

[0081] In the context of the subject matter described herein, a machine-readable medium may be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium include, but are not limited to, an electrical connection with one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0077]

[0082] Further, although operations are described in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, while some specific implementation details are included in the above description, these should not be construed as limitations on the scope of the subject matter described herein, but rather as descriptions of features that may be specific to particular implementations. The particular features described in the context of separate implementations may be implemented in combination within a single implementation. Rather, the various features described in the context of a single implementation may be implemented separately in multiple implementations, or in any suitable sub-combination.

[0078]

[0083] The subject matter is described in terms specific to structural features and / or methodological acts, but it should be understood that the subject matter defined by the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features or acts described above are disclosed as exemplary forms for carrying out the claims.

Claims

1. A method executed by a computer, comprising: determining an encoding state of a video encoder, wherein the encoding state is associated with encoding a first video unit by the video encoder; determining, by a reinforcement learning model, encoding parameters related to rate control of the video encoder based on the encoding state of the video encoder; and encoding a second video unit different from the first video unit based on the encoding parameters; wherein the reinforcement learning model is trained by determining a reward related to the encoding parameters based on the encoding of the second video unit, and the reward is obtained by scaling a basic reward by a scaling factor such that a penalty is imposed on an overshoot of a buffer and the reward increases when the encoding parameters result in higher visual quality, and the scaling factor is based on a ratio of a bandwidth associated with encoding the second video unit to a maximum bandwidth of a transmission channel.

2. The method according to claim 1, wherein the step of determining the encoding parameters comprises: determining, by the reinforcement learning model, an action based on the encoding state of the video encoder; and mapping the action to the encoding parameters.

3. The method according to claim 1, wherein the encoding state associated with encoding the first video unit is: a state representing at least an outcome related to encoding the first video unit; a state of a buffer configured to buffer a video unit encoded by the video encoder before transmission; and a state associated with a network situation for transmitting the encoded video unit.

4. ​ ​ In the method according to claim 3, the results related to encoding at least the first video unit include the encoding parameters of the encoded first video unit and the size of the encoded first video unit, the state of the buffer includes the usage of the buffer, and the state associated with the network situation includes the target bits per pixel. Method.

5. In the method according to claim 4, the usage of the buffer is as follows: The ratio of the occupied space to the maximum space of the buffer; and The remaining space of the buffer measured in the video unit; Method including at least one of the above.

6. In the method according to claim 1, in the reinforcement learning model, when there is no overshoot in the buffer, a smaller value of the encoding parameter results in a basic reward corresponding to higher visual quality, while a smaller value of the encoding parameter corresponding to the case of overshoot in the buffer results in a negative basic reward. Method.

7. In the method according to claim 6, the step of determining the reward is: When there is no overshoot in the buffer, determining the basic reward in such a way that the basic reward is proportional to the encoding parameter with a negative coefficient; Method including the above.

8. A method executed by a computer, comprising: Determining the encoding state of a video encoder, the encoding state being associated with encoding a first video unit by the video encoder; Determining, by a reinforcement learning model, an encoding parameter related to rate control of the video encoder based on the encoding state of the video encoder; and Encoding a second video unit different from the first video unit based on the encoding parameter; Including, the reinforcement learning model is trained by performing a step of determining a reward related to the encoding parameter based on the encoding of the second video unit, the reward imposing a penalty on buffer overshoot and increasing when the encoding parameter results in higher visual quality, and the reinforcement learning model Determining an action associated with the encoding parameter based on the encoding state of the video encoder; Determining an evaluation value of the encoding state regarding encoding the second video unit; Determining a value loss based on the reward and the evaluation value; Determining a policy loss based on the action; and Updating the reinforcement learning model based on the value loss and the policy loss; A method trained by the above.

9. In the method according to claim 1, the reinforcement learning model includes a neural network of an agent, and the neural network: At least one input fully-connected layer configured to extract features from the encoding state; At least one recurrent neural network coupled to receive the extracted features; and At least one output fully-connected layer configured to determine the action of the agent; A method including the above.

10. In the method according to claim 9, the neural network is trained based on an actor-critic architecture, the actor is configured to generate the action based on the encoding state, and the critic is configured to generate an evaluation value regarding the encoding state; and The actor and the critic share a common part by the neural network including the at least one input fully-connected layer and the at least one recurrent neural network. A method.

11. In the method according to claim 1, the encoding parameter includes at least one of a quantization parameter and a lambda parameter. A method.

12. In the method according to claim 1, the video encoder is configured to encode screen content of real-time communication. A method.

13. A processor; and A memory storing instructions for execution by the processor; A device including the above, wherein when the instructions are executed by the processor, the device is caused to perform an operation, and the operation is: A step of determining an encoding state of a video encoder, wherein the encoding state is associated with encoding a first video unit by the video encoder; A step of determining, by a reinforcement learning model, an encoding parameter related to rate control of the video encoder based on the encoding state of the video encoder; and A step of encoding a second video unit different from the first video unit based on the encoding parameter; Including, the reinforcement learning model is trained by performing a step of determining a reward related to the encoding parameter based on the encoding of the second video unit, the reward imposing a penalty on buffer overshoot and increasing when the encoding parameter results in higher visual quality, obtained by scaling a basic reward by a scaling factor, the scaling factor being based on the ratio of the bandwidth associated with encoding the second video unit to the maximum bandwidth of the transmission channel, a device.

14. The device according to claim 13, wherein the encoding state associated with encoding the first video unit is: A state representing at least the result of encoding the first video unit; The state of a buffer configured to buffer the video unit encoded by the video encoder before transmission; and A state associated with the network situation for transmitting the encoded video unit; Including, a device.

15. The device according to claim 14, wherein the result related to encoding at least the first video unit includes the encoding parameter of the encoded first video unit and the size of the encoded first video unit, the state of the buffer includes the usage of the buffer, and the state associated with the network situation includes the target bits per pixel, a device.

16. In the device according to claim 13, in the reinforcement learning model, when there is no overshoot in the buffer, a smaller-valued encoding parameter results in a basic reward corresponding to higher visual quality, while a smaller-valued encoding parameter corresponding to the case where an overshoot in the buffer occurs results in a basic reward with a negative value.

17. In the device according to claim 16, the step of determining the reward is: Determining the basic reward in such a way that the basic reward is proportional to the encoding parameter with a negative coefficient when there is no overshoot in the buffer; A device comprising.

18. In the device according to claim 13, the reinforcement learning model includes a neural network of an agent, and the neural network: At least one input fully connected layer configured to extract features from the encoded state; At least one recurrent neural network coupled to receive the extracted features; and At least one output fully connected layer configured to determine the action of the agent; A device comprising.

19. In the device according to claim 13, the video encoder is configured to encode screen content of real-time communication.

20. A computer program incorporating program instructions, the program instructions being executable by a processor to cause the processor to perform operations, the operations being: Determining an encoding state of a video encoder, the encoding state being associated with encoding a first video unit by the video encoder; Determining, by a reinforcement learning model, an encoding parameter regarding rate control of the video encoder based on the encoding state of the video encoder; and Encoding a second video unit different from the first video unit based on the encoding parameter. including, the reinforcement learning model is trained by performing a step of determining a reward for the encoding parameter based on the encoding of the second video unit, the reward imposing a penalty on overshoot of the buffer and being obtained by scaling a basic reward by a scaling factor such that the reward increases when the encoding parameter results in higher visual quality, the scaling factor being based on a ratio of a bandwidth associated with encoding the second video unit to a maximum bandwidth of a transmission channel, a computer program.

Citation Information

Patent Citations

  • Image coder

    JP2001231039A

  • Method and system for setting media frame output quality

    JP2006524461A

  • Method of and system to set an output quality of a media frame

    US20060192850A1