Space tier rate allocation method and system
By determining the spatial rate factor and frame sample allocation bit rate in the data processing hardware, the bandwidth constraint problem of different devices and applications in video coding is solved, the coding bitstream distortion of multiple spatial layers is optimized, and the consistency of user experience and video transmission efficiency are improved.
Patent Information
- Application Number
- CN202310024039.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-07-26
- Filing Date
- 2019-06-23
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2039-06-23
AI Technical Summary
Existing video coding technologies, faced with bandwidth or resource constraints across different devices and applications, struggle to effectively allocate bitrates to ensure video quality across all spatial layers, resulting in inconsistent user experience.
By receiving the transform coefficients of the scaled video input signal in data processing hardware, determining the spatial rate factor, and allocating bit rate based on the factor and frame samples, the bit rate allocation is adjusted using the spatial rate factor threshold and exponential moving average to optimize the coding bit flow distortion of multiple spatial layers.
It achieves video quality optimization of each spatial layer under the total bit rate limit, improves the consistency of user experience and the transmission efficiency of video content.
Smart Images

Figure CN116016935B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to spatial layer rate allocation in the context of scalable video coding. BACKGROUND
[0002] As video becomes more prevalent in a wide range of applications, video streams can need to be encoded and / or decoded several times depending on the situation of the application. For example, different applications and / or devices can need to adhere to bandwidth or resource constraints. To meet these needs without being overly expensive, several combinations of settings, efficient codecs have been developed that compress video into several resolutions. With codecs such as scalable VP9 and H.264, a video bitstream can contain multiple spatial layers that allow users to reconstruct the original video at different resolutions (i.e., the resolution of each spatial layer). By having scalable capabilities, video content can be transferred from device to device with limited further processing. SUMMARY
[0003] One aspect of the present invention provides a method for allocating bitrates. The method includes receiving, at data processing hardware, transform coefficients corresponding to a scalable video input signal, the scalable video input signal including a plurality of spatial layers, the plurality of spatial layers including a base layer. The method also includes determining, by the data processing hardware, a spatial rate factor based on frame samples from the scalable video input signal. The spatial rate factor defines a factor for bitrate allocation at each spatial layer of an encoded bitstream formed from the scalable video input signal. The spatial rate factor is represented by a difference between a bitrate of each transform coefficient of the base layer and an average bitrate of each transform coefficient of the plurality of spatial layers. The method further includes reducing distortion of the plurality of spatial layers of the encoded bitstream by allocating bitrates to each spatial layer based on the spatial rate factor and the frame samples.
[0004] Embodiments of the invention can include one or more of the following optional features. In some embodiments, the method further includes receiving, at the data processing hardware, a second frame sample from the scalable video input signal; modifying, by the data processing hardware, the spatial rate factor based on the second frame sample from the scalable video input signal; and allocating, by the data processing hardware, a modified bitrate to each spatial layer based on the modified spatial rate factor and the second frame sample. In further embodiments, the method further includes receiving, at the data processing hardware, a second frame sample from the scalable video input signal; modifying, by the data processing hardware, the spatial rate factor frame-by-frame based on an exponential moving average, the exponential moving average corresponding to at least the frame sample and the second frame sample; and allocating, by the data processing hardware, a modified bitrate to each spatial layer based on the modified spatial rate factor.
[0005] In some examples, receiving the scaled video input signal includes receiving the video input signal, scaling the video input signal into a plurality of spatial layers, partitioning each spatial layer into sub-blocks, transforming each sub-block into transform coefficients, and scalar quantizing the transform coefficients corresponding to each sub-block. Determining the spatial rate factor based on the frame samples from the scaled video input signal can include determining a variance estimate for each scalar quantized transform coefficient based on an average over all transform blocks of a frame of the video input signal. Here, the transform coefficients for each sub-block can be distributed identically over all sub-blocks.
[0006] In some implementations, the method further includes determining, by the data processing hardware, that the spatial rate factor satisfies a spatial rate factor threshold. In these implementations, the value corresponding to the spatial rate factor threshold can satisfy the spatial rate factor threshold when the value is less than about 1.0 and greater than about 0.5. The spatial rate factor can include a single parameter configured to allocate a bit rate to each layer of an encoded bitstream. In some examples, the spatial rate factor includes a weighted sum corresponding to a ratio of a variance product, where the ratio includes a numerator based on an estimated variance of scalar quantized transform coefficients from a first spatial layer and a denominator based on an estimated variance of scalar quantized transform coefficients from a second spatial layer.
[0007] Another aspect of the present disclosure provides a system for allocating a bit rate. The system includes data processing hardware and memory hardware in communication with the data processing hardware. The memory hardware stores instructions that, when executed by the data processing hardware, cause it to perform operations. The operations include receiving transform coefficients corresponding to a scaled video input signal, the scaled video input signal including a plurality of spatial layers, the plurality of spatial layers including a base layer. The operations further include determining a spatial rate factor based on frame samples from the scaled video input signal. The spatial rate factor defines a factor for bit rate allocation at each spatial layer of an encoded bitstream formed from the scaled video input signal. The spatial rate factor is represented by a difference between a bit rate of each transform coefficient of the base layer and an average bit rate of each transform coefficient of the plurality of spatial layers. The operations further include reducing distortion of the plurality of spatial layers of the encoded bitstream by allocating a bit rate to each spatial layer based on the spatial rate factor and the frame samples.
[0008] This aspect can include one or more of the following optional features. In some implementations, the operations further include receiving a second frame sample from the scaled video input signal; modifying the spatial rate factor based on the second frame sample from the scaled video input signal; and assigning a modified bit rate to each spatial layer based on the modified spatial rate factor and the second frame sample. In further implementations, the operations further include receiving a second frame sample from the scaled video input signal; modifying the spatial rate factor on a frame-by-frame basis based on an exponentially moving average, the exponentially moving average corresponding to at least the frame sample and the second frame sample; and assigning a modified bit rate to each spatial layer based on the modified spatial rate factor.
[0009] In some examples, receiving the scaled video input signal includes receiving a video input signal, scaling the video input signal into a plurality of spatial layers, partitioning each spatial layer into sub-blocks, transforming each sub-block into transform coefficients, and scalar quantizing the transform coefficients corresponding to each sub-block. Determining the spatial rate factor based on a frame sample from the scaled video input signal can include determining a variance estimate for each scalar quantized transform coefficient based on an average over all transform blocks of a frame of the video input signal. Here, the transform coefficients for each sub-block can be identically distributed over all sub-blocks.
[0010] In some implementations, the operations further include determining that the spatial rate factor satisfies a spatial rate factor threshold. In these implementations, the value corresponding to the spatial rate factor threshold can satisfy the spatial rate factor threshold when the value is less than about 1.0 and greater than about 0.5. The spatial rate factor can include a single parameter configured to assign a bit rate to each layer of an encoded bitstream. In some examples, the spatial rate factor includes a weighted sum corresponding to a ratio of a variance product, where the ratio includes a numerator based on an estimated variance of scalar quantized transform coefficients from a first spatial layer and a coefficient based on an estimated variance of scalar quantized transform coefficients from a second spatial layer.
[0011] The details of one or more implementations of the application are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims. BRIEF DESCRIPTION OF DRAWINGS
[0012] Figure 1 is a schematic diagram of an example rate allocation system.
[0013] Figure 2 is Figure 1 is a schematic diagram of an example encoder within the rate allocation system of
[0014] Figure 3 is Figure 1 is a schematic diagram of an example allocator within the rate allocation system of
[0015] Figure 4 is a flowchart of an example method for implementing a rate allocation system.
[0016] Figure 5 is a schematic diagram of an example computing device that can be used to implement the systems and methods described herein.
[0017] The same reference numbers in different drawings represent the same element. DETAILED DESCRIPTION
[0018] Figure 1 is an example of a rate allocation system 100. The rate allocation system 100 generally includes a video source device 110 that transmits captured video as a video input signal 120 to a remote system 140 via a network 130. At the remote system 140, an encoder 200 and an allocator 300 convert the video input signal 120 into an encoded bitstream 204. The encoded bitstream 204 includes more than one spatial layer L 0-i , where i denotes the number of spatial layers L 0-i . Each spatial layer L is a scalable form of the encoded bitstream 204. A scalable video bitstream refers to a video bitstream in which portions of the bitstream can be removed by way of producing a substream (e.g., a spatial layer L) that forms a valid bitstream for certain target decoders. More specifically, the substream represents the source content (e.g., captured video) of the original video input signal 120 at a lower reconstruction quality than the original captured video. For example, a first spatial layer L1 has a 1280 x 720 720p high definition (HD) resolution, while a base layer L0 scales to a 640 x 360 resolution as an extended form of a video graphics adapter (VGA) resolution. In terms of scalability, a video is generally scalable in time (e.g., by frame rate), in space (e.g., by spatial resolution), and / or in quality (e.g., by fidelity, often referred to as signal-to-noise ratio, SNR).
[0019] The rate allocation system 100 is an example environment in which a user 10, 10a captures video at a video source device 110 and transmits the captured video to other users 10, 10b-c. Here, the encoder 200 and the allocator 300 convert the captured video into an encoded bitstream 204 at an allocated bitstream rate before the users 10b, 10c receive the captured video via video receiving devices 150, 150b-c. Each video receiving device 150 can be configured to receive and / or process a different video resolution. Here, a spatial layer L with a larger layer number i refers to a layer L with a larger resolution, so i = 0 refers to a base layer L0 with the lowest scalable resolution within the bitstream of more than one spatial layer L 0-i Figure 1 The encoded video bitstream 204 includes two spatial layers L0, L1. As such, one video receiving device 150 can receive video content as the lower resolution spatial layer L0, while another video receiving device 150 can receive video content as the higher resolution spatial layer L1. For example, Figure 1 The first video receiving device 150a belonging to user 10b is described as receiving the lower spatial resolution layer L0 as a cell phone, while user 10c, who owns the second receiving device 150b as a laptop, receives the higher spatial resolution layer L1.
[0020] When different video receiving devices 150a-b receive different spatial layers L 0-i , the video quality of each spatial layer L can depend on the bit rate B R and / or allocation factor A F of the received spatial layer L. Here, the bit rate B R corresponds to the number of bits per second, and the allocation factor A F corresponds to the number of bits per sample (i.e., transform coefficient). In the case of a scalable bitstream (e.g., the encoded bitstream 204), the total bit rate B Rtot of the scalable bitstream is typically limited such that each spatial layer L of the scalable bitstream receives a similar bit rate limitation. Due to these limitations, the bit rate B R associated with one spatial layer L can compromise or trade-off the quality of another spatial layer L. More specifically, if the quality on a spatial layer L received by a user 10 via a video receiving device 150 is compromised, the quality can negatively impact the user experience. For example, as real-time communication (RTC) applications become more prevalent for transmitting video content as a form of communication. Users 10 using RTC applications can typically select an application for communication based on the subjective quality of the application. As such, as an application user, the user 10 typically expects to have a positive communication experience without quality issues that can arise due to insufficient bit rate allocated to the spatial layer L received by the application user 10. To help ensure a positive user experience, the allocator 300 is configured to adaptively communicate an allocation factor A F to determine a bit rate B 0-i for each spatial layer L of the plurality of spatial layers L R . By apportioning the allocation factor A 0-i among the plurality of spatial layers L F , the allocator 300 seeks to achieve the highest video quality across all spatial layers L Rtot for a given total bit rate B 0-i .
[0021] Video source device 110 can be any computing device or data processing hardware capable of transmitting captured video and / or video input signal 120 to network 130 and / or remote system 140. In some examples, video source device 110 includes data processing hardware 112, memory hardware 114, and video capture device 116. In some implementations, video capture device 116 is actually an image capture device, which can transmit captured image sequences as video content. For example, some digital cameras and / or webcams are configured to capture images at a particular frequency to form a visual video content. In other examples, video source device 110 captures video in a continuous analog format, which can then be converted to a digital format. In some configurations, video source device 110 includes an encoder to preliminarily encode or compress captured data (e.g., analog or digital) into a format for further processing by encoder 200. In other examples, video source device 110 is configured to access encoder 200 at video source device 110. For example, encoder 200 is a web application hosted on remote system 140, but is accessible by video source device 110 via a network connection. In other examples, portions or all of encoder 200 and / or distributor 300 are hosted on video source device 110. For example, encoder 200 and distributor 300 are hosted on video source device 110, but remote system 140 is used as a backend system that relays bitstreams including spatial layers L 0-i to video sink device 150 according to decoding capabilities of video sink device 150 and connection capabilities of network 130 between video sink device 150 and remote system 140. Additionally or alternatively, video source device 110 is configured such that user 10a can use video capture device 116 to communicate with another user 10b-c over network 130.
[0022] Video input signal 120 is a video signal corresponding to captured video content. Here, video source device 110 captures video content. For example, Figure 1 depicts video source device 110 capturing video content via webcam 116. In some examples, video input signal 120 is an analog signal that is processed into a digital format by encoder 200. In other examples, video input signal 120 undergoes some degree of encoding or digital formatting prior to encoder 200, such that encoder 200 performs a re-quantization process.
[0023] Similar to the video source device 110, the video sink device 150 can be any computing device or data processing hardware capable of receiving transmitted captured video via the network 130 and / or the remote system 140. In some examples, the video source device 110 and the video sink device 150 are configured with identical functionality such that the video sink device 150 can become the video source device 110 and the video source device 110 can become the video sink device 150. In either case, the video sink device 150 includes at least data processing hardware 152 and memory hardware 154. Additionally, the video sink device 150 includes a display 156 configured to display received video content (e.g., at least one layer L of the encoded bitstream 204). As shown, the user 10b, 10c is receiving the encoded bitstream 204 as a spatial layer L and decoding the encoded bitstream 204 and displaying as video on the display 156. In some examples, the video sink device 150 contains a decoder or is configured to access a decoder (e.g., via the network 130) to allow the video sink device 150 to display the content of the encoded bitstream 204. Figure 1 As shown, the user 10b, 10c is receiving the encoded bitstream 204 as a spatial layer L and decoding the encoded bitstream 204 and displaying as video on the display 156. In some examples, the video sink device 150 contains a decoder or is configured to access a decoder (e.g., via the network 130) to allow the video sink device 150 to display the content of the encoded bitstream 204. R As shown, the user 10b, 10c is receiving the encoded bitstream 204 as a spatial layer L and decoding the encoded bitstream 204 and displaying as video on the display 156. In some examples, the video sink device 150 contains a decoder or is configured to access a decoder (e.g., via the network 130) to allow the video sink device 150 to display the content of the encoded bitstream 204.
[0024] In some examples, the encoder 200 and / or the distributor 300 are applications hosted by the remote system 140 (e.g., a distributed system of a cloud environment) that are accessed via the video source device 110 and / or the video sink device 150. In some implementations, the encoder 200 and / or the distributor 300 are applications downloaded to the memory hardware 114, 154 of the video source device 110 and / or the video sink device 150. Regardless of the point of access for the encoder 200 and / or the distributor 300, the encoder 200 and / or the distributor 300 can be configured to communicate with the remote system 140 to access resources 142 (e.g., data processing hardware 144, memory hardware 146, or software resources 148). Access to the resources 142 of the remote system 140 can allow the encoder 200 and / or the distributor 300 to encode the video input signal 120 into the encoded bitstream 204 and / or to assign the bit rate B R to more than one spatial layer L of the encoded bitstream 204. Optionally, as a software resource 148 of the remote system 140 for communicating between the users 10, 10a-c, a real-time communication (RTC) application includes the encoder 200 and / or the distributor 300 as a built-in functionality. 0-i As shown, the user 10b, 10c is receiving the encoded bitstream 204 as a spatial layer L and decoding the encoded bitstream 204 and displaying as video on the display 156. In some examples, the video sink device 150 contains a decoder or is configured to access a decoder (e.g., via the network 130) to allow the video sink device 150 to display the content of the encoded bitstream 204.
[0025] Referring in more detail Figure 1Three users 10, 10a-c are communicating via an RTC application (e.g., a WebRTC video application hosted by a cloud) hosted by a remote system 140. In this example, a first user 10a is in a group video chat with a second user 10b and a third user 10c. As the video capture device 116 captures video of the first user 10a speaking, the video captured via the video input signal 120 is processed by the encoder 200 and the distributor 300 and transmitted via the network 130. Here, the encoder 200 and the distributor 300 operate in conjunction with the RTC application to generate an encoded bitstream 204 having more than one spatial layer L0, L1, where each spatial layer L has an allocated bitrate B R0 , B R1 The allocated bitrate is determined based on an allocation factor A F0 , A F1 based on the video input signal 120. As each video receiving device 150a, 150b has different capabilities, each user 10b, 10c receiving the first user 10a video chat receives a different scaled version of the original video corresponding to the video input signal 120. For example, the second user 10b receives the base spatial layer L0, while the third user 10c receives the first spatial layer L1. Each user 10b, 10c continues to display the video content received in communication with the RTC application on the display 156a, 156b. Although an RTC communication application is shown, the encoder 200 and / or the distributor 300 can be used in other applications involving an encoded bitstream 204 having more than one spatial layer L 0-i .
[0026] Figure 2 is an example of an encoder 200. The encoder 200 is configured to convert the video input signal 120 as input 202 to an encoded bitstream as output 204. Although shown separately, the encoder 200 and the distributor 300 can be integrated into a single device (e.g., as shown by the dashed line in Figure 1 The encoder 200 generally includes a sealer 210, a transformer 220, a quantizer 230, and an entropy encoder 240. Although not shown, the encoder 200 can include additional components for generating the encoded bitstream 204, such as a prediction component (e.g., motion estimation and intra prediction) and / or a loop filter. The prediction component can produce a residual to be passed to the transformer 220 for transformation, where the residual is based on a difference of an original input frame minus a frame prediction (e.g., motion compensation or intra prediction).
[0027] The sealer 210 is configured to scale the video input signal 120 into multiple spatial layers L 0-iIn some implementations, the scaler 210 scales the video input signal 120 by determining portions of the video input signal 120 that can be removed to reduce the spatial resolution. By removing one or more portions, the scaler 210 forms multiple versions of the video input signal 120 to form multiple spatial layers (e.g., substreams). The scaler 210 can repeat the process until the scaler 210 forms a base spatial layer L0. In some examples, the scaler 210 scales the video input signal 120 to form a set number of spatial layers L 0-i In other examples, the scaler 210 is configured to scale the video input signal 120 until the scaler 210 determines that there are no decoders to decode the substreams. When the scaler 210 determines that there are no decoders to decode the substreams corresponding to the scaled versions of the video input signal 120, the scaler 210 identifies the previous version (e.g., spatial layer L) as the base spatial layer L0. Some examples of the scaler 210 include codecs corresponding to the Scalable Video Coding (SVC) extension, such as an extension of the H.264 video compression standard or an extension of the VP9 encoding format.
[0028] The transformer 220 is configured to receive each spatial layer L corresponding to the video input signal 120 from the scaler 210. For each spatial layer L, the transformer 220 partitions each spatial layer L into subblocks at operation 222. With each subblock, the transformer 220 transforms each subblock at operation 224 to generate transform coefficients 226 (e.g., by a discrete cosine transform (DCT)). By producing the transform coefficients 226, the transformer 220 can correlate redundant video data and non-redundant video data to help the encoder 200 remove the redundant video data. In some implementations, the transform coefficients also allow the distributor 300 to easily determine the number of coefficients for each transform block in a spatial layer L that have a non-zero variance.
[0029] The quantizer 230 is configured to perform a quantization or re-quantization process 232 (i.e., scalar quantization). A quantization process generally converts an input parameter (e.g., from a continuous analog data set) to a smaller output value data set. Although a quantization process can convert an analog signal to a digital signal, here the quantization process 232 (sometimes also referred to as a re-quantization process) generally further processes a digital signal. Depending on the form of the video input signal 120, either process can be used interchangeably. By applying a quantization or re-quantization process, data can be compressed, but at the cost of some aspect of data loss, as the smaller data set is a reduction of the larger or continuous data set. Here, the quantization process 232 converts a digital signal. In some examples, the quantizer 230 facilitates the formation of the encoded bitstream 204 by scalar quantizing the transform coefficients 226 from each sub-block of the transformer 220 into quantization indices 234. Here, scalar quantizing the transform coefficients 226 can allow lossy encoding to scale each transform coefficient 226 in order to contrast redundant video data (e.g., data that can be removed during encoding) with valuable video data (e.g., data that should not be removed).
[0030] The entropy encoder 240 is configured to convert the quantization indices 234 (i.e., quantized transform coefficients) and side information into bits. Through this conversion, the entropy encoder 240 forms the encoded bitstream 204. In some implementations, the entropy encoder 240, along with the quantizer 230, enables the encoder 200 to form the encoded bitstream 204, where each layer L 0-i has a bit rate B F0-i based on an allocation factor A R0-i determined by the allocator 300.
[0031] Figure 3 The allocator 300 is an example of an allocator. The allocator 300 is configured to receive non-quantized transform coefficients 226 related to more than one spatial layer L 0-i and determine an allocation factor A 0-i for each received spatial layer L F In some implementations, the allocator 300 determines each allocation factor A F based on a square error based high-rate approximation for scalar quantization. The square error high-rate approximation allows a system to determine an optimal (in the case of a high-rate approximation) bit rate to allocate to N scalar quantizers. Generally, the optimal bit rate to allocate to N scalar quantizers is determined by rate-distortion optimized quantization. Rate-distortion optimization attempts to minimize a distortion subject to a bit rate constraint (e.g., a total bit rate B Rtotto improve the video quality during video compression. Here, the allocator 300 applies the principle of determining the optimal bit rate for N scalar quantizers to determine the optimal allocation factors to allocate the bit rate to more than one spatial layer L 0-i in the encoded bitstream 204.
[0032] In general, the squared error high-rate approximation of scalar quantization can be represented by the following equation:
[0033]
[0034] where, depends on the source distribution of the input signal (e.g., transform coefficients) to the i-th quantizer, is the variance of the signal, and r i is the bit rate for the i-th quantizer in bits per input symbol. The following is the expression for the optimal rate allocation for two scalar quantizers derived using the squared error high-rate approximation.
[0035] The average distortion D2 of the two-quantizer problem, D2, is equal to Similarly, the average rate R2 of the two-quantizer problem is equal to Here, d i is the squared error distortion due to the i-th quantizer, and r i is the bit rate allocated to the i-th quantizer in bits per sample. Although, the parameter d i is a function of the rate r i such that an equation like d i (r i ) is appropriate, for convenience, d i is simply substituted for d i . Substituting the high-rate approximations of d0and di into the equation for D2yields:
[0036]
[0037] With equation (2), one can substitute r1with 2R2- r0to get:
[0038]
[0039] By further taking the derivative of D2with respect to r0, equation (3) yields the following expression:
[0040]
[0041] Setting the above expression (equation (4)) to zero and solving for r0yields the expression for the optimal rate r* of the zero quantizer, denoted as follows:
[0042]
[0043] Because the expression for the high-rate distortion is convex, the minimum found by setting the derivative to zero is global. Similarly, the optimal rate r* of the first quantizer can be denoted as follows:
[0044]
[0045] To find the optimal quantizer distortion and Substitute equations (5) and (6) into the high-rate expression for the scalar quantizer distortion as follows:
[0046]
[0047] A simplified form of equation (7) yields the following equation:
[0048] For all i (8)
[0049] The same two-quantizer analysis can be extended to three quantizers by combining the zero quantizer and the first quantizer into a single quantization system (i.e., a nested system), where the combined quantizers have been solved according to equations (1)-(8). Using a similar approach to the rate allocation for the two-quantizer system, the three-quantizer system is derived as follows.
[0050] Because the average single quantizer distortion of the two-quantizer system is denoted as Substitute d avg into the expression for the average three-quantizer distortion yields the following equation:
[0051]
[0052] Similarly, the average rate of the three-quantizer system is denoted as follows:
[0053] where
[0054] Utilizing the optimal distortion results from the two-quantizer analysis as shown in equation (8), it follows that the three-quantizer distortion can be denoted by the following equation:
[0055]
[0056] Therefore, when equation (11) is simplified and substituted into At equation (11), equation (11) is transformed into the following expression:
[0057]
[0058] Using equation (12), the derivative with respect to r2 can be set to zero and solved for r2 to yield the following equation:
[0059]
[0060] For three quantizers, equation (13) can be more fully expressed as follows:
[0061]
[0062] Based on the first and second quantizers, an expression for the optimal rate allocation r* for N quantizers can be derived. The expression for the optimal rate for the i-th quantizer is as follows:
[0063]
[0064] By substituting the expression for the optimal rate into the expression for the high rate of distortion and performing similar simplifications as for the two quantizer expression, the resulting expression for the optimal distortion for N quantizers is as follows.
[0065]
[0066] Based on the expressions derived from equations (1)-(16), the allocator 300 can apply these expressions to the optimal distortion to determine the optimal allocation factor A 0-i (i.e., that contributes to the optimal bit rate B R ) for each layer L of the plurality of spatial layers L F . Similar to the N quantizer expressions that have been derived, the multi-spatial layer bit rate can be deduced from the expressions associated with the two and three layer rate allocation systems. In some examples, it is assumed that although the spatial layers L 0-i typically have different spatial dimensions, the spatial layers L 0-i originate from the same video source (e.g., the video source device 110). In some implementations, the scalar quantizers that encode the first spatial layer L0 and the second spatial layer L1 are assumed to be identical in structure, even though the values of these scalar quantizers can be different. Furthermore, for each spatial layer L, the number of samples S is typically equal to the number of transform coefficients 226 (i.e., also equal to the number of quantizers).
[0067] In the case of the two spatial layer rate allocation system, the average distortion D2 for the two spatial layers can be expressed as a weighted sum of the average distortions do and di, and corresponds to the first and second spatial layers L0, L1 (i.e., spatial layers 0 and 1) as follows:
[0068]
[0069] where s i Equal to the i-th spatial layer L i The number of samples in , and S = s0 + s1. Similarly, the average bit rate of the two spatial layers can be expressed as follows:
[0070]
[0071] Where r0 and r1 are the average bit rates of the first and second spatial layers L0 and L1, respectively. By substituting the expression for the optimal distortion of N quantizers (i.e., equation (16)) into equation (17) above for D2, D2 can be expressed as follows:
[0072]
[0073] in is the i-th spatial layer L i The variance of the input signal to the jth scalar quantizer in . Solving equation (18) for r1 and substituting the result into equation (19) yields:
[0074]
[0075] Furthermore, by setting the derivative of D2 with respect to r0 to zero and solving for r0, r0 can be expressed by the following equation:
[0076]
[0077] Simplifying equation (21) for ease of expression, The P i Substitute the expression (21) and rearranging the resulting terms to form the following expression similar to the N quantizer allocation expression:
[0078]
[0079] Alternatively, you can use Represent equation (22) to achieve the following equation:
[0080]
[0081] Based on equations (17)-(23), the optimal two-spatial layer distortion can be expressed as follows:
[0082]
[0083] A similar approach can be applied to the three spatial layers L 0-2the best allocation factor for the i-th spatial layer L1. Very similarly to the two spatial layers L0, L1, s i is equal to the number of samples in the i-th spatial layer L1, such that S = s0+ s1+ s2. The three spatial layers L 0-2 The average rate and distortion, R3and D3, can be expressed as a weighted sum of the average rates and distortions r0, r1and r2and d0, d1and d2of the spatial layers 0, 1 and 2 (e.g., the three spatial layers L 0-2 ) as follows:
[0084] and
[0085]
[0086] When similar techniques are applied from the two quantizer result to the three quantizers, R3can be expressed as a combination of the average two-layer rate R2using the following equation:
[0087]
[0088] where
[0089] Similarly, for three quantizers, the distortion can be expressed as follows:
[0090]
[0091] where Using the equation (24) for the two-layer optimal distortion and the equation (8) for the optimal N-quantizer distortion D3can be solved in equation (29) to give the following expression:
[0092]
[0093] where R2can be solved in equation (27) to give the following expression:
[0094]
[0095] Further, equations (31) and (32) can be combined by substituting equation (32) into equation (31) for D3to form the following equation:
[0096]
[0097] An expression for r2can be formed by taking the derivative of D3with respect to r2and setting the result to zero. This expression can be represented by the following equation:
[0098]
[0099] When the terms are rearranged, Equation (34) may look similar to the N-quantizer allocation expression as follows:
[0100]
[0101] Applying this equation (36) to the first layer L0 and the second layer L1, the allocation factor of each layer can be expressed as follows:
[0102]
[0103] and
[0104] Two spatial layers L 0-1 and three spatial layers L 0-2 The two derivations of illustrate that the method can be extended to multiple spatial layers to optimize the rate allocation at the allocator 300 (e.g., for determining the bit rate B allocated to each spatial layer L). R The distribution factor A F Here, the above results are extended to L spatial layers L 0-L The general expression is obtained as shown in the following equation:
[0105]
[0106] where R L is the average rate corresponding to L spatial layers L 0-i The number of bits per sample on L spatial layers; the total number of samples S on L spatial layers, where s i is the number of samples in the i-th spatial layer; Among them, h j,i depends on the source distribution of the signal quantized by the jth quantizer in the i-th spatial layer; and corresponds to the variance of the j-th transform coefficient in the i-th spatial layer.
[0107] In some embodiments, due to various assumptions, equation (39) has different forms. Two different forms of equation (39) are shown below.
[0108]
[0109]
[0110] For example, h j,i The value depends on the i-th spatial layer L i The source distribution of the video input signal 120 quantized by the jth quantizer in . In the example with similar source distribution, h j,ithe values do not change between different quantizers, and thus cancel out due to the ratio of the product terms in equation (39). In other words, h j,0 = h j,1 = h j,2 = h. Thus, when this cancellation occurs, the term
[0111] This effectively eliminates consideration of this parameter, as P i always occurs in the ratio, where h in the numerator is used to cancel out a similar term in the denominator. In practice, h j,0 may be different from h j,1 and h j,2 , as the base spatial layer L0 uses only temporal prediction, while the other spatial layers can use both temporal and spatial prediction. In some configurations, this difference does not significantly affect the allocation factor A F determined by the allocator 300.
[0112] In other embodiments, the encoder 200 introduces a transform block that produces the transform coefficients 226. When this occurs, the combination of transform coefficients 226 can change, which introduces a variable s i '. The variable s i ' corresponds to the average number of transform coefficients 226 per transform block in the i-th spatial layer L i that has a non-zero variance, as shown in equation (39a). Unlike this variable s i ', s i in equation (39b) corresponds to the number of samples S in the i-th spatial layer L i . Furthermore, in equation (39a), the term where σ 2 k,i is the variance of the k-th coefficient in a transform block in the i-th spatial layer L i . In effect, equation (39a) represents the optimal bit rate allocation for the i-th spatial layer L i as an expression of a weighted sum of the ratios of variance products (e.g., ).
[0113] Referring to Figure 3 , in some embodiments, the allocator 300 includes a sampler 310, an estimator 320, and a rate determiner 330. The sampler 310 receives the non-quantized transform coefficients 226 having a plurality of spatial layers L 0-i as an input 302 to the allocator 300. For example, Figure 2The transform coefficients 226 generated by the transformer 220 are shown and are transmitted to the distributor 300 via a dotted line. Using the received unquantized transform coefficients 226, the sampler 310 identifies a frame of the video input signal 120 as samples S F Based on the sample S identified by the sampler 310 F , the allocator 300 determines the allocation factor A for each spatial layer L F In some embodiments, the allocator 300 is configured to dynamically determine the allocation factor A for each spatial layer L. F In these embodiments, the sampler 310 may be configured to iteratively identify frame samples S F The set of allocators 300 can allocate the factor A F Adapted to each sample S identified by the sampler 310 F For example, the distributor 310 is based on the first sample S of the frame of the video input signal 120. F1 To determine the allocation factor A to each spatial layer L F Then (eg, if necessary) based on the second sample S of the frame of the video input signal 120 identified by the sampler 310 F2 , continue to adjust or modify the allocation factor A applied to each spatial layer L F (For example, Figure 3 As shown, from the first sample S F1 The first distribution factor A F1 Become the second sample S F2 The second distribution factor A F2 ). This process may continue iteratively for the duration that the distributor 300 receives the video input signal 120. In these examples, the distributor 300 modifies the distribution factor A F , distribution factor A F Based on the first sample S F1 and the second sample S F2 332 (eg, from the first spatial rate factor 3321 to the second spatial rate factor 3322). Additionally or alternatively, the allocator 300 may use an exponential moving average to modify the allocation factor A frame by frame. F The exponential moving average is usually a weighted moving average that uses a distribution factor A from the previous frame. F The weighted average of the distribution factor A determined by the current frame F In other words, here, the allocation factor A F Each modification of is with the current and previous allocation factors A F The weighted average of .
[0114] The estimator 320 is configured to determine a variance estimate 322 for each transform coefficient from the encoder 200. In some configurations, the estimator 320 assumes that the transform coefficients 226 in each block from the transformer 220 are similarly distributed. Based on this assumption, the variance of the transform coefficients 226 can be estimated by averaging over all transform blocks in a sample frame S F of the video input signal 120. For example, the following expression models the kth transform coefficient 226 in the ith spatial layer L k,i .
[0115]
[0116] where ε b,k,i,t is the kth transform coefficient 226 in the bth transform block in the ith spatial layer L i in the tth frame, B i denotes the number of blocks in the ith spatial layer L i , and S F denotes the number of sample frames used to estimate the variance. In some examples, the value of the ith spatial layer L i , the estimate of the variance of the kth transform coefficient 226 in the ith spatial layer L F is independent of the transform block, when all transform blocks are assumed to have the same statistical quantities. However, in practice, the statistical quantities of the transform blocks can vary across the frame. This is especially true for video conferencing content, where blocks at the frame edges can have less activity than blocks in the center. Thus, estimating the variance based on blocks located in the center of the frame can mitigate the negative effects if these different statistical quantities negatively affect the accuracy of the rate allocation results. In some configurations, the sub-blocks used to estimate the transform coefficient variance represent a subset of all sub-blocks in the video image (e.g., sub-blocks located in the most central portion of the video image or sub-blocks located in the video image that have changed relative to a previous image).
[0117] The rate determiner 330 is configured to determine a spatial rate factor 332 based on the frame samples S F identified by the sampler 310 of the video input signal 120. In some examples, the spatial rate factor 332 defines a factor used to determine the bit rate B 0-i at each spatial layer L R of the encoded bitstream 204. The spatial rate factor 332 is assigned to the spatial layer L i-1the ratio between the bit rate allocated to the spatial layer L0 and the bit rate allocated to the spatial layer L1. In the two spatial example with spatial layers L0 and L1, a spatial rate factor equal to 0.5, and a bit rate allocated to the spatial layer L1 equal to 500 kbps, the bit rate allocated to the spatial layer L0 is equal to 250 kbps (i.e. 0.5 times 500 kbps). In these embodiments, the value of the spatial rate factor 332 is set to the allocation factor A F the difference between the average rate R L and the rate r*0 of the basic layer L0 (e.g. the expression r*0 - R L of equation (39)). Here, the allocation factor A F corresponds to the number of bits per transform coefficient of the basic layer L0 (also denoted r*0), while the average rate R L corresponds to the number of bits per transform coefficient of more than one spatial layer L 0-i . In some configurations, experimental results with two spatial layers show that the spatial rate factor 332 corresponds to the expression As a single parameter, the spatial rate factor 332 can allow the allocator 300 to easily adjust or modify the bit rate B R of each layer L of the encoded bitstream 204.
[0118] Although explained with respect to two spatial layers, the allocator 300 can apply the spatial rate factor 332 and / or the allocation factor A F to any number of spatial layers L 0-i . For example, the allocator 300 determines the allocation factor A F and / or the spatial rate factor 332 with respect to each group of two spatial layers. To illustrate with three layers L 0-2 , the allocator 300 first determines the allocation factor A F for the basic layer L0 and the first layer L1, and then determines the allocation factor A F for the first layer L1 and the second layer L2. Each allocation factor A F can be used to determine a spatial rate factor 332, one spatial rate factor 332 for the basic layer L0 and the first layer L1, and a second spatial rate factor 332 for the first layer L1 and the second layer L2. With the spatial rate factor 332 of each group of two spatial layers, the allocator 300 can average (e.g. weighted average, arithmetic average, geometric average, etc.) the spatial rate factors 332 and / or the allocation factors A F to generate an average spatial rate factor and / or an average allocation factor for any number of spatial layers L 0-i .
[0119] In some examples, the spatial rate factor 332 must satisfy a spatial rate factor threshold 334 (e.g., be within a range of values) for the allocator 300 to help determine a bit rate B R In some implementations, the value satisfies the spatial rate factor threshold 334 when the value is within an interval less than about 1.0 and greater than about 0.5. In other implementations, the spatial rate factor threshold 334 corresponds to a narrower range of values (e.g., 0.55-0.95, 0.65-0.85, 0.51-0.99, 0.65-1.0, 0.75-1.0, etc.) or a wider range of values (e.g., 0.40-1.20, 0.35-0.95,
[0120] 0.49-1.05, 0.42-1.17, 0.75-1.38, etc.). In some configurations, when the spatial rate factor 332 is outside the range of values for the spatial rate factor threshold 334, the allocator 300 adjusts the spatial rate factor 332 to satisfy the spatial rate factor threshold 334. For example, when the spatial rate factor threshold 334 takes an interval of 0.45-0.95, the spatial rate factor 332 outside this interval is adjusted to the nearest maximum value of the interval (e.g., a spatial rate factor 332 of 0.3 is adjusted to a spatial rate factor 332 of 0.45, while a spatial rate factor 332 of 1.82 is adjusted to a spatial rate factor 332 of 0.95).
[0121] Based on the determined spatial rate factor 332, the allocator 300 is configured to optimize video quality by reducing a distortion of one or more spatial layers L 0-i subject to a constraint of a total bit rate B Rtot . To reduce the distortion, the allocator 300 influences (e.g., helps the encoder 200 determine) a bit rate B F for each spatial layer L based on the spatial rate factor 332 calculated for the frame sample S R . For example, when the encoded bitstream 204 includes two spatial layers L0, L1, the allocator 300 determines an allocation factor A F , the allocation factor A F is in turn used to determine the spatial rate factor 332 to generate a first bit rate B corresponding to the equation R1 and a second bit rate B corresponding to the equation R0 , where B Rtot corresponds to the total bit rate available to encode the total bitstream (i.e., all spatial layers L0, L1).
[0122] Figure 4is an example of a method 400 for implementing the rate distribution system 100. In operation 402, the method 400 receives transform coefficients 226 (e.g., non-quantized transform coefficients) corresponding to the video input signal 120 at the data processing hardware 510. The video input signal 120 includes a plurality of spatial layers L 0-i , where multiple spatial layers L 0-i In operation 404, the method 400 performs the processing of the frame samples S from the video input signal 120 by the data processing hardware 510. F The spatial rate factor 332 is determined. The spatial rate factor 332 defines a factor for rate allocation on each spatial layer L of the coded bitstream 204 and is a function of the bit rate of each transform coefficient of the base layer L0 and the bit rate of the multiple spatial layers L. 0-i The average bit rate per transform coefficient R L In operation 406, method 400 is performed by data processing hardware 510 by performing a calculation based on spatial rate factor 332 and frame sample S F Set the bit rate B R The number of spatial layers L allocated to each spatial layer L to reduce the coded bit stream 204 0-i The distortion d.
[0123] Figure 5 is a schematic diagram of an example computing device 500 that can be used to implement the systems and methods described in this document (e.g., encoder 200 and / or distributor 300). Computing device 500 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The components shown here, their connections and relationships, and their functions are merely exemplary and are not intended to limit the embodiments of the inventions described and / or claimed in this document.
[0124] Computing device 500 includes data processing hardware 510, memory hardware 520, storage hardware 530, high-speed interface / controller 540 connecting to memory 520 and high-speed expansion ports 550, and low-speed interface / controller 560 connecting to low-speed bus 570 and storage 530. Each of the components 510, 520, 530, 540, 550, and 560 are interconnected using various busses, and can be mounted on a common motherboard or in other manners as appropriate. Processor 510 can process instructions for execution within the computing device 500, including instructions stored in the memory 520 or on the storage 530 to display graphical information for a graphical user interface (GUI) on an external input / output device, such as display 580 coupled to high-speed interface 540. In other implementations, multiple processors and / or multiple buses can be used, as appropriate, along with multiple memories and types of memory. Also, multiple computing devices 500 can be connected, with each device providing portions of the necessary operations (e.g., as a server bank, a group of blade servers, or a multi-processor system).
[0125] Memory 520 stores information non-transitorily within computing device 500. Memory 520 can be a computer-readable medium, a volatile memory unit(s) or non-volatile memory unit(s). The non-transitory memory 520 can be a physical device that is temporarily or permanently stores programs (e.g., sequences of instructions) or data (e.g., program state information) for use by or singularly computing device 500. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as, for example, in a
[0126] Storage 530 can provide mass storage for computing device 500. In some implementations, storage 530 is a computer-readable medium. In various implementations, storage 530 can be a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid state memory device, or an array of devices, including devices in a storage area network or other configurations. In additional implementations, a computer program product is tangibly embodied in an information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer- or machine-readable medium, such as the memory 520, the storage device 530, or memory on processor 510.
[0127] The high-speed controller 540 manages bandwidth-intensive operations for the computing device 500, while the low-speed controller 560 manages lower bandwidth-intensive operations. Such allocation of functions is exemplary only. In some embodiments, the high-speed controller 540 is coupled to memory 520, display 580 (e.g., through a graphics processor or accelerator), and to high-speed expansion ports 550, which can accept various expansion cards (not shown). In some embodiments, the low-speed controller 560 is coupled to storage device 530 and low-speed expansion port 590. The low-speed expansion port 590 can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet) and can be coupled to one or more input / output devices, such as keyboard, pointing devices, scanners, or networking devices, e.g., switches or routers, through a network adapter (not shown).
[0128] The computing device 500 can be implemented in a number of different forms, as shown in the figure. For example, it can be implemented as a standard server 500a or multiple times in a group of such servers 500a, as a laptop computer 500b, or as part of a rack server system 500c.
[0129] Various implementations of the systems and techniques described here can be realized in digital electronic and / or optical circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0130] These computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer readable medium, apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0131] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for performing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0132] To provide for interaction with a user, one or more aspects of the disclosure can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube), LCD (liquid crystal display), or touch screen for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user's client device in response to requests received from the web browser.
[0133] A variety of implementations have been described. However, it should be understood that various modifications can be made without departing from the spirit and scope of the disclosure. Accordingly, other implementations are within the scope of the following claims.
Claims
1. A computer-implemented method, characterized by, When executed by data processing hardware, the operations for causing the data processing hardware to perform include: receiving non-quantized transform coefficients corresponding to a scaled video input signal, the scaled video input signal including a plurality of spatial layers; and for each spatial layer of the plurality of spatial layers: determining an allocation factor based on frame samples from the scaled video input signal, the allocation factor corresponding to a variance estimate of the received non-quantized transform coefficients; allocating a bit rate based on the allocation factor and the frame samples; and after allocating the bit rate based on the allocation factor and the frame samples, dynamically adjusting the allocation factor by: iteratively identifying a set of frame samples from the scaled video input signal; and adjusting the allocation factor based on each identified set of samples from the scaled video input signal.
2. The method of claim 1, wherein, Determining an allocation factor includes using a high-rate approximation based squared error for scalar quantization.
3. The method of claim 1, wherein, Allocating a bit rate includes determining the bit rate using rate-distortion optimized quantization.
4. The method of claim 1, wherein, The allocated bit rate minimizes an amount of distortion subject to a bit rate constraint in the spatial layer.
5. The method of claim 1, wherein, Each spatial layer of the plurality of spatial layers originates from a same video source device.
6. The method of claim 1, wherein, Determining an allocation factor includes determining a weighted sum of average distortion rates for each spatial layer of the plurality of spatial layers.
7. The method of claim 1, wherein, Dynamically adjusting the allocation factor includes using an exponentially moving average to modify the allocation factor for each frame from the scaled video input signal.
8. The method of claim 1, wherein, The variance estimate of the received non-quantized transform coefficients includes a mean value for each transform block of a plurality of transform blocks of the frame samples.
9. A system for allocating bit rates, characterized by The system includes: data processing hardware; and memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations including: receiving non-quantized transform coefficients corresponding to a scaled video input signal, the scaled video input signal including a plurality of spatial layers; and for each spatial layer of the plurality of spatial layers: determining an allocation factor based on frame samples from the scaled video input signal, the allocation factor corresponding to a variance estimate of the received non-quantized transform coefficients; allocating a bit rate based on the allocation factor and the frame samples; and after allocating the bit rate based on the allocation factor and the frame samples, dynamically adjusting the allocation factor by: iteratively identifying a set of frame samples from the scaled video input signal; and adjusting the allocation factor based on each identified set of samples from the scaled video input signal.
10. The system of claim 9, wherein, Determining an allocation factor includes using a high-rate approximation based squared error for scalar quantization.
11. The system of claim 9, wherein, Allocating a bit rate includes determining the bit rate using rate-distortion optimized quantization.
12. The system of claim 9, wherein, The allocated bit rate minimizes an amount of distortion subject to a bit rate constraint in the spatial layer.
13. The system of claim 9, wherein, Each spatial layer of the plurality of spatial layers originates from a same video source device.
14. The system of claim 9, wherein, Determining an allocation factor includes determining a weighted sum of average distortion rates for each spatial layer of the plurality of spatial layers. Dynamically adjusting the allocation factor includes using an exponentially moving average to modify the allocation factor for each frame from the scaled video input signal.
15. The system of claim 9, wherein, Dynamically adjusting the allocation factor includes modifying the allocation factor for each frame from the scaled video input signal using an exponential moving average.
16. The system of claim 9, wherein, wherein, The received variance estimate of the non-quantized transform coefficients includes a mean value for each transform block of a plurality of transform blocks of the frame samples.
Citation Information
Patent Citations
Video rate control based on transform-coefficients histogram
CN102948147A
System and method for improved fine granular scalable video using base layer coding information
CN1321399A
Techniques for managing output bandwidth for a conferencing server
US20080101410A1