System and method for RGB video coding enhancement
Adaptive residual color space conversion using YCgCo color space enhances video coding efficiency and image quality by reducing redundancy and preserving high-frequency information in screen content.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- INTERDIGITAL VC HOLDINGS INC
- Filing Date
- 2023-12-22
- Publication Date
- 2026-05-11
AI Technical Summary
Existing video compression methods fail to fully characterize the features of screen content, leading to reduced compression performance and image/video quality issues such as blurriness in curves and text.
Implement adaptive residual color space conversion using YCgCo color space for encoding and decoding video content, enabling flexible conversion between RGB and YCgCo color spaces based on rate-distortion metrics, and applying different interpolation filters for motion compensation.
Improves video coding efficiency by reducing redundancy and preserving high-frequency information, resulting in sharper curves and text in reconstructed images.
Smart Images

Figure 0007856624000020 
Figure 0007856624000021 
Figure 0007856624000022
Abstract
Description
Background Art
[0001] The present invention relates to systems and methods for RGB video coding enhancement.
[0002] This application claims priority based on U.S. Provisional Patent Application No. 61 / 953,185, filed on March 14, 2014, U.S. Provisional Patent Application No. 61 / 994,071, filed on May 15, 2014, and U.S. Provisional Patent Application No. 62 / 040,317, filed on August 21, 2014, each of which is entitled "RGB VIDEO CODING ENHANCEMENT" and is hereby incorporated by reference in its entirety.
[0003] Screen content sharing applications have become more popular as device and network capabilities have improved. Examples of popular screen content sharing applications include remote desktop applications, video conferencing applications, and mobile media presentation applications. Screen Contents can include numerous video and / or image elements having one or more (one or more) primary colors and / or sharp edges. Such image and video elements may include relatively sharp curves and / or text within such elements.
Summary of the Invention
Problems to be Solved by the Invention
[0004] Various video compression means and methods can be used to encode screen content and / or transmit such content to a receiver, but such methods and means cannot fully characterize the features of the screen content. Such lack of characterization may result in reduced compression performance in the reconstructed image or video content. In such implementations, the reconstructed image or video content may be adversely affected by image or video quality issues. For example, such curves and / or text may be blurry, indistinct, or in other states that make them difficult to recognize within the screen content. [Means for solving the problem]
[0005] Systems, methods, and devices for encoding and decoding video content are disclosed. In embodiments, the systems and methods can be implemented to perform adaptive residual color space conversion. A video bitstream can be received, and a first flag can be determined based on the video bitstream. A residual can also be generated based on the video bitstream. The residual can be converted from a first color space to a second color space in response to the first flag.
[0006] In an embodiment, determining a first flag may include receiving the first flag at the coding unit level. The first flag may be received only if a second flag at the coding unit level indicates that at least one residual having a non-zero value exists in the coding unit. Converting the residuals from a first color space to a second color space may be done by applying a color space conversion matrix. This color space conversion matrix may correspond to a YCgCo to RGB lossy conversion matrix that can be applied in lossy coding. In another embodiment, the color space conversion matrix may correspond to a YCgCo to RGB lossy conversion matrix that can be applied in lossless coding. Converting the residuals from a first color space to a second color space may include applying a scale factor matrix, in which case the color space conversion matrix is unnormalized, and each row of the scale factor matrix may include a scale factor corresponding to the norm of the corresponding row of the unnormalized color space conversion matrix. The color space conversion matrix may include at least one fixed-point precision coefficient. A second flag based on the video bitstream can be communicated at the sequence level, picture level, or slice level, and the second flag can indicate whether the process of converting residuals from the first color space to the second color space is enabled at the sequence level, picture level, or slice level, respectively.
[0007] In the embodiment, the residuals of the coding unit can be coded in a first color space. The best mode for coding such residuals can be determined based on the cost of coding the residuals in the available color spaces. A flag can be determined based on the determined best mode and can be included in the output bitstream. The above and other aspects of the disclosed invention are described below. [Effects of the Invention]
[0008] A system, method, and device for encoding and decoding video content are provided. [Brief explanation of the drawing]
[0009] [Figure 1] This is a block diagram illustrating an exemplary screen content sharing system according to an embodiment. [Figure 2] This block shows an exemplary video encoding system according to an embodiment. [Figure 3] An exemplary video decoding system according to an embodiment is shown in the block diagram. [Figure 4] This figure shows an exemplary prediction unit mode according to an embodiment. [Figure 5] This figure shows an exemplary color image according to the embodiment. [Figure 6] This figure shows an exemplary method for carrying out the disclosed embodiments of the present invention. [Figure 7] This figure shows another exemplary method for carrying out the disclosed embodiments of the present invention. [Figure 8] This block shows an exemplary video encoding system according to an embodiment. [Figure 9] An exemplary video decoding system according to an embodiment is shown in the block diagram. [Figure 10] This is a block diagram illustrating an exemplary subdivision of a prediction unit into a conversion unit according to an embodiment. [Figure 11A] This is a system diagram of an exemplary communication system that can implement the present invention. [Figure 11B] Figure 11A is a system diagram of an exemplary wireless transceiver unit (WTRU) that can be used in the communication system shown. [Figure 11C] Figure 11A is a system diagram of an exemplary wireless access network and an exemplary core network that can be used within the communication system shown. [Figure 11D]Figure 11A is a system diagram of another exemplary radio access network and an exemplary core network that can be used within the communication system shown. [Figure 11E] Figure 11A is a system diagram of another exemplary radio access network and an exemplary core network that can be used within the communication system shown. [Modes for carrying out the invention]
[0010] The following provides a detailed explanation with reference to various diagrams illustrating the concepts. While this explanation provides detailed examples of possible implementations, it should be noted that the details are intended to be illustrative only and are not intended to limit the scope of this application.
[0011] Screen content compression methods are becoming increasingly important as more people share device content for use, for example, in media presentation and remote desktop applications. The display capabilities of mobile devices have increased to high-definition or ultra-high-definition resolutions in some embodiments. Video coding tools, such as block coding modes and conversions, may not be optimized for higher-definition screen content coding. Such tools can increase the bandwidth available for transmitting screen content in content sharing applications.
[0012] Figure 1 shows a block diagram of an exemplary screen content sharing system 191. The system 191 may include a receiver 192, a decoder 194, and a display 198 (sometimes called a “renderer”). The receiver 192 can provide an input bitstream 193 to the decoder 194, which can decode the bitstream to produce a decoded picture 195, which can be provided to one or more display picture buffers 196. The display picture buffers 196 can provide the decoded picture 197 to the display 198 for presentation on the device’s display.
[0013] Figure 2 shows a block diagram of a block-based single-layer video encoder 200, which can be implemented, for example, to provide a bitstream to the receiver 192 of system 191 in Figure 1. As shown in Figure 2, the encoder 200 predicts the input video signal 201 using techniques such as spatial prediction (sometimes called “intra prediction”) and temporal prediction (sometimes called “inter prediction”), in an effort to improve compression efficiency. The encoder 200 may include other encoder control logic 240 that can determine mode determination and / or the form of prediction. Such determination may be at least partially based on criteria such as rate-based criteria, distortion-based criteria, and / or a combination thereof. The encoder 200 may provide one or more prediction blocks 206 to element 204, which may generate a prediction residual 205 (which may be a difference signal between the input signal and the predicted signal) and provide it to a transform element 210. The encoder 200 can transform the prediction residual 205 in the transformation element 210 and quantize the prediction residual 205 in the quantization element 215. The quantized residual can be provided to the entropy coding element 230 as a residual coefficient block 222, along with mode information (e.g., intra-prediction or inter-prediction) and prediction information (motion vector, reference picture index, intra-prediction mode, etc.). The entropy coding element 230 can compress the quantized residual and provide it along with the output video bitstream 235. The entropy coding element 230 can also, or instead, use the coding mode, prediction mode, and / or motion information 208 when generating the output video bitstream 235.
[0014] In an embodiment, the encoder 200 can also, or instead, apply inverse quantization in the inverse quantization element 225 to the residual coefficient block 222, and also apply an inverse transform in the inverse transform element 220 to generate a reconstructed residual that can be added back to the prediction signal 206 in the element 209, thereby generating a reconstructed video signal. The resulting reconstructed video signal can, in some embodiments, be processed using a loop filter process implemented in the loop filter element 250 (e.g., by using one or more of a deblocking filter, a sample adaptive offset, and / or an adaptive loop filter). The resulting reconstructed video signal can, in some embodiments, be stored in the reference picture store 270 in the form of the reconstructed block 255, whereupon it can be used, for example, by the motion prediction (estimation and compensation) element 280 and / or the spatial prediction element 260 to predict future video signals. Note that in some embodiments, the resulting reconstructed video signal generated by the element 209 can be provided to the spatial prediction element 260 without being processed by an element such as the loop filter element 250.
[0015] Figure 3 shows a block diagram of a block-based single-layer decoder 300 that can receive a video bitstream 335, which can be a bitstream such as the bitstream 235 that can be generated by the encoder 200 in Figure 2. The decoder 300 can reconstruct the bitstream 335 for display on a device. The decoder 300 can analyze the bitstream 335 in an entropy decoder element 330 to generate residual coefficients 326. The residual coefficients 326 can be dequantized in a dequantization element 325 and / or inverse transformed in an inverse transform element 320 to obtain a reconstructed residual that can be provided to element 309. To obtain a prediction signal, coding mode, prediction mode, and / or motion information 327 can be used, and in some embodiments, one or both of spatial prediction information provided by a spatial prediction element 360 and / or temporal prediction information provided by a temporal prediction element 390 are used. Such a prediction signal can be provided as a prediction block 329. The predicted signal and the reconstructed residual can be added in element 309 to generate a reconstructed video signal, which can be provided to the loop filter element 350 for loop filtering and can also be stored in the reference picture store 370 for use when displaying the picture and / or decoding the video signal. Note that the prediction mode 328 can be provided to element 309 by the entropy decoding element 330 for use when generating a reconstructed video signal that can be provided to the loop filter element 350 for loop filtering.
[0016] Video coding (encoding) standards such as High Efficiency Video Coding (HEVC) can reduce the transmission bandwidth and / or storage. In some embodiments, the HEVC implementation can operate as a block-based hybrid video coding (encoding), in which case the encoder and decoder implemented generally operate as described herein with reference to FIGS. 2 and 3. HEVC can enable the use of larger video blocks and can use quadtree partitioning to convey block coding (encoding) information. In such embodiments, a picture or a slice of a picture can be divided into coding tree blocks (CTBs) each having the same size (e.g., 64×64). Each CTB can be divided into coding units (CUs) using quadtree partitioning, and each CU can be further divided into a prediction unit (PU) and a transform unit (TU), each of which can also be divided using quadtree partitioning.
[0017] In embodiments, for each encoded CU, the associated PU can be split using one of eight exemplary splitting modes, examples of which are shown in Figure 4 as modes 410, 420, 430, 440, 460, 470, 480, and 490. In some embodiments, time prediction can be applied to reconstruct the encoded PU. Linear filters can be applied to obtain pixel values at fractional positions. Interpolation filters used in some such embodiments may have seven or eight taps for lumens and / or four taps for chromens. Content-based deblocking filters can be used so that different deblocking filter behaviors can be applied at each of the TU and PU boundaries depending on a number of factors, including one or more of the following: differences in encoding modes, differences in motion, differences in reference pictures, differences in pixel values, etc. In embodiments of entropy coding, context-adaptive binary arithmetic coding (CABAC) can be used for one or more block-level syntax elements. In some embodiments, CABAC may not be used for high-level parameters. The bins that can be used in CABAC coding may include regular bins with context-based coding and bins with bypass coding that does not use context.
[0018] Screen content video can be captured in red-green-blue (RGB) format. RGB signals may contain redundancy between the three color components. While such redundancy can be less efficient in embodiments implementing video compression, the use of the RGB color space may be chosen for applications where high fidelity is desired for the decoded screen content video, as color space conversion (e.g., from RGB coding to YCbCr coding) can introduce losses into the original video signal due to rounding and clipping operations sometimes used to convert color components between different spaces. In some embodiments, video compression efficiency can be improved by leveraging the correlation between the three color components of the color space. For example, a component-predictive coding tool can use the residual of the G component to predict the residuals of the B component and / or R component. In a YCbCr embodiment, the residual of the Y component can be used to predict the residuals of the Cb component and / or Cr component.
[0019] In some embodiments, motion compensation prediction techniques can be used to leverage redundancy between temporally adjacent pictures. Such embodiments can support motion vectors with a precision of one-quarter pixel for the Y component and one-eighth pixel for the Cb and / or Cr components. In some embodiments, fractional sample interpolation can be used, which may include a separable 8-tap filter for half-pixel positions and a 7-tap filter for quarter-pixel positions. Table 1 below shows exemplary filter coefficients for fractional interpolation of the Y component. Fractional interpolation of the Cb and / or Cr components can be performed using similar filter coefficients, except that in the 4:2:0 video format implementation, the motion vector may have a precision of one-eighth of a pixel. In the 4:2:0 video format implementation, the Cb and Cr components may contain less information than the Y component, and the 4-tap interpolation filter can reduce the complexity of fractional interpolation filtering without sacrificing the efficiency that can be gained in motion compensation prediction for the Cb and Cr components compared to the 8-tap interpolation filter implementation. Table 2 below shows exemplary filter coefficients that can be used for fractional interpolation of the Cb and Cr components.
[0020] [Table 1]
[0021] [Table 2]
[0022] In some embodiments, a video signal initially captured in RGB color format can be encoded in the RGB domain, for example, if high fidelity is desired for the decoded video signal. Inter-component prediction tools can improve the efficiency of encoding RGB signals. In some embodiments, redundancy that may exist between the three color components may not be fully utilized, because in some such embodiments, the G component can be used to predict the B and / or R components, but the correlation between the B and R components may not be used. Decorrelation of such color components can improve the encoding performance of RGB video coding.
[0023] Fractional interpolation filters can be used to encode RGB video signals. An interpolation filter design that focuses on encoding YCbCr video signals in a 4:2:0 color format may not be suitable for encoding RGB video signals. For example, the B and R components of RGB video may represent richer color information and possess higher frequency characteristics than the chrominance components of the converted color space, such as the Cb and Cr components in the YCbCr color space. A 4-tap fractional filter that can be used for the Cb and / or Cr components may not be sufficiently accurate for motion compensation prediction of the B and R components when encoding RGB video. In a lossless coding embodiment, a reference picture may be used for motion compensation prediction, which may be mathematically identical to the original picture associated with such a reference picture. In such an embodiment, such a reference picture may contain more edges (i.e., high-frequency signals) compared to a lossy coding embodiment using the same original picture, and the high-frequency information in such a reference picture may be reduced and / or distorted due to the quantization process. In such embodiments, shorter-tap interpolation filters can be used for the B and R components, which can preserve higher-frequency information within the original picture.
[0024] In some embodiments, a residual color conversion method can be used to adaptively select either the RGB color space or the YCgCo color space for coding residual information associated with RGB video. Such a residual color space conversion method can be applied to either or both lossless and lossy coding without incurring excessive computational complexity overhead during the coding and / or decoding process. In another embodiment, interpolation filters can be adaptively selected for use in motion-compensated prediction of different color components. Such a method can allow for the flexibility to use different fractional interpolation filters at the sequence, picture, and / or CU levels, thereby improving the efficiency of motion-compensated predictive coding.
[0025] In some embodiments, residual coding can be performed in a color space different from the original color space to remove redundancy in the original color space. Coding in the YCbCr color space can provide a more compact representation of the original video signal than coding in the RGB color space (for example, inter-component correlation can be lower in the YCbCr color space than in the RGB color space), and the coding efficiency of YCbCr can be higher than that of RGB. Therefore, video coding of natural content (e.g., camera-captured video content) can be performed in the YCbCr color space instead of the RGB color space. In most cases, the source video can be captured in RGB format, and high fidelity of the reconstructed video may be desired.
[0026] Color space conversions are not always reversible, and the output color space may have the same dynamic range as that of the input color space. For example, when RGB video is converted to the ITU-R BT.709 YCbCr color space with the same bit depth, some loss may occur due to rounding and truncation operations that may be performed during such a color space conversion. YCgCo can be a color space that can have properties similar to the YCbCr color space, but the conversion process between RGB and YCgCo (i.e., RGB to YCgCo and YCgCo to RGB) can be computationally simpler than the conversion process between RGB and YCbCr, since only shift and addition operations can be used during such conversions. YCgCo can also support fully reversible conversions by increasing the bit depth of the intermediate operations by 1 (i.e., the color values derived after the inverse conversion can be numerically the same as the original color values). This aspect can be desirable because it can be applicable to both irreversible and reversible embodiments.
[0027] Due to the coding efficiency and ability to perform reversible transformations provided by the YCgCo color space, in embodiments, the residuals can be converted from RGB to YCgCo before residual coding. The decision of whether to apply the RGB to YCgCo transformation process can be performed adaptively at the sequence and / or slice and / or block level (e.g., CU level). For example, the decision can be based on whether applying the transformation provides an improvement to the rate-distortion (RD) metric (e.g., a weighted combination of rate and distortion). Figure 5 shows an exemplary image 510 which can be an RGB picture. Image 510 can be decomposed into three color components of YCgCo. In such embodiments, both reversible and reversible versions of the transformation matrix can be specified for reversible coding and reversible coding, respectively. When the residuals are coded in the RGB domain, the encoder can treat the G component as the Y component, and the B and R components as the Cb and Cr components, respectively. In this disclosure, the order G, B, R may be used to represent RGB video, rather than the order R, G, B. While embodiments described herein may be illustrated using examples where the conversion is performed from RGB to YCgCo, it should be noted that those skilled in the art will understand that conversions between RGB and other color spaces (e.g., YCbCr) can also be performed using the disclosed embodiments. All such embodiments are intended to be within the scope of this disclosure.
[0028] A reversible conversion from the GBR color space to the YCgCo color space can be performed using equations (1) and (2) shown below. These equations can be used for both reversible and reversible coding. Equation (1) shows a means, according to an embodiment, for performing a reversible conversion from the GBR color space to YCgCo.
[0029]
number
[0030] This can be done using shifts without multiplication or division, and the reason for this is... Co=R·B t = B + (Co >> 1) Cg = G·t Y = t + (Cg >> 1) Therefore.
[0031] In such embodiments, the inverse conversion from YCgCo to GBR can be performed using equation (2).
[0032]
number
[0033] This can be done using Shift, and the reason is... t = Y - (Cg >> 1) G = Cg + t B = t - (Co >> 1) R=Co+B Therefore.
[0034] In some embodiments, the irreversible transformation can be performed using equations (3) and (4) shown below. Such irreversible transformations can be used for irreversible coding, but in some embodiments they cannot be used for reversible coding. Equation (3) shows a means, according to an embodiment, for performing an irreversible transformation from the GBR color space to YCgCo.
[0035]
number
[0036] The inverse conversion from YCgCo to GBR can be performed using equation (4) according to the embodiment.
[0037]
number
[0038] As shown in equation (3), the forward color space transformation matrix that can be used for lossy coding may not be normalized. The magnitude and / or energy of the residual signal in the YCgCo domain may be reduced compared to that of the original residual in the RGB domain. Since the YCgCo residual coefficients may be overquantized by using the same quantization parameter (QP) that could be used in the RGB domain, this reduction of the residual signal in the YCgCo domain may impair the lossy coding performance of the YCgCo domain. In embodiments, a QP adjustment method can be used in which a delta QP can be added to the original QP value to compensate for the change in the magnitude of the YCgCo residual signal when a color space transformation can be applied. The same delta QP can be applied to both the Y component and the Cg and / or Co components. In embodiments that implement equation (3), different rows of the forward transformation matrix may not have the same norm. The same QP adjustment may not guarantee that both the Y component and the Cg and / or Co components have amplitude levels similar to those of the G component and the B and / or R components.
[0039] To ensure that the YCgCo residual signal converted from the RGB residual signal has an amplitude similar to that of the RGB residual signal, in one embodiment, a pair of scaled forward and inverse transformation matrices can be used to transform the residual signal between the RGB domain and the YCgCo domain. More specifically, the forward transformation matrix from the RGB domain to the YCgCo domain can be defined by equation (5).
[0040]
number
[0041] Here,
[0042]
number
[0043] This can represent matrix multiplication of elements of two matrices that can be in the same position, where a and b can be scaling factors to compensate for the norms of different rows in the original forward color space transformation matrix, such as those used in equation (3), which can be derived using equations (6) and (7).
[0044]
number
[0045]
number
[0046] In such embodiments, the reverse conversion from the YCgCo domain to the RGB domain can be performed using equation (8).
[0047]
number
[0048] In equations (5) and (8), the scaling factor can be a real number, which may require floating-point multiplication when converting color spaces between RGB and YCgCo. To reduce implementation complexity, in embodiments, the multiplication of the scaling factor can be approximated by a computationally efficient multiplication using an integer M performed by an N-bit right shift.
[0049] The disclosed color space conversion methods and systems can be enabled and / or disabled at the sequence, picture, or block (e.g., CU, TU) level. For example, in the embodiments, predictive residual color space conversion can be adaptively enabled and / or disabled at the coding unit level. The encoder can select the optimal color conversion space between GBR and YCgCo for each CU.
[0050] Figure 6 shows an exemplary method 600 for an RD optimization process using adaptive residual color transformation in an encoder as described herein. In block 605, the residual of the CU can be coded using the “best mode” of coding for its implementation (e.g., intra-prediction for intra-coding, motion vector and reference picture index for inter-coding), which can be a pre-configured coding mode, a coding mode previously determined to be the best among those available, or another pre-determined coding mode that has been determined to have the lowest or relatively lower RD cost at the time of performing at least the function of block 605. In block 610, a flag, which in this example is called “CU_YCgCo_residual_flag,” but can be called using any term or combination of terms, can be set to “false” (or to any other indicator that indicates false, zero, etc.) that the coding of the residual of the coding unit should not be performed using the YCgCo color space. In response to a flag evaluated in block 610 as false or equivalent, in block 615, the encoder performs residual coding in the GBR color space and the RD cost for such coding (in Figure 6, "RDCost") GBR It is called "cost calculation," but here again, any label or term can be used to refer to such costs.
[0051] In block 620, a decision can be made as to whether the RD cost for GBR color space coding is lower than the RD cost for best mode coding. If the RD cost for GBR color space coding is lower than the RD cost for best mode coding, in block 625, CU_YCgCo_residual_flag for best mode can be set to false or equivalent (or left set to false or equivalent), and the RD cost for best mode can be set to the RD cost for residual coding in the GBR color space. Method 600 can then proceed to block 630, where CU_YCgCo_residual_flag can be set to a true or equivalent indicator.
[0052] If, in block 620, it is determined that the RD cost for the GBR color space is greater than or equal to the RD cost for best-mode coding, the RD cost for best-mode coding can be left at the value it was set before the evaluation in block 620, and block 625 can be bypassed. Method 600 can proceed to block 630, where CU_YCgCo_residual_flag can be set to true or an equivalent indicator. Setting CU_YCgCo_residual_flag to true (or an equivalent indicator) in block 630 facilitates the coding of residuals of coding units using the YCgCo color space, and thus facilitates the evaluation of the RD cost of coding using the YCgCo color space compared to the RD cost of best-mode coding, as described below.
[0053] In block 635, the residuals of the coding unit can be coded using the YCgCo color space, and the RD cost of such coding can be determined (such cost is shown in Figure 6 as "RDCost"). YCgCoThis is called "costs," but again, any label or term can be used to refer to such costs.
[0054] In block 640, a decision can be made as to whether the RD cost for YCgCo color space coding is lower than the RD cost for best mode coding. If the RD cost for YCgCo color space coding is lower than the RD cost for best mode coding, in block 645, the CU_YCgCo_residual_flag for best mode can be set to true or equivalent (or left set to true or equivalent), and the RD cost for best mode can be set to the RD cost for residual coding in the YCgCo color space. Method 600 can end in block 650.
[0055] If, in block 640, it is determined that the RD cost for the YCgCo color space is higher than the RD cost for best-mode coding, the RD cost for best-mode coding can be left at the value it was set before the evaluation in block 640, and block 645 can be bypassed. Method 600 can terminate in block 650.
[0056] As those skilled in the art will understand, the disclosed embodiments, including Method 600 and any subset thereof, can enable a comparison of GBR and YCgCo color space codings and their respective RD costs, which can enable the selection of a color space coding having a lower RD cost.
[0057] Figure 7 shows another exemplary method 700 for an RD optimization process using adaptive residual color transformation in an encoder as described herein. In an embodiment, the encoder may attempt to use the YCgCo color space for residual coding if at least one of the reconstructed GBR residuals in the current coding unit is non-zero. If all of the reconstructed residuals are zero, it can indicate that the prediction in the GBR color space is sufficient and that the transformation to the YCgCo color space cannot further improve the efficiency of residual coding. In such an embodiment, the number of cases examined for RD optimization can be reduced and the coding process can be performed more efficiently. Such embodiments can be implemented in systems using large quantization parameters, such as a large quantization step size.
[0058] In block 705, the residual of a CU can be coded using the “best mode” of coding for its implementation (e.g., intra-prediction for intra-coding, motion vector and reference picture index for inter-coding), which can be a pre-configured coding mode, a coding mode previously determined to be the best available, or another pre-determined coding mode that has the lowest or relatively lower RD cost at the time of performing at least the function of block 705. In block 710, a flag called “CU_YCgCo_residual_flag” in this example can be set to “false” (or to any other arbitrary indicator indicating false, zero, etc.) to indicate that coding of the residual of a coding unit should not be performed using the YCgCo color space. Again, note that such a flag can be referred to using any term or combination of terms. In response to a flag evaluated in block 710 as false or equivalent, in block 715, the encoder performs residual coding in the GBR color space and the RD cost for such coding (in Figure 7, "RDCost") GBR It is called "cost calculation," but here again, any label or term can be used to refer to such costs.
[0059] In block 720, a decision can be made as to whether the RD cost for GBR color space coding is lower than the RD cost for best mode coding. If the RD cost for GBR color space coding is lower than the RD cost for best mode coding, in block 725, the CU_YCgCo_residual_flag for best mode can be set to false or equivalent (or left set to false or equivalent), and the RD cost for best mode can be set to the RD cost for residual coding in the GBR color space.
[0060] If, in block 720, it is determined that the RD cost for the GBR color space is greater than or equal to the RD cost for best-mode coding, the RD cost for best-mode coding can be left at the value it was set before the evaluation of block 720, and block 725 can be bypassed.
[0061] In block 730, a decision can be made as to whether at least one of the reconstructed GBR coefficients is non-zero (i.e., whether all reconstructed GBR coefficients are equal to zero). If at least one non-zero reconstructed GBR coefficient exists, in block 735, CU_YCgCo_residual_flag can be set to true or an equivalent indicator. Setting CU_YCgCo_residual_flag to true (or an equivalent indicator) in block 735 facilitates the coding of residuals of coding units using the YCgCo color space, and therefore facilitates the evaluation of the RD cost of coding using the YCgCo color space compared to the RD cost of best-mode coding, as described below.
[0062] If at least one reconstructed GBR coefficient is non-zero, then in block 740, the residuals of the coding unit can be coded using the YCgCo color space, and the RD cost of such coding can be determined (such cost is shown in Figure 7 as "RDCost"). YCgCo This is called "costs," but again, any label or term can be used to refer to such costs.
[0063] In block 745, a decision can be made as to whether the RD cost for YCgCo color space coding is lower than the RD cost for best mode coding. If the RD cost for YCgCo color space coding is lower than the RD cost for best mode coding, in block 750, the CU_YCgCo_residual_flag for best mode can be set to true or equivalent (or left set to true or equivalent), and the RD cost for best mode can be set to the RD cost for residual coding in the YCgCo color space. Method 700 can end in block 755.
[0064] If, in block 745, it is determined that the RD cost for the YCgCo color space is greater than or equal to the RD cost for best-mode coding, the RD cost for best-mode coding can remain at the value it was set before the evaluation in block 745, and block 750 can be bypassed. Method 700 can terminate in block 755.
[0065] As those skilled in the art will understand, the disclosed embodiments, including Method 700 and any subset thereof, can enable a comparison of GBR and YCgCo color space coding and their respective RD costs, which can enable the selection of a color space coding having a lower RD cost. Method 700 in Figure 7 can provide a more efficient means of determining an appropriate setting for a flag such as the exemplary CU_YCgCo_residual_coding_flag described herein, while Method 600 in Figure 6 can provide a more complete means of determining an appropriate setting for a flag such as the exemplary CU_YCgCo_residual_coding_flag described herein. Any embodiment, or any variation, subset, or implementation using any one or more aspects thereof, all of which are intended to be within the scope of this disclosure, and the values of such flags can be transmitted in an encoded bitstream, such as those described with respect to Figure 2 and with respect to any other encoder described herein. Figure 8 shows a block diagram of a block-based single-layer video encoder 800, which can be implemented according to an embodiment to provide a bitstream to a receiver 192 of system 191 in Figure 1, for example. As shown in Figure 8, an encoder such as encoder 800 predicts an input video signal 801 using techniques such as spatial prediction (sometimes called “intra prediction”) and temporal prediction (sometimes called “inter prediction”). Encoder 800 may include other encoder control logic 840 that can determine mode determination and / or the form of prediction. Such determination may be at least partially based on criteria such as rate-based criteria, distortion-based criteria, and / or a combination thereof. Encoder 800 may provide one or more prediction blocks 806 to an adder element 804, which may generate a prediction residual 805 (which may be a difference signal between the input signal and the prediction signal) and provide it to a transform element 810. The encoder 800 can transform the prediction residual 805 in the transformation element 810 and quantize the prediction residual 805 in the quantization element 815. The quantized residual can be provided to the entropy coding element 830 as a residual coefficient block 822, along with mode information (e.g., intra-prediction or inter-prediction) and prediction information (motion vector, reference picture index, intra-prediction mode, etc.). The entropy coding element 830 can compress the quantized residual and provide it along with the output video bitstream 835. The entropy coding element 830 may also, or alternatively, use the coding mode, prediction mode, and / or motion information 808 when generating the output video bitstream 835.
[0066] In an embodiment, the encoder 800 may, in addition to or instead, generate a reconstructed video signal by applying inverse quantization to the residual coefficient block 822 in the inverse quantization element 825 and applying inverse transform in the inverse transform element 820 to generate a reconstructed residual that can be added back to the predicted signal 806 in the adder element 809. In an embodiment, the inverse residual transform of such a reconstructed residual may be generated by the inverse residual transform element 827 and provided to the adder element 809. In such an embodiment, the residual coding element 826 may provide a control switch 817 via the control signal 823 with a display of the value of CU_YCgCo_residual_coding_flag 891 (or CU_YCgCo_residual_flag, or any other one or more flags or indicators that provide the performance or display of the functions described herein with respect to the described CU_YCgCo_residual_coding_flag and / or the described CU_YCgCo_residual_flag). In response to receiving a control signal 823 indicating the reception of such a flag, the control switch 817 can direct the reconstructed residual to the inverse residual transform element 827 for the generation of the inverse residual transform of the reconstructed residual. The values of flag 891 and / or control signal 823 can indicate the encoder's decision on whether to apply a residual transform process that includes both a forward residual transform 824 and a reverse residual transform 827. In some embodiments, the encoder evaluates the cost and benefit of applying or not applying the residual transform process, so the control signal 823 can take different values. For example, the encoder may evaluate the rate-distortion cost of applying the residual transform process to a portion of the video signal.
[0067] The resulting reconstructed video signal generated by the adder 809 can, in some embodiments, be processed using a loop filtering process performed in the loop filter element 850 (for example, by using one or more of a deblocking filter, a sample-adaptive offset, and / or an adaptive loop filter). The resulting reconstructed video signal can, in some embodiments, be stored in the reference picture store 870 in the form of a reconstructed block 855, in which case it can be used, for example, by a motion prediction (estimation and compensation) element 880 and / or a spatial prediction element 860 to predict future video signals. It should be noted that, in some embodiments, the resulting reconstructed video signal generated by the adder element 809 can be provided to the spatial prediction element 860 without being processed by elements such as the loop filter element 850.
[0068] As shown in Figure 8, in the embodiment, an encoder such as encoder 800 can determine the value of CU_YCgCo_residual_coding_flag891 (or CU_YCgCo_residual_flag, or any other one or more flags or indicators that perform or indicate the functions described herein relating to the described CU_YCgCo_residual_coding_flag and / or CU_YCgCo_residual_flag) in the color space determination for residual coding element 826. The color space determination for residual coding element 826 can provide indication of such flag to control switch 807 via control signal 823. When control switch 807 receives control signal 823 indicating the receipt of such flag, it can respond by directing the predicted residual 805 to residual conversion element 824 so that the RGB to YCgCo conversion process can be adaptively applied to the predicted residual 805 in residual conversion element 824. In some embodiments, this transformation process may be performed before transformation and quantization are performed on the coding units processed by the transformation element 810 and the quantization element 815. In some embodiments, this transformation process may be performed, in addition or alternatively, before inverse transformation and inverse quantization are performed on the coding units processed by the inverse transformation element 820 and the inverse quantization element 825. In some embodiments, CU_YCgCo_residual_coding_flag891 may be provided to the entropy coding element 830 for inclusion in the bitstream, in addition or alternatively.
[0069] Figure 9 shows a block diagram of a block-based single-layer decoder 900 that can receive a video bitstream 935, which can be a bitstream such as bitstream 835 that can be generated by the encoder 800 of Figure 8. The decoder 900 can reconstruct the bitstream 935 for display on a device. The decoder 900 can analyze the bitstream 935 in an entropy decoder element 930 to generate residual coefficients 926. The residual coefficients 926 can be dequantized in a dequantization element 925 and / or inverse transformed in an inverse transform element 920 to obtain a reconstructed residual that can be provided to an adder element 909. To obtain a prediction signal, coding mode, prediction mode, and / or motion information 927 can be used, and in some embodiments, one or both of spatial prediction information provided by a spatial prediction element 960 and / or temporal prediction information provided by a temporal prediction element 990 are used. Such a prediction signal can be provided as a prediction block 929. The predicted signal and the reconstructed residual can be added in the adder element 909 to generate a reconstructed video signal, which can be provided to the loop filter element 950 for loop filtering and can also be stored in the reference picture store 970 for use when displaying the picture and / or decoding the video signal. Note that the prediction mode 928 can be provided to the adder element 909 by the entropy decoding element 930 for use when generating a reconstructed video signal that can be provided to the loop filter element 950 for loop filtering.
[0070] In the embodiment, the decoder 900 can decode the bitstream 935 in the entropy decoding element 930 and determine CU_YCgCo_residual_coding_flag991 (or CU_YCgCo_residual_flag, or any other one or more flags or indicators that provide the performance or indication of the functions described herein relating to the described CU_YCgCo_residual_coding_flag and / or CU_YCgCo_residual_flag) which could have been encoded into the bitstream 935 by an encoder such as the encoder 800 in Figure 8. The value of CU_YCgCo_residual_coding_flag991 can be used to determine whether the YCgCo to RGB inverse conversion process can be performed in the residual inverse conversion element 999 on the reconstructed residuals generated by the inverse conversion element 920 and provided to the adder element 909. In this embodiment, a control signal indicating the flag 991 or its reception can be provided to the control switch 917, which in turn can direct the reconstructed residuals to the inverse residual transform element 999 in order to generate the inverse residual transform of the reconstructed residuals.
[0071] In some embodiments, the complexity of the video coding system can be reduced by performing adaptive color space conversion on the predicted residual, rather than as part of motion-compensated prediction or intra-prediction, because such embodiments may not require the encoder and / or decoder to store the predicted signals in two different color spaces.
[0072] To improve residual coding efficiency, the residual block can be divided into multiple square transform units to perform transform coding of the predicted residual, and possible TU sizes can be 4×4, 8×8, 16×16, and / or 32×32. Figure 10 shows an exemplary division of a PU into TUs, where PU1010 in the lower left can represent an embodiment in which the TU size can be equal to the PU size, and PU1020, 1030, and 1040 can represent embodiments in which each respective exemplary PU can be divided into multiple TUs.
[0073] In some embodiments, the color space conversion of the predicted residual can be adaptively enabled and / or disabled at the TU level. Such embodiments can provide finer granularity for switching between different color spaces compared to enabling and / or disabling adaptive color conversion at the CU level. Such embodiments can improve the coding gain that adaptive color space conversion can achieve.
[0074] Referring again to the exemplary encoder 800 in Figure 8, an encoder such as the exemplary encoder 800 can test each coding mode (e.g., intra-coding mode, inter-coding mode, intra-block copy mode) twice, once with a color space conversion and once without, in order to select a color space for residual coding of the CU. In some embodiments, various “fast” or more efficient coding logics, such as those described herein, can be used to improve the efficiency of such coding complexity.
[0075] In some embodiments, YCgCo can provide a more compact representation of the original color signals than RGB, so the RD cost with color space conversion enabled can be determined and compared to the RD cost with color space conversion disabled. In some such embodiments, the calculation of the RD cost with color space conversion disabled can be performed if there is at least one non-zero coefficient when color space conversion is enabled.
[0076] To reduce the number of coding modes to be tested, in some embodiments, the same coding mode can be used for both the RGB and YCbCr color spaces. In intra-mode, selected lumane and chromane-intra predictions can be shared between the RGB and YCgCo spaces. In inter-mode, selected motion vectors, reference pictures, and motion vector predictors can be shared between the RGB and YCgCo color spaces. In intra-block copy mode, selected block vectors and block vector predictors can be shared between the RGB and YCgCo color spaces. To further reduce coding complexity, in some embodiments, TU partitioning can be shared between the RGB and YCgCo color spaces.
[0077] Since correlations can exist between the three color components (Y, Cg, and Co in the YCgCo domain, and G, B, and R in the RGB domain), in some embodiments, the same intra-prediction direction can be selected for all three color components. In each of the two color spaces, the same intra-prediction mode can be used for all three color components.
[0078] Since correlations can exist between CUs within the same domain, one CU may select the same color space (e.g., RGB or YCgCo) as its parent CU to encode its residual signal. Alternatively, a child CU may derive a color space from information associated with its parent, such as the selected color space and / or the RD cost of each color space. In some embodiments, coding complexity can be reduced by not checking the RD cost of the residual coding in the RGB domain for one CU if the residual of the parent CU is coded in the YCgCo domain. In addition, or instead, the check for the RD cost of the residual coding in the YCgCo domain can be skipped if the residual of the child CU's parent CU is coded in the RGB domain. In some embodiments, the RD costs of the child CU's parent CU in two color spaces can be used for the child CU if the two color spaces are tested in the parent CU's coding. If the parent CU of a child CU selects the YCgCo color space, and the RD cost of YCgCo is less than that of RGB, then for the child CU, the RGB color space can be skipped, and vice versa.
[0079] Many prediction modes, including many intra-angle prediction modes, one or more DC modes, and / or one or more planar prediction modes, can be supported by several embodiments. Testing residual coding with color space conversion for all such intra-prediction modes may increase the complexity of the encoder. In embodiments, instead of calculating the full RD cost for all supported intra-prediction modes, a subset of N intra-prediction candidates can be selected from the supported modes without considering the bits of residual coding. The N selected intra-prediction candidates can be tested in the converted color space by applying residual coding and then calculating the RD cost. The best mode with the lowest RD cost among the supported modes can be selected as the intra-prediction mode in the converted color space.
[0080] As referred to herein, the disclosed color space conversion systems and methods can be enabled and / or disabled at the sequence level, and / or at the picture and / or block level. In the exemplary embodiments shown in Table 3 below, syntax elements (examples of which are highlighted in bold in Table 3, but which can take any form, label, term, or combination thereof, all of which are intended to be within the scope of this disclosure) can be used within a sequence parameter set (SPS) to indicate whether the residual color space conversion coding tool is enabled. In some such embodiments, the color space conversion is applied to video content having the same resolution for the lumen and chroma components, so the disclosed adaptive color space conversion systems and methods can be enabled for the "444" chroma format. In such embodiments, the color space conversion to the 444 chroma format may be constrained at a relatively high level. In such embodiments, if a non-444 color format can be used, bitstream conformance constraints can be applied to enforce the disabling of the color space conversion.
[0081] [Table 3]
[0082] In an embodiment, the exemplary syntax element "sps_residual_csc_flag" may, if equal to 1, indicate that the residual color space conversion coding tool can be enabled. The exemplary syntax element sps_residual_csc_flag may, if equal to 0, indicate that the residual color space conversion can be disabled and that the flag CU_YCgCo_residual_flag at the CU level is presumed to be 0. In such an embodiment, if the ChromaArrayType syntax element is not equal to 3, the value of the exemplary sps_residual_csc_flag syntax element (or its equivalent) may be equal to 0 in order to maintain bitstream compatibility.
[0083] In another embodiment, as shown in Table 4 below, an exemplary sps_residual_csc_flag syntax element (the example is highlighted in bold in Table 4, but it can take any form, label, term, or combination thereof, all of which are intended to be within the scope of this disclosure) can be transmitted depending on the value of the ChromaArrayType syntax element. In such an embodiment, if the input video is in 444 color format (i.e., ChromaArrayType is equal to 3, for example, "ChromaArrayType==3" in the table), the exemplary sps_residual_csc_flag syntax element can be transmitted to indicate whether color space conversion is enabled. If such input video is not in 44 color format (i.e., ChromaArrayType is not equal to 3), the exemplary sps_residual_csc_flag syntax element may not be transmitted and can be set to equal to 0.
[0084] [Table 4]
[0085] When the residual color space conversion coding tool is enabled, in the embodiment, additional flags may be added at the CU level and / or TU level, as described herein, to enable color space conversion between the GBR color space and the YCgCo color space.
[0086] In embodiments where examples are shown in Table 5 below, the exemplary coding unit syntax element “cu_ycgco_residue_flag” (the example is highlighted in bold in Table 5, but it can take any form, label, term, or combination thereof, all of which are intended to be within the scope of this disclosure) can, when equal to 1, indicate that the residual of the coding unit can be coded and / or decoded in the YCgCo color space. In such embodiments, the cu_ycgco_residue_flag syntax element or its equivalent can, when equal to 0, indicate that the residual of the coding unit can be coded in the GBR color space.
[0087] [Table 5]
[0088] In another embodiment, an example of which is shown in Table 6 below, the exemplary conversion unit syntax element “tu_ycgco_residue_flag” (the example is highlighted in bold in Table 6, but it can take any form, label, term, or combination thereof, all of which are intended to be within the scope of this disclosure) can, if equal to 1, indicate that the residual of the conversion unit can be encoded and / or decoded in the YCgCo color space. In such embodiments, the tu_ycgco_residue_flag syntax element or its equivalent can, if equal to 0, indicate that the residual of the conversion unit can be encoded in the GBR color space.
[0089] [Table 6]
[0090] Some interpolation filters may be less efficient when interpolating fractional pixels for motion compensation prediction, which can be used in some embodiments of screen content coding. For example, a 4-tap filter may not be accurate when interpolating the B and R components at fractional positions when encoding RGB video. In embodiments of lossless coding, an 8-tap lumen filter may not be the most efficient means of preserving useful high-frequency texture information contained within the original lumen components. In embodiments, separate representations of the interpolation filter can be used for different color components.
[0091] In one such embodiment, one or more default interpolation filters (e.g., a set of 8-tap filters, a set of 4-tap filters) can be used as candidate filters for the fractional pixel interpolation process. In another embodiment, a set of interpolation filters different from the default interpolation filters can be explicitly communicated in the bitstream. To enable adaptive filter selection for different color components, it is possible to communicate syntax elements that specify the interpolation filter to be selected for each color component. The disclosed filter selection systems and methods can be used at various coding levels, such as the sequence level, picture and / or slice level, and the CU level. The choice of the operational coding level can be made based on the coding efficiency and / or computation and / or operational complexity of the available implementations.
[0092] In embodiments where the default interpolation filter is used, a flag can be used to indicate whether a set of 8-tap filters or a set of 4-tap filters can be used for fractional pixel interpolation of color components. One such flag may indicate a filter selection for the Y component (or the G component in embodiments of the RGB color space), and another such flag may be used for the Cb and Cr components (or the B and R components in embodiments of the RGB color space). The table below provides examples of such flags that can be communicated at the sequence level, picture and / or slice level, as well as the CU level.
[0093] Table 7 below illustrates embodiments in which such flags are propagated to enable the selection of a default interpolation filter at the sequence level. The disclosed syntax can be applied to any parameter set, including video parameter sets (VPS), sequence parameter sets (SPS), and picture parameter sets (PPS). Table 7 illustrates embodiments in which exemplary syntax elements can be propagated in an SPS.
[0094] [Table 7]
[0095] In such embodiments, the exemplary syntax element “sps_luma_use_default_filter_flag” (the example is highlighted in bold in Table 7, but it can take any form, label, term, or combination thereof, all of which are intended to be within the scope of this disclosure) can, if equal to 1, indicate that the luma components of all pictures associated with the current sequence parameter set can use the same set of luma interpolation filters (e.g., the default luma filter set) for fractional pixel interpolation. In such embodiments, the exemplary syntax element sps_luma_use_default_filter_flag can, if equal to 0, indicate that the luma components of all pictures associated with the current sequence parameter set can use the same set of chroma interpolation filters (e.g., the default chroma filter set) for fractional pixel interpolation.
[0096] In such embodiments, the exemplary syntax element “sps_chroma_use_default_filter_flag” (the example is highlighted in bold in Table 7, but it can take any form, label, term, or combination thereof, all of which are intended to be within the scope of this disclosure) can, if equal to 1, indicate that the chroma components of all pictures associated with the current sequence parameter set can use the same set of chroma interpolation filters (e.g., the default chroma filter set) for fractional pixel interpolation. In such embodiments, the exemplary syntax element sps_chroma_use_default_filter_flag can, if equal to 0, indicate that the chroma components of all pictures associated with the current sequence parameter set can use the same set of chroma interpolation filters (e.g., the default chroma filter set) for fractional pixel interpolation.
[0097] In the embodiment, flags can be transmitted at the picture and / or slice level to facilitate the selection of interpolation filters at the picture and / or slice level (i.e., for a given color component, all CUs in the picture and / or slice can use the same interpolation filter). Table 8 below shows an example of transmission using syntax elements in the slice segment header according to the embodiment.
[0098] [Table 8]
[0099] In such embodiments, the exemplary syntax element “slice_luma_use_default_filter_flag” (the example is highlighted in bold in Table 8, but it can take any form, label, term, or combination thereof, all of which are intended to be within the scope of this disclosure) can, if equal to 1, indicate that the luma components of the current slice can use the same set of luma interpolation filters (e.g., the default luma filter set) for interpolation of fractional pixels. In such embodiments, the exemplary slice_luma_use_default_filter_flag syntax element can, if equal to 0, indicate that the luma components of the current slice can use the same set of chroma interpolation filters (e.g., the default chroma filter set) for interpolation of fractional pixels.
[0100] In such embodiments, the exemplary syntax element “slice_chroma_use_default_filter_flag” (the example is highlighted in bold in Table 8, but it can take any form, label, term, or combination thereof, all of which are intended to be within the scope of this disclosure) can, if equal to 1, indicate that the chroma components of the current slice can use the same set of chroma interpolation filters (e.g., the default chroma filter set) for interpolation of fractional pixels. In such embodiments, the exemplary syntax element slice_chroma_use_default_filter_flag can, if equal to 0, indicate that the chroma components of the current slice can use the same set of lumen interpolation filters (e.g., the default lumen filter set) for interpolation of fractional pixels.
[0101] In embodiments where a flag can be transmitted at the CU level to facilitate the selection of interpolation filters at the CU level, such a flag can be transmitted using a coding unit syntax as shown in Figure 9. In such embodiments, the color component of a CU can adaptively select one or more interpolation filters that can provide a predictive signal for that CU. Such a selection can represent a coding improvement that can be achieved by adaptive interpolation filter selection.
[0102] [Table 9]
[0103] In such embodiments, the exemplary syntax element “cu_use_default_filter_flag” (its example is highlighted in bold in Table 9, but it can take any form, label, term, or combination thereof, all of which are intended to be within the scope of this disclosure) can, when equal to 1, indicate that both the lumer and chroma can use the default interpolation filter for fractional pixel interpolation. In such embodiments, the exemplary cu_use_default_filter_flag syntax element or its equivalent can, when equal to 0, indicate that either the lumer or chroma component of the current CU can use a different set of interpolation filters for fractional pixel interpolation.
[0104] In such embodiments, the exemplary syntax element “cu_luma_use_default_filter_flag” (the example is highlighted in bold in Table 9, but it can take any form, label, term, or combination thereof, all of which are intended to be within the scope of this disclosure) can, if equal to 1, indicate that the luma component of the current cu will use the same set of luma interpolation filters (e.g., the default luma filter set) for interpolation of fractional pixels. In such embodiments, the exemplary syntax element cu_luma_use_default_filter_flag can, if equal to 0, indicate that the luma component of the current cu will be able to use the same set of chroma interpolation filters (e.g., the default chroma filter set) for interpolation of fractional pixels.
[0105] In such embodiments, the exemplary syntax element “cu_chroma_use_default_filter_flag” (the example is highlighted in bold in Table 9, but it can take any form, label, term, or combination thereof, all of which are intended to be within the scope of this disclosure) can, if equal to 1, indicate that the chroma components of the current cu can use the same set of chroma interpolation filters (e.g., the default chroma filter set) for interpolation of fractional pixels. In such embodiments, the exemplary syntax element cu_chroma_use_default_filter_flag can, if equal to 0, indicate that the chroma components of the current cu can use the same set of lumen interpolation filters (e.g., the default lumen filter set) for interpolation of fractional pixels.
[0106] In embodiments, the coefficients of candidate interpolation filters can be explicitly transmitted in the bitstream. Any interpolation filter that may differ from the default interpolation filter can be used for fractional pixel interpolation of a video sequence. In such embodiments, to facilitate the delivery of filter coefficients from the encoder to the decoder, the filter coefficients can be carried in the bitstream using an exemplary syntax element "interp_filter_coef_set()" (its example is highlighted in bold in Table 10, but it can take any form, label, term, or combination thereof, all of which are intended to be within the scope of this disclosure). Table 10 shows the syntax structure for transmitting such coefficients of candidate interpolation filters.
[0107] [Table 10]
[0108] In such embodiments, the exemplary syntax element “arbitrary_interp_filter_used_flag” (its example is highlighted in bold in Table 10, but it can take any form, label, term, or combination thereof, all of which are intended to be within the scope of this disclosure) can specify whether any interpolation filter exists. If the exemplary syntax element arbitrarary_interp_filter_used_flag is set to 1, any interpolation filter can be used for the interpolation process.
[0109] In such embodiments, the exemplary syntax element “num_interp_filter_set” (the example is highlighted in bold in Table 10, but it can take any form, label, term, or combination thereof, all of which are intended to be within the scope of this disclosure) or its equivalent can specify the number of interpolation filter sets presented in the bitstream.
[0110] Furthermore, in such embodiments, the exemplary syntax element “interp_filter_coeff_shifting” (the example is highlighted in bold in Table 10, but it can take any form, label, term, or combination thereof, all of which are intended to be within the scope of this disclosure) or its equivalent may specify the number of right-shift operations used for pixel interpolation.
[0111] In such embodiments, the exemplary syntax element “num_interp_filter[i]” (the example is highlighted in bold in Table 10, but it can take any form, label, term, or combination thereof, all of which are intended to be within the scope of this disclosure) or its equivalent can specify the number of interpolation filters in the i-th set of interpolation filters.
[0112] Here again, in such embodiments, the exemplary syntax element “num_interp_filter_coeff[i]” (the example is highlighted in bold in Table 10, but it can take any form, label, term, or combination thereof, all of which are intended to be within the scope of this disclosure) or its equivalent can specify the number of taps used for the interpolation filter in the i-th set of interpolation filters.
[0113] Here again, in such embodiments, the exemplary syntax element “interp_filter_coeff_abs[i][j][l]” (the example is highlighted in bold in Table 10, but it can take any form, label, term, or combination thereof, all of which are intended to be within the scope of this disclosure) or its equivalent can specify the absolute value of the l-th coefficient of the j-th interpolation filter in the i-th set of interpolation filters.
[0114] Also in such embodiments, the exemplary syntax element “interp_filter_coeff_sign[i][j][l]” (the example is highlighted in bold in Table 10, but it can take any form, label, term, or combination thereof, all of which are intended to be within the scope of this disclosure) or its equivalent can specify the sign of the l-th coefficient of the j-th interpolation filter in the i-th set of interpolation filters.
[0115] The disclosed syntax elements may be shown in any high-level parameter sets such as VPS, SPS, PPS, and slice segment headers. It should also be noted that additional syntax elements may be used at the sequence level, picture level, and / or CU level to facilitate the selection of interpolation filters for the operational coding level. It should also be noted that the disclosed flags may be replaced by variables that can indicate the selected filter set. In the intended embodiments, it should be noted that any number of interpolation filter sets (e.g., two, three, or more) may be transmitted in the bitstream.
[0116] The disclosed embodiments allow interpolation of pixels at fractional positions during the motion compensation prediction process using any combination of interpolation filters. For example, in an embodiment capable of performing lossy coding of a 4:4:4 video signal (in RGB or YCbCr format), the default 8-tap filter can be used to generate fractional pixels for three color components (i.e., R, G, and B components). In another embodiment capable of performing lossless coding of a video signal, the default 4-tap filter can be used to generate fractional pixels for three color components (i.e., Y, Cb, and Cr components in the YCbCr color space, and R, G, and B components in the RGB color space).
[0117] Figure 11A is a diagram of an exemplary communication system 100 that can implement one or more disclosed embodiments. The communication system 100 may be a multiple access system that provides content such as voice, data, video, messaging, and broadcast to multiple wireless users. The communication system 100 may enable multiple wireless users to access such content through the sharing of system resources, including wireless bandwidth. For example, the communication system 100 may utilize one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), quadrature FDMA (OFDMA), and single-carrier FDMA (SC-FDMA).
[0118] As shown in Figure 11A, the communication system 100 may include radio transceiver units (WTRUs) 102a, 102b, 102c, and / or 102d (sometimes referred to generally or collectively as WTRUs 102), radio access networks (RANs) 103 / 104 / 105, core networks 106 / 107 / 109, public switched telephone networks (PSTNs) 108, the internet 110, and other networks 112, but it will be understood that the disclosed systems and methods intend any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 102a, 102b, 102c, and 102d can be any type of device configured to operate and / or communicate in a radio environment. For example, WTRU102a, 102b, 102c, and 102d can be configured to transmit and / or receive wireless signals and may include user equipment (UEs), mobile stations, fixed or mobile subscriber units, pagers, cellular phones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, and consumer electronics.
[0119] The communication system 100 may also include base stations 114a and 114b. Each of the base stations 114a and 114b can be any type of device configured to wirelessly interface with at least one of the WTRUs 102a, 102b, 102c, and 102d to facilitate access to one or more communication networks such as the core network 106 / 107 / 109, the Internet 110, and / or network 112. For example, base stations 114a and 114b could be base transceiver stations (BTS), node B, enode B, home node B, home enode B, site controller, access point (AP), and wireless router. Although base stations 114a and 114b are shown as single elements, it will be understood that base stations 114a and 114b can include any number of interconnected base stations and / or network elements.
[0120] Base station 114a can be part of RAN 103 / 104 / 105, which may also include other base stations and / or network elements (not shown) such as base station controllers (BSCs), radio network controllers (RNCs), and relay nodes. Base station 114a and / or base station 114b can be configured to transmit and / or receive radio signals within a specific geographic area, which may be called a cell (not shown). A cell can be further divided into cell sectors. For example, a cell associated with base station 114a can be divided into three sectors. Thus, in one embodiment, base station 114a may include three transceivers, for example, one per sector of the cell. In another embodiment, base station 114a can utilize multiple-input multiple-output (MIMO) technology, and thus multiple transceivers can be utilized per sector of the cell.
[0121] Base stations 114a and 114b can communicate with one or more WTRUs 102a, 102b, 102c, and 102d over air interfaces 115 / 116 / 117, where air interfaces 115 / 116 / 117 can be any suitable radio communication link (e.g., radio frequency (RF), microwave, infrared (IR), ultraviolet (UV), visible light, etc.). Air interfaces 115 / 116 / 117 can be established using any suitable radio access technology (RAT).
[0122] More specifically, as mentioned above, the communication system 100 can be a multiple access system and can utilize one or more channel access schemes such as CDMA, TDMA, FDMA, OFDMA, and SC-FDMA. For example, base stations 114a and WTRU 102a, 102b, 102c in RAN 103 / 104 / 105 can implement radio technologies such as Universal Mobile Communications System (UMTS) Terrestrial Radio Access (UTRA), which can establish air interfaces 115 / 116 / 117 using broadband CDMA (WCDMA). WCDMA can include communication protocols such as High Speed Packet Access (HSPA) and / or Advanced HSPA (HSPA+). HSPA can include High Speed Downlink Packet Access (HSDPA) and / or High Speed Uplink Packet Access (HSUPA).
[0123] In another embodiment, base stations 114a and WTRUs 102a, 102b, 102c can implement radio technologies such as Advanced UMTS Terrestrial Radio Access (E-UTRA), which can establish air interfaces 115 / 116 / 117 using Long-Term Evolution (LTE) and / or LTE Advanced (LTE-A).
[0124] In other embodiments, base stations 114a and WTRUs 102a, 102b, and 102c can implement radio technologies such as IEEE 802.16 (i.e., Global Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Provisional Standard 2000 (IS-2000), Provisional Standard 95 (IS-95), Provisional Standard 856 (IS-856), Global System for Mobile Communications (GSM), High Speed Data Rate for GSM Evolution (EDGE), and GSM EDGE (GERAN). Base station 114b in Figure 11A can be, for example, a wireless router, home node B, home e-node B, or access point, and can utilize any suitable RAT to facilitate wireless connectivity in localized areas such as workplaces, homes, vehicles, and campuses. In one embodiment, base stations 114b and WTRUs 102c and 102d can establish a wireless local area network (WLAN) by implementing wireless technologies such as IEEE 802.11. In another embodiment, base stations 114b and WTRUs 102c and 102d can establish a wireless personal area network (WPAN) by implementing wireless technologies such as IEEE 802.15. In yet another embodiment, base stations 114b and WTRUs 102c and 102d can establish a picocell or femtocell using cellular-based RATs (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, etc.). As shown in Figure 11A, base station 114b may have a direct connection to the internet 110. Therefore, base station 114b may not need to access the internet 110 via the core network 106 / 107 / 109.
[0125] RAN103 / 104 / 105 can communicate with core networks 106 / 107 / 109, which can be any type of network configured to provide voice, data, applications, and / or Voice over Internet Protocol (VoIP) services to one or more WTRU102a, 102b, 102c, and 102d. For example, core networks 106 / 107 / 109 can provide call control, billing services, mobile location-based services, prepaid calls, internet connectivity, video distribution, etc., and / or perform high-level security functions such as user authentication. Although not shown in Figure 11A, it will be understood that RAN103 / 104 / 105 and / or core networks 106 / 107 / 109 can communicate directly or indirectly with other RANs that utilize the same RAT or a different RAT as RAN103 / 104 / 105. For example, in addition to connecting to RANs 103 / 104 / 105 which can utilize E-UTRA radio technology, core networks 106 / 107 / 109 can also communicate with other RANs (not shown) that utilize GSM radio technology.
[0126] Core networks 106 / 107 / 109 can also serve as gateways for WTRUs 102a, 102b, 102c, and 102d to access PSTN 108, the Internet 110, and / or other networks 112. PSTN 108 may include a circuit-switched telephone network providing basic telephone services (POTS). The Internet 110 may include a global system of interconnected computer networks and devices using common communication protocols such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and Internet Protocol (IP) within the TCP / IP Internet Protocol Suite. Network 112 may include wired or wireless networks owned and / or operated by other service providers. For example, network 112 may include another core network connected to one or more RANs that can utilize the same or different RATs as RAN 103 / 104 / 105.
[0127] Some or all of the WTRUs 102a, 102b, 102c, and 102d in the communication system 100 may include multimode functionality. For example, WTRUs 102a, 102b, 102c, and 102d may include multiple transceivers for communicating with different radio networks on different radio links. For example, WTRU 102c, shown in Figure 11A, may be configured to communicate with base station 114a, which can utilize cellular-based radio technology, and also with base station 114b, which can utilize IEEE 802 radio technology.
[0128] Figure 11B is a system diagram of an exemplary WTRU 102. As shown in Figure 11B, the WTRU 102 may include a processor 118, a transceiver 120, a transmit / receive element 122, a speaker / microphone 124, a keypad 126, a display / touchpad 128, a non-removable memory 130, a removable memory 132, a power supply 134, a Global Positioning System (GPS) chipset 136, and other peripherals 138. It will be understood that the WTRU 102 may include any subcombinations of the above elements while maintaining consistency with the embodiment. Furthermore, the embodiments are intended to include, but are not limited to, base stations 114a, 114b, and / or, in particular, base stations (BTS), node B, site controller, access point (AP), home node B, evolved node B (enode B), home evolved node B (HeNB), home evolved node B gateway, and proxy node, which may be represented by base stations 114a, 114b, as shown in Figure 11B and may include some or all of the elements described herein.
[0129] The processor 118 can be a general-purpose processor, a dedicated processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors working with a DSP core, a controller, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), and a state machine, etc. The processor 118 can perform signal coding, data processing, power control, input / output processing, and / or any other functions that enable the WTRU 102 to operate in a wireless environment. The processor 118 can be coupled to the transceiver 120, and the transceiver 120 can be coupled to the transmit / receive element 122. Although Figure 11B shows the processor 118 and the transceiver 120 as separate components, it will be understood that the processor 118 and the transceiver 120 can be integrated together in an electronic package or chip.
[0130] The transmit / receive element 122 can be configured to transmit signals to or receive signals from a base station (e.g., base station 114a) on the air interface 115 / 116 / 117. For example, in one embodiment, the transmit / receive element 122 may be an antenna configured to transmit and / or receive RF signals. In another embodiment, the transmit / receive element 122 may be a radiator / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, the transmit / receive element 122 may be configured to transmit and / or receive both RF and optical signals. It will be understood that the transmit / receive element 122 can be configured to transmit and / or receive any combination of radio signals.
[0131] In addition, although the transmit / receive element 122 is shown as a single element in Figure 11B, the WTRU 102 can include any number of transmit / receive elements 122. More specifically, the WTRU 102 can utilize MIMO technology. Therefore, in one embodiment, the WTRU 102 can include two or more transmit / receive elements 122 (e.g., multiple antennas) for transmitting and receiving radio signals over the air interfaces 115 / 116 / 117.
[0132] The transceiver 120 can be configured to modulate the signal transmitted by the transmit / receive element 122 and demodulate the signal received by the transmit / receive element 122. As mentioned above, the WTRU 102 can have multimode capabilities. Therefore, the transceiver 120 can include multiple transceivers to enable the WTRU 102 to communicate via multiple RATs, such as UTRA and IEEE 802.11.
[0133] The processor 118 of the WTRU102 can be coupled to a speaker / microphone 124, a keypad 126, and / or a display / touchpad 128 (e.g., a liquid crystal display (LCD) display unit or an organic light-emitting diode (OLED) display unit) and can receive user input data from them. The processor 118 can also output user data to the speaker / microphone 124, the keypad 126, and / or the display / touchpad 128. In addition, the processor 118 can obtain information from any type of suitable memory, such as non-removable memory 130 and / or removable memory 132, and store data in them. The non-removable memory 130 may include random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 132 may include a subscriber identification module (SIM) card, a memory stick, and a secure digital (SD) memory card, etc. In other embodiments, the processor 118 can obtain information from memory located on a server or home computer (not shown), rather than from memory physically located on the WTRU 102, and can store data in such memory.
[0134] The processor 118 can receive power from the power supply 134 and can be configured to distribute and / or control power to other components within the WTRU 102. The power supply 134 can be any suitable device for supplying power to the WTRU 102. For example, the power supply 134 may include one or more dry cell batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel-metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, and fuel cells.
[0135] The processor 118 can also be coupled to a GPS chipset 136, which can be configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 102. In addition to, or instead of, the information from the GPS chipset 136, the WTRU 102 can receive location information from base stations (e.g., base stations 114a, 114b) on the air interfaces 115 / 116 / 117, and / or determine its own position based on the timing of signals received from two or more nearby base stations. It will be understood that the WTRU 102 can acquire location information using any suitable location determination method while maintaining consistency with the embodiments.
[0136] The processor 118 can be further coupled to other peripherals 138, which may include one or more software modules and / or hardware modules that provide additional features, functions, and / or wired or wireless connectivity. For example, peripherals 138 may include an accelerometer, e-compass, satellite transceiver, digital camera (for photos or videos), Universal Serial Bus (USB) port, vibration device, TV transceiver, hands-free headset, Bluetooth® module, frequency modulation (FM) radio unit, digital music player, media player, video game player module, and internet browser.
[0137] Figure 11C is a system diagram of RAN103 and core network 106 according to an embodiment. As mentioned above, RAN103 can communicate with WTRU102a, 102b, and 102c over air interface 115 using UTRA radio technology. RAN103 can also communicate with core network 106. As shown in Figure 11C, RAN103 may include nodes B140a, 140b, and 140c, each of which may include one or more transceivers for communication with WTRU102a, 102b, and 102c over air interface 115. Nodes B140a, 140b, and 140c may each be associated with a specific cell (not shown) within RAN103. RAN103 may also include RNC142a and 142b. It will be understood that RAN103 may include any number of nodes B and RNC while maintaining consistency with the embodiment.
[0138] As shown in Figure 11C, nodes B140a and B140b can communicate with RNC142a. In addition, node B140c can communicate with RNC142b. Nodes B140a, B140b, and B140c can communicate with their respective RNC142a and B142b via the Iub interface. RNC142a and B142b can communicate with each other via the Iur interface. Each of RNC142a and B142b can be configured to control their respective connected nodes B140a, B140b, and B140c. In addition, each of RNC142a and B142b can be configured to implement or support other functions such as outer loop power control, load control, admission control, packet scheduling, handover control, macro diversity, security functions, and data encryption.
[0139] The core network 106 shown in Figure 11C may include a media gateway (MGW) 144, a mobile switching center (MSC) 146, a serving GPRS support node (SGSN) 148, and / or a gateway GPRS support node (GGSN) 150. Although each of the above elements is shown as part of the core network 106, it will be understood that any one of these elements may be owned and / or operated by an entity different from the core network operator.
[0140] RNC142a in RAN103 can connect to MSC146 in core network 106 via the IuCS interface. MSC146 can connect to MGW144. MSC146 and MGW144 provide WTRU102a, 102b, and 102c with access to circuit-switched networks such as PSTN108, facilitating communication between WTRU102a, 102b, and 102c and conventional land-line communication devices.
[0141] RNC142a in RAN103 can also connect to SGSN148 in core network 106 via the IuPS interface. SGSN148 can connect to GGSN150. SGSN148 and GGSN150 provide WTRU102a, 102b, and 102c with access to packet-switched networks such as the Internet 110, facilitating communication between WTRU102a, 102b, and 102c and IP-enabled devices.
[0142] As mentioned above, the core network 106 may also be connected to network 112, which may include other wired or wireless networks owned and / or operated by other service providers.
[0143] Figure 11D is a system diagram of RAN 104 and core network 107 according to an embodiment. As mentioned above, RAN 104 can communicate with WTRU 102a, 102b, and 102c over air interface 116 using E-UTRA radio technology. RAN 104 can also communicate with core network 107.
[0144] RAN104 may include e-nodes B160a, 160b, and 160c, but it will be understood that RAN104 may include any number of e-nodes B while maintaining consistency with the embodiment. Each of the e-nodes B160a, 160b, and 160c may include one or more transceivers for communicating with WTRU102a, 102b, and 102c on the air interface 116. In one embodiment, e-nodes B160a, 160b, and 160c can implement MIMO technology. Thus, e-node B160a can, for example, use multiple antennas to transmit radio signals to and receive radio signals from WTRU102a.
[0145] Each of the e-nodes B160a, 160b, and 160c can be associated with a specific cell (not shown) and configured to handle wireless resource management decisions, handover decisions, and user scheduling on the uplink and / or downlink. As shown in Figure 11D, the e-nodes B160a, 160b, and 160c can communicate with each other over the X2 interface.
[0146] The core network 107 shown in Figure 11D may include a Mobility Management Gateway (MME) 162, a Serving Gateway 164, and a Packet Data Network (PDN) Gateway 166. Although each of the above elements is shown as part of the core network 107, it will be understood that any one of these elements may be owned and / or operated by an entity different from the core network operator.
[0147] The MME162 can connect to each of the e-nodes B160a, 160b, and 160c within RAN104 via the S1 interface and can act as a control node. For example, the MME162 can be responsible for user authentication of WTRU102a, 102b, and 102c, bearer activation / deactivation, and selection of a specific serving gateway during the initial connection of WTRU102a, 102b, and 102c. The MME162 can also provide control plane functionality for exchanges between RAN104 and other RANs (not shown) utilizing other radio technologies such as GSM or WCDMA.
[0148] The serving gateway 164 can connect to each of the e-nodes B160a, 160b, and 160c in RAN104 via the S1 interface. The serving gateway 164 can generally perform route selection and forwarding of user data packets to and from WTRU102a, 102b, and 102c. The serving gateway 164 can also perform other functions, such as anchoring the user plane during e-node B handover, triggering paging when downlink data is available to WTRU102a, 102b, and 102c, and managing and remembering the context of WTRU102a, 102b, and 102c.
[0149] The serving gateway 164 can also be connected to the PDN gateway 166, which provides WTRU 102a, 102b, and 102c with access to a packet-switched network such as the Internet 110, thereby facilitating communication between WTRU 102a, 102b, and 102c and IP-enabled devices.
[0150] The core network 107 can facilitate communication with other networks. For example, the core network 107 can provide WTRU 102a, 102b, and 102c with access to a circuit-switched network such as PSTN 108, thereby facilitating communication between WTRU 102a, 102b, and 102c and conventional land-line communication devices. For example, the core network 107 may include an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that acts as an interface between the core network 107 and PSTN 108, or can communicate with the IP gateway. In addition, the core network 107 can provide WTRU 102a, 102b, and 102c with access to network 112, which may include other wired or wireless networks owned and / or operated by other service providers.
[0151] Figure 11E is a system diagram of RAN105 and core network 109 according to an embodiment. RAN105 can be an access service network (ASN) that communicates with WTRU102a, 102b, and 102c over air interface 117 using IEEE 802.16 wireless technology. As will be further described below, communication links between different functional entities of WTRU102a, 102b, 102c, RAN105, and core network 109 can be defined as reference points.
[0152] As shown in Figure 11E, RAN105 may include base stations 180a, 180b, 180c and an ASN gateway 182, but it will be understood that RAN105 may include any number of base stations and ASN gateways while maintaining consistency with the embodiment. Each of the base stations 180a, 180b, 180c may be associated with a specific cell (not shown) within RAN105, and each may include one or more transceivers for communicating with WTRU102a, 102b, 102c on the air interface 117. In one embodiment, the base stations 180a, 180b, 180c may implement MIMO technology. Thus, base station 180a may, for example, use multiple antennas to transmit radio signals to and receive radio signals from WTRU102a. Base stations 180a, 180b, and 180c can also provide mobility management functions such as handoff triggering, tunnel establishment, radio resource management, traffic classification, and quality of service (QoS) policy implementation. The ASN gateway 182 can act as a traffic aggregation point and is responsible for paging, subscriber profile caching, and route selection to the core network 109.
[0153] The air interface 117 between WTRU102a, 102b, 102c and RAN105 can be defined as the R1 reference point, implementing the IEEE 802.16 specification. In addition, each of WTRU102a, 102b, and 102c can establish a logical interface (not shown) with the core network 109. The logical interface between WTRU102a, 102b, 102c and the core network 109 can be defined as the R2 reference point, which can be used for authentication, authorization, IP host configuration management, and / or mobility management.
[0154] The communication links between base stations 180a, 180b, and 180c can be defined as R8 reference points, including protocols to facilitate WTRU handover and data transfer between base stations. The communication links between base stations 180a, 180b, and 180c and the ASN gateway 182 can be defined as R6 reference points. R6 reference points may include protocols to facilitate mobility management based on mobility events associated with each of WTRU 102a, 102b, and 102c.
[0155] As shown in Figure 11E, RAN 105 can connect to core network 109. The communication link between RAN 105 and core network 109 can be defined as an R3 reference point, including, for example, protocols to facilitate data transfer and mobility management functions. Core network 109 may include a Mobile IP Home Agent (MIP-HA) 184, an Authentication Authorization and Billing (AAA) server 186, and a gateway 188. Although each of the above elements is shown as part of core network 109, it will be understood that any one of these elements may be owned and / or operated by an entity different from the core network operator.
[0156] The MIP-HA can handle IP address management and enable WTRU102a, 102b, and 102c to roam between different ASNs and / or between different core networks. The MIP-HA184 can provide WTRU102a, 102b, and 102c with access to packet-switched networks such as the Internet 110, facilitating communication between WTRU102a, 102b, and 102c and IP-enabled devices. The AAA server 186 can handle user authentication and support user services. The gateway 188 can facilitate inter-network connectivity with other networks. For example, the gateway 188 can provide WTRU102a, 102b, and 102c with access to circuit-switched networks such as the PSTN 108, facilitating communication between WTRU102a, 102b, and 102c and conventional landline communication devices. In addition, gateway 188 provides access to network 112 to WTRU 102a, 102b, and 102c, and network 112 may include other wired or wireless networks owned and / or operated by other service providers.
[0157] Although not shown in Figure 11E, it will be understood that RAN105 can connect to other ASNs, and core network 109 can connect to other core networks. The communication link between RAN105 and other ASNs can be defined as an R4 reference point, and the R4 reference point can include protocols for coordinating the mobility of WTRU102a, 102b, and 102c between RAN105 and other ASNs. The communication link between core network 109 and other core networks can be defined as an R5 reference point, and the R5 reference point can include protocols for facilitating inter-network connectivity between home core networks and local core networks.
[0158] While features and elements have been described above in specific combinations, those skilled in the art will understand that each feature or element can be used alone or in any combination with other features and elements. In addition, the methods described herein can be implemented in computer programs, software, or firmware contained within a computer-readable medium for execution by a computer or processor. Examples of computer-readable mediums include electronic signals (transmitted over wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital multi-purpose disks (DVDs). A processor in conjunction with software can be used to implement a radio frequency transceiver for use in a WTRU, UE, terminal, base station, RNC, or any host computer. [Industrial applicability]
[0159] The present invention can be used in systems, methods, and devices for encoding and decoding video content. [Explanation of symbols]
[0160] 102, 102a-102d, WTRU 103, 104, 105 RAN 106, 107, 109 Core Network 108 PSTN 110 Internet
Claims
1. A method of video encoding content, The steps include obtaining the residuals of coding blocks among multiple coding blocks of different sizes in an image sequence, A step of determining whether to apply a color space conversion to the residual of the coding block based on a rate distortion cost comparison, Based on the decision that the color space conversion is applied to the coding block, the steps include: applying the color space conversion to the coding block; The steps include including in the bitstream an encoding unit adaptive color space conversion indication configured to indicate whether a color space conversion is applied to one of the plurality of encoding blocks. Equipped with, A step of determining whether there is at least one non-zero residual coefficient among the residual coefficients associated with the coding block among the plurality of coding blocks, wherein the step of including a coding unit adaptive color space conversion indication for the coding block among the plurality of coding blocks in the bitstream is based on determining whether there is at least one non-zero residual coefficient among the residual coefficients associated with the coding block among the plurality of coding blocks. How to prepare even more.
2. The method of claim 1, comprising the step of determining whether there is at least one non-zero residual coefficient among the residual coefficients associated with the coding block among the plurality of coding blocks, wherein the step of determining whether there is at least one non-zero coefficient among the lumar residual coefficients.
3. The method of claim 1, wherein the step of determining whether there is at least one non-zero residual coefficient among the residual coefficients associated with the coding block among the plurality of coding blocks is to determine whether there is at least one non-zero coefficient among the chroma residual coefficients.
4. A step of deciding to enable adaptive color space conversion for the sequence of images, The steps include including an adaptive color space conversion enablement indication in the sequence parameter set associated with the sequence of images to indicate that adaptive color space conversion is permitted to be used for the sequence of images. The method of claim 1, further comprising the following:
5. A step of determining whether to enable adaptive color space conversion for the sequence of images, wherein the step of determining whether to apply the color space conversion to the residuals of the coding block is based on the step of determining whether to enable adaptive color space conversion for the sequence of images. The method of claim 1, further comprising the following:
6. A step of calculating the rate distortion cost associated with performing residual coding in the GBR color space, A step of calculating the rate distortion cost associated with performing residual coding in the YCgCo color space, wherein the decision to apply the color space conversion to the coding block among the plurality of coding blocks is based on a rate distortion cost associated with performing residual coding in the YCgCo color space that is lower than the rate distortion cost associated with performing residual coding in the GBR color space. The method of claim 1, further comprising the following:
7. In an image sequence, the residuals of the coded blocks are obtained from multiple coded blocks of different sizes. Based on the rate distortion cost comparison, it is determined whether to apply a color space conversion to the residual of the coding block. Based on the decision that the color space conversion is applied to the coding block, the color space conversion is applied to the coding block. The bitstream includes an encoding unit adaptive color space conversion indication configured to indicate whether a color space conversion is applied to one of the encoding blocks among the plurality of encoding blocks. Processor configured in such a way Equipped with, The aforementioned processor, The system is further configured to determine whether there is at least one non-zero residual coefficient among the residual coefficients associated with the coding block among the plurality of coding blocks, and including coding unit adaptive color space conversion indications for the coding block among the plurality of coding blocks in the bitstream is based on determining whether there is at least one non-zero residual coefficient among the residual coefficients associated with the coding block among the plurality of coding blocks. Video encoding device.
8. The video encoding apparatus of claim 7, wherein determining whether there is at least one non-zero residual coefficient among the residual coefficients associated with the encoding block among the plurality of encoding blocks is equivalent to determining whether there is at least one non-zero coefficient among the lumar residual coefficients.
9. The video encoding apparatus of claim 7, wherein determining whether there is at least one non-zero residual coefficient among the residual coefficients associated with the encoding block among the plurality of encoding blocks is equivalent to determining whether there is at least one non-zero coefficient among the chroma residual coefficients.
10. The aforementioned processor, It was decided to enable adaptive color space conversion for the aforementioned sequence of images, The sequence parameter set associated with the sequence of images includes an adaptive color space conversion enablement indication, indicating that adaptive color space conversion is permitted for use with the sequence of images. A video encoding apparatus according to claim 7, further configured as follows.
11. The aforementioned processor, The system is further configured to determine whether to enable adaptive color space conversion for the sequence of images, and the determination of whether to apply color space conversion to the residuals of the coding blocks is based on the determination to enable adaptive color space conversion for the sequence of images. A video encoding apparatus according to claim 7.
12. The aforementioned processor, We calculate the rate distortion cost associated with performing residual coding in the GBR color space. The calculation of the rate distortion cost associated with performing residual coding in the YCgCo color space and the determination that the color space conversion is applied to the coding block among the plurality of coding blocks is based on a rate distortion cost associated with performing residual coding in the YCgCo color space that is lower than the rate distortion cost associated with performing residual coding in the GBR color space. A video encoding apparatus according to claim 7, further configured as follows.
13. A computer-readable medium that includes instructions causing a processor to perform any of the methods of claims 1 to 6.