Cross-plane filtering for chroma signal enhancement in video coding

Cross-plane filtering enhances chroma signals in video coding by using adaptive filters from the luma plane, addressing blurring and texture issues while minimizing overhead and performance degradation.

JP7796856B2Active Publication Date: 2026-01-09INTERDIGITAL MADISON PATENT HLDG
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024218775
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2013-07-12
Filing Date
2024-12-13
Publication Date
2026-01-09
Estimated Expiration
2033-09-27

AI Technical Summary

Technical Problem

Existing video coding techniques result in blurred edges and textures in the chroma plane, which are not adequately addressed by current chroma prediction methods.

Method used

Implement cross-plane filtering using adaptive filters that leverage information from the corresponding luma plane to enhance chroma signals, with quantized filter coefficients to minimize overhead and performance degradation, applicable to various color subsampling formats and regions of the video image.

Benefits of technology

Improves chroma signal quality by reducing blurring and texture artifacts while maintaining reasonable computational and bandwidth efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007796856000021
    Figure 0007796856000021
  • Figure 0007796856000022
    Figure 0007796856000022
  • Figure 0007796856000023
    Figure 0007796856000023
Patent Text Reader

Abstract

To solve the problem that video coding using a known chroma prediction technique may provide a result of a video image having an extremely unclear edge and / or texture in a chroma plane.SOLUTION: Cross-plane filtering is used for recovering an unclear edge and / or texture in one or both chroma planes while using information from a corresponding luma plane. An adaptive cross-plane filter is implemented. A cross-plane filter coefficient is quantized in such a manner that performance deterioration caused by overhead in a bitstream is minimized, and / or transferred. Cross-plane filtering is applied to a selective area (e.g., an edge area) of a video image. The cross-plane filter is implemented in a single-layer video coding system and / or a multi-layer video coding system.SELECTED DRAWING: Figure 9A
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] The present invention relates to a method and apparatus for performing cross-plane filtering for chroma signal enhancement in video coding.

[0002] This application claims priority to U.S. Provisional Patent Application No. 61 / 707,682, filed September 28, 2012, U.S. Provisional Patent Application No. 61 / 762,611, filed February 8, 2013, U.S. Provisional Patent Application No. 61 / 778,218, filed March 12, 2013, and U.S. Provisional Patent Application No. 61 / 845,792, filed July 12, 2013, which applications are incorporated herein by reference in their entireties.

[0003] Video encoding (coding) systems are often used to compress digital video signals, e.g., to reduce the memory space consumed and / or the transmission bandwidth consumption associated with such signals. For example, block-based hybrid video encoding (coding) systems are widely deployed and frequently used.

[0004] A digital video signal typically has three color planes, including a luma plane, a blue-difference chroma plane, and a red-difference chroma plane. Pixels in the chroma plane typically have a smaller dynamic range than pixels in the luma plane, and the chroma planes of a video image typically are smoother and / or have less detail than the luma plane. Therefore, chroma blocks of a video image are easier to accurately predict, e.g., consuming fewer resources and / or resulting in smaller prediction errors. Summary of the Invention [Problem to be solved by the invention]

[0005] However, video coding using known chroma prediction techniques results in video images with significantly blurred edges and / or textures in the chroma plane. The present invention provides a method and apparatus for improved cross-plane filtering for chroma signal enhancement in video coding. [Means for solving the problem]

[0006] Cross-plane filtering is used to recover blurred edges and / or textures in one or both chroma planes using information from the corresponding luma plane. An adaptive cross-plane filter is implemented. Cross-plane filter coefficients are quantized and / or signaled to ensure reasonable (e.g., reduced and / or minimized) overhead in the bitstream without causing performance degradation. One or more characteristics of the cross-plane filter (e.g., size, separability, symmetry, etc.) are determined to ensure reasonable (e.g., reduced and / or minimized) overhead in the bitstream without causing performance degradation. Cross-plane filtering is applied to videos having various color subsampling formats (e.g., 4:4:4, 4:2:2, and 4:2:0). Cross-plane filtering is applied to selected regions of the video image, such as edge areas and / or areas specified by one or more parameters signaled in the bitstream. Cross-plane filtering is implemented in single-layer video coding systems and / or multi-layer video coding systems.

[0007] An exemplary video decoding process according to cross-plane filtering includes receiving a video signal and a cross-plane filter associated with the video signal. The video decoding process includes applying the cross-plane filter to luma plane pixels of the video signal to determine a chroma offset. The video decoding process includes adding the chroma offset to corresponding chroma plane pixels of the video signal.

[0008] A video coding device is configured for cross-plane filtering. The video coding device includes a network interface configured to receive a video signal and a cross-plane filter associated with the video signal. The video coding device includes a processor configured to apply the cross-plane filter to luma plane pixels of the video signal to determine a chroma offset. The processor is configured to add the chroma offset to corresponding chroma plane pixels of the video signal.

[0009] An exemplary video encoding process according to cross-plane filtering includes receiving a video signal, generating a cross-plane filter using components of the video signal, quantizing filter coefficients associated with the cross-plane filter, encoding the filter coefficients into a bitstream representing the video signal, and transmitting the bitstream.

[0010] A video coding device is configured for cross-plane filtering. The video coding device includes a network interface configured to receive a video signal. The video coding device includes a processor configured to generate a cross-plane filter using components of the video signal. The processor is configured to quantize filter coefficients associated with the cross-plane filter. The processor is configured to encode the filter coefficients into a bitstream representing the video signal. The processor is configured to transmit the bitstream, for example, via the network interface. [Effects of the Invention]

[0011] SUMMARY OF THE INVENTION The present invention provides a method and apparatus for improved cross-plane filtering for chroma signal enhancement in video coding. [Brief explanation of the drawings]

[0012] [Figure 1] 1 is a block diagram illustrating an example block-based video encoder. [Figure 2] 1 is a block diagram illustrating an example block-based video decoder. [Figure 3] FIG. 1 is a block diagram illustrating an example two-layer spatial scalable video encoder. [Figure 4] FIG. 2 is a block diagram illustrating an example two-layer spatial scalable video decoder. [Figure 5] FIG. 2 is a block diagram of an example inter-layer prediction processing and management unit. [Figure 6A] FIG. 1 illustrates an exemplary 4:4:4 color subsampling format. [Figure 6B] FIG. 1 illustrates an exemplary 4:2:2 color subsampling format. [Figure 6C] FIG. 1 illustrates an exemplary 4:2:0 color subsampling format. [Figure 7]FIG. 1 is a block diagram illustrating an example of cross-plane filtering. [Figure 8A] FIG. 10 is a block diagram illustrating another example of cross-plane filtering. [Figure 8B] FIG. 10 is a block diagram illustrating another example of cross-plane filtering. [Figure 9A] FIG. 10 is a block diagram illustrating another example of cross-plane filtering. [Figure 9B] FIG. 10 is a block diagram illustrating another example of cross-plane filtering. [Figure 10A] FIG. 10 illustrates exemplary sizes and / or support regions of cross-plane filters (filter_Y4Cb and filter_Y4Cr) for selected chroma pixels in 4:4:4. [Figure 10B] A diagram showing example sizes and / or support regions of cross-plane filters (filter_Y4Cb and filter_Y4Cr) for selected chroma pixels in 4:2:2. [Figure 10C] A diagram showing example sizes and / or support regions of cross-plane filters (filter_Y4Cb and filter_Y4Cr) for selected chroma pixels in 4:2:0. [Figure 11A] A diagram showing exemplary unified sizes and / or support regions of cross-plane filters (filter_Y4Cb and filter_Y4Cr) for selected chroma pixels in 4:4:4. [Figure 11B] A diagram showing exemplary unified sizes and / or support regions of cross-plane filters (filter_Y4Cb and filter_Y4Cr) for selected chroma pixels in 4:2:2. [Figure 11C] A diagram showing exemplary unified sizes and / or support regions of cross-plane filters (filter_Y4Cb and filter_Y4Cr) for selected chroma pixels in 4:2:0. [Figure 12A]10A-10C illustrate an exemplary lack of symmetry properties of exemplary cross-plane filtering. [Figure 12B] FIG. 10 illustrates exemplary horizontal and vertical symmetry properties of exemplary cross-plane filtering. [Figure 12C] FIG. 10 illustrates an exemplary vertical symmetry property of exemplary cross-plane filtering. [Figure 12D] FIG. 10 illustrates an exemplary horizontal symmetry property of exemplary cross-plane filtering. [Figure 12E] FIG. 1 illustrates an exemplary point-symmetry property of exemplary cross-plane filtering. [Figure 13A] FIG. 1 illustrates exemplary horizontal and vertical one-dimensional filters with no symmetry. [Figure 13B] FIG. 1 illustrates exemplary horizontal and vertical one-dimensional filters with symmetry. [Figure 14] FIG. 10 illustrates an exemplary syntax table illustrating an example of signaling a set of cross-plane filter coefficients. [Figure 15A] FIG. 10 illustrates an exemplary arrangement of cross-plane filter coefficients. [Figure 15B] FIG. 10 illustrates an exemplary arrangement of cross-plane filter coefficients. [Figure 16] FIG. 10 illustrates an example syntax table illustrating an example of signaling multiple sets of cross-plane filter coefficients. [Figure 17] FIG. 10 illustrates an exemplary syntax table illustrating an example of conveying information specifying a region for cross-plane filtering. [Figure 18] 1A and 1B are diagrams illustrating examples of multiple image regions detected according to an implementation of region-based cross-plane filtering. [Figure 19] FIG. 10 illustrates an example syntax table illustrating an example of conveying information about multiple regions along with multiple sets of cross-plane filter coefficients. [Figure 20]FIG. 1 illustrates an exemplary picture level selection algorithm for cross-plane filtering. [Figure 21A] FIG. 1 is a system diagram of an example communication system in which one or more disclosed embodiments may be implemented. [Figure 21B] FIG. 21B is a system diagram of an example wireless transmit / receive unit (WTRU) used within the communication system shown in FIG. 21A. [Figure 21C] 21B is a system diagram of an example radio access network and an example core network used within the communication system shown in FIG. 21A. [Figure 21D] 21B is a system diagram of an example radio access network and an example core network used within the communication system shown in FIG. 21A. [Figure 21E] 21B is a system diagram of an example radio access network and an example core network used within the communication system shown in FIG. 21A. DETAILED DESCRIPTION OF THE INVENTION

[0013] 1 shows an exemplary block-based video encoder. An input video signal 102 is processed, for example, block-by-block. A video block unit includes 16x16 pixels. Such a block unit is called a macroblock (MB). The video block unit size is extended, for example, to 64x64 pixels. The extended-size video block is used to compress high-resolution video signals (e.g., video signals above 1080p). The extended block size is called a coding unit (CU). The CU is divided into one or more prediction units (PUs), to which separate prediction methods are applied.

[0014] For one or more input video blocks (e.g., each input video block), such as a MB or CU, spatial prediction 160 and / or temporal prediction 162 are performed. Spatial prediction 160, referred to as intra prediction, predicts, for example, a video block using pixels from one or more previously coded neighboring blocks within a video picture and / or slice. Spatial prediction 160 reduces spatial redundancy inherent in a video signal. Temporal prediction 162, referred to as inter prediction and / or motion-compensated prediction, predicts, for example, a video block using pixels from one or more previously coded video pictures. Temporal prediction reduces temporal redundancy inherent in a video signal. A temporal prediction signal for a video block includes one or more motion vectors and / or one or more reference picture indexes to identify, for example, which reference picture in reference picture store 164 the temporal prediction signal comes from, e.g., if multiple reference pictures are used.

[0015] After spatial prediction and / or temporal prediction are performed, a mode decision block 180 (e.g., in an encoder) selects a prediction mode, e.g., based on a rate-distortion optimization method. The prediction block is subtracted (116) from the video block. The prediction residual is transformed (104) and / or quantized (106). One or more quantized residual coefficients are inverse quantized (110) and / or inverse transformed (112), e.g., to form a reconstructed residual. The reconstructed residual is added (126) to the prediction block, e.g., to form a reconstructed video block.

[0016] Further in-loop filtering, such as one or more deblocking filters and / or adaptive loop filter 166, is applied to the reconstructed video block, for example, before it is stored in reference picture store 164 and / or before it is used to code a subsequent video block. To form output video bitstream 120, the coding mode (e.g., inter or intra), prediction mode information, motion information, and / or quantized residual coefficients are sent to entropy coding unit 108, where they are further compressed and / or packed, for example, to form bitstream 120.

[0017] 2 shows an exemplary block-based video decoder corresponding to the block-based encoder shown in FIG. 1. A video bitstream 202 is unpacked and / or entropy decoded, e.g., in an entropy decoding unit 208. Coding mode and / or prediction information is sent to a spatial prediction unit 260 (e.g., for intra-coding) or a temporal prediction unit 262 (e.g., for inter-coding) to form, e.g., a prediction block. One or more residual transform coefficients are sent to an inverse quantization unit 210 and / or an inverse transform unit 212 to reconstruct, e.g., a residual block. The prediction block and the residual block are summed at 226 to form, e.g., a reconstructed block. The reconstructed block is processed through in-loop filtering (e.g., using a loop filter 266), e.g., before being added to a reconstructed output video 220 and transmitted (e.g., to a display device) and / or stored in a reference picture store 264, e.g., for use in predicting one or more subsequent video blocks.

[0018] Video is consumed on devices having varying capabilities in terms of computing power, memory and / or storage size, display resolution, display frame rate, etc., e.g., by smartphones and / or tablets. Networks and / or transmission channels have varying characteristics in terms of packet loss rates, available channel bandwidth, burst error rates, etc. Video data is transmitted over a combination of wired and / or wireless networks, which complicates one or more underlying video transmission channel characteristics. In such scenarios, scalable video coding improves the video quality provided by video applications, e.g., by video applications running on devices with different capabilities over heterogeneous networks.

[0019] Scalable video coding encodes a video signal according to the highest representation (e.g., temporal resolution, spatial resolution, quality, etc.) but allows decoding from respective subsets of one or more video streams according to specified rates and / or representations utilized by, for example, one or more applications running on a client device. Scalable video coding allows for bandwidth and / or storage savings.

[0020] 3 illustrates an exemplary two-layer scalable video coding system having one base layer (BL) and one enhancement layer (EL). The spatial resolution between the two layers is different, and spatial scalability is applied. A base layer encoder (e.g., a High Efficiency Video Coding (HEVC) encoder) encodes the base layer video input, for example, block-by-block, to generate a base layer bitstream (e.g., according to the block diagram shown in FIG. 1). An enhancement layer encoder encodes the enhancement layer video input, for example, block-by-block, to generate an enhancement layer bitstream (e.g., according to the block diagram shown in FIG. 1). The coding efficiency of the scalable video coding system (e.g., the coding efficiency of the enhancement layer coding) is improved. For example, signal correlation from the reconstructed video of the base layer is used to improve prediction accuracy.

[0021] The reconstructed video of the base layer is processed, and at least a portion of one or more processed base layer pictures is inserted into an enhancement layer decoded picture buffer (EL DPB) and / or used to predict the enhancement layer video input. The base layer video and the enhancement layer video are substantially the same video source represented at different spatial resolutions, and they correspond to each other, for example, through a downsampling process. Inter-layer prediction (ILP) processing is performed by an inter-layer processing and / or management unit, such as an upsampling operation used to align the spatial resolution of the base layer reconstruction with the spatial resolution of the enhancement layer video. The scalable video coding bitstream includes a base layer bitstream generated by a base layer encoder, an enhancement layer bitstream generated by an enhancement layer encoder, and / or inter-layer prediction information.

[0022] The inter-layer prediction information is generated by an ILP processing and management unit. For example, the ILP information includes one or more of the type of inter-layer processing applied, one or more parameters used in the processing (e.g., which upsampling filter is used), which of one or more processed base layer pictures should be inserted into the EL DPB, etc. The base layer bitstream and the enhancement layer bitstream and / or the ILP information are multiplexed together to form, for example, a scalable bitstream (e.g., an SHVC bitstream).

[0023] Figure 4 shows an exemplary two-layer scalable video decoder corresponding to the scalable encoder shown in Figure 3. The decoder performs one or more operations, e.g., in the reverse order of the encoder. The scalable bitstream is demultiplexed into a base layer bitstream, an enhancement layer bitstream, and / or ILP information. The base layer decoder decodes the base layer bitstream and / or generates a base layer reconstruction.

[0024] The ILP processing and management unit receives the ILP information and / or processes the base layer reconstruction, for example, according to the received ILP information. The ILP processing and management unit selectively inserts one or more processed base layer pictures into the EL DPB, for example, according to the received ILP information. The enhancement layer decoder decodes the enhancement layer bitstream using a combination of temporal reference pictures and / or inter-layer reference pictures (e.g., one or more processed base layer pictures) to reconstruct the enhancement layer video. For purposes of this disclosure, the terms “inter-layer reference picture” and “processed base layer picture” are used interchangeably.

[0025] 5 illustrates an exemplary inter-layer prediction and processing management unit, implemented, for example, in the exemplary two-layer spatial scalable video encoder illustrated in FIG. 3 and / or the exemplary two-layer spatial scalable video decoder illustrated in FIG. 4. The inter-layer prediction and processing management unit includes one or more stages (e.g., three stages as illustrated in FIG. 5). In the first stage (e.g., stage 1), the reconstructed picture of the BL is enhanced (e.g., before being upsampled). In the second stage (e.g., stage 2), upsampling is performed (e.g., in spatial scalability, when the resolution of the BL is lower than that of the EL). The output of the second stage has substantially the same resolution as the resolution of the EL with the sampling grid adjusted. In the third stage (e.g., stage 3), enhancement is performed, for example, before the upsampled picture is placed in the EL DPB, which improves inter-layer reference picture quality.

[0026] It should be noted that one or more of the three stages described above are performed by the inter-layer prediction and processing management unit. For example, in signal-to-noise ratio (SNR) scalability, where the BL picture has substantially the same resolution as the EL picture but lower quality, one or more (e.g., all stages) of the three stages described above are not performed, and, for example, the BL reconstructed picture is directly inserted into the EL DPB for inter-layer prediction. For spatial scalability, the second stage is performed, for example, to make the upsampled BL reconstructed picture have a sampling grid adjusted with respect to the EL picture. The first and third stages are performed to improve inter-layer reference picture quality, which helps, for example, to achieve higher efficiency in EL coding.

[0027] Execution of picture-level ILP in a scalable video coding system (as shown in FIGS. 3 and 4) reduces implementation complexity, e.g., because the encoder and / or decoder logic of the base layer and / or enhancement layer, respectively, is reused, e.g., at least partially unchanged, at the block level. Higher-level (e.g., picture- and / or slice-level) configuration implements insertion of one or more respective processed base layer pictures into the enhancement layer DPB. To improve coding efficiency, in a scalable system, one or more block-level modifications are allowed, e.g., to facilitate block-level inter-layer prediction in addition to picture-level inter-layer prediction.

[0028] The single and / or multi-layer video coding systems described herein are used to code color video. In color video, each pixel, carrying luma and chroma information, is created from a combination of the respective intensities of primary colors (e.g., YCbCr, RGB, or YUV). Each video frame of color video is composed of three rectangular arrays corresponding to the three color channels. One or more samples within a color channel (e.g., each color channel) have discrete and / or finite magnitudes, which are represented in digital video applications using 8-bit values. In video capture and / or display systems, red, green, and blue (RGB) primaries are used.

[0029] In video coding and / or transmission, video signals in RGB space are converted to one or more other color spaces (e.g., having luminance and / or chromaticity coordinates), such as YUV for PAL and SECAM TV systems and YIQ for NTSC TV systems, for example, to reduce bandwidth consumption and / or for compatibility with monochrome video applications. The value of the Y component represents the brightness of the pixel, while the other two components (e.g., Cb and Cr) carry chromaticity information. Digital color spaces (e.g., YCbCr) are scaled and / or shifted versions of analog color spaces (e.g., YUV). The transformation matrix for deriving YCbCr coordinates from RGB coordinates is expressed as Equation (1).

[0030]

number

[0031] Because the human visual system (HVS) is less sensitive to color than to brightness, the chrominance components Cb and Cr can be subsampled with only a slight degradation in perceived video quality. Color subsampling formats are indicated by triplet numbers separated by colons. For example, when following a 4:2:2 color subsampling format, the horizontal sampling rate for the chrominance components is reduced by half, while the vertical sampling rate remains unchanged. When following a 4:2:0 color subsampling format, the sampling rate for the chrominance components is reduced by half both horizontally and vertically to reduce the associated data rate. When following a 4:4:4 color subsampling format, which is used for applications using very high video quality, the chrominance components have substantially the same sampling rate as that used for the luma component. Exemplary sampling grids showing luma and chrominance samples for the color subsampling formats described above are shown in Figures 6A through 6C, respectively.

[0032] The Y, Cb, and Cr color planes of a frame in a video sequence are correlated (e.g., highly correlated) in content, but the two chroma planes exhibit less texture and / or edges than the luma plane. The three color planes share the same motion. When a block-based hybrid video coding system (according to Figures 1 and 2) is applied to a color block, the three planes in the block are not coded separately. When the color block is coded by inter prediction, the two chroma blocks reuse the motion information of the luma block, such as motion vectors and / or reference indices. When the color block is coded by intra prediction, for example, the luma block has more prediction directions to choose from than one or both of the two chroma blocks because the luma block has more diverse and / or stronger edges.

[0033] For example, according to H.264 / AVC intra prediction, a luma block has nine candidate directions, while a chroma block has four candidate directions. According to HEVC intra prediction, a chroma block has four candidate directions, while a luma block has five or more candidate directions (e.g., 35 candidate directions). The respective transform and / or quantization processes for the luma and / or chroma prediction errors are performed separately, e.g., after intra prediction or inter prediction. At low bit rates (e.g., when QP for luma is greater than 34), chroma has lighter quantization (e.g., smaller quantization step size) than the corresponding luma, for example, because edges and / or textures in the chroma planes are more delicate and adversely affected by heavy quantization, which causes visible artifacts such as color bleeding.

[0034] A device configured to perform video coding (e.g., to encode and / or decode a video signal) is called a video coding device. Such video coding devices include video-enabled devices such as televisions, digital media players, DVD players, Blu-ray players, network-connected media player devices, desktop computers, laptop personal computers, tablet devices, mobile phones, video conferencing systems, or hardware- and / or software-based video encoding systems. Such video coding devices include wireless communication network elements such as wireless transmit / receive units (WTRUs), base stations, gateways, or other network elements.

[0035] The video coding device is configured to receive a video signal (e.g., a video bitstream) via a network interface. The video coding device has a wireless network interface, a wired network interface, or a combination thereof. For example, if the video coding device is a wireless communication network element (e.g., a wireless transmit / receive unit (WTRU)), the network interface is a transceiver of the WTRU. In another example, if the video coding device is a video-enabled device not configured for wireless communication (e.g., a back-end rack encoder), the network interface is a wired network connection (e.g., an optical fiber connection). In another example, the network interface is an interface configured to communicate with a physical storage medium (e.g., an optical disk drive, a memory card interface, or a direct connection to a video camera, etc.). It should be understood that the network interface is not limited to these examples and includes other interfaces that enable the video coding device to receive a video signal.

[0036] The video coding device is configured to perform cross-plane filtering on one or more video signals (eg, source video signals received by a network interface of the video coding device).

[0037] Cross-plane filtering is used, for example, to recover blurred edges and / or texture in one or both chroma planes using information from the corresponding luma plane. An adaptive cross-plane filter is implemented. For example, according to a threshold level of transmission performance of a bitstream associated with the video signal, the cross-plane filter coefficients are quantized and / or signaled such that overhead in the bitstream reduces (e.g., minimizes) performance degradation. The cross-plane filter coefficients are transmitted within the bitstream (output video bitstream) and / or out-of-band with respect to the bitstream.

[0038] One or more characteristics of the cross-plane filter (e.g., size, separability, symmetry, etc.) are determined to incur a reasonable overhead in the bitstream without causing performance degradation. Cross-plane filtering is applied to video having various color subsampling formats (e.g., including 4:4:4, 4:2:2, and 4:2:0). Cross-plane filtering is applied to selected regions of the video image (e.g., edge areas and / or one or more signals conveyed in the bitstream). Cross-plane filters are implemented in single-layer video coding systems. Cross-plane filters are implemented in multi-layer video coding systems.

[0039] The luma plane is used as guidance for improving the quality of one or both chroma planes. For example, one or more portions of information related to the luma plane are mixed into the corresponding chroma plane. For purposes of this disclosure, the three color planes of an original (e.g., uncoded) video image are represented by Y_org, Cb_org, and Cr_org, respectively, and the three color planes of a coded version of the original video image are represented by Y_rec, Cb_rec, and Cr_rec, respectively.

[0040] 7 illustrates an example of cross-plane filtering, where Y_rec, Cb_rec, and Cr_rec are used to convert back to RGB space, where the three planes are represented by R_rec, G_rec, and B_rec, respectively, using an inverse process (e.g., of process (1) shown above). Y_org, Cb_org, and Cr_org are then converted back to RGB space (e.g., substantially simultaneously) to obtain the original RGB planes represented by R_org, G_org, and B_org. A least-squares (LS) training method obtains plane pairs (R_org, R_rec), (G_org, G_rec), and (B_org, B_rec) as a training data set to train three filters for the R plane, G plane, and B plane, represented by filter_R, filter_G, and filter_B, respectively. By filtering R_rec, G_rec, and B_rec using filter_R, filter_G, and filter_B, respectively, three improved RGB planes, denoted as R_imp, G_imp, and B_imp, are obtained, and / or distortion between R_org and R_imp, G_org and G_imp, and B_org and B_imp is reduced (e.g., minimized) compared to the distortion between R_org and R_rec, G_org and G_rec, and B_org and B_rec, respectively. R_imp, G_imp, and B_imp are transformed into YCbCr space, and Y_imp, Cb_imp, and Cr_imp are obtained, where Cb_imp and Cr_imp are the outputs of the cross-plane filtering process.

[0041] 7 consumes computational resources (e.g., an undesirably large amount of computational resources) on one or both of the encoder and / or decoder sides. Because both the spatial conversion process and the filtering process are linear, at least a portion of the illustrated cross-plane filtering procedure may be approximated using a simplified process, for example, in which one or more of the operations (e.g., all of the operations) are performed in YCbCr space.

[0042] As shown in FIG. 8A, to improve the quality of Cb_rec, the LS training module takes Y_rec, Cb_rec, Cr_rec, and Cb_org as a training data set, and jointly derived optimal filters filter_Y4Cb, filter_Cb4Cb, and filter_Cr4Cb are applied to Y_rec, Cb_rec, and Cr_rec, respectively. The respective outputs of filtering on the three planes are summed together to obtain an improved Cb plane, e.g., denoted as Cb_imp. The three optimal filters are trained by the LS method, and the distortion between Cb_imp and Cb_org is minimized, e.g., according to Equation (2):

[0043]

number

[0044] where:

[0045]

number

[0046] denotes two-dimensional (2-D) convolution, + and - denote matrix addition and subtraction, respectively, and E[(X) 2 ] represents the average of the squares of each element of matrix X.

[0047] As shown in FIG. 8B, to improve the quality of Cr_rec, the LS training module takes Y_rec, Cb_rec, Cr_rec, and Cr_org as a training data set, and jointly derived optimal filters filter_Y4Cr, filter_Cb4Cr, and filter_Cr4Cr are applied to Y_rec, Cb_rec, and Cr_rec, respectively. The outputs of each of the filtering on the three planes are summed together to obtain an improved Cr plane, e.g., denoted as Cr_imp. The three optimal filters are trained by the LS method, and the distortion between Cr_imp and Cr_org is minimized, e.g., according to Equation (3).

[0048]

number

[0049] Cr contributes only slightly to improving Cb. Cb contributes only slightly to improving Cr.

[0050] The cross-plane filtering technique shown in Figures 8A and 8B may be simplified. For example, as shown in Figure 9A, in LS training, the quality of the Cb plane is improved by utilizing the Y plane and the Cb plane, but not by utilizing the Cr plane, and two filters, filter_Y4Cb and filter_Cb4Cb, are derived together and applied to Y and Cb, respectively. The outputs of each of the filters are summed together to obtain an improved Cb plane, for example, denoted as Cb_imp.

[0051] 9B, in LS training, the quality of the Cr plane is improved by utilizing the Y and Cr planes, but not by utilizing the Cb plane, and two filters, filter_Y4Cr and filter_Cr4Cr, are derived together and applied to Y and Cr, respectively. The outputs of each of the filters are summed together to obtain an improved Cr plane, e.g., denoted as Cr_imp.

[0052] The cross-plane filtering technique shown in Figures 9A and 9B reduces the computational complexity of training and / or filtering, respectively, and / or reduces the overhead bits of transmitting cross-plane filter coefficients to the decoder side, with only a slight performance degradation.

[0053] To perform cross-plane filtering in a video coding system, one or more of determining the cross-plane filter size, quantizing and / or transmitting (e.g., signaling) the cross-plane filter coefficients, or adapting the cross-plane filtering to one or more local areas are addressed.

[0054] To train an optimal cross-plane filter, an appropriate filter size is determined. The filter size is roughly proportional to the overhead associated with the filter and / or the computational complexity of the filter. For example, a 3×3 filter has 9 transmitted filter coefficients and utilizes 9 multiplications and 8 additions to filter one pixel. A 5×5 filter has 25 transmitted filter coefficients and utilizes 25 multiplications and 24 additions to filter one pixel. Larger size filters achieve lower minimum distortion and / or provide better performance (e.g., as seen in Equations (2) and (3)). The filter size is selected, for example, to balance computational complexity, overhead, and / or performance.

[0055] Trained filters applied to the planes themselves, such as filter_Cb4Cb and filter_Cr4Cr, are implemented as low-pass filters. Trained filters used for cross-planes, such as filter_Y4Cb, filter_Y4Cr, filter_Cb4Cr, and filter_Cr4Cb, are implemented as high-pass filters. The use of different filters of different sizes has little impact on the performance of the corresponding video coding system. The size of the cross-plane filters is kept small (e.g., as small as possible), e.g., so that the performance penalty is negligible. For example, the cross-plane filter size is selected so that substantially no performance degradation is observed. A large-sized cross-plane filter (e.g., an M×N cross-plane filter, where M and N are integers) is implemented.

[0056] For example, for low-pass filters such as filter_Cb4Cb and filter_Cr4Cr, the filter size is implemented as 1×1, and the filter has one coefficient that is multiplied with each pixel being filtered. The filter coefficients of 1×1 filter_Cb4Cb and filter_Cr4Cr are fixed at 1.0, and filter_Cb4Cb and filter_Cr4Cr are omitted (e.g., not applied and / or not transmitted).

[0057] For high-pass filters such as filter_Y4Cb and filter_Y4Cr, the filter size depends on the color sampling format or is independent of the color sampling format. The cross-plane filter size depends on the color sampling format. For example, the size and / or support area of ​​the cross-plane filters (e.g., filter_Y4Cb and filter_Y4Cr) are implemented for selected chroma pixels, for example, as shown in Figures 10A to 10C, where circles represent the respective positions of luma samples, filled triangles represent the respective positions of chroma samples, and luma samples used to filter the selected chroma samples (e.g., represented by outline-only triangles) are represented by gray circles. As shown, the filter size of filter_Y4Cb and filter_Y4Cr is 3x3 for 4:4:4 and 4:2:2 color formats and 4x3 for 4:2:0 color format. The filter size is independent of the color format, for example, as shown in Figures 11A to 11C. The filter size is, for example, 4x3, according to the size for the 4:2:0 format.

[0058] For example, according to equations (4) and (5), the cross-plane filtering process applies the trained high-pass filter to the Y plane and obtains the filtering results represented by Y_offset4Cb and Y_offset4Cr as offsets to be added to the corresponding pixels in the chroma planes.

[0059]

number

[0060]

number

[0061] The cross-plane filter coefficients are quantized. The trained cross-plane filter has real-valued coefficients that are, for example, quantized before transmission. For example, filter_Y4Cb is roughly approximated by an integer filter, denoted filter_int. The elements of filter_int have a small dynamic range (e.g., from -8 to 7 in a 4-bit representation). A second coefficient, denoted coeff., is used to more accurately approximate filter_int to filter_Y4Cb, for example, according to equation (6).

[0062]

number

[0063] In equation (6), coeff. is a real value, for example, M / 2 according to equation (7). N where M and N are integers.

[0064]

number

[0065] To transmit filter_Y4Cb, for example, the coefficients in filter_int are coded into the bitstream together with M and N. The quantization techniques described above are extended to quantize filter_Y4Cr, for example.

[0066] The cross-plane filters (e.g., filter_Y4Cb and / or filter_Y4Cr) have flexible separability and / or symmetry. The cross-plane filter characteristics introduced herein are described with respect to an exemplary 4×3 cross-plane filter (e.g., according to FIGS. 10A-10C or 11A-11C), but are applicable to other filter sizes.

[0067] Cross-plane filters have various symmetry properties, for example, as shown in Figures 12A through 12E. Cross-plane filters have no symmetry, for example, as shown in Figure 12A. Each square represents one filter coefficient and is labeled with a unique index indicating that its value is different from that of the remaining filter coefficients. Cross-plane filters have horizontal and vertical symmetry, for example, as shown in Figure 12B, where a coefficient has the same value as one or more corresponding coefficients in one or more other quadrants. Cross-plane filters have vertical symmetry, for example, as shown in Figure 12C. Cross-plane filters have horizontal symmetry, for example, as shown in Figure 12D. Cross-plane filters have point symmetry, for example, as shown in Figure 12E.

[0068] 12A-12E, the cross-plane filter is not limited to the symmetry shown in Figure 12A-12E, and may have one or more other symmetries. A cross-plane filter has symmetry if at least two coefficients in the filter have the same value (e.g., at least two coefficients are labeled with the same index). For example, for a high-pass cross-plane filter (e.g., filter_Y4Cb and filter_Y4Cr), it is beneficial to implement no symmetry for one or more (e.g., all) coefficients along the boundary of the filter support region, but to implement some symmetry (e.g., horizontal and vertical, horizontal, vertical, or point symmetry) for one or more (e.g., all) of the coefficients within the filter support region.

[0069] Cross-plane filters are separable. For example, cross-plane filtering using a 4x3 two-dimensional filter is equivalent to applying a 1x3 horizontal filter to the rows (e.g., during the first stage) and a 4x1 vertical filter to the columns of the output of the first stage (e.g., during the second stage). The order of the first and second stages is reversed. Symmetry is applied to the 1x3 horizontal filter and / or the 4x1 vertical filter. Figures 13A and 13B show two one-dimensional filters with and without symmetry, respectively.

[0070] Regardless of whether a cross-plane filter is separable and / or symmetric, the coding of filter coefficients into the bitstream is limited to filter coefficients with unique values. For example, according to the cross-plane filter shown in Figure 12A, 12 filter coefficients (indexed from 0 to 11) are coded. According to the cross-plane filter shown in Figure 12B, 4 filter coefficients (indexed from 0 to 3) are coded. Enforcing symmetry in the cross-plane filter reduces the amount of overhead (e.g., in the video signal bitstream).

[0071] For example, if the cross-plane filters (e.g., filter_Y4Cb and filter_Y4Cr) are high-pass filters, the sum of the filter coefficients of the cross-plane filters is equal to zero. According to this property of being constant, a coefficient (e.g., at least one coefficient) in the cross-plane filter has a magnitude equal to the sum of the other coefficients but has an opposite sign. If the cross-plane filter has X coefficients to be transmitted (e.g., X is equal to 12, as shown in FIG. 12A), X−1 coefficients are coded (e.g., explicitly coded) in the bitstream. The decoder receives the X−1 coefficients and derives (e.g., implicitly derives) values ​​for the remaining coefficients, for example, based on a zero-sum constraint.

[0072] The cross-plane filtering coefficients are signaled, for example, within a video bitstream. The exemplary syntax table of Figure 14 shows an example of signaling a set of two-dimensional non-separable asymmetric cross-plane filter coefficients for a chroma plane (e.g., Cb or Cr). The following applies to the entries in the exemplary syntax table: The entry num_coeff_hori_minus1 plus 1 (+1) indicates the number of coefficients in the horizontal direction of the cross-plane filter. The entry num_coeff_vert_minus1 plus 1 (+1) indicates the number of coefficients in the vertical direction of the cross-plane filter. The entry num_coeff_reduced_flag equal to 0 indicates that the number of cross-plane filter coefficients is equal to (num_coeff_hori_minus1 + 1) × (num_coeff_vert_minus1 + 1), for example, as shown in Figure 15A. As shown in Figure 15A, num_coeff_hori_minus1 is equal to 2, and num_coeff_vert_minus1 is equal to 3.

[0073] The entry num_coeff_reduced_flag equal to 1 indicates that the number of cross-plane filter coefficients, which is generally equal to (num_coeff_hori_minus1+1)×(num_coeff_vert_minus1+1), is reduced to (num_coeff_hori_minus1+1)×(num_coeff_vert_minus1+1)−4, e.g., by removing the four corner coefficients, as shown in FIG. 15B. The support region of the cross-plane filter is reduced, e.g., by removing the four corner coefficients. The use of the num_coeff_reduced_flag entry provides, e.g., increased flexibility as to whether filter coefficients are reduced.

[0074] The entry filter_coeff_plus8[i] minus 8 corresponds to the i-th cross-plane filter coefficient. The value of the filter coefficient is, for example, in the range of -8 to 7. In such a case, the entry filter_coeff_plus8[i] is in the range of 0 to 15 and is, for example, coded according to 4-bit fixed-length coding (FLC). The entries scaling_factor_abs_minus1 and scaling_factor_sign together specify the value of the scaling factor (e.g., M in equation (7)), as follows:

[0075] M=(1-2×scaling_factor_sign)×(scaling_factor_abs_minus1+1) The entry bit_shifting specifies the number of bits to be right-shifted after the scaling process. This entry represents N in equation (7).

[0076] Different regions of a picture have different statistical characteristics. Deriving cross-plane filter coefficients for one or more such regions (e.g., each such region) improves chroma coding performance. To illustrate, different sets of cross-plane filter coefficients are applied to different regions of a picture or slice, for which multiple sets of cross-plane filter coefficients are transmitted at the picture level (e.g., in an adaptive picture set (APS)) and / or slice level (e.g., in a slice header).

[0077] If cross-plane filtering is used in a post-processing implementation applied to the reconstructed video before the video is displayed, for example, one or more sets of filter coefficients are transmitted as a supplemental enhancement information (SEI) message. For each color plane, the total number of filters is signaled. If the number is one or more, one or more sets of cross-plane filtering coefficients are transmitted, for example, sequentially.

[0078] The example syntax table of Figure 16 shows an example of conveying multiple sets of cross-plane filter coefficients in an SEI message named cross_plane_filter(). The following applies to the entries in the example syntax table: An entry cross_plane_filter_enabled_flag equal to 1 specifies that cross-plane filtering is enabled. In contrast, an entry cross_plane_filter_enabled_flag equal to 0 specifies that cross-plane filtering is disabled.

[0079] The entry cb_num_of_filter_sets specifies the number of cross-plane filter coefficient sets used to code the Cb plane of the current picture. An entry cb_num_of_filter_sets equal to 0 indicates that cross-plane filtering is not applied to the Cb plane of the current picture. The entry cb_filter_coeff[i] is the i-th set of cross-plane filter coefficients for the Cb plane. The entry cb_filter_coeff is a data structure and includes one or more of num_coeff_hori_minus1, num_coeff_vert_minus1, num_coeff_reduced_flag, filter_coeff_plus8, scaling_factor_abs_minus1, scaling_factor_sign, or bit_shifting.

[0080] The entry cr_num_of_filter_sets specifies the number of cross-plane filter coefficient sets used to code the Cr plane of the current picture. An entry cr_num_of_filter_sets equal to 0 indicates that cross-plane filtering is not applied to the Cr plane of the current picture. The entry cr_filter_coeff[i] is the i-th set of cross-plane filter coefficients for the Cr plane. The entry cr_filter_coeff is a data structure and includes one or more of num_coeff_hori_minus1, num_coeff_vert_minus1, num_coeff_reduced_flag, filter_coeff_plus8, scaling_factor_abs_minus1, scaling_factor_sign, or bit_shifting.

[0081] Region-based cross-plane filtering is performed. For example, if it is desired to recover the loss of high-frequency information in the associated chroma plane (e.g., following the guidance of the luma plane), cross-plane filtering is adapted to filter one or more local areas within a video image. For example, cross-plane filtering is applied to areas rich in edges and / or textures. For example, edge detection is first performed to find one or more regions to which the cross-plane filter is applied. A high-pass filter, such as filter_Y4Cb and / or filter_Y4Cr, is first applied to the Y plane.

[0082] The magnitude of the filtering result indicates whether the filtered pixel is in a high-frequency area. A large magnitude indicates a sharp edge in the filtered pixel's region. A magnitude close to 0 indicates the filtered pixel is in a uniform region. A threshold is used to measure the filtering output by filter_Y4Cb and / or filter_Y4Cr. If the filtering output is greater than the threshold, it is added to the corresponding pixel in the chroma plane. For example, each chroma pixel in a smooth region is left unchanged, which avoids random filtering noise. Region-based cross-plane filtering reduces the complexity of video coding while maintaining coding performance. For example, region information including one or more regions is transmitted to the decoder.

[0083] In a region-based cross-plane filtering implementation, one or more regions with different statistical characteristics (e.g., smooth regions, colorful regions, texture-rich regions, and / or edge-rich regions) are detected, for example, at the encoder side. Multiple cross-plane filters are derived and applied to corresponding ones of the one or more regions. Information about each of the one or more regions is transmitted to the decoder side. Such information includes, for example, the area of ​​the region, the position of the region, and / or the specific cross-plane filter to be applied to the region.

[0084] The exemplary syntax table of Figure 17 shows an example of conveying information about a particular region. The following applies to entries in the exemplary syntax table: The entries top_offset, lef_offset, righ_offset, and bottom_offset specify the area and / or location of the current region. The entries represent the distance, e.g., in pixels, from the top, left, right, and bottom edges of the current region to the corresponding four edges of the associated picture, as shown in Figure 18, for example.

[0085] cross_plane_filtering_region_info() contains information about cross-plane filtering of a specified region in the Cb plane, cross-plane filtering of a specified region in the Cr plane, or cross-plane filtering of a specified region in each of the Cb and Cr planes.

[0086] The entry cb_filtering_enabled_flag equal to 1 indicates that cross-plane filtering for the current region of the Cb plane is enabled. The entry cb_filtering_enabled_flag equal to 0 indicates that cross-plane filtering for the current region of the Cb plane is disabled. The entry cb_filter_idx specifies that the cross-plane filter cb_filter_coeff[cb_filter_idx] (which conveys, for example, cb_filter_coeff as shown in Figure 16) is applied to the current region of the Cb plane.

[0087] The entry cr_filtering_enabled_flag equal to 1 indicates that cross-plane filtering for the current region of the Cr plane is enabled. The entry cr_filtering_enabled_flag equal to 0 indicates that cross-plane filtering for the current region of the Cr plane is disabled. The entry cr_filter_idx specifies that the cross-plane filter cr_filter_coeff[cr_filter_idx] (which conveys, for example, cr_filter_coeff as shown in Figure 16) is applied to the current region of the Cr plane.

[0088] Information about one or more regions is transmitted at the picture level (e.g., in an APS or SEI message) or at the slice level (e.g., in a slice header). The example syntax table in Figure 19 shows an example of conveying multiple regions along with multiple cross-plane filters in an SEI message named cross_plane_filter(). Information about the regions is written in italics.

[0089] The following applies to entries in an example syntax table: The entry cb_num_of_regions_minus1 plus 1 (+1) specifies the number of regions in the Cb plane. Each region is filtered by the corresponding cross-plane filter. The entry cb_num_of_regions_minus1 equal to 0 indicates that the entire Cb plane is filtered by one cross-plane filter. The entry cb_region_info[i] is the information of the i-th region in the Cb plane. The entry cb_region_info is a data structure and includes one or more of top_offset, left_offset, right_offset, bottom_offset, cb_filtering_enabled_flag, or cb_filter_idx.

[0090] The entry cr_num_of_regions_minus1 plus 1 (+1) specifies the number of regions in the Cr plane. Each region is filtered by the corresponding cross-plane filter. An entry cr_num_of_regions_minus1 equal to 0 indicates that the entire Cr plane is filtered by one cross-plane filter. The entry cr_region_info[i] is the information of the i-th region in the Cr plane. The entry cr_region_info is a data structure and includes one or more of top_offset, left_offset, right_offset, bottom_offset, cr_filtering_enabled_flag, or cr_filter_idx.

[0091] Cross-plane filtering is used in single-layer video coding systems and / or multi-layer video coding systems. In a single-layer video coding system (e.g., as shown in FIGS. 1 and 2), cross-plane filtering is applied, for example, to improve reference pictures (e.g., pictures stored in reference picture store 164 and / or 264) so ​​that one or more subsequent frames are better predicted (e.g., with respect to chroma planes).

[0092] Cross-plane filtering is used as a post-processing method. For example, cross-plane filtering is applied to the reconstructed output video 220 (e.g., before it is displayed). Although such filtering is not part of the MCP loop and therefore does not affect the coding of subsequent pictures, the post-processing improves (e.g., directly) the quality of the video for display. For example, cross-plane filtering is applied in HEVC post-processing using supplemental enhancement information (SEI) signaling. Cross-plane filter information estimated at the encoder side is delivered, for example, in the SEI message.

[0093] According to an example using multi-layer video coding (e.g., as shown in Figures 3 and 4), cross-plane filtering is applied to one or more upsampled BL pictures to predict pictures of higher layers, e.g., before the one or more pictures are placed in an EL DPB buffer (e.g., a reference picture list). As shown in Figure 5, cross-plane filtering is performed in a third stage. To improve the quality of one or both chroma planes in an upsampled base layer reconstructed picture (e.g., an ILP picture), the corresponding luma planes involved in training and / or filtering are from the same ILP picture, where the training and / or filtering process is the same as that used in single-layer video coding.

[0094] According to another example using multi-layer video coding, to support cross-plane training and / or filtering, for example, to enhance a chroma plane in an ILP picture, the corresponding luma plane is used (e.g., directly) in the base layer reconstructed picture without upsampling. For example, according to 2X spatial SVC using a 4:2:0 video source, the size of the base layer luma plane is substantially the same (e.g., exactly the same) as the size of one or both corresponding chroma planes in the ILP picture. The sampling grids of the two types of planes are different. For example, the luma plane in the base layer picture is filtered by a phase correction filter to align (e.g., precisely align) with the sampling grid of the chroma plane in the ILP picture. One or more of the following operations are the same as those described elsewhere herein, for example, for single-layer video coding. The color format is assumed to be 4:4:4 (e.g., according to FIG. 10A or FIG. 11A). The use of the base layer luma plane to support cross-plane filtering for chroma planes in ILP pictures is extended to other ratios of spatial scalability and / or other color formats, for example, by simple derivations.

[0095] According to another example using multi-layer video coding, cross-plane filtering is applied to a reconstructed base layer picture that has not been upsampled. The output of the cross-plane filtering is upsampled. As shown in Figure 5, cross-plane filtering is performed in the first stage. In the case of spatial scalability (e.g., BL has a lower resolution than EL), cross-plane filtering is applied to fewer pixels, which involves lower computational complexity than one or more of the other multi-layer video coding examples described herein. For example, referring to equation (2),

[0096]

number

[0097] Equations (2) and (3) do not apply directly because Y have different dimensions and cannot be directly subtracted. rec , Cb rec , Cr rec has the same resolution as the base layer picture. org has the same resolution as the enhancement layer picture. The derivation of the cross-plane filter coefficients according to this example of multi-layer video coding is achieved using equations (8) and (9),

[0098]

number

[0099]

number

[0100] where U is an upsampling function that takes a base layer picture as input and outputs an upsampled picture with the enhancement layer resolution.

[0101] According to the cross-plane filtering technique shown in Figures 9A and 9B, a chroma plane is enhanced by a luma plane and by itself (e.g., excluding the other chroma plane), and equations (8) and (9) are simplified, for example, as shown in equations (10) and (11).

[0102]

number

[0103]

number

[0104] 9A and 9B, the size of filter_Cb4Cb and / or filter_Cr4Cr is reduced to 1×1, and the value of the filter coefficient is set to 1.0. Equations (10) and (11) are simplified, for example, as shown in Equations (12) and (13).

[0105]

number

[0106]

number

[0107] Cross-plane filtering is adaptively applied, for example, when applied to multi-layer video coding, in the first stage and / or the third stage as shown in FIG.

[0108] Cross-plane filtering may be adaptively applied at one or more coding levels, including, for example, one or more of the sequence level, picture level, slice level, or block level. According to sequence-level adaptation, for example, the encoder determines to utilize cross-plane filtering in the first stage and / or the third stage to code a portion of a video sequence (e.g., the entire video sequence). Such a decision may be represented, for example, as a binary flag included in a sequence header and / or one or more sequence-level parameter sets, such as a video parameter set (VPS) and / or a sequence parameter set (SPS).

[0109] According to picture-level adaptation, for example, the encoder decides to utilize cross-plane filtering in the first stage and / or the third stage to code one or more EL pictures (e.g., each EL picture in a video sequence), and such decision is represented, for example, as a binary flag included in the picture header and / or included in one or more picture-level parameter sets, such as an adaptation parameter set (APS) and / or a picture parameter set (PPS).

[0110] According to slice-level adaptation, for example, the encoder may decide to utilize cross-plane filtering in the first stage and / or the third stage to code one or more EL video slices (e.g., each EL slice). Such a decision may be represented, for example, as a binary flag included in the slice header. The signaling mechanism described above may be implemented (e.g., extended) according to one or more other levels of adaptation.

[0111] Picture-based cross-plane filtering is implemented, for example, for multi-layer video coding. Information regarding such cross-plane filtering is signaled. For example, one or more flags, such as uplane_filtering_flag and / or vplane_filtering_flag, are coded and transmitted to a decoder, for example, once per picture. The flags uplane_filtering_flag and / or vplane_filtering_flag indicate, for example, whether cross-plane filtering should be applied to the Cb plane and / or whether cross-plane filtering should be applied to the Cr plane, respectively. The encoder determines (e.g., per picture) for which chroma planes of one or more pictures cross-plane filtering should be enabled or disabled. The encoder is configured to make such a decision, for example, to improve coding performance and / or according to a desired level of coding performance and complexity (e.g., turning on cross-plane filtering increases decoding complexity).

[0112] The encoder is configured to utilize one or more techniques to determine whether to apply picture-based cross-plane filtering to one or more chroma planes. For example, according to an example of performing picture-level selection, the Cb plane before and after filtering, e.g., Cb_rec and Cb_imp, is compared with the original Cb plane in the EL picture, e.g., Cb_org. Mean square error (MSE) values ​​before and after filtering, denoted as MSE_rec and MSE_imp, respectively, are calculated and compared. In the example, MSE_imp is smaller than MSE_rec, indicating that applying cross-plane filtering reduces distortion, and cross-plane filtering is enabled on the Cb plane. If MSE_imp is not smaller than MSE_rec, cross-plane filtering is disabled on the Cb plane. According to this technique, MSE is calculated on a whole-picture basis, which means that a single weighting factor is applied to one or more pixels (e.g., each pixel) in the MSE calculation.

[0113] According to another example of performing picture-level selection, the MSE is calculated based on one or more pixels involved in ILP, e.g., based only on pixels involved in ILP. When the encoder determines whether to apply cross-plane filtering to the Cb plane, the ILP map for the picture is not yet available. For example, the decision is made before coding the EL picture, but the ILP map is not available until the EL picture is coded.

[0114] Another example of picture-level selection is to use a multi-pass coding strategy. In a first pass, an EL picture is coded and an ILP map is recorded. In a second pass, a decision is made as to whether cross-plane filtering should be used, for example, according to an MSE calculation limited to the ILP blocks indicated by the ILP map. The picture is coded according to this decision. Such multi-pass coding is time-consuming and requires greater computational complexity than single-pass coding.

[0115] One or more moving objects in each picture (e.g., each picture of a video sequence) are more likely to be coded by an ILP picture than non-moving objects. The ILP maps of consecutive pictures (e.g., consecutive pictures of a video sequence) are correlated (e.g., exhibit a high degree of correlation). Such consecutive ILP maps exhibit one or more displacements (e.g., relatively small displacements) between each other. Such displacements may be attributed, for example, to different time instances of the pictures.

[0116] According to another example of performing picture-level selection, the ILP maps of one or more previously coded EL pictures are used to predict the ILP map of the current EL picture being coded. The predicted ILP map is used to find one or more blocks that are likely to be used for ILP in coding the current EL picture. Such likely-to-be-used blocks are referred to as potential ILP blocks. The one or more potential ILP blocks are included in the MSE calculation (e.g., as described above) and / or used in determining whether to apply cross-plane filtering, for example, based on the calculated MSE.

[0117] The dimensions of the ILP map depend, for example, on the granularity selected by the encoder. If the dimensions of a picture are W×H (e.g., in terms of luma resolution), then, for example, the dimensions of the ILP map are W×H, and entries therein indicate whether corresponding pixels are used for ILP. The dimensions of the ILP map are (W / M)×(H / N), and entries therein indicate whether corresponding blocks of size M×N are used for ILP. According to an exemplary implementation, M=N=4 is selected.

[0118] For example, a precise ILP map, recorded after an EL picture is coded, is a binary map in which entries (e.g., each entry) are restricted to one of two possible values ​​(e.g., 0 or 1) that indicate whether the entry is used for ILP. Values ​​of 0 and 1 indicate, for example, that the entry is used for ILP or is not used for ILP, respectively.

[0119] The predicted ILP map is a multi-level map. According to such an ILP map, each entry has multiple possible values ​​that represent a multi-level confidence in predicting the block used for ILP. A larger value indicates a higher confidence. According to an exemplary implementation, possible values ​​of the predicted ILP map ranging from 0 to 128 are used, with 128 representing the highest confidence and 0 representing the lowest confidence.

[0120] 20 shows an example picture-level selection algorithm 2000 for cross-plane filtering. The illustrated picture-level selection algorithm is applied, for example, to the Cb plane and / or the Cr plane. At 2010, for example, before encoding the first picture, a predicted ILP map, represented by PredILPMap, is initialized. According to the illustrated algorithm, it is assumed that each block has an equal probability of being used for ILP, and the value of each entry in PredILPMap is set to 128.

[0121] At 2020, the encoder determines whether to apply cross-plane filtering. An enhanced Cb plane, Cb_imp, is generated by cross-plane filtering. The weighted MSE is calculated using, for example, equations (14) and (15).

[0122]

number

[0123]

number

[0124] In equations (14) and (15), Cb_rec and Cb_imp represent the Cb plane before and after cross-plane filtering, Cb_org represents the original Cb plane of the current EL picture being coded, and (x, y) represents the location of a pixel in the grid of the luma plane. As shown, equations (14) and (15) assume 4:2:0 color subsampling, and the entries in the ILP map represent a 4x4 block size, so the corresponding locations in the Cb plane and PredILPMap are (x / 2, y / 2) and (x / 4, y / 4), respectively. For each pixel, the squared error (Cb_imp(x / 2, y / 2) - Cb_org(x / 2, y / 2)) 2 or (Cb_rec(x / 2,y / 2)-Cb_org(x / 2,y / 2)) 2 For example, the error is weighted by the corresponding factor in PredILPMap before being accumulated in Weighted_MSE_imp or Weighted_MSE_rec. This means that distortions on one or more pixels that are more likely to be used for ILP have a higher weight in the weighted MSE.

[0125] Alternatively or additionally, at 2020, an enhanced Cr plane Cr_imp is generated by cross-plane filtering. The weighted MSE is calculated using, for example, equations (16) and (17).

[0126]

number

[0127]

number

[0128] In equations (16) and (17), Cr_rec and Cr_imp represent the Cr plane before and after cross-plane filtering, Cr_org represents the original Cr plane of the current EL picture being coded, and (x, y) represents the location of a pixel in the grid of the luma plane. As shown, equations (16) and (17) assume 4:2:0 color subsampling, and the entries in the ILP map represent a 4x4 block size, so the corresponding locations in the Cr plane and PredILPMap are (x / 2, y / 2) and (x / 4, y / 4), respectively. For each pixel, the squared error (Cr_imp(x / 2, y / 2) - Cr_org(x / 2, y / 2)) 2 or (Cr_rec(x / 2,y / 2)-Cr_org(x / 2,y / 2)) 2 For example, the error is weighted by the corresponding factor in PredILPMap before being accumulated in Weighted_MSE_imp or Weighted_MSE_rec. This means that distortions on one or more pixels that are more likely to be used for ILP have a higher weight in the weighted MSE.

[0129] Weighted_MSE_imp and Weighted_MSE_rec are compared to each other. If Weighted_MSE_imp is smaller than Weighted_MSE_rec, it indicates that cross-plane filtering reduces distortion (e.g., distortion of one or more potential ILP blocks), and cross-plane filtering is enabled. If Weighted_MSE_imp is not smaller than Weighted_MSE_rec, cross-plane filtering is disabled.

[0130] Once the determination is made in 2020, the current EL picture is coded in 2030, and the current ILP map, represented by CurrILPMap, is recorded in 2040. The current ILP map is used, for example, with EL pictures that follow the current EL picture. The current ILP map is accurate rather than predicted and is binary. If the corresponding block is used for ILP, the value of the entry for that block is set to 128. If the corresponding block is not used for ILP, the value of the entry for that block is set to 0.

[0131] At 2050, the current ILP map is used to update the predicted ILP map, for example, as shown in equation 18. According to an exemplary update process, the sum of the previously predicted ILP map (e.g., PredILPMap(x,y)) and the current ILP map (e.g., CurrILPMap(x,y)) is divided by 2, which means that ILP maps associated with other pictures have a relatively small influence on the updated predicted ILP map.

[0132]

number

[0133] At 2060, it is determined whether the end of the video sequence has been reached. If the end of the video sequence has not been reached, one or more of the operations described above (e.g., 2020 through 2060) are repeated, e.g., to code successive EL pictures. If the end of the video sequence has been reached, at 2070, the exemplary picture level selection algorithm 2000 ends.

[0134] For example, the video coding techniques described herein that utilize cross-plane filtering may be implemented by transporting video in a wireless communication system such as the exemplary wireless communication system 2100 and its components shown in Figures 21A through 21E.

[0135] 21A is a diagram of an example communication system 2100 in which one or more disclosed embodiments may be implemented. For example, a wireless network (e.g., a wireless network comprising one or more components of communication system 2100) may be configured such that bearers extending beyond the wireless network (e.g., beyond a walled garden associated with the wireless network) are assigned QoS characteristics.

[0136] The communications system 2100 is a multiple-access system that provides content, such as voice, data, video, messaging, and broadcast, to multiple wireless users. The communications system 2100 enables the multiple wireless users to access such content through sharing of system resources, including wireless bandwidth. For example, the communications system 2100 may utilize one or more channel access methods, such as code division multiple access (CDMA), time division multiple access (TDMA), frequency division multiple access (FDMA), orthogonal FDMA (OFDMA), and single-carrier FDMA (SC-FDMA).

[0137] 21A, the communications system 2100 includes multiple wireless transmit / receive units (WTRUs), such as at least one WTRU, e.g., WTRUs 2102a, 2102b, 2102c, 2102d, a radio access network (RAN) 2104, a core network 2106, a public switched telephone network (PSTN) 2108, the Internet 2110, and other networks 2112, although it should be understood that the disclosed embodiments contemplate any number of WTRUs, base stations, networks, and / or network elements. Each of the WTRUs 2102a, 2102b, 2102c, 2102d is any type of device configured to operate and / or communicate in a wireless environment. For example, WTRUs 2102a, 2102b, 2102c, 2102d are configured to transmit and / or receive wireless signals and include user equipment (UE), mobile stations, fixed or mobile subscriber units, pagers, cellular telephones, personal digital assistants (PDAs), smartphones, laptops, netbooks, personal computers, wireless sensors, and home appliances.

[0138] The communications system 2100 also includes a base station 2114a and a base station 2114b. Each of the base stations 2114a, 2114b is any type of device configured to wirelessly interface with at least one of the WTRUs 2102a, 2102b, 2102c, 2102d to facilitate access to one or more communications networks, such as the core network 2106, the Internet 2110, and / or the network 2112. By way of example, the base stations 2114a, 2114b may be a base transceiver station (BTS), a Node B, an eNodeB, a Home Node B, a Home eNodeB, a site controller, an access point (AP), a wireless router, etc. While the base stations 2114a, 2114b are each shown as a single element, it should be understood that the base stations 2114a, 2114b may include any number of interconnected base stations and / or network elements.

[0139] The base station 2114a is part of the RAN 2104, which also includes other base stations and / or network elements (not shown), such as base station controllers (BSCs), radio network controllers (RNCs), relay nodes, etc. The base station 2114a and / or base station 2114b are configured to transmit and / or receive radio signals within a particular geographic area called a cell (not shown). A cell is further divided into cell sectors. For example, the cell associated with the base station 2114a is divided into three sectors. Thus, in one embodiment, the base station 2114a includes three transceivers, i.e., one for each sector of the cell. In another embodiment, the base station 2114a utilizes multiple-input multiple-output (MIMO) technology and thus utilizes multiple transceivers for each sector of the cell.

[0140] The base stations 2114a, 2114b communicate with one or more of the WTRUs 2102a, 2102b, 2102c, 2102d over an air interface 2116, which may be any suitable wireless communication link (e.g., radio frequency (RF), microwave, infrared (IR), ultraviolet (UV), visible light, etc.). The air interface 2116 may be established using any suitable radio access technology (RAT).

[0141] More specifically, as mentioned above, the communication system 2100 is a multiple-access system and utilizes one or more channel access schemes, such as CDMA, TDMA, FDMA, OFDMA, and SC-FDMA. For example, the base station 2114a and the WTRUs 2102a, 2102b, and 2102c in the RAN 2104 implement a radio technology such as Universal Mobile Telecommunications System (UMTS) Terrestrial Radio Access (UTRA), which establishes the air interface 2116 using Wideband CDMA (WCDMA). WCDMA includes communication protocols such as High Speed ​​Packet Access (HSPA) and / or Evolved HSPA (HSPA+). HSPA includes High Speed ​​Downlink Packet Access (HSDPA) and / or High Speed ​​Uplink Packet Access (HSUPA).

[0142] In another embodiment, the base station 2114a and the WTRUs 2102a, 2102b, 2102c implement a radio technology such as Evolved UMTS Terrestrial Radio Access (E-UTRA), which establishes the air interface 2116 using Long Term Evolution (LTE) and / or LTE-Advanced (LTE-A).

[0143] In other embodiments, the base station 2114a and the WTRUs 2102a, 2102b, 2102c implement a radio technology such as IEEE 802.16 (i.e., Worldwide Interoperability for Microwave Access (WiMAX)), CDMA2000, CDMA2000 1X, CDMA2000 EV-DO, Interim Standard 2000 (IS-2000), Interim Standard 95 (IS-95), Interim Standard 856 (IS-856), Global System for Mobile Communications (GSM), Enhanced Data Rates for GSM Evolution (EDGE), and GSM EDGE (GERAN).

[0144] 21A may be, for example, a wireless router, a Home NodeB, a Home eNodeB, or an access point, and may utilize any suitable RAT to facilitate wireless connectivity in a local area, such as a workplace, home, vehicle, campus, etc. In one embodiment, the base station 2114b and the WTRUs 2102c, 2102d implement a radio technology such as IEEE 802.11 to establish a wireless local area network (WLAN). In another embodiment, the base station 2114b and the WTRUs 2102c, 2102d implement a radio technology such as IEEE 802.15 to establish a wireless personal area network (WPAN). In yet another embodiment, the base station 2114b and the WTRUs 2102c, 2102d utilize a cellular-based RAT (e.g., WCDMA, CDMA2000, GSM, LTE, LTE-A, etc.) to establish a picocell or femtocell. 21A, the base station 2114b has a direct connection to the Internet 2110. Therefore, the base station 2114b does not need to access the Internet 2110 via the core network 2106.

[0145] The RAN 2104 communicates with a core network 2106, which is any type of network configured to provide voice, data, application, and / or Voice over Internet Protocol (VoIP) services to one or more of the WTRUs 2102a, 2102b, 2102c, 2102d. For example, the core network 2106 may provide call control, billing services, mobile location-based services, prepaid calling, Internet connectivity, video distribution, etc., and / or perform high-level security functions, such as user authentication. Although not shown in FIG. 21A , it is understood that the RAN 2104 and / or core network 2106 communicate directly or indirectly with other RANs that employ the same RAT as the RAN 2104 or a different RAT. For example, in addition to connecting to the RAN 2104 that employs E-UTRA radio technology, the core network 2106 may also communicate with another RAN (not shown) that employs GSM radio technology.

[0146] The core network 2106 also serves as a gateway for the WTRUs 2102a, 2102b, 2102c, 2102d to access the PSTN 2108, the Internet 2110, and / or other networks 2112. The PSTN 2108 includes a circuit-switched telephone network providing plain old telephone service (POTS). The Internet 2110 includes a global system of interconnected computer networks and devices that use common communication protocols, such as Transmission Control Protocol (TCP), User Datagram Protocol (UDP), and Internet Protocol (IP) in the TCP / IP Internet protocol suite. The networks 2112 include wired or wireless communication networks owned and / or operated by other service providers. For example, the network 2112 includes another core network connected to one or more RANs that utilize the same RAT as the RAN 2104 or a different RAT.

[0147] Some or all of the WTRUs 2102a, 2102b, 2102c, 2102d in the communication system 2100 include multi-mode capability, i.e., the WTRUs 2102a, 2102b, 2102c, 2102d include multiple transceivers for communicating with different wireless networks over different wireless links. For example, the WTRU 2102c shown in FIG. 21A is configured to communicate with a base station 2114a that utilizes cellular-based wireless technology and also configured to communicate with a base station 2114b that utilizes IEEE 802.11 wireless technology.

[0148] 21B is a system diagram of an example WTRU 2102. As shown in FIG. 21B, the WTRU 2102 includes a processor 2118, a transceiver 2120, a transmit / receive element 2122, a speaker / microphone 2124, a keypad 2126, a display / touchpad 2128, non-removable memory 2130, removable memory 2132, a power source 2134, a global positioning system (GPS) chipset 2136, and other peripherals 2138. It will be understood that the WTRU 2102 may include any subcombination of the above elements while remaining consistent with an embodiment.

[0149] The processor 2118 may be a general-purpose processor, a special-purpose processor, a conventional processor, a digital signal processor (DSP), multiple microprocessors, one or more microprocessors in conjunction with a DSP core, a controller, a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) circuit, any other type of integrated circuit (IC), a state machine, etc. The processor 2118 performs signal coding, data processing, power control, input / output processing, and / or any other functionality that enables the WTRU 2102 to operate in a wireless environment. The processor 2118 is coupled to the transceiver 2120, which is coupled to the transmit / receive element 2122. While FIG. 21B depicts the processor 2118 and the transceiver 2120 as separate components, it will be understood that the processor 2118 and the transceiver 2120 may be integrated together in an electronic package or chip.

[0150] The transmit / receive element 2122 is configured to transmit signals to or receive signals from a base station (e.g., base station 2114a) over the air interface 2116. For example, in one embodiment, the transmit / receive element 2122 is an antenna configured to transmit and / or receive RF signals. In another embodiment, the transmit / receive element 2122 is an emitter / detector configured to transmit and / or receive, for example, IR, UV, or visible light signals. In yet another embodiment, the transmit / receive element 2122 is configured to transmit and receive both RF and light signals. It should be understood that the transmit / receive element 2122 may be configured to transmit and / or receive any combination of wireless signals.

[0151] 21B, the transmit / receive element 2122 is shown as a single element, the WTRU 2102 may include any number of transmit / receive elements 2122. More specifically, the WTRU 2102 utilizes MIMO technology. Thus, in one embodiment, the WTRU 2102 includes two or more transmit / receive elements 2122 (e.g., multiple antennas) for transmitting and receiving wireless signals over the air interface 2116.

[0152] The transceiver 2120 is configured to modulate signals to be transmitted by the transmit / receive element 2122 and demodulate signals received by the transmit / receive element 2122. As mentioned above, the WTRU 2102 has multi-mode capabilities. Thus, the transceiver 2120 includes multiple transceivers to enable the WTRU 2102 to communicate via multiple RATs, such as, for example, UTRA and IEEE 802.11.

[0153] The processor 2118 of the WTRU 2102 is coupled to and receives user input data from a speaker / microphone 2124, a keypad 2126, and / or a display / touchpad 2128 (e.g., a liquid crystal display (LCD) display unit or an organic light emitting diode (OLED) display unit). The processor 2118 also outputs user data to the speaker / microphone 2124, the keypad 2126, and / or the display / touchpad 2128. In addition, the processor 2118 obtains information from and stores data in any type of suitable memory, such as non-removable memory 2130 and / or removable memory 2132. The non-removable memory 2130 includes random access memory (RAM), read-only memory (ROM), a hard disk, or any other type of memory storage device. The removable memory 2132 includes a subscriber identity module (SIM) card, a memory stick, a secure digital (SD) memory card, etc. In other embodiments, the processor 2118 obtains information from and stores data in memory that is not physically located on the WTRU 2102, but is located on a server or home computer (not shown), or the like.

[0154] The processor 2118 is configured to receive power from the power source 2134 and to distribute and / or control the power to other components within the WTRU 2102. The power source 2134 is any suitable device for powering the WTRU 2102. For example, the power source 2134 may include one or more dry batteries (e.g., nickel-cadmium (NiCd), nickel-zinc (NiZn), nickel-metal hydride (NiMH), lithium-ion (Li-ion), etc.), solar cells, fuel cells, etc.

[0155] The processor 2118 is also coupled to a GPS chipset 2136, which is configured to provide location information (e.g., longitude and latitude) regarding the current location of the WTRU 2102. In addition to or instead of information from the GPS chipset 2136, the WTRU 2102 receives location information over the air interface 2116 from base stations (e.g., base stations 2114a, 2114b) and / or determines its location based on the timing of signals received from two or more nearby base stations. It will be understood that the WTRU 2102 acquires location information using any suitable location-determination method while remaining consistent with an embodiment.

[0156] The processor 2118 is further coupled to other peripherals 2138, which include one or more software and / or hardware modules that provide additional features, functionality, and / or wired or wireless connectivity. For example, the peripherals 2138 include an accelerometer, an e-compass, a satellite transceiver, a digital camera (for photos or videos), a Universal Serial Bus (USB) port, a vibration device, a television transceiver, a hands-free headset, a Bluetooth® module, a frequency modulation (FM) radio unit, a digital music player, a media player, a video game player module, an internet browser, etc.

[0157] 21C is a system diagram of an embodiment of a communications system 2100 including a RAN 2104a and a core network 2106a, which comprise exemplary implementations of the RAN 2104 and the core network 2106, respectively. As mentioned above, the RAN 2104, e.g., the RAN 2104a, communicates with the WTRUs 2102a, 2102b, and 2102c over the air interface 2116 utilizing UTRA radio technology. The RAN 2104a also communicates with the core network 2106a. As shown in FIG. 21C, the RAN 2104a includes Node Bs 2140a, 2140b, and 2140c, which each include one or more transceivers for communicating with the WTRUs 2102a, 2102b, and 2102c over the air interface 2116. The Node Bs 2140a, 2140b, 2140c are each associated with a particular cell (not shown) within the RAN 2104a. The RAN 2104a also includes RNCs 2142a, 2142b. It will be appreciated that the RAN 2104a may include any number of Node Bs and RNCs while remaining consistent with an embodiment.

[0158] As shown in Figure 21C, Node Bs 2140a, 2140b communicate with RNC 2142a. Additionally, Node B 140c communicates with RNC 2142b. Node Bs 2140a, 2140b, 2140c communicate with their respective RNCs 2142a, 2142b via an Iub interface. RNCs 2142a, 2142b communicate with each other via an Iur interface. Each of RNCs 2142a, 2142b is configured to control its respective Node B 2140a, 2140b, 2140c to which it is connected. Additionally, each of RNCs 2142a, 2142b is configured to perform or support other functions, such as outer loop power control, load control, admission control, packet scheduling, handover control, macro diversity, security functions, and data encryption.

[0159] The core network 2106a shown in Figure 21C includes a media gateway (MGW) 2144, a mobile switching center (MSC) 2146, a serving GPRS support node (SGSN) 2148, and / or a gateway GPRS support node (GGSN) 2150. While each of the above elements is shown as part of the core network 2106a, it should be understood that any one of these elements may be owned and / or operated by an entity different from the core network operator.

[0160] The RNC 2142a in the RAN 2104a is connected to the MSC 2146 in the core network 2106a via an IuCS interface. The MSC 2146 is connected to the MGW 2144. The MSC 2146 and MGW 2144 provide the WTRUs 2102a, 2102b, 2102c with access to circuit-switched networks, such as the PSTN 2108, to facilitate communications between the WTRUs 2102a, 2102b, 2102c and traditional landline communications devices.

[0161] The RNC 2142a in the RAN 2104a is also connected to an SGSN 2148 in the core network 2106a via an IuPS interface. The SGSN 2148 is connected to a GGSN 2150. The SGSN 2148 and GGSN 2150 provide the WTRUs 2102a, 2102b, 2102c with access to packet-switched networks, such as the Internet 2110, to facilitate communications between the WTRUs 2102a, 2102b, 2102c and IP-enabled devices.

[0162] As mentioned above, the core network 2106a is also connected to the network 2112, which includes other wired or wireless networks owned and / or operated by other service providers.

[0163] 21D is a system diagram of an embodiment of a communications system 2100 including a RAN 2104b and a core network 2106b, which respectively comprise exemplary implementations of the RAN 2104 and the core network 2106. As mentioned above, the RAN 2104, e.g., the RAN 2104b, utilizes E-UTRA radio technology to communicate with the WTRUs 2102a, 2102b, 2102c over the air interface 2116. The RAN 2104b also communicates with the core network 2106b.

[0164] The RAN 2104b includes eNodeBs 2140d, 2140e, and 2140f, although it will be understood that the RAN 2104b may include any number of eNodeBs while remaining consistent with an embodiment. The eNodeBs 2140d, 2140e, and 2140f each include one or more transceivers for communicating with the WTRUs 2102a, 2102b, and 2102c over the air interface 2116. In one embodiment, the eNodeBs 2140d, 2140e, and 2140f implement MIMO technology. Thus, the eNodeB 2140d, for example, uses multiple antennas to transmit wireless signals to, and receive wireless signals from, the WTRU 2102a.

[0165] Each of the eNodeBs 2140d, 2140e, 2140f is associated with a particular cell (not shown) and is configured to handle radio resource management decisions, handover decisions, and scheduling of users on the uplink and / or downlink, etc. As shown in Figure 21D, the eNodeBs 2140d, 2140e, 2140f communicate with each other over the X2 interface.

[0166] The core network 2106b shown in Figure 21D includes a mobility management gateway (MME) 2143, a serving gateway 2145, and a packet data network (PDN) gateway 2147. Although each of the above elements is shown as part of the core network 2106b, it will be understood that any one of these elements may be owned and / or operated by an entity different from the core network operator.

[0167] The MME 2143 is connected to each of the eNodeBs 2140d, 2140e, 2140f in the RAN 2104b via an S1 interface and acts as a control node. For example, the MME 2143 is responsible for authenticating users of the WTRUs 2102a, 2102b, 2102c, bearer activation / deactivation, selecting a specific serving gateway during initial attach of the WTRUs 2102a, 2102b, 2102c, etc. The MME 2143 also provides a control plane function for switching between the RAN 2104b and other RANs (not shown) that employ other radio technologies, such as GSM or WCDMA.

[0168] The serving gateway 2145 is connected to each of the eNodeBs 2140d, 2140e, 2140f in the RAN 2104b via an S1 interface. The serving gateway 2145 generally routes and forwards user data packets to and from the WTRUs 2102a, 2102b, 2102c. The serving gateway 2145 also performs other functions, such as anchoring the user plane during inter-eNodeB handovers, triggering paging when downlink data is available for the WTRUs 2102a, 2102b, 2102c, and managing and storing the context of the WTRUs 2102a, 2102b, 2102c.

[0169] The serving gateway 2145 is also connected to a PDN gateway 2147, which provides the WTRUs 2102a, 2102b, 2102c with access to packet-switched networks, such as the Internet 2110, to facilitate communications between the WTRUs 2102a, 2102b, 2102c and IP-enabled devices.

[0170] The core network 2106b facilitates communication with other networks. For example, the core network 2106b may provide the WTRUs 2102a, 2102b, 2102c with access to circuit-switched networks, such as the PSTN 2108, to facilitate communication between the WTRUs 2102a, 2102b, 2102c and traditional landline communication devices. For example, the core network 2106b may include or communicate with an IP gateway (e.g., an IP Multimedia Subsystem (IMS) server) that serves as an interface between the core network 2106b and the PSTN 2108. In addition, the core network 2106b may provide the WTRUs 2102a, 2102b, 2102c with access to the network 2112, which may include other wired or wireless networks owned and / or operated by other service providers.

[0171] 21E is a system diagram of an embodiment of a communications system 2100 including a RAN 2104c and a core network 2106c that comprise exemplary implementations of the RAN 2104 and the core network 2106, respectively. The RAN 2104, e.g., the RAN 2104c, is an access service network (ASN) that communicates with the WTRUs 2102a, 2102b, 2102c over the air interface 2116 utilizing IEEE 802.16 radio technology. As described herein, communication links between different functional entities of the WTRUs 2102a, 2102b, 2102c, the RAN 2104c, and the core network 2106c are defined as reference points.

[0172] 21E, the RAN 2104c includes base stations 2140g, 2140h, and 2140i and an ASN gateway 2141, although it should be understood that the RAN 2104c may include any number of base stations and ASN gateways while remaining consistent with an embodiment. The base stations 2140g, 2140h, and 2140i are each associated with a particular cell (not shown) within the RAN 2104c and each include one or more transceivers for communicating with the WTRUs 2102a, 2102b, and 2102c over the air interface 2116. In one embodiment, the base stations 2140g, 2140h, and 2140i implement MIMO technology. Thus, the base station 2140g, for example, uses multiple antennas to transmit wireless signals to and receive wireless signals from the WTRU 2102a. The base stations 2140g, 2140h, 2140i also provide mobility management functions such as handoff triggering, tunnel establishment, radio resource management, traffic classification, and quality of service (QoS) policy enforcement. The ASN gateway 2141 acts as a traffic aggregation point and is responsible for paging, caching of subscriber profiles, and routing to the core network 2106c.

[0173] The air interface 2116 between the WTRUs 2102a, 2102b, 2102c and the RAN 2104c is defined as an R1 reference point, which implements the IEEE 802.16 specification. In addition, each of the WTRUs 2102a, 2102b, 2102c establishes a logical interface (not shown) with the core network 2106c. The logical interface between the WTRUs 2102a, 2102b, 2102c and the core network 2106c is defined as an R2 reference point, which is used for authentication, authorization, IP host configuration management, and / or mobility management.

[0174] The communication link between each of the base stations 2140g, 2140h, 2140i is defined as an R8 reference point that includes protocols for facilitating WTRU handovers and the transfer of data between base stations. The communication link between the base stations 2140g, 2140h, 2140i and the ASN gateway 2141 is defined as an R6 reference point that includes protocols for facilitating mobility management based on mobility events associated with each of the WTRUs 2102a, 2102b, 2102c.

[0175] As shown in Figure 21E, the RAN 2104c is connected to a core network 2106c. The communication link between the RAN 2104c and the core network 2106c is defined as an R3 reference point, which includes, for example, protocols for facilitating data forwarding and mobility management functions. The core network 2106c includes a Mobile IP Home Agent (MIP-HA) 2154, an Authentication, Authorization, and Accounting (AAA) server 2156, and a gateway 2158. While each of the above elements is shown as part of the core network 2106c, it should be understood that any one of these elements may be owned and / or operated by an entity different from the core network operator.

[0176] The MIP-HA is responsible for IP address management, allowing the WTRUs 2102a, 2102b, 2102c to roam between different ASNs and / or different core networks. The MIP-HA 2154 provides the WTRUs 2102a, 2102b, 2102c with access to packet-switched networks, such as the Internet 2110, to facilitate communication between the WTRUs 2102a, 2102b, 2102c and IP-enabled devices. The AAA server 2156 is responsible for user authentication and user service support. The gateway 2158 facilitates interworking with other networks. For example, the gateway 2158 provides the WTRUs 2102a, 2102b, 2102c with access to circuit-switched networks, such as the PSTN 2108, to facilitate communication between the WTRUs 2102a, 2102b, 2102c and traditional land-line communication devices. Additionally, the gateway 2158 provides the WTRUs 2102a, 2102b, 2102c with access to the network 2112, which may include other wired or wireless networks owned and / or operated by other service providers.

[0177] Although not shown in Figure 21E, it should be understood that the RAN 2104c is connected to other ASNs, and the core network 2106c is connected to other core networks. The communication link between the RAN 2104c and the other ASNs is defined as the R4 reference point, which includes protocols for coordinating mobility of the WTRUs 2102a, 2102b, 2102c between the RAN 2104c and the other ASNs. The communication link between the core network 2106c and the other core networks is defined as the R5 reference point, which includes protocols for facilitating interworking between a home core network and a visited core network.

[0178] Although features and elements have been described above in particular combinations, those skilled in the art will understand that each feature or element may be used alone or in any combination with the other features and elements. Additionally, the methods described herein may be implemented in a computer program, software, or firmware embodied in a computer-readable medium, executed by a computer or processor. Examples of computer-readable media include electronic signals (transmitted over wired or wireless connections) and computer-readable storage media. Examples of computer-readable storage media include, but are not limited to, read-only memory (ROM), random-access memory (RAM), registers, cache memory, semiconductor memory devices, magnetic media such as internal hard disks and removable disks, magneto-optical media, and optical media such as CD-ROM disks and digital versatile disks (DVDs). A processor in conjunction with software may be used to implement a radio frequency transceiver for use in a WTRU, terminal, base station, RNC, or any host computer. Features and / or elements described herein according to one or more exemplary embodiments may be used in combination with features and / or elements described herein according to one or more other exemplary embodiments.

[0179] Embodiment 1. A method of video decoding, comprising: receiving a video signal and a cross-plane filter associated with said video signal; applying the cross-plane filter to luma plane pixels of the video signal to determine a chroma offset; adding the chroma offset to a corresponding chroma plane pixel of the video signal; A method comprising: 2. receiving an indication of a region of the video signal to which the cross-plane filter is to be applied; applying the cross-plane filter to the region; 2. The method of embodiment 1, further comprising: 3. The method of embodiment 1, further comprising receiving an indication to apply the cross-plane filter to at least one of a sequence level, a picture level, a slice level, or a block level of the video signal. 4. The method of embodiment 1, wherein the step of applying the cross-plane filter is performed as a post-process in a single-layer video coding process. 5. The method of embodiment 1, wherein the step of applying the cross-plane filter is performed in a multi-layer video coding process, and the luma plane pixels are upsampled base layer luma plane pixels and the chroma plane pixels are upsampled base layer chroma plane pixels. 6. The method of embodiment 1, wherein the step of applying the cross-plane filter is performed in a multi-layer video coding process, the luma plane pixels are base layer luma plane pixels that are not upsampled, and the chroma plane pixels are base layer chroma plane pixels that are upsampled. 7. A network interface configured to receive a video signal and a cross-plane filter associated with said video signal; applying the cross-plane filter to luma plane pixels of the video signal to determine a chroma offset; and a processor configured to add the chroma offset to a corresponding chroma plane pixel of the video signal; A video coding device comprising: 8. The processor: receiving, via the network interface, an indication of a region of the video signal to which the cross-plane filter is to be applied; and 8. The video coding device of embodiment 7, further configured to apply the cross-plane filter to the region. 9. The video coding device of embodiment 7, wherein the processor is further configured to receive, via the network interface, an indication to apply the cross-plane filter to at least one of a sequence level, a picture level, a slice level, or a block level of the video signal. 10. The video coding device of embodiment 7, wherein the processor is configured to apply the cross-plane filter as a post-process in a single-layer video coding process. 11. The video coding device of embodiment 7, wherein the processor is configured to apply the cross plane in a multi-layer video coding process, the luma plane pixels being upsampled base layer luma plane pixels, and the chroma plane pixels being upsampled base layer chroma plane pixels. 12. The video coding device of embodiment 7, wherein the processor is configured to apply the cross-plane filter in a multi-layer video coding process, the luma plane pixels being non-upsampled base layer luma plane pixels, and the chroma plane pixels being upsampled base layer chroma plane pixels. 13. A method of encoding a video signal, comprising: generating a cross-plane filter using components of the video signal; quantizing filter coefficients associated with the cross-plane filter; encoding the filter coefficients into a bitstream representing the video signal; transmitting the bitstream; A method comprising: 14. The method of embodiment 13, wherein the cross-plane filter is designed for application to a luma plane component of the video signal, and application of the cross-plane filter to the luma plane component produces an output that is applicable to a chroma plane component of the video signal. 15. The method of embodiment 13, wherein the cross-plane filter is generated according to a training set. 16. The method of embodiment 15, wherein the training set includes coded luma components of the video signal, coded chroma components of the video signal, and original chroma components of the video signal. 17. The method of embodiment 13, further comprising determining the characteristics of the cross-plane filter according to at least one of coding performance or color subsampling format. 18. The method of embodiment 17, wherein the characteristic is at least one of the size of the cross-plane filter, the separation of the cross-plane filter, or the symmetry of the cross-plane filter. 19. Identifying a region within an encoded picture of the video signal; transmitting an indication to apply the cross-plane filter to the region; 14. The method of embodiment 13, further comprising: 20. A network interface configured to receive a video signal; generating a cross-plane filter using components of the video signal; quantizing filter coefficients associated with the cross-plane filter; encoding the filter coefficients into a bitstream representing the video signal; and a processor configured to transmit the bitstream via the network interface; A video coding device comprising: 21. The video coding device of embodiment 20, wherein the processor is configured to design the cross-plane filter for application to a luma plane component of the video signal, and application of the cross-plane filter to the luma plane component generates an output that is applicable to a chroma plane component of the video signal. 22. The video coding device of embodiment 20, wherein the processor is configured to generate the cross-plane filter according to a training set. 23. The video coding device of embodiment 22, wherein the training set includes coded luma components of the video signal, coded chroma components of the video signal, and original chroma components of the video signal. 24. The video coding device of embodiment 20, wherein the processor is further configured to determine characteristics of the cross-plane filter according to at least one of coding performance or color subsampling format. 25. The video coding device of embodiment 24, wherein the characteristic is at least one of the size of the cross-plane filter, the separability of the cross-plane filter, or the symmetry of the cross-plane filter. 26. The processor: Identifying a region within an encoded picture of said video signal; and 21. The video coding device of embodiment 20, further configured to send, via the network interface, an indication to apply the cross-plane filter to the region. [Industrial Applicability]

[0180] The present invention can be generally applied to wireless communication systems. [Explanation of symbols]

[0181] 102 Input Video Signal 108 Entropy Coding Units 110 Inverse Quantization Block 112 Inverse Transform Block 120 output video bitstreams 160 spatial prediction blocks 162 Temporal Prediction Blocks 164 Reference Picture Store Store 166 Loop Filter 2102a~2102d WTRU 2104 RAN 2106 Core Network 2108 PSTN 2110 Internet

Claims

1. 1. A method for decoding a video signal, comprising: receiving the video signal including an encoded first color component and an encoded second color component; reconstructing a first color component from the encoded first color component and a second color component from the encoded second color component; obtaining N coefficients of a cross-plane filter associated with the first color component; applying the cross-plane filter associated with the first color component to the reconstructed second color component to determine a color component offset associated with the first color component; adding the determined color component offset associated with the first color component to the reconstructed first color component; Equipped with The step of obtaining the N coefficients of the cross-plane filter includes decoding only N-1 coefficients and deriving the remaining coefficients from the N-1 coefficients. method.

2. 2. The method of claim 1, wherein deriving the remaining coefficient comprises deriving the remaining coefficient as the reciprocal of a sum of the N-1 coefficients.

3. A video decoding device, comprising: receiving a video signal including an encoded first color component and an encoded second color component; reconstructing a first color component from the encoded first color component and a second color component from the encoded second color component; obtaining N coefficients of a cross-plane filter associated with the first color component; applying the cross-plane filter associated with the first color component to the reconstructed second color component to determine a color component offset associated with the first color component; adding the determined color component offset associated with the first color component to the reconstructed first color component. a processor configured to: The apparatus, wherein obtaining the N coefficients of the cross-plane filter includes decoding only N-1 coefficients and deriving the remaining coefficients from the N-1 coefficients.

4. 4. The apparatus of claim 3, wherein deriving the remaining coefficient comprises deriving the remaining coefficient as the reciprocal of a sum of the N-1 coefficients.

5. 1. A method for encoding a video signal, comprising: receiving the video signal including a first color component and a second color component; encoding the first color component and the second color component; reconstructing a first color component from the encoded first color component and a second color component from the encoded second color component; obtaining N coefficients of a cross-plane filter associated with the first color component, wherein obtaining the N coefficients includes obtaining one coefficient among the N coefficients from N−1 other coefficients; applying the cross-plane filter associated with the first color component to the reconstructed second color component to determine a color component offset associated with the first color component; adding the determined color component offset associated with the first color component to the reconstructed first color component to obtain a modified reconstructed first color component; encoding the N-1 other coefficients of a cross-plane filter; A method for providing

6. 6. The method of claim 5, wherein obtaining one of the N coefficients comprises deriving the one coefficient as the reciprocal of a sum of the N-1 other coefficients.

7. 1. A video encoding device, comprising: receiving a video signal including a first color component and a second color component; encoding the first color component and the second color component; reconstructing a first color component from the encoded first color component and a second color component from the encoded second color component; obtaining N coefficients of a cross-plane filter associated with the first color component, wherein obtaining the N coefficients includes obtaining one coefficient among the N coefficients from N−1 other coefficients; applying the cross-plane filter associated with the first color component to the reconstructed second color component to determine a color component offset associated with the first color component; adding the determined color component offset associated with the first color component to the reconstructed first color component to obtain a modified reconstructed first color component; Encode the N-1 other coefficients of the cross-plane filter Processor configured to A device comprising:

8. 8. The apparatus of claim 7, wherein obtaining one coefficient among the N coefficients comprises deriving the one coefficient as the reciprocal of a sum of the N-1 other coefficients.

Citation Information

Patent Citations

  • Manufacture of element for type variable resistor

    JP1986075505A

  • JPP7433019B

  • Dynamic image encoding / decoding method and device

    WO2010001999A1

  • Dynamic image encoding method and dynamic image decoding method

    WO2011033643A1

  • Image processing device and image processing method

    WO2013164922A1