Region-based cross-component prediction

By adopting a region-based cross-component prediction method in video encoding and decoding, the problems of resource-intensive computing and long-term delay in the prior art are solved, and efficient video encoding and decoding in hardware code processors are realized.

CN119968840APending Publication Date: 2025-05-09GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280101018.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-13
Filing Date
2022-12-16
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In the video encoding and decoding process, the problem of resource-intensive computing and long delays occurs in cross-component prediction, especially in hardware code processors.

Method used

Using a region-based cross-component prediction method, the need to perform highly resource-intensive computing in the hardware code processor is reduced by derive the CCCM prediction filter coefficients for the entire area of ​​the frame rather than the individual CU.

Benefits of technology

The delay of the code processing process is greatly reduced, so that CCCM prediction can be effectively executed in the hardware code processor, and the efficiency of video encoding and decoding is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119968840A_ABST
    Figure CN119968840A_ABST
Patent Text Reader

Abstract

Region-based cross-component prediction improves convolutional cross-component mode (CCCM) prediction by enabling derivation of filter coefficients for predicting chroma samples from luma samples for an entire region of a frame of a video stream, such as code processing tree units (CTUs), rather than requiring derivation of such filter coefficients for each individual code processing unit (CU). Deriving filter coefficients for the entire region rather than for each individual CU being processed significantly reduces the latency of video code processing and thus enables CCCM prediction for hardware code processor implementations.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] A digital video stream may represent a video using a sequence of frames or still images. Digital video may be used for a variety of applications, including, for example, video conferencing, high-definition video entertainment, video advertising, or sharing of user-generated videos. A digital video stream may contain a large amount of data and consume a considerable amount of computing or communication resources of a computing device for processing, transmission, or storage of video data. Various methods have been proposed to reduce the amount of data in a video stream, including encoding or decoding techniques. Summary of the invention

[0002] Disclosed herein are, among other things, systems and techniques for region-based cross-component prediction.

[0003] A method for region-based cross-component prediction according to an implementation of the present disclosure includes: identifying a region within a frame to be encoded or decoded; determining regional filter coefficients for the region; determining an input value for a current luma sample within a portion of the region; determining a predicted chroma sample for the current luma sample based on the input value and the regional filter coefficients; and encoding or decoding the predicted chroma sample.

[0004] In some implementations of the method, determining the regional filter coefficients includes deriving at least a portion of the regional filter coefficients based on one or both of the region or neighboring regions.

[0005] In some implementations of the method, deriving at least the portion of the region filter coefficients based on one or both of the region or the neighboring region includes minimizing a mean square error between predicted chroma samples and reconstructed chroma samples within a reference region of the frame.

[0006] In some implementations of the method, the mean square error is performed using chrominance samples from a padded area outside the region.

[0007] In some implementations of the method, determining the region filter coefficients includes decoding one or more syntax elements for signaling the region filter coefficients from a bitstream associated with the frame.

[0008] In some implementations of the method, the method includes determining to use the region filter coefficients for determining the predicted chrominance sample based on a classification of the current luma sample.

[0009] In some implementations of the method, the portion of the region is a code processing unit, and different regional filter coefficients are used to determine the second predicted chrominance sample based on a classification of a second current luma sample within the code processing unit.

[0010] In some implementations of the method, identifying the region includes decoding one or more syntax elements signaled within the bitstream associated with the region.

[0011] In some implementations of the method, determining the predicted chroma samples includes determining a spatial weight value for the region according to a prediction method for the region of the portion of the area, and determining the predicted chroma samples using the spatial weight value.

[0012] In some implementations of the method, the portion of the region is a code processing unit, and the region filter coefficients are determined for use with a plurality of code processing units of the region.

[0013] In some implementations of the method, the size of the region is larger than the minimum chroma cell size.

[0014] In some implementations of the method, the region is a code processing tree unit of size 128x128 or 64x64.

[0015] A device for region-based cross-component prediction according to an implementation of the present disclosure includes a memory and a processor, the processor being configured to execute instructions stored in the memory to: determine a regional filter coefficient for a region within a frame to be encoded or decoded; determine a first predicted chroma sample of a first luminance sample based on an input value of the first luminance sample within a first part of the region and based on the regional filter coefficient; determine a second predicted chroma sample of the second luminance sample based on an input value of a second luminance sample within a second part of the region and based on the regional filter coefficient; and encode or decode the first predicted chroma sample and the second predicted chroma sample.

[0016] In some implementations of the device, a first portion of the regional filter coefficients is signaled within a bitstream associated with the frame, and a second portion of the regional filter coefficients is derived based on video data within the frame.

[0017] In some implementations of the device, the region is a current code processing tree unit, and the region filter coefficients are derived using reconstructed chroma samples from one or more neighboring code processing tree units of the current code processing tree unit.

[0018] In some implementations of the device, the regional filter coefficients are used for both the first predicted chroma samples and the second predicted chroma samples based on the classification of the first luma samples and the second luma samples.

[0019] In some implementations of the device, the classification is based on one or more of gradient, direction, or pixel value bands.

[0020] A non-transitory computer-readable storage device according to an implementation of the present disclosure includes program instructions executable by one or more processors, which, when executed, cause the one or more processors to perform operations for region-based cross-component prediction, wherein the operations include: determining filter coefficients for chroma samples within a plurality of code processing units of a code processing tree unit for predicting a frame to be encoded or decoded; determining a current luma sample within a code processing unit of the plurality of code processing units; determining a predicted chroma sample of the current luma sample based on an input value and the filter coefficients; and encoding or decoding the predicted chroma sample.

[0021] In some implementations of a non-transitory computer-readable storage device, determining the filter coefficients includes one of: deriving the filter coefficients based on one or both of the code processing tree unit or a neighboring code processing tree unit of the code processing tree unit; decoding one or more syntax elements for signaling the filter coefficients from a bitstream associated with the frame; or deriving a first portion of the filter coefficients and decoding a second portion of the filter coefficients from the bitstream.

[0022] In some implementations of a non-transitory computer-readable storage device, determining the predicted chroma sample includes determining a spatial weight value of the region according to a prediction method for the region of the code processing unit, and determining the predicted chroma sample using the spatial weight value.

[0023] These and other aspects of the present disclosure are disclosed in the following detailed description of implementations, the accompanying claims and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The description herein refers to the drawings described below, wherein like reference numerals refer to like parts throughout the several views.

[0025] Figure 1 is a schematic diagram of an example of a video encoding and decoding system.

[0026] Figure 2 is a block diagram of an example of a computing device that may implement a transmitting station or a receiving station.

[0027] Figure 3 is a diagram of an example of a video stream to be encoded and decoded.

[0028] Figure 4 is a block diagram of an example of an encoder.

[0029] Figure 5 is a block diagram of an example of a decoder.

[0030] Figure 6is a diagram of examples of portions of a video frame.

[0031] Figure 7 An example of a reference region for region-based cross-component prediction is shown.

[0032] Figure 8 An example of a neighborhood of luma samples used to predict chroma samples is shown.

[0033] Fig. 9 Example resolutions for luma and chroma blocks are shown.

[0034] Fig.10 is a flow diagram of an example of a technique for region-based cross-component prediction. DETAILED DESCRIPTION

[0035] The video compression scheme may include decomposing a corresponding image or frame of a video stream into smaller parts such as blocks or code processing tree units (CTUs), and generating a coded bitstream using a technique for limiting the information included in its corresponding CTU. The bitstream may be decoded to recreate the source frame from the limited information. Encoding a CTU into a bitstream or decoding a CTU from a bitstream may include predicting the value of a pixel or CTU based on similarities with other pixels or CTUs that have been processed by the code in the same frame. Those similarities may be determined using intra-frame prediction, which attempts to predict the pixel values ​​of a code processing unit (CU) of the CTU using pixels outside the CU (e.g., pixels in the same frame as the CU but outside the CU). During encoding, the result of an intra-frame prediction mode performed on a CU is a prediction unit (PU). A prediction residual may be determined based on the difference between the pixel values ​​of the CU and the pixel values ​​of the PU. The prediction residual and the intra-frame prediction mode used to ultimately obtain the prediction residual may then be encoded into the bitstream. During decoding, the prediction residual is reconstructed into a CU using the PU generated based on the intra prediction mode, and is thereafter included in the output video stream.

[0036] A CU includes a luminance (also referred to as luma) component and two chrominance (also referred to as chroma) components. In some cases, these luminance components and chrominance components may be referred to as luminance blocks and chrominance blocks. For example, the luminance component of a CU may be expressed in the Y plane of the CU, and the chrominance components may be expressed in the U plane and V plane or the Cr plane and Cb plane of the CU. The luminance component is understood to include a certain number of luminance samples, and each chrominance component is understood to include a certain number of chrominance samples. Typically, the luminance samples provide a measure of the brightness of the entire target CU, and therefore represent the structural quality of the video content of the target CU, while the chrominance samples provide a measure of the color of the entire target CU. Therefore, conventional video compression schemes typically use a more sophisticated prediction method to predict the luminance component of a CU than the method used to predict the chrominance component of a CU. Such schemes may also utilize methods for predicting those chrominance components based on the predicted luminance components.

[0037] An example of such a chroma prediction method based on luma is the cross-component linear model (CCLM) prediction, which is proposed to be used in conjunction with the H.266 codec, which is also called Versatile Video Coding (VVC), which is used in an intra-predicted CU to predict the chroma signal based on the weighted luma signal. In the case of CCLM prediction, the chroma samples of the CU are predicted based on the reconstructed luma samples of the same CU using a linear model represented as pred_C(i, j) = α * rec_L` (i, j) + β, where pred_C (i, j) represents the predicted chroma samples in the CU and rec_L` (i, j) represents the downsampled reconstructed luma samples of the same CU. The CCLM prediction parameters α and β are weights derived from up to four adjacent chroma samples and their corresponding downsampled luma samples using one or more lookup tables. Downsampling is to align the resolution of the luma component and the chroma component of the CU. Specifically, when the resolution of the luma component and the chroma component are already equal (e.g., 4:4:4), the downsampling operation can be omitted; however, when the resolution of the luma component and the chroma component are not equal (e.g., 4:2:0) so that the chroma component is generally smaller than the luma component, one or more downsampling filters can be applied to the luma samples within the luma component in the horizontal and vertical directions. Examples of downsampling filters may include: type 0, in which each chroma sample in the entire CU exists between two vertical luma samples; and type 2, in which a chroma sample exists for each luma sample in the entire CU. Due to the high correlation between luma values ​​and chroma values, CCLM prediction is generally more efficient than conventional chroma space prediction when the CU has rich textures, especially chroma textures.

[0038] Although CCLM prediction provides advantages over historical methods for chrominance prediction based on brightness, there is still an opportunity to further improve the accuracy and / or efficiency of CCLM prediction. One such opportunity is related to a newer method for chrominance prediction based on brightness based on CCLM prediction, which is called convolution cross-component model (CCCM) prediction. CCCM prediction uses a seven-tap filter, which includes five-tap spatial components, a tap nonlinear term, and a tap deviation term. The spatial component includes the current brightness sample C and four neighbor samples, referred to as N, S, E, and W (e.g., arranged in a plus sign, X, diamond, or other shape, where C is located in the middle in any such case). The nonlinear term P is represented as a power of 2 of C and scaled to the sample value range of the content, represented as P = (C * C + midVal) >> bitDepth, where bitDepth represents the bit precision of the video content, and midVal is the intermediate chrominance value within the bit precision. For example, for 10-bit video content, bitDepth will be equal to 10 and midVal will be equal to 512. The bias term B represents a scalar offset between the input and output, similar to the offset term in CCLM prediction, and is set to the mid-chroma value of the bit precision (e.g., 512 for 10-bit video content) – therefore, B is equal to midVal.

[0039] The output of CCCM prediction—the predicted chrominance values ​​based on C—is given as the filter coefficients c i The filter coefficient c is calculated by convolution between the input value and the predicted chroma value predChromaVal, where i ranges from 0 to 6 (inclusive). The predicted chroma value predChromaVal is expressed as predChromaVal = c0C + c1N + c2S +c3E + c4W + c5P + c6B. i The MSE is determined by minimizing the mean square error (MSE) between the predicted chroma samples and the reconstructed chroma samples in the reference area corresponding to one or more CTUs, the one or more CTUs including the current CTU, the current CTU including the CU being predicted. In one example, the reference area may include N (e.g., 6) chroma sample rows above and to the left of the CU, and the reference area may be extended by a CU width to the right and a CU height below the CU boundary accordingly. The reference area is adjusted to include only available chroma samples. An extension of the reference area may be provided, represented as a sample around the periphery of the actual reference area, to support chroma samples along the side of the reference area when such side samples are not otherwise available. MSE minimization is performed by calculating the autocorrelation matrix of the luma input samples and the cross-correlation vector between the luma input samples and the predicted chroma output samples.

[0040] While CCCM prediction offers many improvements over CCLM prediction alone, it is not without its drawbacks. Specifically, CCCM prediction requires performing multiple 64-bit division operations—with arbitrary denominators—to derive the filter coefficients c. i . Due to the nature of function solving, these division operations must be performed sequentially, and each filter coefficient value is expressed accordingly using a relatively high number of bits (e.g., using a bit precision of 22). Therefore, CCCM prediction of derived filter coefficients generally introduces long delays. This delay is particularly evident in hardware code processors (coders) (i.e., combined hardware encoders and decoders or separate hardware encoders and hardware decoders), which are limited to only a certain amount of processing per cycle and generally have a limited number of cycle budgets for small CUs. Because hardware code processors must be designed to handle the worst-case scenario (e.g., CCCM prediction is required for each sample within the entire CU), these limitations inevitably prevent CCCM prediction from being implemented in hardware code processors. Specifically, in such a worst-case scenario, given that there is not enough time to process the chroma samples within each CU of each frame, it will not be possible to play the video at the desired frame rate (e.g., 30 frames per second). Therefore, it will be desirable to modify CCCM prediction to make it available for hardware code processor implementations.

[0041] The implementation of the present disclosure uses a region-based approach to cross-component prediction to solve problems such as these, wherein CCCM prediction filter coefficients are determined for a relatively large region (e.g., CTU) and used in all CUs of the relatively large region. By deriving filter coefficients for the entire region of the frame instead of individual CUs at one time, it is no longer necessary to perform a highly resource-intensive filter coefficient derivation calculation sequence for individual CUs, thereby greatly reducing the latency of the code processing process so that CCCM prediction can be performed in a hardware code processor. In general, the region corresponds to a single CTU within the frame, but in any case should be larger than the minimum chroma unit size allowed by the target video codec. Accordingly, the size of a given region can be signaled within the bitstream. The filter coefficients for a given region can be derived based on spatially adjacent regions within the frame, signaled within the bitstream (e.g., within an adaptive parameter set (APS) or a slice header), or both, such as where some of the filter coefficients for the region are derived and other filter coefficients are signaled. In some cases, multiple filter coefficient sets can be used within a single region. For example, different filter sets may be used based on the classification of the reconstructed luma samples used to predict the target chroma samples, in which case a first set of filter coefficients may be used to predict a first chroma sample in a given region, and a second set of filter coefficients may be used to predict a second chroma sample in the same region. In some cases, different filter shapes may be used for cross-component prediction. In some cases, a region-based cross-component prediction approach as disclosed herein may be combined with a CCLM prediction approach, e.g., as explained above, to improve prediction accuracy for regions of certain types and / or sizes.

[0042] Although reference is made herein by way of example to CTUs, CUs, PUs, etc. as commonly used in video codecs such as H.265, known as High Efficiency Video Coding (HEVC), and H.266, implementations of the present disclosure may be used in conjunction with other video codec structures. In one specific but non-limiting example, implementations of the present disclosure may be used in conjunction with superblocks, macroblocks, blocks, etc. as commonly used in video codecs such as VP9, ​​AV1, and AV2 currently under development. Accordingly, references herein to specific video codec structures such as CTUs, CUs, PUs, etc. should be viewed as expressions of non-limiting example video codec structures that may be used in conjunction with implementations of the present disclosure.

[0043] Further details of techniques for region-based cross-component prediction are first described herein with reference to systems in which such techniques may be implemented. Figure 1 1 is a schematic diagram of an example of a video encoding and decoding system 100. The transmitting station 102 may be, for example, Figure 2The computer with internal hardware configuration is described. However, other implementations of the sending station 102 are possible. For example, the processing of the sending station 102 can be distributed among multiple devices.

[0044] The network 104 can connect the sending station 102 and the receiving station 106 to encode and decode the video stream. Specifically, the video stream can be encoded in the sending station 102, and the encoded video stream can be decoded in the receiving station 106. The network 104 can be, for example, the Internet. The network 104 can also be a local area network (LAN), a wide area network (WAN), a virtual private network (VPN), a cellular telephone network, or any other means of transferring the video stream from the sending station 102 to the receiving station 106 (in this example).

[0045] In one example, receiving station 106 may be a Figure 2 A computer with an internal hardware configuration is described. However, other suitable implementations of receiving station 106 are possible. For example, the processing of receiving station 106 may be distributed among multiple devices.

[0046] Other implementations of the video encoding and decoding system 100 are possible. For example, one implementation may omit the network 104. In another implementation, the video stream may be encoded and then stored for transmission to a receiving station 106 or any other device with memory at a later time. In one implementation, the receiving station 106 receives (e.g., via the network 104, a computer bus, and / or some communication pathway) the encoded video stream and stores the video stream for later decoding. In an example implementation, a real-time transport protocol (RTP) is used to send the encoded video over the network 104. In another implementation, a transport protocol other than RTP may be used, such as a hypertext transfer protocol (HTTP) video streaming protocol.

[0047] When used in a video conferencing system, for example, the sending station 102 and / or the receiving station 106 may include the capability to both encode and decode video streams as described below. For example, the receiving station 106 may be a video conference participant that receives an encoded video bitstream from a video conference server (e.g., the sending station 102) for decoding and viewing, and further encodes his or her own video bitstream and sends it to the video conference server for decoding and viewing by other participants.

[0048] In some implementations, the video encoding and decoding system 100 may alternatively be used to encode and decode data other than video data. For example, the video encoding and decoding system 100 may be used to process image data. The image data may include blocks of data from an image (e.g., CTUs of frames of a video stream). In such implementations, the sending station 102 may be used to encode the image data, and the receiving station 106 may be used to decode the image data.

[0049] Alternatively, receiving station 106 may represent a computing device that stores encoded image data for later use, such as after receiving the encoded or pre-encoded image data from sending station 102. As a further alternative, sending station 102 may represent a computing device that decodes the image data, such as before sending the decoded image data to receiving station 106 for display.

[0050] Figure 2 is a block diagram of an example of a computing device 200 that can implement a transmitting station or a receiving station. For example, the computing device 200 can implement Figure 1 The computing device 200 may be in the form of a computing system including a plurality of computing devices, or in the form of a computing device (e.g., a mobile phone, a tablet computer, a laptop computer, a notebook computer, a desktop computer, etc.).

[0051] The processor 202 in the computing device 200 may be a conventional central processing unit. Alternatively, the processor 202 may be another type of device or devices now existing or later developed that can manipulate or process information. For example, although the disclosed implementations may be practiced with one processor (e.g., processor 202) as shown, more than one processor may be used to achieve advantages in speed and efficiency.

[0052] In an implementation, the memory 204 in the computing device 200 may be a read-only memory (ROM) device or a random access memory (RAM) device. However, other suitable types of storage devices may be used as the memory 204. The memory 204 may include code and data 206 accessed by the processor 202 using a bus 212. The memory 204 may also include an operating system 208 and application programs 210, which include at least one program that allows the processor 202 to perform the methods described herein. For example, the application programs 210 may include applications 1 to N, which further include encoding and / or decoding software that performs enhanced multi-stage intra prediction, etc., as described herein.

[0053] The computing device 200 may also include an auxiliary storage 214, which may be, for example, a memory card used with a mobile computing device. Since a video communication session may contain a considerable amount of information, they may be stored in full or in part in the auxiliary storage 214 and loaded into the memory 204 for processing as needed.

[0054] The computing device 200 may also include one or more output devices, such as a display 218. In one example, the display 218 may be a touch-sensitive display that combines a display with a touch-sensitive element operable to sense touch input. The display 218 may be coupled to the processor 202 via the bus 212. In addition to or in lieu of the display 218, other output devices may be provided that allow a user to program or otherwise use the computing device 200. When the output device is or includes a display, the display may be implemented in various ways, including by a liquid crystal display (LCD), a cathode ray tube (CRT) display, or a light emitting diode (LED) display, such as an organic LED (OLED) display.

[0055] The computing device 200 may also include or be in communication with an image sensing device 220, such as a camera or any other image sensing device 220 now existing or later developed that can sense an image, such as an image of a user operating the computing device 200. The image sensing device 220 may be positioned so that it is pointed at the user operating the computing device 200. In one example, the position and optical axis of the image sensing device 220 may be configured so that the field of view includes an area directly adjacent to the display 218 and from which the display 218 is visible.

[0056] The computing device 200 may also include or communicate with a sound sensing device 222, such as a microphone or any other sound sensing device now existing or later developed that can sense sounds near the computing device 200. The sound sensing device 222 may be positioned so that it is directed toward a user operating the computing device 200, and may be configured to receive sounds, such as voice or other utterances, uttered by the user while the user is operating the computing device 200.

[0057] although Figure 2 The processor 202 and memory 204 of the computing device 200 are depicted as being integrated into one unit, but other configurations may be utilized. The operations of the processor 202 may be distributed across multiple machines (where individual machines may have one or more processors), which may be coupled directly or across a local area network or other network. The memory 204 may be distributed across multiple machines, such as network-based storage or storage in multiple machines that perform the operations of the computing device 200.

[0058] Although depicted here as one bus, the bus 212 of the computing device 200 may be comprised of multiple buses. Further, the secondary storage 214 may be directly coupled to other components of the computing device 200 or may be accessed via a network, and may include an integrated unit (such as a memory card) or multiple units (such as multiple memory cards). Thus, the computing device 200 may be implemented in a wide variety of configurations.

[0059] Figure 3 is a diagram of an example of a video stream 300 to be encoded and decoded. Video stream 300 includes a video sequence 302. At the next level, video sequence 302 includes a plurality of adjacent video frames 304. Although three frames are depicted as adjacent frames 304, video sequence 302 may include any number of adjacent frames 304. Adjacent frames 304 may then be further subdivided into individual video frames, e.g., frame 306.

[0060] At the next level, frame 306 can be divided into a series of planes, slices or segments 308. For example, segment 308 can be a subset of a frame that allows parallel processing. Segment 308 can also be a subset of a frame that can separate video data into separate colors. For example, a frame 306 of color video data can include a luminance plane and two chrominance planes. Segment 308 can be sampled at different resolutions.

[0061] Regardless of whether the frame 306 is divided into segments 308, the frame 306 may be further subdivided into CTUs 310, which may contain data corresponding to, for example, NxM pixels in the frame 306, where N and M may refer to the same integer value or different integer values. The CTU 310 may also be arranged to include data from one or more slices 308 of pixel data. The CTU 310 may have any suitable size, such as 4x4 pixels, 8x8 pixels, 16x8 pixels, 8x16 pixels, 16x16 pixels, or larger up to a maximum size, which may be 128x128 pixels or another NxM pixel size.

[0062] Figure 4 4 is a block diagram of an example of encoder 400. As described above, encoder 400 may be implemented in transmission station 102, such as by providing a computer software program stored in a memory (e.g., memory 204). The computer software program may include machine instructions that, when executed by a processor (such as processor 202), cause transmission station 102 to generate a signal in a manner that is consistent with the present invention. Figure 4 The encoder 400 may also be implemented as dedicated hardware included in, for example, the transmission station 102. In some implementations, the encoder 400 is a hardware encoder.

[0063] The encoder 400 has the following stages for performing various functions in a forward path (shown by solid connecting lines) to produce an encoded or compressed bitstream 420 using the video stream 300 as input: an intra / inter prediction stage 402, a transform stage 404, a quantization stage 406, and an entropy code processing stage 408. The encoder 400 may also include a reconstruction path (shown by dashed connecting lines) for reconstructing frames for encoding of future CTUs. Figure 4 4 , the encoder 400 has the following stages for performing various functions in the reconstruction path: a dequantization stage 410, an inverse transform stage 412, a reconstruction stage 414, and a loop filter stage 416. Other structural variations of the encoder 400 may be used to encode the video stream 300.

[0064] In some cases, the functions performed by encoder 400 may occur after filtering of video stream 300. That is, before encoder 400 receives video stream 300, video stream 300 may be pre-processed according to one or more implementations of the present disclosure. Alternatively, encoder 400 itself may continue to perform the filtering of video stream 300. Figure 4 Such pre-processing is performed on the video stream 300 prior to the described functionality, such as prior to processing the video stream 300 at the intra / inter prediction stage 402 .

[0065] When the video stream 300 is submitted for encoding after performing preprocessing, the corresponding adjacent frames 304, such as frame 306, can be processed in units of CTUs. At the intra / inter prediction stage 402, the corresponding CU of the CTU can be encoded using intra-frame prediction (also referred to as intra-prediction) or inter-frame prediction (also referred to as inter-prediction). In any case, PUs can be formed. In the case of intra-frame prediction, PUs can be formed from samples that have been previously encoded and reconstructed in the current frame. In the case of inter-frame prediction, PUs can be formed from samples in one or more previously constructed reference frames.

[0066] Next, the PU may be subtracted from the CU at the intra / inter prediction stage 402 to produce a prediction residual, also referred to as a residual. The transform stage 404 transforms the residual into transform coefficients, for example, in the frequency domain, using a block-based transform. The quantization stage 406 converts the transform coefficients into discrete quantum values, referred to as quantized transform coefficients, using a quantizer value or quantization level. For example, the transform coefficients may be divided by a quantizer value and truncated.

[0067] The quantized transform coefficients are then entropy encoded by the entropy encoding stage 408. The entropy encoded coefficients are then output to a compressed bitstream 420 along with other information used to decode the CU (which may include, for example, syntax elements such as to indicate the prediction type, transform type, motion vectors, or quantizer values ​​used). The compressed bitstream 420 may be formatted using various techniques such as variable length code processing or arithmetic code processing. The compressed bitstream 420 may also be referred to as a coded video stream or a coded video bitstream, and the terms will be used interchangeably herein.

[0068] The reconstruction path (shown by the dashed connecting line) can be used to ensure that the encoder 400 and (hereinafter Figure 5 The decoder 500 (described below) uses the same reference frame to decode the compressed bitstream 420. The reconstruction path performs the same Figure 5 The decoder 400 may be configured to decode the CU in the intra / inter prediction stage 402. The decoder 400 may be configured to decode the CU in the intra / inter prediction stage 402. The decoder 400 may be configured to decode the CU in the intra / inter prediction stage 402. The decoder 400 may be configured to decode the CU in the intra / inter prediction stage 402. The decoder 400 may be configured to decode the CU in the inter ...

[0069] Other variations of the encoder 400 may be used to encode the compressed bitstream 420. In some implementations, for certain CUs, CTUs, or frames, a non-transform based encoder may directly quantize the residual signal without the transform stage 404. In some implementations, the encoder may have the quantization stage 406 and the dequantization stage 410 combined in a common stage.

[0070] Figure 5 5 is a block diagram of an example of a decoder 500. The decoder 500 may be implemented in the receiving station 106, for example, by providing a computer software program stored in the memory 204. The computer software program may include machine instructions that, when executed by a processor such as the processor 202, cause the receiving station 106 to receive the received signal in a manner that is consistent with the exemplary embodiment of the present invention. Figure 5 The decoder 500 may also be implemented in hardware included in, for example, the sending station 102 or the receiving station 106. In some implementations, the decoder 500 is a hardware decoder.

[0071] Similar to the reconstruction path of the encoder 400 discussed above, in one example, the decoder 500 includes the following stages for performing various functions to produce an output video stream 516 from the compressed bitstream 420: an entropy decoding stage 502, a dequantization stage 504, an inverse transform stage 506, an intra / inter prediction stage 508, a reconstruction stage 510, a loop filtering stage 512, and a post-filter stage 514. Other structural variations of the decoder 500 may be used to decode the compressed bitstream 420.

[0072] When the compressed bitstream 420 is submitted for decoding, the data elements within the compressed bitstream 420 may be decoded by the entropy decoding stage 502 to produce a set of quantized transform coefficients. The dequantization stage 504 dequantizes the quantized transform coefficients (e.g., by multiplying the quantized transform coefficients by a quantizer value), and the inverse transform stage 506 inverse transforms the dequantized transform coefficients to produce a derived residual, which may be the same as the derived residual created by the inverse transform stage 412 in the encoder 400. Using the header information decoded from the compressed bitstream 420, the decoder 500 may use the intra / inter prediction stage 508 to create a PU that is the same as the predicted block created in the encoder 400 (e.g., at the intra / inter prediction stage 402).

[0073] At the reconstruction stage 510, the PU can be added to the derived residual to create a reconstructed CU. A loop filtering stage 512 can be applied to the reconstructed CU to reduce blocking artifacts. Examples of filters that can be applied at the loop filtering stage 512 include, but are not limited to, deblocking filters, directional enhancement filters, and loop recovery filters. Other filtering can be applied to the reconstructed CU. In this example, a post-filter stage 514 is applied to the reconstructed CU to reduce blocking distortion, and the result is output as an output video stream 516. The output video stream 516 may also be referred to as a decoded video stream, and the terms will be used interchangeably herein.

[0074] Other variations of decoder 500 may be used to decode compressed bitstream 420. In some implementations, decoder 500 may produce output video stream 516 without post-filter stage 514, or otherwise omit post-filter stage 514.

[0075] Figure 6 is a diagram of examples of portions of a video frame 600, which may be, for example, Figure 3306 shown in FIG. The video frame 600 includes a plurality of 64×64 CTUs, such as four 64x64 CTUs 610 in two rows and two columns in a matrix or Cartesian plane, as shown. Each 64×64 CTU 610 may include up to four 32×32 CUs 620. Each 32×32 CU 620 may include up to four 16×16 CUs 630. Each 16×16 CU 630 may include up to four 8×8 CUs 640. Each 8×8 CU 640 may include up to four 4×4 CUs 650. Each 4×4 CU 650 may include 16 pixels, which may be represented in four rows and four columns in each corresponding CU in a Cartesian plane or matrix.

[0076] In some implementations, the video frame 600 may include CTUs larger than 64x64 and / or CUs smaller than 4x4. The video frame 600 may be partitioned into various arrangements based on features within the video frame 600 and / or other criteria. Although one arrangement of CUs is shown, any arrangement may be used. Figure 6 N×N CTUs and CUs are shown, but in some implementations, N×M CTUs and / or CUs may be used, where N and M are different numbers. For example, 32×64 CTUs, 64×32 CTUs, 16×32 CUs, 32×16 CUs, or any other size may be used. In some implementations, N×2N CTUs or CUs, 2N×N CTUs or CUs, or combinations thereof may be used.

[0077] Pixels may include information representing an image captured in the video frame 600, such as brightness information, color information, and position information. In some implementations, a block, such as the 16×16 pixel block shown, may include: a luma block 660, which may include luma pixels 662; and two chroma blocks 670, 680, such as a U or Cb chroma block 670, and a V or Cr chroma block 680. The chroma blocks 670, 680 may include chroma pixels 690. For example, the luma block 660 may include 16×16 luma pixels 662, and each chroma block 670, 680 may include 8×8 chroma pixels 690, as shown.

[0078] In some implementations, code processing the video frame 600 may include ordered block-level code processing. Ordered block-level code processing may include code processing the CUs of the video frame 600 in an order such as a raster scan order, where the CU may be identified and processed starting with the CTU in the upper left corner of the video frame 600 or a portion of the video frame 600, and then proceeding from left to right and from the top row to the bottom row along the rows, thereby identifying each CU for processing in turn. For example, the 64×64 CTU in the left column of the top row of the video frame 600 may be the first code-processed CTU, and the 64×64 CTU immediately to the right of the first CTU may be the second code-processed CTU. The second row from the top may be the second code-processed row, such that the 64×64 CTU in the left column of the second row may be code-processed after the 64×64 CTU in the rightmost column of the first row.

[0079] In some implementations, coding the CTUs of the video frame 600 may include using quadtree coding, which may include coding smaller CUs within the CTU in a raster scan order. For example, the 64×64 CTU shown in the lower left corner of the portion of the video frame 600 may be coded using quadtree coding, where the upper left 32×32 CU may be coded, then the upper right 32×32 CU may be coded, then the lower left 32×32 CU may be coded, and then the lower right 32×32 CU may be coded. Each 32×32 CU may be coded using quadtree coding, where the upper left 16×16 CU may be coded, then the upper right 16×16 CU may be coded, then the lower left 16×16 CU may be coded, and then the lower right 16×16 CU may be coded. Each 16×16 CU may be coded using quadtree coding, where the top-left 8×8 CU may be coded, then the top-right 8×8 CU may be coded, then the bottom-left 8×8 CU may be coded, and then the bottom-right 8×8 CU may be coded. Each 8×8 CU may be coded using quadtree coding, where the top-left 4×4 CU may be coded, then the top-right 4×4 CU may be coded, then the bottom-left 4×4 CU may be coded, and then the bottom-right 4×4 CU may be coded. In some implementations, the 8x8 CU may be omitted for the 16x16 CU, and the 16x16 CU may be coded using quadtree coding, where the top-left 4x4 CU may be coded, and then the other 4x4 CUs in the 16x16 CU may be coded in raster scan order.

[0080] In some implementations, transcoding the video frame 600 may include transcoding information included in an original version of the image or video frame by, for example, omitting some information from the original version of the image or video frame from the corresponding encoded image or encoded video frame. For example, transcoding may include reducing spectral redundancy, reducing spatial redundancy, or a combination thereof. Reducing spectral redundancy may include using a color model based on a luma component (Y) and two chroma components (U and V or Cb and Cr), which may be referred to as a YUV or YCbCr color model or color space. Using the YUV color model may include using a relatively large amount of information to represent the luma component of a portion of the video frame 600, and using a relatively small amount of information to represent each corresponding chroma component of the portion of the video frame 600. For example, a portion of the video frame 600 may be represented by a high-resolution luma component that may include a 16×16 luma sample block, and two lower-resolution chroma components, where each chroma component represents a portion of the image as an 8×8 chroma sample block. A sample may indicate a value, for example, a value in a range from 0 to 255, and may be stored or transmitted using, for example, eight bits. Although the present disclosure is described with reference to a YUV color model, another color model may also be used. Reducing spatial redundancy may include transforming the CU into a frequency domain using, for example, a discrete cosine transform. For example, a unit of an encoder may perform a discrete cosine transform using transform coefficient values ​​based on spatial frequencies.

[0081] Although described herein with reference to a matrix or Cartesian representation of video frame 600 for clarity, video frame 600 may be stored, sent, processed, or a combination thereof in a data structure so that pixel values ​​and / or luminance samples and chrominance samples of video frame 600 may be efficiently represented. For example, video frame 600 may be stored, sent, processed, or any combination thereof in a two-dimensional data structure such as the matrix shown or in a one-dimensional data structure such as a vector matrix. In addition, although described herein as showing an image subsampled by chrominance, where U and V have half the resolution of Y, video frame 600 may have different configurations of its color channels. For example, still with reference to the YUV color space, full resolution may be used for all color channels of video frame 600. In another example, a color space other than the YUV color space may be used to represent the resolution of the color channels of video frame 600.

[0082] Figure 7An example of a reference region 700 for region-based cross-component prediction is shown. Reference region 700 shows chroma samples of a CTU, where some of those chroma samples are filled with patterns 702, 704, and 706. Specifically, the chroma samples filled with pattern 702 correspond to a current PU 708 being predicted, the chroma samples filled with pattern 704 are reconstructed chroma samples that can be used to predict the chroma samples filled with pattern 702, and the chroma samples filled with pattern 706 represent a filling region that is used to expand the reference region to accommodate prediction of chroma samples located along the edge of the chroma samples filled with pattern 704. The filling region surrounds some or all of the perimeter of reference region 700 and is one or more chroma samples wide. Figure 7 In the example shown, the padding regions are one chroma sample wide, which is indicated based on that they are a single chroma sample with pattern 706 adjacent to each outermost chroma sample padded with pattern 704. Given that—as will be described below—the CCCM filter coefficients for the current luma sample are determined using four neighboring samples (e.g., N, S, E, and W), the padding regions ensure that all four neighboring sample regions are available even for samples along the edge of the portion of the reference region 700 padded with pattern 704. Since the chroma samples padded with pattern 706 are not available within the CTU itself, it can be understood that they contain (i.e., are set to) padding values. Although PU 708 is shown as having size 8×4, the present disclosure is not limited to a particular PU size.

[0083] The reference area 700 may include a top area 710, which may include 1 to N (where N>1) rows of pixels. The reference area 700 may include an upper right area 712, which may include 1 to N rows of pixels. The reference area 700 may include a left area 714 having 1 to M (where M>1) columns of pixels. The reference area 700 may include a lower left area 716 having 1 to M (where M>1) columns of pixels. In one example, N=M. The reference area 700 may be based on a chroma color format. For example, for 4:4:4 content, the reference area 700 may also be 4 samples wide; and for 4:2:0 or 4:2:2 color formats, the reference area 700 may be 2 samples wide. In one example, when the upper right area 712 is available, only the 4×4 luminance block in the upper right corner is included in the reference area 700. Similarly, if the lower left area 716 is available, only the 4×4 luminance block in the lower right corner is included in the reference area 700. Reference region 700 may be adjusted accordingly based on the chroma color format.In another example, top region 710 may always be 1 sample wide for both luma and chroma, while left region 714 may be 4 samples wide for luma.

[0084] While conventional approaches to CCCM prediction require deriving filter coefficients for each PU, such as PU 708, individually and therefore only for a small portion of the reference region 700, the region-based cross-component prediction disclosed herein includes deriving filter coefficients for the entire reference region 700. In this manner, the reference region 700 corresponds to a region of a frame being predicted, and more specifically to a CTU including PU 708 within the frame. However, in some cases, the reference region 700 may correspond in whole or in part to multiple CTUs, such as a CTU including PU 708 and one or more neighboring CTUs of the CTU.

[0085] Figure 8 An example of a neighborhood 800 of a luma sample 802 used to predict chroma samples is shown. Neighborhood 800 shows a 3×3 neighborhood by way of example. In some cases, neighborhood 800 may be larger or smaller than 3x3 and / or neighborhood 800 may be a shape other than a square, such as a non-square rectangle or diamond. Luma sample 802 is located in the middle of neighborhood 800. Luma sample 802—labeled C to indicate that it is the current luma sample being processed—is surrounded by neighboring luma samples 804, 806, 808, and 810 that will be used to predict chroma samples for luma sample 802. In the example shown, luma samples 804, 806, 808, and 810 are labeled with directional names N, S, E, and W (i.e., north, south, east, and west), respectively, relative to the location of luma sample 802. Luma sample 802 and neighboring luma samples 804, 806, 808, and 810 together constitute the value of the five-tap spatial component used in CCCM prediction and used to calculate the predicted chroma sample of luma sample 802, which is represented as predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P +c6B, where the filter coefficients c i is derived for the entire region rather than just for the CU that includes luma samples 802 using region-based cross-component prediction as disclosed herein.

[0086] Fig. 9 Example resolutions for luma blocks and chroma blocks are shown. As described above, and in order to ensure that the appropriate luma samples are used to predict the chroma samples of a given CU 900, it may be desirable to downsample (i.e., reduce the resolution of) the luma block of the CU being processed so that the final resolution of the luma block is the same as the resolution of the chroma blocks of the CU. For example, downsampling may be performed where the resolutions of the luma blocks and chroma blocks are originally provided in a format such as 4:2:0. However, where the resolutions of the luma blocks and chroma blocks of a given CU are already the same (e.g., 4:4:4), the downsampling operation on the CU may be skipped.

[0087] Further details of the techniques for region-based cross-component prediction are now described. Fig.10 1 is a flow chart of an example of a technique 1000 for region-based cross-component prediction. For example, the technique 1000 may be performed in whole or in part at a prediction stage (e.g., intra / inter prediction stage 402) of an encoder for encoding a video stream or a prediction stage (e.g., intra / inter prediction stage 508) of a decoder for decoding a bitstream.

[0088] Technique 1000 may be implemented as a software program that can be executed, for example, by a computing device such as sending station 102 or receiving station 106. For example, the software program may include machine-readable instructions that may be stored in a memory such as memory 204 or auxiliary storage 214, and when executed by a processor such as processor 202, may cause a computing device to perform technique 1000. Technique 1000 may be implemented using specialized hardware or firmware. For example, a hardware component such as a hardware code processor may be configured to perform technique 1000. As explained above, some computing devices may have multiple memories or processors, and the operations described in technique 1000 may be distributed using multiple processors, memories, or both. For ease of explanation, technique 1000 is depicted and described herein as a series of steps or operations. However, the steps or operations according to the present disclosure may occur in various orders and / or concurrently. In addition, other steps or operations not presented and described herein may be used. In addition, all of the steps or operations shown may not be required to implement the techniques according to the disclosed subject matter.

[0089] At 1002, a region within a current frame being processed (i.e., encoded or decoded) is identified. For example, the region may be a CTU. During encoding, the region may be identified as a single CTU during frame segmentation. During decoding, the region may be identified using one or more syntax elements signaled within a bitstream. The region has a size greater than the minimum chroma block size. For example, the region may be 128x128 or 64x64.

[0090] At 1004, the regional filter coefficients for the region are determined. The regional filter coefficients are CCCM prediction filter coefficients (ie, filter coefficients c i ). Determining the regional filter coefficients may include deriving the regional filter coefficients based on one or more previously coded and spatially adjacent regions, identifying the regional filter coefficients using one or more syntax elements signaled within the bitstream, or both. Deriving the regional filter coefficients includes minimizing the reference region—for example, Figure 7700 - the MSE between the predicted chroma samples and the reconstructed chroma samples in the reference region 700 shown in FIG. Thus, previous CCCM prediction approaches derive filter coefficients for individual CUs and therefore use predicted chroma samples and reconstructed chroma samples limited to a portion of the reference region corresponding to the given CU, while deriving regional filter coefficients includes using the entire reference region to minimize the MSE. However, in some cases, the padded portion of the reference region (e.g., having Figure 7 Some or all of the samples of the pattern 706 shown in FIG. 1 may be excluded from the regional filter coefficient determination process. Fig. 9 In the case of downsampling as described, the downsampling may be performed before determining the regional filter coefficients.

[0091] In some implementations, the regional filter coefficients for the identified region may be derived using reconstructed chroma samples from one or more other regions. For example, where the identified region is the current CTU, the regional filter coefficients may be derived using reconstructed chroma samples from one or both of the left neighboring CTU of the current CTU or the upper neighboring CTU of the current CTU. In another example, where the identified region is the current CTU, the regional filter coefficients may be derived using reconstructed chroma samples from one or more of the upper left neighboring CTU of the current CTU, the upper right neighboring CTU of the current CTU, the lower left neighboring CTU of the current CTU, or the lower right neighboring CTU of the current CTU.

[0092] During encoding, the regional filter coefficients for the region are derived; however, during decoding, the regional filter coefficients for the region may be derived and / or signaled. For example, signaling the regional filter coefficients may include signaling the regional filter coefficients explicitly or implicitly within a bitstream, such as within an adaptive parameter set, a slice header, or another structure that may be used to store syntax elements for decoding the encoded video data from a bitstream. In some cases, some but not all of the regional filter coefficients for the region may be signaled. In this case, the remaining regional filter coefficients may be derived, as described above. For example, in this case, a first subset of the regional filter coefficients may be signaled, and a second subset of the regional filter coefficients may be derived.

[0093] In addition, in some cases, one or more regional filter coefficients signaled within the bitstream may be refined as part of a process for determining regional filter coefficients. For example, refining the regional filter coefficients may include deriving regional filter coefficients as described above and comparing the derived regional filter coefficients with the signaled regional filter coefficients. In some such cases, where the comparison indicates that the derived regional filter coefficient is within a first threshold range of the signaled regional filter coefficients, the signaled regional filter coefficient or the derived regional filter coefficient may be used as the refined regional filter coefficient. In other such cases, where the comparison indicates that the derived regional filter coefficient is outside the first threshold range of the signaled regional filter coefficients but within a second threshold range thereof, the signaled regional filter coefficient and the derived regional filter coefficient may be combined (e.g., averaged) to produce the refined regional filter coefficient. In still other such cases, where the comparison indicates that the derived regional filter coefficient is outside both the first threshold range and the second threshold range of the signaled regional filter coefficients, the derived regional filter coefficient may be used as the refined regional filter coefficient. Other examples are also possible.

[0094] In some implementations, determining the regional filter coefficients may include determining multiple sets of regional filter coefficients for the region. For example, different sets of regional filter coefficients may be determined based on different classifications of the reconstructed luminance samples of the region. In this case, luminance samples corresponding to the same classification may be understood as sharing the same set of regional filter coefficients. The classifications may be derived in parallel, and therefore there is no dependency between them. The classifications may be based, for example, on gradients, directions, pixel value bands, etc. For example, gradient-based classifications may be derived at the 4x4 luminance block level. In another example, band-based classifications may be derived at the 2x2 luminance block level and based on the average value of a given 2x2 luminance block. In some cases, overlapping classifications may be used, in which samples may be counted in more than one classification. In the case of overlapping classifications, different weights may be used. For example, compared to when the sample is not directly classified into the target classification, when the sample is directly classified into the target classification, the sample may have a greater weight. In some cases, when a luminance sample is needed but the luminance sample has not yet been reconstructed, filling (e.g., pixel repetition) may be used for classification and prediction.

[0095] At 1006, an input value of a current luma sample in the region is determined. Specifically, the current luma sample is located in a sub-portion of the region, such as a CU or PU being predicted. The input value includes the current luma sample, a plurality of adjacent luma samples of the current luma sample, and the bit precision of the video data being encoded or decoded. For example, the input value may correspond to seven taps of a seven-tap filter for CCCM prediction, the taps including the current luma sample C, four adjacent luma samples N, S, E, and W, a nonlinear term P (expressed as a power of 2 of C and scaled to the sample value range of the content based on the bit precision (e.g., expressed as P = (C * C + midVal) >> bitDepth, where bitDepth represents the bit precision of the video content, and midVal is an intermediate chroma value within the bit precision)) and a bias term B (expressed as a scalar offset between input and output, similar to the offset term in CCLM prediction, and set to the intermediate chroma value of the bit precision).

[0096] A filter is used to identify the current luma sample and a plurality of neighboring luma samples. In one example, the filter can be applied to a 3x3 neighborhood within the CU including the current luma sample, such as Figure 8 The filter can have a Figure 8 The example uses a plus sign shape so that the plurality of adjacent brightness samples includes four adjacent brightness samples labeled N, S, E, and W, such as Figure 8 ; however, other shape examples may be used, and neighborhoods of other sizes may be used. For example, an x-shaped filter may be used in a 3x3 neighborhood, a diamond-shaped filter may be used in a 5x5 neighborhood, and so on. In some implementations, filters with a number of coefficients below a threshold may be used for CUs below a specified size (e.g., 8x8), and / or filters with a number of coefficients above a threshold may be used for CUs above the specified size.

[0097] At 1008, predicted chroma samples are determined based on the input value of the current luma sample and the regional filter coefficients. For example, the predicted chroma samples, denoted as predChromaVal, may be determined by calculating predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P + c6B, where C, N, S, E, W, P, and B are the input values ​​of the current luma sample and c0, c1, c2, c3, c4, c5, and c6 are the regional filter coefficients.

[0098] In some implementations, the predicted chroma sample may be determined by a weighted combination (e.g., an average) of a first predicted chroma sample determined as explained above (i.e., using region-based cross-component prediction as disclosed herein) and a second predicted chroma sample determined using CCLM prediction. Thus, the predicted chroma sample may be determined by a weighted combination (e.g., an average) of sample values ​​using a spatial weight value determined for the region according to a prediction method for the region of a portion of the region (e.g., a CU including the current luma sample). For example, the weighting values ​​used may depend on the sample position relative to the CU, such that a larger weighting value is used for region-based cross-component prediction at the bottom and / or right portion of the CU, while a larger weighting is used for CCLM prediction at the top and / or left portion of the CU. This may be desirable because CCLM prediction is generally well adapted to local textures, while region-based cross-component prediction disclosed herein is generally well adapted to large regions. The weighting values ​​may be predefined (e.g., for use during encoding) or signaled in the bitstream (e.g., for use during decoding).

[0099] In some implementations, the regional filter coefficients used to determine the predicted chroma samples can be a first set of regional filter coefficients, and a second set of regional filter coefficients can be used to determine a second predicted chroma sample for the identified region. For example, the specific regional filter coefficients used for the predicted chroma samples and the second predicted chroma samples can be based on classification information of luma samples corresponding to those predicted chroma samples.

[0100] At 1010, the predicted chroma samples are encoded (e.g., encoded into a bitstream) or decoded (e.g., for output within an output video stream) based on whether the technique 1000 is performed during encoding or decoding. In some implementations, the predicted chroma samples can be reconstructed for use in predicting one or more other chroma samples in the region (e.g., within the same CU or PU where the current luma sample is located and that corresponds to the predicted chroma sample).

[0101] Technique 1000 explains a method for predicting chroma samples corresponding to luma samples. While some cases may involve predicting chroma samples for each luma sample accordingly (e.g., in a CU, CTU, or elsewhere), in some cases, technique 1000 can be used to predict chroma samples for some but not all luma samples.

[0102] The aspects of encoding and decoding described above illustrate some examples of encoding and decoding techniques. However, it should be understood that when those terms are used in the claims, encoding and decoding may mean compressing data, decompressing data, transforming data, or any other processing or change of data.

[0103] The word "example" is used herein to mean serving as an example, instance or illustration. Any aspect or design described herein as an "example" is not necessarily to be interpreted as being preferred or advantageous over other aspects or designs. Rather, the use of the word "example" is intended to present concepts in a specific manner. As used in this application, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clearly indicated in the context, the statement "X includes A or B" is intended to mean any one of its natural inclusive arrangements. That is, if X includes A; X includes B; or X includes both A and B, then "X includes A or B" is satisfied under any of the above examples. In addition, unless otherwise specified or the context clearly indicates a singular form, the article "one" used in this application and the appended claims should generally be interpreted as meaning "one or more". In addition, the use of the term "implementation" or "an implementation" throughout this disclosure is not intended to mean the same implementation, unless so described.

[0104] The implementation of the sending station 102 and / or the receiving station 106 (and the algorithms, methods, instructions, etc. stored thereon and / or executed by it (including by the encoder 400 and the decoder 500 or another encoder or decoder as disclosed herein)) can be implemented in hardware, software, or any combination thereof. The hardware may include, for example, a computer, an intellectual property (IP) core, an application specific integrated circuit (ASIC), a programmable logic array, an optical processor, a programmable logic controller, a microcode, a microcontroller, a server, a microprocessor, a digital signal processor, or any other suitable circuit. In the claims, the term "processor" should be understood to cover any of the aforementioned hardware, either individually or in combination. The terms "signal" and "data" are used interchangeably. Further, the parts of the sending station 102 and the receiving station 106 do not necessarily have to be implemented in the same manner.

[0105] Furthermore, in one aspect, for example, the sending station 102 or the receiving station 106 can be implemented using a general purpose computer or general purpose processor with a computer program that, when executed, performs any of the corresponding methods, algorithms and / or instructions described herein. Additionally or alternatively, for example, a special purpose computer / processor can be utilized that can include other hardware for performing any of the methods, algorithms or instructions described herein.

[0106] The sending station 102 and the receiving station 106 can be implemented, for example, on a computer in a video conference system. Alternatively, the sending station 102 can be implemented on a server, and the receiving station 106 can be implemented on a device (such as a handheld communication device) separated from the server. In this example, the sending station 102 can encode the content into an encoded video signal and send the encoded video signal to the communication device. Then, the communication device can decode the encoded video signal then. Alternatively, the communication device can decode the content stored locally on the communication device, for example, the content not sent by the sending station 102. Other suitable transmission and reception implementation schemes are available. For example, the receiving station 106 can be a generally fixed personal computer instead of a portable communication device.

[0107] In addition, all or part of the implementation of the present disclosure may take the form of a computer program product that can be accessed from, for example, a computer-usable or computer-readable medium. A computer-usable or computer-readable medium can be any device that can, for example, tangibly contain, store, communicate, or transmit a program for use by or in conjunction with any processor. The medium can be, for example, an electronic, magnetic, optical, electromagnetic, or semiconductor device. Other suitable media are also available.

[0108] The above implementations and other aspects have been described to facilitate easy understanding of the present disclosure and are not intended to limit the present disclosure. On the contrary, the present disclosure is intended to cover various modifications and equivalent arrangements included within the scope of the appended claims, which scope should be given the broadest interpretation allowed under the law to cover all such modifications and equivalent structures.

Claims

1. A method for region-based cross-component prediction, the method comprising: Identifying a region within a frame to be encoded or decoded; determining regional filter coefficients for the region; determining an input value of a current brightness sample within a portion of the region; determining a predicted chroma sample for the current luma sample based on the input value and the regional filter coefficients; as well as The predicted chroma samples are encoded or decoded.

2. The method of claim 1 , wherein determining the regional filter coefficients comprises: At least a portion of the regional filter coefficients are derived based on one or both of the region or neighboring regions.

3. The method of claim 2, wherein deriving at least the portion of the regional filter coefficients based on one or both of the region or the neighboring regions comprises: A mean square error between predicted chroma samples and reconstructed chroma samples within a reference region of the frame is minimized.

4. The method of claim 3, wherein the mean square error is performed using chrominance samples from a fill area outside the region.

5. The method of claim 1 , wherein determining the regional filter coefficients comprises: One or more syntax elements for signaling the regional filter coefficients are decoded from a bitstream associated with the frame.

6. The method of any one of claims 1, 2, 3, 4 or 5, comprising: Determining the regional filter coefficients for use in determining the predicted chroma samples based on a classification of the current luma sample.

7. The method of claim 6, wherein the portion of the region is a code processing unit, and wherein different regional filter coefficients are used to determine a second predicted chrominance sample based on a classification of a second current luma sample within the code processing unit.

8. The method of any one of claims 1, 2, 3, 4 or 5, wherein identifying the region comprises: One or more syntax elements signaled within the bitstream associated with the region are decoded.

9. The method of any one of claims 1, 2, 3, 4 or 5, wherein determining the predicted chroma samples comprises: determining a spatial weight value for the zone based on a prediction mode for the zone of the portion of the area; as well as The predicted chroma samples are determined using the spatial weight values.

10. The method of any one of claims 1, 2, 3, 4 or 5, wherein the portion of the region is a code processing unit and the regional filter coefficients are determined for use with multiple code processing units of the region.

11. The method of any one of claims 1, 2, 3, 4 or 5, wherein the size of the region is larger than a minimum chroma unit size.

12. The method of claim 11, wherein the region is a code processing tree unit having a size of 128x128 or 64x64.

13. An apparatus for region-based cross-component prediction, the apparatus comprising: Memory; as well as a processor configured to execute instructions stored in the memory to: determining regional filter coefficients for a region within a frame to be encoded or decoded; determining, based on an input value of a first luma sample within a first portion of the region and based on the regional filter coefficients, a first predicted chroma sample for the first luma sample; determining a second predicted chroma sample for a second luma sample within a second portion of the region based on an input value of the second luma sample and based on the regional filter coefficients; as well as The first predicted chroma samples and the second predicted chroma samples are encoded or decoded.

14. The apparatus of claim 13, wherein a first portion of the regional filter coefficients are signaled within a bitstream associated with the frame, and a second portion of the regional filter coefficients are derived based on video data within the frame.

15. The apparatus of claim 13, wherein the region is a current code processing tree unit, and the regional filter coefficients are derived using reconstructed chroma samples from one or more neighboring code processing tree units of the current code processing tree unit.

16. The apparatus of any one of claims 13, 14 or 15, wherein the regional filter coefficients are used for both the first predicted chroma samples and the second predicted chroma samples based on the classification of the first luma samples and the second luma samples.

17. The apparatus of claim 16, wherein the classification is based on one or more of gradient, direction, or pixel value bands.

18. A non-transitory computer-readable storage device comprising program instructions executable by one or more processors, the program instructions when executed causing the one or more processors to perform operations for region-based cross-component prediction, the operations comprising: determining filter coefficients for predicting chrominance samples within a plurality of code processing units of a code processing tree unit within a frame to be encoded or decoded; determining a current brightness sample within a code processing unit of the plurality of code processing units; determining a predicted chroma sample for the current luma sample based on an input value and the filter coefficients; as well as The predicted chroma samples are encoded or decoded.

19. The non-transitory computer-readable storage device of claim 18, wherein determining the filter coefficients comprises one of: deriving the filter coefficients based on one or both of the code processing tree unit or a neighboring code processing tree unit of the code processing tree unit; decoding, from a bitstream associated with the frame, one or more syntax elements for signaling the filter coefficients; or A first portion of the filter coefficients is derived and a second portion of the filter coefficients is decoded from the bitstream.

20. The non-transitory computer-readable storage device of claim 18, wherein determining the predicted chroma samples comprises: Determining a spatial weight value of the region according to a prediction mode for the region of the code processing unit; as well as The predicted chroma samples are determined using the spatial weight values.