Simplifying the cross-component model

By classifying reference samples into multiple models and applying them based on classification, the method reduces computational complexity in video encoding and decoding, enhancing efficiency and accuracy in chroma sample prediction.

JP2026507623APending Publication Date: 2026-03-04INTERDIGITALCE PATENT HLDG SAS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-16
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

Conventional cross-component-based intra prediction techniques for video coding increase computational complexity due to the use of multiple reference samples and model parameters, which complicates the prediction process.

Method used

A method and apparatus for video encoding and decoding that reduce computational complexity by simultaneously accumulating data for two models using a first loop through reference samples classified into different classes, and applying these models in a second loop to predict chroma samples based on their classification, with the option to rescale parameters to a lower bit-size representation.

Benefits of technology

This approach reduces computational complexity while maintaining accurate chroma sample prediction, improving efficiency in video encoding and decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026507623000001_ABST
    Figure 2026507623000001_ABST
Patent Text Reader

Abstract

Disclosed are apparatuses and methods including techniques for encoding and decoding video data. The disclosed techniques include obtaining video data including data representing a video data domain, and then calculating models used for cross-component-based prediction of chroma samples from the video data domain. The calculation of the models includes simultaneously accumulating data for multiple models using a single loop through reference samples selected from the video data. The accumulated data for the models is generated based on the reference samples according to their classification into respective classes. The models are then applied using a single loop through samples from the video domain, and the models are applied to predict chroma samples according to their classification into respective classes.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS [1] This application claims the benefit of European Patent Application No. 23305278.6, filed March 2, 2023, which is incorporated herein by reference in its entirety. [Background technology]

[0002] background [2] Predictive video coding employs prediction to exploit spatial and temporal redundancy in video content. Typically, to encode a video block, intra or inter prediction is applied to the video block to exploit spatial or temporal correlation, and then the difference between the original and predicted video block is transformed, quantized, and entropy coded. To reconstruct the video block, the inverse process corresponding to entropy coding, quantization, transformation, and prediction is applied. Cross-component-based intra prediction techniques are sometimes used by encoders, in which chroma samples are predicted based on reconstructed luma samples according to a prediction model derived from reference samples spatially associated with the chroma samples. The use of multiple reference samples and model parameters can increase the accuracy of chroma sample prediction. However, improved prediction accuracy comes at the cost of increased computational complexity associated with both deriving and applying the prediction model. Summary of the Invention

[0003] overview [3] Aspects disclosed in this disclosure describe a method for encoding and decoding video data. The method includes obtaining video data including data representing a video data domain and then calculating a model used for cross-component-based prediction of chroma samples from the video data domain. The calculation of the model includes simultaneously accumulating data for a first model and a second model using a first loop through reference samples selected from the video data. In one aspect, the accumulated data for the first model and the second model are generated based on the reference samples according to their respective classification into a first class and a second class. In another aspect, parameters of the first model and the second model may be rescaled to a lower bit-size representation. Then, using a second loop through samples of the video domain, the first model and the second model are applied, and the models are applied to predict chroma samples according to their classification into the first class or the second class.

[0004] [4] Aspects disclosed in this disclosure describe an apparatus for encoding and decoding video data. The apparatus includes at least one processor and a memory storing instructions. When executed by the at least one processor, the instructions cause the apparatus to obtain video data including data representing a video data region and then calculate a model used for cross-component-based prediction of chroma samples from the video data region. The calculation of the model includes simultaneously accumulating data for a first model and a second model using a first loop through reference samples selected from the video data. In one aspect, the accumulated data for the first model and the second model are generated based on the reference samples according to their respective classification into a first class and a second class. In another aspect, parameters of the first model and the second model may be rescaled to a lower bit-size representation. The instructions further cause the system to apply the first model and the second model using a second loop through samples of the video region, where the models are applied to predict chroma samples according to their classification into the first class or the second class.

[0005] [5] A further aspect disclosed in this disclosure describes a non-transitory computer-readable medium including instructions executable by at least one processor to implement a method for encoding and decoding video data. The method includes obtaining video data including data representing a video data domain and then calculating a model used for cross-component-based prediction of chroma samples from the video data domain. The calculation of the model includes simultaneously accumulating data for a first model and a second model using a first loop through reference samples selected from the video data. In one aspect, the accumulated data for the first model and the second model are generated based on the reference samples according to their respective classification into a first class and a second class. In another aspect, parameters of the first model and the second model may be rescaled to a lower bit-size representation. Then, using a second loop through samples of the video domain, the first model and the second model are applied, and the models are applied to predict chroma samples according to their classification into the first class or the second class.

[0006] [6] This Summary section is provided to introduce some concepts in a simplified form that are further described below in the Detailed Description section. This Summary section is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Moreover, the claimed subject matter is not limited to limitations that solve any or all disadvantages noted in any part of this disclosure. [Brief explanation of the drawings]

[0007] BRIEF DESCRIPTION OF THE DRAWINGS [Figure 1] [7] FIG. 1 is a block diagram of an exemplary system in which aspects of the present embodiments may be implemented. [Figure 2] [8] FIG. 1 is a functional block diagram of an exemplary video encoder in which aspects of the present embodiments may be implemented. [Figure 3][9] FIG. 1 is a functional block diagram of an exemplary video decoder in which aspects of the present embodiments may be implemented. [Figure 4]

[10] A diagram showing a reference area for CC-based prediction in which aspects of the present embodiment can be implemented. [Figure 5]

[11] A diagram illustrating multi-model CC-based prediction, in which aspects of the present embodiments can be implemented. [Figure 6]

[12] FIG. 1 is a flowchart of an exemplary method for CC-based prediction in which aspects of the present embodiments may be implemented. [Figure 7]

[13] is a flowchart illustrating the derivation of a model for CC-based prediction, which can implement aspects of the present embodiments. [Figure 8]

[14] is a flowchart illustrating the application of multi-model CC-based prediction, which can implement aspects of the present embodiments. [Figure 9]

[15] is a flowchart illustrating the joint derivation of a model for CC-based prediction, which can implement aspects of the present embodiments. [Figure 10]

[16] is a flowchart illustrating the joint application of multi-model CC-based prediction, which can realize aspects of the present embodiments. [Figure 11]

[17] FIG. 1 is a flowchart of an exemplary method in which aspects of the present embodiments may be implemented. DETAILED DESCRIPTION OF THE INVENTION

[0008] Detailed Description

[18] Apparatuses and methods for video encoding and video decoding are disclosed herein. Aspects of the disclosure describe techniques for reducing the computational complexity of deriving and applying models for cross-component-based intra prediction. Conventional systems and methods for predictive video coding are described below with reference to Figures 1-3, followed by a description of aspects of the disclosure with reference to Figures 4-11.

[0009]

[19] Figure 1 shows a block diagram of an exemplary system 100. System 100 may be embodied as a device and configured to implement one or more of the aspects described herein. Examples of such devices include, but are not limited to, various electronic devices, such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 100, singly or in combination, may be embodied as an integrated circuit, multiple integrated circuits, and / or discrete components. For example, in at least one embodiment, processing element 110 and encoder / decoder element 130 of system 100 are distributed across multiple integrated circuits and / or discrete components. In various embodiments, system 100 is communicatively coupled to other systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports.

[0010]

[20] System 100 includes at least one processor 110, which may be configured to execute loaded instructions, for example, to implement various aspects described herein. Processor 110 may include embedded memory, input / output interfaces, and various other circuitry known in the art. System 100 includes at least one memory 120, such as a volatile memory device and / or a non-volatile memory device. System 100 includes storage device 140, which may include non-volatile memory and / or volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drives, and / or optical disk drives. Storage device 140 may include, for example, an internal storage device, an external storage device, and / or a network-accessible storage device.

[0011]

[21] System 100 includes an encoder / decoder module 130 configured to process data and provide encoded or decoded video data. Encoder / decoder module 130 may include its own processor and memory. Encoder / decoder module 130 may be implemented as a separate element of system 100 or may be incorporated within processor 110 as a combination of hardware and / or software, as known to those skilled in the art. Furthermore, encoder / decoder module 130 represents a module that may be implemented in a separate device to perform encoding and / or decoding functions.

[0012]

[22] Program code to be loaded into the processor 110 or the encoder / decoder 130 to implement various aspects described herein may be stored in the storage device 140 and later loaded into the memory 120 for execution by the processor 110. According to various embodiments, one or more of the processor 110, the memory 120, the storage device 140, and the encoder / decoder module 130 may store one or more of various items during the performance of the processes described herein. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, arithmetic logic, and intermediate or final results from processing equations, formulas, and operations.

[0013]

[23] In some embodiments, memory internal to processor 110 and / or encoder / decoder module 130 is used to store instructions and provide working memory for processing functions required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be either processor 110 or encoder / decoder module 130) may be used for one or more of these functions. The external memory may be memory 120 and / or storage device 140 and may comprise, for example, dynamic volatile memory and / or non-volatile flash memory. In some embodiments, external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, high-speed external dynamic volatile memory, such as RAM, is used as working memory for video coding and decoding operations.

[0014]

[24] Input to the elements of system 100 may be provided through various input devices, as shown in block 105. Such input devices may include, but are not limited to, (i) an RF section that receives RF signals transmitted over the air, for example by a broadcast station, (ii) a composite input (COMP), (iii) a USB input, and / or (iv) an HDMI input.

[0015]

[25] In various embodiments, the input devices of block 105 have respective associated input processing elements known in the art. For example, the RF section may be associated with elements suitable for (i) selecting a desired frequency (also referred to as signal selection or band-limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) band-limiting again to a narrower frequency band, e.g., to select a signal frequency band, which may be referred to as a channel in certain embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF section of various embodiments includes one or more elements for performing these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF section may include, for example, a tuner for performing some of these functions, including down-converting a received signal to a lower frequency (e.g., an intermediate frequency or a frequency near baseband) or to baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted over a wired (e.g., cable) medium and perform frequency filtering, down-conversion, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.

[0016]

[26] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 100 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented, for example, within a separate input processing integrated circuit or within processor 110, as desired. Similarly, aspects of USB or HDMI interface processing may be implemented, for example, within a separate interface integrated circuit or within processor 110, as desired. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 110 operating in combination with memory and storage elements, and encoder / decoder 130, to process the data stream as desired for presentation to an output device.

[0017]

[27] The various elements of the system 100 may be provided within a unitary housing in which the various elements may be interconnected and data may be transmitted between them using a suitable connection arrangement 115, such as an internal bus known in the art, including an I2C bus, wiring, and a printed circuit board.

[0018]

[28] System 100 includes a communication interface 150 that enables communication with other devices over a communication channel 190. Communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data over communication channel 190. Communication interface 150 may include, but is not limited to, a modem or a network card. Communication channel 190 may be implemented in a wired and / or wireless medium, for example.

[0019]

[29] In various embodiments, data may be streamed to system 100 using a Wi-Fi network, such as IEEE 802.11. The Wi-Fi signal in these embodiments is received via communication channel 190 and communication interface 150, which may be adapted for Wi-Fi communication. Communication channel 190 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to enable streaming applications and other over-the-top communications. In other embodiments, data may be streamed to system 100 using a set-top box that delivers data via an HDMI connection in input block 105, or data may be streamed to system 100 using an RF connection in input block 105.

[0020]

[30] System 100 can provide output signals to various output devices, including display device 165, audio device (e.g., speakers) 175, and other peripheral devices 185. Other peripheral devices 185, in various example embodiments, include one or more of a standalone DVR, a disc player, a stereo system, a lighting system, and other equipment that provides functionality based on the output of system 100. In various embodiments, control signals are communicated between system 100 and display 165, audio device 175, or other peripheral devices 185 using signaling such as AV.Link, CEC, or other communication protocols that enable inter-device control with or without user intervention. Output devices may be communicatively coupled to system 100 via dedicated connections via respective interfaces 160, 170, and 180. Alternatively, output devices may be connected to system 100 using communication channel 190 via communication interface 150. Display 165 and audio device 175 may be integrated as a single unit with other elements of system 100 within an electronic device, such as a television. In various embodiments, display interface 160 includes a display driver, such as a timing controller (TCon) chip.

[0021]

[31] Alternatively, display device 165 and audio device 175 may be separate from one or more of the other components, for example, if the RF portion of input 105 is part of a separate set-top box. In various embodiments in which display device 165 and audio device 175 are external components, output signals may be provided via dedicated output connections, including, for example, an HDMI port, a USB port, or a COMP output.

[0022]

[32] Figure 2 shows a functional block diagram of an exemplary video encoder 200. Video encoder 200 may be employed by system 100 described with reference to Figure 1. For example, video encoder 200 may be an encoder operating according to a coding standard such as Advanced Video Coding (AVC, H.264 / MPEG-4 | ISO / IEC 14496-10), High Efficiency Video Coding (HEVC, ITU-T H.265 | ISO / IEC 23008-2), or Versatile Video Coding (VVC, Standard ITU-T H.266, ISO / IEC 23090-3, 2020).

[0023]

[33] Prior to encoding, the video data may be preprocessed by a precoding processor (not shown). Such preprocessing may include applying a color model transformation to the color components of the input video frames (e.g., converting from RGB 4:4:4 to YCbCr 4:2:0) or mapping the color components of the input video frames to obtain a signal distribution that is more resistant to compression (e.g., applying a histogram equalizer and / or a denoising filter to one or more of the color components of the video frames). Preprocessing may also include associating metadata with the video data, which may be attached to the encoded video bitstream.

[0024]

[34] In encoder 200, video frames are encoded by encoder elements, generally as described below. An original video picture (frame) to be encoded is partitioned into coding units (i.e., original blocks) by image partitioner 202. Typically, a coding unit (CU) includes a luma block and respective chroma blocks; therefore, operations described herein that apply to a CU generally apply to the luma block and respective chroma blocks. After partitioning 202, each CU may be coded using an intra-prediction mode or an inter-prediction mode. In intra-prediction mode, prediction of the CU is performed by intra predictor 260. In intra-prediction mode, the content of a CU within a frame is predicted based on content from one or more other CUs of the same frame using reconstructed versions of the other CUs (available from the output of adder 255). In inter-prediction mode, motion estimation and motion compensation are performed by motion estimator 275 and motion compensator 270, respectively. In inter-prediction mode, the content of a CU in a frame is predicted based on the content from one or more other CUs in neighboring frames using reconstructed versions of the other CUs (available from a reference picture buffer 280). The encoder determines 205 the prediction result to use for encoding the CU (either obtained by operating in intra-prediction mode 260 or by operating in inter-prediction mode 270, 275) and indicates the selected prediction mode, for example, by a prediction mode flag. The selected prediction result may then be enhanced (e.g., filtered) by a prediction enhancer 285 to output a respective prediction block. Once a prediction block is generated for each CU, a respective residual block is calculated, for example, by subtracting 210 the predicted CU (i.e., the prediction block) from the CU (i.e., the original block).

[0025]

[35] Each residual block or partition thereof (i.e., transform block) of a CU is then transformed into a coefficient block by a transformer 220. That is, the residual samples of the transform block are transformed into transform coefficients of the coefficient block. The resulting coefficient block is quantized by a quantizer 230. An entropy encoder 245 is then employed to entropy encode the quantized coefficient block and respective coding parameters (e.g., syntax elements including motion vectors and other control data). Thus, the entropy-coded quantized coefficient block and respective coding parameters associated with each video frame of the original video are packed into a bitstream of coded video data.

[0026]

[36] Following coding of an original block (CU) as described above, the encoder 200 reconstructs the coded original block to provide a reference for future prediction. Accordingly, the quantized coefficient blocks (provided by the quantizer 230) are inversely quantized by the inverse quantizer 240 and then inversely transformed by the inverse transformer 250 to reconstruct (decode) a residual block of each original block. Adding 255 the reconstructed residual block to the respective prediction block results in a respective reconstructed original block. An in-loop filter 265 may then be applied to the reconstructed picture (formed by the reconstructed original block) to perform, for example, deblocking filtering and / or sample adaptive offset (SAO) filtering to reduce coding artifacts. The filtered reconstructed picture may then be stored in a reference picture buffer 280 and available for future prediction in inter-prediction mode. Accordingly, the encoder 200 also performs decoding operations 240, 250, in which the coded picture (frame) is reconstructed. The reconstructed picture may then be stored in a reference picture buffer 280 and used to facilitate motion estimation 275 and compensation 270 as described above.

[0027]

[37] Figure 3 shows a functional block diagram of an exemplary video decoder 300. The video decoder 300 may be employed by the system 100 described with reference to Figure 1. In general, the operation of the video decoder 300 is the reverse of that of the video encoder 200. In the decoder 300, the bitstream of coded video data generated by the video encoder 200 is first entropy decoded by an entropy decoder 330, which decodes quantized coefficient blocks and various coding parameters from the bitstream. The quantized coefficient blocks are inversely quantized by an inverse quantizer 340 and then inversely transformed by an inverse transformer 350 to decode (reconstruct) respective residual blocks. Adding 355 the reconstructed residual blocks to respective predictive blocks results in respective reconstructed original blocks. Depending on the selected prediction mode, a predicted original block may be obtained 370 from an intra predictor 360 or a motion compensator 375 and then enhanced (e.g., filtered) by a prediction enhancer 390 to generate a predicted block. An in-loop filter 365 may be applied to the reconstructed picture (formed by the reconstructed original block) to output a reconstructed (decoded) video frame. The filtered reconstructed picture is also stored in a reference picture buffer 380 to facilitate motion compensation 375.

[0028]

[38] A post-decoding processor (not shown) can further process the reconstructed video. For example, the post-decoding processing can include an inverse color model conversion (e.g., YCbCr 4:2:0 to RGB 4:4:4) or an inverse mapping that reverses the mapping process performed by the pre-encoding processor. The post-decoding processor can use metadata derived by the pre-encoding processor and / or signaled in the video bitstream.

[0029]

[39] Although the aspects described herein are described with reference to CUs, the aspects described are equally applicable to any region of a video frame (i.e., a region of video data) to which intra prediction may be applied by encoder 260 or decoder 360. In general, the aspects described herein may be applied to a region of video data formed by a video partition of any shape or size.

[0030]

[40] A CU includes a luma component Y and chroma components Cr and Cb (any one of which is also referred to herein as C). Typically, the chroma component C is subsampled and therefore has reduced resolution relative to the corresponding luma component Y. In general, the image content of the chroma component C is correlated with the image content of the corresponding luma component Y and its nearby spatial neighborhood. To take advantage of such cross-component correlation, approaches exist for predicting chroma samples from the chroma component C based on corresponding luma samples derived from the reconstructed corresponding luma component Y. Some of these approaches, generally referred to herein as cross-component (CC)-based prediction, apply models described below, including the cross-component linear model (CCLM), the convolutional cross-component model (CCCM), and the gradient- and position-based convolutional cross-component model (GL-CCCM).

[0031]

[41] CCLM is a model for CC-based prediction, where chroma samples from the C component of a CU (coded in intra-prediction mode) are used to predict Y rs The chroma sample at pixel location (i,j), i.e., the linear prediction of the chroma sample C(i,j), can be expressed as follows:

number

number

number

[0032]

[42] There are several variations on CCLM. The variations vary: 1) the location and / or number N of reference pairs used to estimate the model parameters α and β; 2) the method for estimating the model parameters; or 3) the conversion of the luminance component Y to its downsampled version Y. rsThe type of filter that can be used when downsampling to 100% may be relevant. For example, when the Enhanced Compression Model (ECM) is used (see M. Coban, et al., "Algorithm description of Enhanced Compression Model 4 (ECM 4)," document JVET-Y2025, 23rd Meeting, by teleconference, 7-16 July 2021), the CCLM included in VVC is extended by adding a multi-model linear model (MMLM) mode (see K. Zhang, et al., "Enhanced Cross-component Linear Model Intra-prediction," document JVET-D0110). Multi-model CC-based prediction is described below with respect to Figure 5.

[0033]

[43] CCCM is another model for CC-based prediction (see P. Astola, et al., “AHG12: Convolutional cross-component model (CCCM) for intra prediction,” document JVET-Z0064, 26th Meeting, by teleconference, 20-29 April 2022). Like CCLM, CCCM can also be used to predict chroma samples based on corresponding subsampled reconstructed luma samples. Also, like CCLM, there is the option of using single-model or multi-model variants of CCCM, as described with respect to FIG. 5. Preferably, multi-model CCCM should be selected when a large number of reference pairs (e.g., N≧128) are available.

[0034]

[44] The parameters of CCCM include a 3 × 3 kernel K, a nonlinear term p, and a bias term b. The kernel K is a positive-signed kernel, and the kernel coefficients k are located at the center, north, south, west, and east of the pixel location where the kernel is convolved. C , kN , k S , k W , and k E For example, the luminance sample Y rs Applying kernel K to (i,j) gives the convolution result Y rs (i,j)*K is as follows: Y rs (i,j)*K=Y rs (i,j)·k c +Y rs (i,j-1)·k N + Y rs (i,j+1)·k s +Y rs (i-1,j)·k w +Y rs (i+1,j)·k E (3) The nonlinear term p is the Y scaled for the bit depth used. rs For example, for a bit depth of 10 bits, the nonlinear term p may be: p≡P(Y rs (i,j) 2 =(Y rs ,(i,j) 2 +512)>>10 (4) The bias term b may be determined as the median saturation value (eg, 512 for 10-bit content).

[0035]

[45] Therefore, the prediction of saturation samples based on the CCCM model can be expressed as follows:

number

number

number

number

[0036]

[46] CCCM saturation sample C p The prediction of (i,j) (formulated in equation (5)) can be expressed as follows:

number

[0037]

[47] As explained above, the parameter vector Φ can be estimated by minimizing the model's sum of squared errors as follows: E(n)=C ref (n)-s ref (n) T Φ,n∈1:N (8) where n is the index of the position (i,j) in the reference area used, and therefore s ref (n)=[y rs (i,j),y rs (i,j-1),y rs (i,j+1),y rs (i-1,j),y rs (i+1,j),p,b], and c ref (n)=c ref (i,j). Therefore, the error sum of squares

number

number

number

[0038]

[48] ​​To derive the model parameter vector Φ, the autocorrelation matrix A must be inverted, as shown in equation (9). For example, the autocorrelation matrix A is TThe parameter vector Φ can then be calculated using backsubstitution. Generally, this process follows adaptive loop filtering (ALF) when ECM is used (see M. Coban, et al., "Algorithm description of Enhanced Compression Model 7 (ECM 7)," document JVET-AB2025, 28th Meeting, by teleconference, October 2022). To eliminate the use of square root operations, LDL decomposition is used instead of Cholesky decomposition. T decomposition is used (hence also known as alternative Cholesky decomposition). In another variant, the autocorrelation matrix can be inverted using a Gaussian elimination technique (see J. Lainema, et al., “AHG12: Simplified linear model solver,” document JVET-AC0053, 29th Meeting, by teleconference, 11-20 January 2023). When ECM is used, the inversion of the autocorrelation matrix uses integer 64-bit arithmetic.

[0039]

[49] Yet another model of CC-based prediction is GL-CCCM, which uses gradient and location information for prediction (see R.G. Youlavari, et al., “EE2-1.12: Gradient and location based convolutional cross-component model (GL-CCCM) for intra prediction,” document JVET-AC0054, 29th Meeting, by teleconference, 11-20 January 2023). Therefore, GL-CCCM can be expressed as follows:

number

[0040]

[50] As explained above, the reconstructed luminance samples Y r is downsampled to match the lower resolution of the chroma samples C, and the downsampled luma samples Y rsIn one aspect, downsampling can be avoided by directly using the reconstructed luma samples. For example, (as in HJJhu, et al., "EE2-1.13 and 1.14: CCCM using non-downsampled luma samples," document JVET-AC0147, 29th Meeting, by teleconference, 11-20 January 2023), Y rs The observation vector s associated with the saturation sample corresponding to position (i,j) in may be: s=[Y r (i,j),L0,L1,L2,L3,L4,L5,p0,p1,p2,p4,b], (15) where L0=Y r (i-2,j-1), L1=Y r (i,j-1), L2=Y r (i+2,j-1), L3=Y r (i-2,j+1), L4=Y r (i,j+1), and L5=Y r (i+2,j+1). The elements p0, p1, p2, and p4 represent nonlinear functions of L0, L1, L2, and L4, respectively. Furthermore, the bias term b can be determined as the median saturation value (e.g., 512 for 10-bit content). This model is based on 12 parameters Φ≡(φ0,φ1,φ2,φ3,φ4,φ5,φ6,φ7,φ8,φ9,φ 10 ,φ 11 )

[0041]

[51] Figure 4 shows a diagram of the reference area used for CC-based prediction. In the example of Figure 4, a luma sample 400A and a corresponding (subsampled) chroma sample 400B are shown. A luma block (e.g., from a CU or its partition) and its corresponding chroma block are indicated by a dark grey square. The reference luma sample y ref (n) containing the reference luminance area and the reference chroma sample c refThe corresponding reference chroma area containing (n) is indicated by a white square. Corresponding samples from the reference luma area and the reference chroma area can be used to calculate the parameters of any of the models mentioned above (e.g., CCLM, CCCM, or GL-CCCM). Thus, the reference luma sample y ref (n) and their corresponding reference saturation samples c ref (n) may be selected from the respective reference areas. Because chroma content 400B is subsampled relative to luma content 400A (e.g., when using a 4:2:0 format), the chroma samples correspond to luma samples that may be derived (or interpolated) from the luma content of the four luma samples (e.g., the four luma samples indicated by the dotted rectangle in 400A are downsampled to one luma sample corresponding to the chroma sample, indicated by the dotted rectangle in 400B).

[0042]

[52] As shown in FIG. 4, three regions may be defined within both the luma and chroma reference areas. These are denoted as R1, R2, and R3. In one aspect, reference samples may be selected from regions explicitly signaled in the bitstream (e.g., at the CU level). For example, reference chroma and luma samples may be selected from regions R1, R2, R3, or a combination thereof. Extended areas (e.g., samples indicated by light gray squares) may also be used to support filtering of reference luma and chroma samples (white squares) located along the boundaries of the luma and chroma reference areas. For each luma and corresponding chroma block, a reference area containing reconstructed samples is determined. However, if reconstructed samples are not available, padding may be applied. For example, the extended area may be used to support convolution with a 3×3 kernel K (see Equation (3)), and if samples within this extended area are not available, they may be extrapolated from available reconstructed samples or padded with zeros.

[0043]

[53] Figure 5 illustrates a multi-model CC-based prediction 500. Multi-model prediction can be used with any (or a combination) of the models mentioned above (e.g., CCLM, CCCM, or GL-CCCM). When a multi-model predictor is applied, first, N reference pairs (y ref (n),c ref (n)) are first classified into multiple classes. The classification is performed by using the reference luminance samples y ref (n) and / or reference saturation sample c ref (n) may be based on a feature space (e.g., including features representing the intensity levels of luminance samples). For example, for two classes, namely class A 510 and class B 520, a reference pair (y ref (n),c ref (n)) can be divided into two groups based on a threshold T 530. This threshold T can be, for example, ref (n) of the pixel intensity levels. As shown in FIG. 5, samples below (or equal to) the threshold are classified into a first class 510 (denoted by black circles), and samples above the threshold are classified into a second class 520 (denoted by white circles). A model associated with each class may be derived (e.g., as described above with respect to Equation (9)) based on reference pairs within each class. Thus, a parameter vector Φ derived based on reference pairs from class A 510 may be A and a parameter vector Φ derived based on a reference pair from class B520. B and a second model having: Accordingly, the first model can be used to predict chroma samples using corresponding luma samples belonging to class A 510, and the second model can be used to predict chroma samples using corresponding luma samples belonging to class B 520. Multi-model CC-based prediction using any of the above-described models for each of any number of classes is generally described with reference to FIG.

[0044] 6 is a flowchart of an example method for CC-based prediction 600. The CC-based prediction method 600 predicts chroma samples of a CU (or a partition thereof) based on corresponding luma samples from the CU's reconstructed and (possibly) subsampled luma samples. The method 600 begins in step 610 with selecting reference samples, e.g., N reference pairs y ref (n) and c ref (n) may be selected from the reference area. In step 620, as described with reference to FIG. 5, if multi-model prediction is applied, the reference samples are classified into multiple classes. If single-model prediction is applied, step 620 can be skipped, and all reference samples are considered to belong to the same class. In step 630, a model is derived for each class. Accordingly, for each model of a class, a parameter vector Φ of the model is derived based on the reference samples from the respective class. Then, in step 640, a chroma sample of a CU may be predicted based on the corresponding luma sample using the model (derived in step 630) associated with the class to which the corresponding luma sample belongs.

[0045]

[55] In general, reference samples in a reference area can be classified into Q classes. Such classification can be based on features extracted from luma samples and / or chroma samples in the reference area. A respective model can be derived for the Q classes, with a respective parameter vector {Φ q :q=1,…,Q} for each parameter vector Φ q is used to predict chroma samples based on corresponding luma samples belonging to its associated class q. When performing CC-based prediction, the encoder 200 may apply a single model prediction (e.g., using parameter vector Φ) or a single model prediction (e.g., using parameter vector {Φ q :q=1,...,Q}) to select whether to apply multi-model prediction.

[0046]

[56] Figure 7 is a flowchart illustrating the derivation of a model for CC-based prediction 700. The encoder 200 may be configured to select between a single model Φ (derived as shown in flowchart portion 700A) and multiple models Φ and Φ (derived as shown in flowchart portions 700B and 700C) to predict chroma samples of a CU based on the model that provides a better prediction. As described with respect to equations (9)-(11), to derive these models, the respective autocorrelation matrices A and cross-correlation vectors B must be calculated. Accordingly, to derive Φ, the autocorrelation matrix A and cross-correlation vector B are initialized 705. N reference samples (e.g., samples selected from reference areas R, R, R, or a combination thereof) are then looped over 725, whereby the respective contributions of these samples to the elements of A and B may be added 710 (e.g., according to equations (10) and (11)). Once all samples have been processed 715, a model Φ is derived 720 (e.g., according to equation (9)). Next, an autocorrelation matrix A and a cross-correlation vector B are initialized 725 to derive a model Φ. Then, looping 745 over the N reference samples, the contributions of samples belonging to a first class (e.g., the class having luminance samples with intensity levels below a threshold T 730) are added 735 to the elements of A and B (e.g., according to equations (10) and (11)). Once all samples have been processed 740, a model Φ is derived 750 (e.g., according to equation (9)). Similarly, an autocorrelation matrix A and a cross-correlation vector B are initialized 755 to derive a model Φ. Then, looping 775 over the N reference samples, the contributions of samples belonging to a second class (e.g., the class having luminance samples with intensity levels above threshold T 760) are added 765 to the elements of A2 and B2 (e.g., according to equations (10) and (11)). Once all samples have been processed 770, a model Φ2 is derived 780 (e.g., according to equation (9)).Once models Φ, Φ, and Φ are derived 720, 750, 780, the encoder 200 may determine whether to predict the chroma samples of the CU based on a single model Φ or based on multiple models Φ and Φ. For example, the encoder may apply both single and multiple models to predict the chroma samples and then select the one that results in a lower cost (e.g., rate-distortion cost).

[0047]

[57] Figure 8 is a flowchart illustrating the application of multi-model CC-based prediction 800. Conventionally, as shown in Figure 8, multi-model prediction (e.g., based on CCCM) may be applied in two steps. First, chroma samples of the CU to be predicted are tested to see if they belong to a first class (e.g., corresponding luma samples with intensity values ​​below threshold T 810), looping 835 through the CU's samples until all samples have been processed 830. Samples found to belong to the first class are predicted 820 based on model Φ1. Next, samples of the CU to be predicted are tested to see if they belong to a second class (e.g., corresponding luma samples with intensity values ​​above threshold T 860), looping 885 through the CU's samples again until all samples have been processed 880. Samples found to belong to the second class are predicted 870 based on model Φ2.

[0048]

[58] The above CC-based prediction model has the following computational drawbacks. The calculation of the autocorrelation matrix and cross-correlation vectors associated with models Φ0, Φ1, and Φ2 involves two tests 730, 760 and three loops 725, 745, 775 through the reference samples, as described with reference to FIG. 7. For the application of a multi-model predictor, prediction is performed using two tests 810, 860 and two loops 835, 885 through the CU samples, as shown in FIG. 8. Furthermore, the calculation of the multi-model threshold T requires an additional loop through the reference samples. Furthermore, when calculating the model's parameter vector, the autocorrelation matrix (a positive definite symmetric matrix) must be inverted as shown in Equation (9). Traditionally, matrix inversion is performed based on successive row permutations with weighted linear combinations of other rows. This calculation is performed using integer arithmetic, and division is performed by tabular division (using lookup tables). At the end of this process, large integer values ​​may be obtained that require representation by 64-bit integers. As a result, when the model is applied to estimate saturation samples (see equations (7) or (12)), this estimation may require costly resources in terms of computational operations applied to 64-bit data, as well as memory access and storage of the 64-bit data.

[0049]

[59] Aspects of the present disclosure address the above-mentioned shortcomings. According to some aspects, calculation of autocorrelation matrices and cross-correlation vectors associated with models Φ, Φ, and Φ may be performed using a single test and a single loop through reference samples, as further described with reference to FIG. 9. Furthermore, application of multiple models may be performed using a single test and a single loop through the samples of a CU, as further described with reference to FIG. 10. Furthermore, as disclosed herein, the threshold T may be calculated simultaneously with the calculation of the autocorrelation matrices and cross-correlation vectors, eliminating the need for a separate loop through the reference samples. Furthermore, to reduce computational complexity, model parameters may be represented by fewer than 64 bits (e.g., 32-bit integers), as further described below.

[0050]

[60] Figure 9 illustrates a joint derivation of models for CC-based prediction 900. In the joint model derivation 900, the autocorrelation matrices and cross-correlation vectors of each model Φ, Φ, and Φ are calculated simultaneously, so that the calculations can be accomplished in a single loop through the reference samples, as will now be described with reference to the example of Figure 9. Thus, following initialization 905 of the autocorrelation matrices A1 and A2 and the cross-correlation vectors B1 and B2, the joint model derivation 900 is performed by first looping 945 through the N reference samples used to predict the chroma samples associated with the CU (e.g., samples selected from reference areas R1, R2, R3, or a combination thereof). If a sample belongs to a first class (e.g., the class representing those luma samples having intensity values ​​below a threshold T in 910), the sample's contribution to the elements of the autocorrelation matrix A1 and the cross-correlation vector B1 is added (or accumulated) 920 to A1 and B1, respectively. For example, for sample n belonging to the first class, its contribution s i (n)·s j (n) is added to the element Ai,j of the autocorrelation matrix A1, and its contribution s i (n C(n) is the element B of the cross-correlation vector B1 i (see equations (10) and (11)). Otherwise, if the sample belongs to a second class (e.g., the class representing those luminance samples having intensity values ​​above the threshold T in 910), then the sample's contribution to the elements of its respective autocorrelation matrix A2 and to its respective cross-correlation vector B2 is added (or accumulated) 930 to A2 and B2, respectively. For example, for sample n belonging to the second class, its contribution s i (n)·s j (n) is added to the element A(i,j) of the autocorrelation matrix A2, and its contribution s i (n)·C(n) is added to the elements B i of the cross-correlation vector B (see equations (10) and (11)). Once all N reference samples have been processed 940, a model Φ can be derived 960 based on the autocorrelation matrix A and the cross-correlation vector B (i.e., Φ = (A) -1B1), a model Φ2 can be derived 970 based on the autocorrelation matrix A2 and the cross-correlation vector B2 (i.e., Φ2=(A1) -1 B2). To derive a model Φ, the autocorrelation matrices A and A can be combined to A = A + A, and the cross-correlation vectors B and B can be combined to B = B + B. A model Φ can then be derived 950 based on the autocorrelation matrix A and the cross-correlation vector B (i.e., Φ = (A) -1 B0).

[0051]

[61] Thus, in the joint model derivation 900, the elements of the autocorrelation matrix and the elements of the cross-correlation vector (used to calculate the respective models Φ0950, Φ1960, Φ2970) are accumulated through the reference samples associated with the CU using one test 910 and a single loop 945. This contrasts with the model derivation described above with respect to Figure 7, where the accumulation of the elements of the autocorrelation matrix and the cross-correlation vector (used to calculate the respective models Φ0720, Φ1750, and Φ2780) is done separately, using two tests 730, 760, and involving three loops 725, 745, 775 through the reference samples.

[0052]

[62] In one aspect, the reference area may be divided into multiple regions (e.g., regions R1, R2, and R3 in Figure 4), and autocorrelation matrices and cross-correlation vectors may be calculated for each region. Similar to the example of Figure 9, a single loop may be used to go through the reference samples, accumulating the contributions of samples from different regions into their respective autocorrelation matrices and cross-correlation vectors (e.g.,

number

number

number

[0053]

[63] As mentioned above, conventionally, multi-model CC-based prediction is applied in two steps, as described with reference to Figure 8, where two loops 835, 885 through the samples of the CU (for which chroma samples are predicted) are used and two different tests 810, 860 are applied. To reduce complexity, in one aspect, the two models of multi-model CC-based prediction are first derived, and then these models are applied using a single test and a single loop, as described with reference to Figure 10.

[0054]

[64] Figure 10 is a flowchart illustrating the joint application of multi-model CC-based prediction 1000. In the example of Figure 10, model Φ11010 and model Φ21020 are first derived, for example, as described with respect to Figure 9. The samples of the CU are then looped 1070 to apply these models. If the sample belongs to a first class (e.g., a corresponding luminance sample having an intensity value less than or equal to threshold T1030), model Φ1 is applied 1040 to the sample (e.g., according to equation (7) or (12)). Otherwise, if the sample belongs to a second class (e.g., a corresponding luminance sample having an intensity value greater than threshold T1030), model Φ2 is applied 1050 to the sample. The application 1040, 1050 of the models to the samples of the CU continues until all samples have been processed 1060. As described above, in this embodiment, models Φ1 and Φ2 are applied in a single loop 1070 using a single test 1030.

[0055]

[65] In a further aspect, the computational complexity of extracting features to facilitate classification of reference samples into different classes (e.g., class A 510 and class B 520) can be reduced. For example, one or more classification features can be extracted from a smaller region (e.g., a sub-region of a reference area or CU). The features can be determined incrementally while looping 945 through the reference samples to calculate the autocorrelation matrix and cross-correlation vectors 920, 930. Thus, in one aspect, while looping 945 through the reference samples, a histogram can be updated based on sequentially obtained reference samples, and one or more features can be recalculated based on the updated histogram.

[0056]

[66] For example, a single line or a single column within the reference area (e.g., regions R1, R2, and / or R3) can be used to extract a threshold T (i.e., a classification feature) that represents the average of the luminance samples along the single line or column. The threshold T may be determined incrementally based on a histogram of the luminance samples, which may be incrementally constructed when looping 945 through the reference samples. The histogram is updated based on sequentially acquired reference samples while looping 945 through the reference samples, and the threshold T is recalculated based on the updated histogram. In this way, one loop 945 can be used to calculate T and to calculate 920, 930 the autocorrelation matrices (A1 and A2) and cross-correlation vectors (B1 and B2). Furthermore, to further reduce complexity (in terms of computation speed and memory storage), a quantized histogram may be used.

[0057]

[67] As disclosed herein, the representation of the model parameters may be simplified. To that end, in one embodiment, the model parameters generated 950, 960, 970 by the process of FIG. 9 may be rescaled 955, 965, 975. p CC-based prediction of (i,j) may be implemented as follows:

number

[0058]

[68] In one aspect, to reduce complexity, the model parameters Φ are rescaled 955, 965, 975 so that the rescaled parameters can be represented with fewer bits. That is, the rescaled parameters can be stored using an n-bit buffer. As a result, applying the parameter vector to predict chroma samples (the operation involved in equation (16)) can be performed using a buffer with a desired maximum n-bit number less than 64 bits. For example, it may be desirable to limit n to a 32-bit or 16-bit signed integer so that the processor can use a SIMD accelerator (such as AVX or SSE).

[0059]

[69] The value of nd is the sum of nb, SIMD code, and s max where s max is the maximum value of the elements in the observation vector s. For example, if the prediction in (16) is performed using repeated sums of weighted samples, the value of nd can be determined so that the following requirement is met: a)s max Φ m should be within the nb-bit range for m={1,…,M}. b)

number

number

[0060]

[70] The above requirements can be met with a SIMD instruction set. For example, adapting a parameter vector to predict chroma samples can be achieved by:

number

number

[0061]

[71] Rescaling of the coefficients can be performed by right shifting as follows: Φ m >>shC or (Φ m +(1<<(shC-1)))>>shC (18) The value of shC should be less than or equal to sh. Furthermore, in equation (17), sh is replaced by (sh-shC). In one aspect, if the value of shC is greater than sh, default parameters may be used, e.g., Φ for m=1,...,M-1. m =0, and Φ M =1, sh=0}. In another embodiment, if the value of shC is greater than sh, sh is replaced by (shC-sh) and the right shift is replaced by a left shift.

[0062] 11 is a flowchart of an exemplary method 1100 in which aspects of the present embodiment may be implemented. The method 1100 may be employed by the encoder 200 or the decoder 300. According to the method 1100, at step 1110, video data including data representing a video data region is obtained. At step 1120, the method calculates a model applicable to CC-based prediction of chroma samples from the video data region. As mentioned above, the model may be, for example, one of CCLM, CCCM, GL-CCCM, or a combination thereof. The calculation of the model includes, at step 1130, processes 920, 930 simultaneously accumulating data for a first model Φ1 and data for a second model Φ2 using a single loop 945 through reference samples selected from the video data. As mentioned above, the accumulated data for the first model may include a first autocorrelation matrix and a first cross-correlation matrix, and the accumulated data for the second model may include a second autocorrelation matrix and a second cross-correlation matrix.

[0063]

[73] Further, the accumulated data of the first model and the second model are calculated based on the reference samples according to their respective classification into the first class or the second class in step 1130. Features may be extracted based on the reference samples using a single loop 945 through the reference samples to classify the reference samples into the first class or the second class. In one aspect, features may be extracted based on a subset of the reference samples. In another aspect, a histogram may be constructed, and the histogram may be updated based on sequentially obtained reference samples while looping 945 through the reference samples, and features may be recalculated based on the updated histogram. For example, the extracted feature may be a threshold T calculated based on the reference samples.

[0064]

[74] Furthermore, the method 1100 can combine 950 the accumulated data of the first model with the accumulated data of the second model into combined data of a single model Φ. In one aspect, the first model Φ and the second model Φ are used for multi-model prediction of chroma samples. In this case, application of the first model Φ and the second model Φ can be performed in a single loop 1070 through samples of the video domain, and the first model and the second model are applied 1040, 1050 to predict chroma samples according to their classification into a first class or a second class, as shown in FIG. 10. In a further aspect, the parameters of the first model Φ, the second model Φ, and the single model Φ can be rescaled to a lower bit-size representation (e.g., 32 bits or 16 bits) to reduce the computational complexity involved in applying the models, as described above.

[0065]

[75] Several aspects and embodiments of this disclosure have been described that provide at least the following outputs and results, including all combinations across different claim categories and types: Encoding syntax elements into the coded video data that may enable a decoder to decode the coded video data according to any of the aspects described herein. A bitstream containing one or more of the described syntax elements or variations thereof. A bitstream can be any set of data, whether transmitted, stored, or otherwise made available. · Creating, transmitting, receiving, and / or decoding bitstreams. An electronic device (e.g., a television, a set-top box, a mobile phone, or a tablet) that tunes to a channel (e.g., using a tuner) to receive the bitstream or receives the bitstream wirelessly (e.g., using an antenna). The electronic device decodes syntax elements from the bitstream and optionally displays the resulting image (e.g., using a monitor, screen, or any other type of display). Various other generalized and detailed outputs, results, implementations, and claims are also supported and contemplated throughout this disclosure.

[0066]

[76] Various methods are described herein, each of which includes one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Additionally, terms such as “first,” “second,” etc. may be used in various embodiments to modify elements, components, steps, operations, etc. (e.g., “first decode” and “second decode,” etc.). The use of such terms does not imply any ordering to the modified operations unless specifically required. Thus, in this example, the first decode need not be performed before the second decode, but may occur before, during, or during an overlapping period with the second decode.

[0067]

[77] Various methods and other aspects described herein can be used to modify modules, such as those of the video encoder 200 and decoder 300 shown in Figures 2 and 3. Furthermore, the aspects are not limited to a particular standard (such as VVC or HEVC) and may, for example, be applied to other standards and recommendations, and extensions of any such standards and recommendations. Unless otherwise indicated or technically precluded, the aspects described herein may be used individually or in combination.

[0068]

[78] Various numerical values ​​are used in this application. The specific values ​​are for illustrative purposes and the described aspects are not limited to these specific values.

[0069]

[79] Various implementations involve decoding. As used herein, "decoding" can encompass all or part of the processes performed on a received encoded sequence to produce a final output suitable for display, for example. In various embodiments, such processes include one or more of the processes typically performed by a decoder (e.g., entropy decoding, inverse quantization, inverse transform, and differential decoding). Whether the phrase "decoding process" is intended to refer specifically to a subset of operations or to the broader decoding process generally should be clear based on the context of the specific description and is believed to be well understood by those skilled in the art.

[0070]

[80] Various implementations involve encoding. Similar to what was discussed above with respect to "decoding," as used in this application, "encoding" can encompass, for example, all or part of a process performed on an input video sequence to produce a coded bitstream. Furthermore, the terms "reconstructed" and "decoded" can be used interchangeably, the terms "encoded" or "coded" can be used interchangeably, and the terms "image," "picture," and "frame" can be used interchangeably. Typically, but not necessarily, the term "reconstructed" is used on the encoder side, and the term "decoded" is used on the decoder side.

[0071]

[81] Please note that the syntactic elements used here are descriptive terms and therefore do not preclude the use of other syntactic element names.

[0072]

[82] This disclosure has described various information, e.g., syntax, that may be transmitted or stored. This information may be packaged or arranged in various ways, including methods common to video standards, such as placing information in an SPS, PPS, NAL unit, header (e.g., a NAL unit header or slice header), or SEI message. Other methods are also possible, including methods common to system-level or application-level standards, such as signaling information in one or more of the following: a. SDP (Session Description Protocol), which is a format for describing multimedia communication sessions for purposes such as session announcements and session invitations, as described in, for example, RFCs, and is used in conjunction with RTP (Real-time Transport Protocol) transport. b. DASH Media Presentation Description (MPD) Descriptor, used for example in DASH and transmitted over HTTP. A descriptor is associated with a Representation or a collection of Representations to provide additional characteristics to the content Representation. c. RTP header extensions used, for example, during RTP streaming. d. ISO Base Media File Format, for example as used by OMAF, which uses boxes (also called "atoms" in some specifications) that are object-oriented building blocks defined by a unique type identifier and a length. e. HLS (HTTP Live Streaming) manifests transmitted over HTTP. A manifest can be associated with a version or collection of versions of content, for example, and can provide characteristics of the version or collection of versions.

[0073]

[83] Implementations and aspects described herein may be embodied in, for example, a method or process, an apparatus, a software program, a data stream, or a signal. Even if discussed only in the context of a single implementation form (e.g., discussed only as a method), the implementation of the discussed features may also be implemented in other forms (e.g., an apparatus or a program). An apparatus may be implemented, for example, in appropriate hardware, software, and firmware. A method may be implemented, for example, in an apparatus, for example, a processor, which generally refers to a processing device including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, for example, computers, mobile phones, portable / personal digital assistants ("PDAs"), and other devices that enable communication of information between end users.

[0074]

[84] References to "in one aspect" or "in one embodiment" or "in one implementation" and other variations thereof mean that a particular feature, structure, characteristic, etc. described in connection with that aspect / embodiment / implementation is included in at least one embodiment. Thus, appearances of the phrases "in one aspect" or "in one embodiment" or "in one implementation" in various places throughout this application, as well as any other variations, do not necessarily all refer to the same embodiment.

[0075]

[85] This application may also refer to “determining” various pieces of information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from memory.

[0076]

[86] Additionally, this application may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, transferring information, replicating information, calculating information, determining information, predicting information, or estimating information.

[0077]

[87] Additionally, this application may refer to "receiving" various information. Receiving, like "accessing," is intended to be a broad term. Receiving information can include, for example, one or more of accessing the information or retrieving the information (e.g., from a memory). Furthermore, "receiving" typically includes, in some manner, during operation, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information.

[0078]

[88] It should be understood that the use of any of the following terms " / ," "and / or," and "at least one of" is intended to encompass the selection of only the first listed alternative (A), or the selection of only the second listed alternative (B), or the selection of both alternatives (A and B), for example, in the case of "A / B," "A and / or B," and "at least one of A and B." As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C," such phrases are intended to encompass the selection of only the first listed alternative (A), or the selection of only the second listed alternative (B), or the selection of only the third listed alternative (C), or the selection of only the first and second listed alternatives (A and B), or the selection of only the first and third listed alternatives (A and C), or the selection of only the second and third listed alternatives (B and C), or the selection of all three alternatives (A, B, and C). This may be expanded as many times as there are listed items, as would be apparent to one of skill in the art.

[0079]

[89] Also, as used herein, the term "signal" refers, among other things, to indicating something to a corresponding decoder. For example, in a particular embodiment, an encoder signals quantization parameters for inverse quantization. In this way, in one embodiment, the same parameters are used at both the encoder and decoder sides. Thus, for example, an encoder may transmit certain parameters to a decoder (explicit signaling) so that the decoder may use the same certain parameters. Conversely, if the decoder already has certain parameters and other parameters, signaling may be used without transmission (implicit signaling) to simply allow the decoder to know and select certain parameters. By avoiding the transmission of any actual function, bit savings are achieved in various embodiments. It should be understood that signaling may be implemented in various ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder. While the above refers to the verb form of the word "signal," the word "signal" may also be used as a noun herein.

[0080]

[90] As will be apparent to those skilled in the art, multiple implementations can produce a variety of signals formatted to carry information that can be, for example, stored or transmitted. The information can include, for example, instructions for performing a method or data produced by one of the described implementations. For example, a signal can be formatted to carry a bitstream of the described embodiments. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is well known. The signal can be stored on a processor-readable medium.

Claims

1. obtaining video data including data representing a video data region; calculating a model used for cross-component based prediction of chroma samples from said video data domain; wherein the calculation comprises: using a first loop through selected reference samples from the video data to simultaneously accumulate data for the first model and the second model; A method comprising:

2. 2. The method of claim 1, further comprising: rescaling at least one parameter of the first model and the second model, wherein the rescaling comprises rescaling the parameter to a lower bit-size representation.

3. 3. The method of claim 1, further comprising applying the first model and the second model using a second loop through samples of the video domain, wherein the first model and the second model are applied to predict the saturation samples according to their classification into a first class or a second class.

4. The method of any one of claims 1 to 3, further comprising combining the accumulated data of the first model and the second model into combined data of a single model.

5. The method of claim 4 , further comprising rescaling at least one parameter of the single model, wherein the rescaling comprises rescaling the parameter to a lower bit-size representation.

6. 6. The method according to claim 1, wherein the accumulated data of the first model and the second model are generated based on the reference samples according to their respective classification into a first class and a second class.

7. extracting, based on the reference samples, one or more features used in the classification of the reference samples into the first class and the second class using the first loop through the reference samples. The method of claim 6 further comprising:

8. The method of claim 7 , wherein the extracted one or more features are extracted based on a subset of the reference samples.

9. said extracting said one or more features further comprising: updating a histogram based on the reference samples sequentially acquired while looping through the reference samples in the first loop; recalculating the one or more features based on the updated histogram; and The method of claim 7, comprising:

10. The method of claim 7 , wherein the one or more features include a threshold calculated based on the reference sample.

11. 11. The method of claim 1, wherein the accumulated data for the first model comprises a first autocorrelation matrix and a first cross-correlation vector, and the accumulated data for the second model comprises a second autocorrelation matrix and a second cross-correlation vector.

12. 12. The method of any one of claims 1 to 11, wherein the model is one of a cross-component linear model (CCLM), a convolutional cross-component model (CCCM), a gradient and position-based convolutional cross-component model (GL-CCCM), or a combination thereof.

13. The method of any one of claims 1 to 12, implemented by a video encoder.

14. The method of any one of claims 1 to 12, implemented by a video decoder.

15. at least one processor; and a memory storing instructions that, when executed by at least one processor, cause the apparatus to: obtaining video data including data representing a video data region; and calculating a model used for cross-component based prediction of chroma samples from the video data domain, said calculation comprising: using a first loop through selected reference samples from the video data to simultaneously accumulate data for the first model and the second model; 1. An apparatus comprising:

16. The instructions may include: further rescaling at least one parameter of the first model and the second model, wherein the rescaling includes rescaling the parameter to a lower bit size representation.

16. The apparatus of claim 15.

17. The instructions may include: and applying the first model and the second model using a second loop through samples of the video domain, the first model and the second model being applied to predict the chroma samples according to their classification into a first class or a second class.

17. Apparatus according to claim 15 or 16.

18. The instructions may include: Combining the accumulated data of the first model and the second model into combined data of a single model. The apparatus according to any one of claims 15 to 17, further comprising:

19. The instructions may include: further rescaling at least one parameter of the single model, wherein the rescaling comprises rescaling the parameter to a lower bit size representation.

20. The apparatus of claim 18.

20. 20. The apparatus according to claim 15, wherein the accumulated data of the first model and the second model are generated based on the reference samples according to their respective classification into a first class and a second class.

21. The instructions may include: extracting, based on the reference samples, one or more features used in the classification of the reference samples into the first class and the second class using the first loop through the reference samples. The apparatus of claim 20 , further comprising:

22. The apparatus of any one of claims 15 to 21, wherein the at least one processor that executes the instructions is of a video encoder.

23. The apparatus of any one of claims 15 to 21, wherein the at least one processor that executes the instructions is of a video decoder.

24. 1. A non-transitory computer-readable medium containing instructions executable by at least one processor to implement a method, the method comprising: obtaining video data including data representing a video data region; and calculating a model used for cross-component based prediction of chroma samples from the video data domain, said calculating comprising: using a first loop through selected reference samples from the video data to simultaneously accumulate data for the first model and the second model; 1. A non-transitory computer-readable medium, comprising: